Expected gain per round, in units of the initial bet, given the count when the bet goes down. The learned policy plays every hand. Whiskers show a 95% interval.
The same playing strategy with bets that follow the count. The 0–3 strategies sit out rounds at a count without a player edge (back-counting). The 1–3 strategies play every round. Bet levels were set on one simulation run and scored on another.
Every player at the table counts and plays the same learned strategy. Seat 1 acts first; later seats see the earlier players' cards before deciding. Units won per 100 rounds dealt.
Q(s,a) is the expected gain of action a, followed by optimal play. V(s) is the best Q. The advantage A(s,a) = Q(s,a) − V(s) is what that action costs relative to the best one. Hover or focus a cell for all actions.
Two-card decisions whose best action at some count differs from the count-blind play (all counts pooled), with at least 20,000 visits on both actions and a gain of at least 0.002 bets. Each line gives the count where the change starts, the learned counterpart of index tables such as the Illustrious 18.