guide contents
Baixada · Study · Field guidePart I · The solve, and how to read it
Chapter viii

How far to trust it

A solution is not a promise that every printed mix is equally polished. The lab now measures the important question at a node: how much can a perfectly adapting opponent gain from the strategy shown here? The answer is the quality badge; the older reach and self-loss numbers explain why a corner went bad.

§ 1

The quality badge is the headline

The badge is a per-node adversarial best-response gap, expressed in match-win percentage points. It compares the exported mix with a genuine opponent best response, rather than asking the solve to grade itself.

solid · ≤ 1 pp
The measured exploit is at most one match-win point. Read the mix normally.
caution · 1–5 pp
There is visible room for an opponent to gain. Use it as a clue, not a prescription.
weak · > 5 pp
The lab replaces the strategy number with ≈. This branch is not reliable strategy.

A BR table is shipped with the solved spot when this measurement is available. Older exports keep the legacy warning as a compatibility fallback.

§ 2

Recognising a weak corner

This is an actual 11×10 pilot continuation from the lab. The pale cells are not merely low probability: each has a measured best-response gap above 5 pp. Hover or tap a cell to see its own measurement.

hi \ lo555532AKJQ764
55010010010010010010099999999100
5506169999999999910010067
550577895989797989865
5506881869596989965
31006172809295999857
210097536575939660
A10010010061769568
K100100100647865
J100100656760
Q1001005856
71005147
610074
459

Hover or tap a cell. Solid cells carry trained mixes; the pale ≈ cells are flagged.

Plate VI Real lab data — 11×10 v4, after the mão de onze is accepted and a low-reach first-round continuation. 0 of the exact two-card grid’s 87 cells are weak by their measured BR gap, so the lab renders ≈ instead of a strategy number.

Don’t infer a tactic from an ≈ cell. Step back to the last solid decision, or explore a different line.

§ 3

Why an off-path mix gets noisy

Before the BR table, the lab's warning came from two diagnostics. They stay here because they explain how a corner went bad — they describe training coverage, not quality.

Lself(h)=50aσ(ah)(maxaq(ah)q(ah))L_{\mathrm{self}}(h) = 50\sum_a \sigma(a\mid h)\bigl(\max_{a'}q(a'\mid h)-q(a\mid h)\bigr)
σ(ah)\sigma(a\mid h)
how often the solve takes action a holding h
q(ah)q(a\mid h)
the solve’s equilibrium value after action a at holding h
5050
converts the solver’s −1…1 value scale into match-win percentage points

Self-loss asks whether the mix is close to its own best action under the solve's q values — an alarm, not an adversarial guarantee.

ρown(h)=downσ(adh)\rho_{\mathrm{own}}(h) = \prod_{d\,\in\,\ell_{\mathrm{own}}}\sigma(a_d\mid h)
own\ell_{\mathrm{own}}
the acting player’s earlier decisions on the line
ρown\rho_{\mathrm{own}}
how often that player’s own exported strategy walks this holding to the node

Near-zero own reach means the averaging almost never collected experience here; the legacy alarm uses ε = 0.001.

The cause is concrete: once a non-owner action's probability reaches zero, CFR stops descending into that branch — correct for a truly dominated action, but it freezes whatever early random values sit below. Low own reach records that history; the BR gap says whether the result is actually exploitable.

§ 4

What the certificate does — and doesn’t — say

The solve’s certificate is a whole-game, reach-weighted bound against exact best responses. It says the strategy profile is close to equilibrium on average across the deal tree. A tiny off-path branch contributes almost nothing to that average, so the certificate alone cannot certify every cell. That is why the per-node BR measurement belongs on the chart.

Trust the node’s own badge first. Use the global certificate to understand the solve as a whole, and self-loss plus own reach to understand how a warning arose.