Per-Agent Calibration Scores: Portugal vs Spain
Result: POR 0-1 ESP (90') β ESP advance in 90' (Merino 90'+1' sub winner) β ESP to QF
Auto-scored baseline: winner_90 + advance + scoreline, tracked per-agent below | who_scores 0/20 for all (separate)
Match shape: Predicted 50/50 coin-flip, modal 1-1 (unanimous). ESP won 1-0 in 90' via a 90'+1' sub winner. The 0-0 draw shape held for 90 minutes, broken at the death. ESP xG 1.77 vs POR 0.60 β ESP were the better side.
mm3
Findings: 20/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | F11 NEUTRALIZED β roof CLOSED β no heat/travel/altitude discount (TOML + Β§6.5 Step 3) | β | β | 5 | Brief: "Roof CLOSED, F11 neutralized. β Validated." Indoor neutral conditions confirmed. |
| 2 | F12 conservation = 0 β R16 + Ronaldo farewell + both attacking (TOML conservation_discount = 0) |
β | β | 5 | Brief: "F12 = 0% both." Both full-throttle, no conservation. Exactly as staked. |
| 3 | SimΓ³n 0 GA = confirmed-form real hot-GK (Γ0.5β0.6 weight β "real draw-forcer"); NOT opponent-inflated (TL;DR + Β§6.5 Step 6) | β | β | 5 | Brief key learning #1: "ESP's 0 GA HELD vs Tier 1! SimΓ³n clean sheet (0 GA holds vs Tier 1!)." mm3 sided with "real," validated. Draw-forcing mechanism held for 90' (0-0 at 90'). |
| 4 | F16 Ronaldo age-decline applies β demoted from ~28β30% β 22% scorer; R32-5g override applied but F16 discount STILL baked in (TOML finding_16_age_decline = "Ronaldo" + Β§6.5 Step 2/Who Scores) |
β | β | 5 | Ronaldo blanked (0 goals). Brief learning #3: "Ronaldo farewell β 0 goals. At 41, farewell milestone β minutes not goals (F16)." Direction right; discount was correct (reality was even more severe β 0). |
Note (minor): mm3 also added a "+1pt POR milestone amplifier" nudge in the synthesis table β a small probability lean that did NOT fire (Ronaldo scored 0). This is a minor wrong probability nudge, not a core staked finding; the modal POR-win outcome it contributed to is already docked in winner_90 (5/15). Not double-counted here.
Model-breaking: 10/10
All structural model-breakers caught and confirmed pre-match: β GK (Costa + SimΓ³n both confirmed, no change), β referee (Taylor pen band 15β25%, correctly regressed HIGH-WIDE β 0 pens, within band), β weather (roof closed, F11 neutralized), β lineup (FΓ©lix-over-LeΓ£o noted, accounted in Who Scores). No unanticipated structural break occurred. The decisive event (Merino 90'+1' sub) is a Finding 19 late-sub tail β anticipated as a category (mm3 flagged Ramos super-sub + Williams-sub, F18/F19) but Merino himself wasn't named; this is a Who Scores tail miss (scored separately, 0/20), NOT a structural model-breaker. Not docked here.
TOTAL: 5 (auto) + 20 (findings) + 10 (MB) = 35 / 80 β 44% β Band F
Honest read: mm3 had the strongest qualitative reasoning of the three (findings 20/20 β the two marquee findings of the match, SimΓ³n-real + F16-Ronaldo, both validated by the brief). The F-band is driven entirely by outcome luck: the modal POR-win (40%) + 1-1 scoreline missed, and the TOML scoreline list didn't contain the exact "0-1" string the auto-matcher needed (so scoreline 0 vs g5's 20 despite near-identical Β§7K tables). The findings layer correctly distinguishes mm3's sound reasoning from the unlucky outcome.
g4
Findings: 20/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | F11 FULLY NEUTRALIZED β roof CLOSED, no heat penalty, no altitude (Β§6.5 Step 3) | β | β | 5 | Brief: "Roof CLOSED, F11 neutralized. β " Exactly as staked. |
| 2 | F12 = 0 for both β knockout + farewell + settled XI (Β§6.5 Step 2) | β | β | 5 | Brief: "F12 = 0% both." Full effort both sides. |
| 3 | SimΓ³n 0 GA = real hot-GK (capped Γ0.25 open shape per R32-5f β +2% draw, +1% ESP); NOT treated as opponent-inflated (TL;DR + Β§6.5 Step 4/6 + F7 note "SimΓ³n 0 GA β real") | β | β | 5 | Brief learning #1: 0 GA held vs Tier 1. g4 gave ESP MORE credit than mm3 (Γ0.25 in open shape, +1% ESP + +2% draw) β this is precisely why g4 was the only agent to lean ESP and nail the winner. |
| 4 | F16 / finishing-quality gate β "Ronaldo farewell = minutes NOT goals (R32-5g)", POR partial finishing gate (+1% draw), Ronaldo β10pt scorer discount (TOML finding_16_age_decline = "Ronaldo" + Β§6.5 Step 5) |
β | β | 5 | Ronaldo blanked (0 goals); POR scored 0. Brief learning #3 confirms farewell β minutes not goals. The finishing-quality gate fired exactly as staked. |
Bonus context: g4 also staked F4a "single-absence, NOT gutting" for ESP wings (Williams+Pino out, β2pt only) β held (ESP won, spine intact, scored). g4 was the most ESP-leaning agent on every material call (SimΓ³n-real credit +1 ESP, finishing-gate against POR, momentum +2 ESP), which is why it was the only agent to get winner + advance + scoreline-margin right.
Model-breaking: 10/10
All structural model-breakers caught: β GK (Costa + SimΓ³n confirmed, no change), β referee (pen band 20β25% β 0 pens, within band), β weather (roof closed), β lineup (FΓ©lix-over-LeΓ£o explicitly profiled in Β§1 + accounted). g4's Match Flow Β§9 even flagged the 75β90' window as "high drama... late subs... golden goal" β the exact window Merino scored in β though Merino wasn't named (Who Scores tail, scored separately). No structural break missed.
TOTAL: 40 (auto) + 20 (findings) + 10 (MB) = 70 / 80 β 88% β Band B
Honest read: g4 was the clear winner on this match β the only agent to land winner + advance + scoreline-margin (auto 40/50, top of the three). Its qualitative reasoning was equally strong (findings 20/20), with the decisive edge being that it gave ESP real credit for the 0 GA block (Γ0.25 open-shape hot-GK β +1% ESP) and applied the POR finishing-quality gate. This is g4's documented strength profile firing: "best on clear favorites + scoreline precision" β here, best on correctly identifying the marginal ESP edge in a coin-flip.
g5
Findings: 10/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | F11 NEUTRALIZED β roof CLOSED, 36Β°C irrelevant, cleanest possible environmental input (Β§6.2 + Β§6.5) | β | β | 5 | Brief: "Roof CLOSED, F11 neutralized. β " Held. |
| 2 | F12 = 0 β knockout + farewell = zero conservation (Β§6.5 Step 2, empirical Γ1.5 β still 0) | β | β | 5 | Brief: "F12 = 0% both." Held. |
| 3 | ESP 0 GA = OPPONENT-INFLATED β "CPV/KSA/AUT, no Tier 1 attack tested yet, the ENG-MEX warning," "if the block bends... POR flips the tie" (g5's MARQUEE finding, staked 5Γ: TL;DR, Β§2 weakness, Β§6.1, Β§6.2, Watch line, match-flow 45β70') | β | β | 0 | DIRECTION WRONG. Brief learning #1 (verbatim): "g5 flagged it as opponent-inflated (the ENG-MEX warning). Wrong β SimΓ³n's defense is genuinely elite. The 0 GA record survived its first real test." POR scored 0. This was the match's central swing variable and g5 staked the wrong side β directly causing the POR-win modal + POR-advance miss. |
| 4 | R32-5g: Ronaldo record-chasing OVERRIDES F16 age-41 β Ronaldo #1 at 22%, "milestoneβgoals... R32 goal validates" (TOML finding_16_age_decline = "" + Β§6.5 Step 2 + Who Scores #1 + R32-5g note) |
β | β | 0 | DIRECTION WRONG. Ronaldo blanked (0 goals). Brief learning #3: farewell milestone β minutes not goals (F16). Age-decline WON over record-chasing β the opposite of g5's staked position. g5 explicitly demurred on F16 (finding_16_age_decline = ""), betting the override. The bet lost. |
Note: g5's two marquee findings (the very ones it leaned on hardest, repeated 5Γ each) were BOTH wrong, and they're exactly the two findings the brief calls out as validated learnings. g5 framed this match as "the 0-GA breaks vs first real attack" β the defining prediction was falsified. g5 also staked "sterile-domination risk returns" (ESP attack) β didn't fire (ESP xG 1.77, scored 1).
Model-breaking: 10/10
Structural model-breakers caught: β GK (Costa + SimΓ³n confirmed, no change), β referee (pen band 20β28% β 0 pens, within band), β weather (roof closed), β lineup (FΓ©lix LW noted). No real structural break missed β the model-breaking category scores whether real breakers were CAUGHT, and none materialized. Important nuance: g5's headline "model-breaker" flag β the ESP 0 GA as a breakable ENG-MEX-style record β was a FALSE POSITIVE (g5 over-flagged a break that didn't happen), not a missed breaker. That reasoning error is scored under findings (#3, 0pts) to avoid double-counting; it does not dock model-breaking, which g5 otherwise completed cleanly (no real breaker unanticipated). g5 did NOT name Merino (Who Scores tail, scored separately).
TOTAL: 30 (auto) + 10 (findings) + 10 (MB) = 50 / 80 β 63% β Band C
Honest read: g5's Band C is propped up by outcome luck β it alone listed the exact "0-1" scoreline string (scoreline 20/25) and had loss as #2 at 35% (winner_90 10/15) β despite staking the WRONG side on both marquee findings. The findings layer (10/20, the lowest of the three) correctly captures that g5's core reasoning was wrong on the two questions that defined this match. g5 is the project's default + "best model-breaker detector," but here its signature call (ESP 0 GA opponent-inflated) was the match's biggest analytical miss. The auto-score's generosity (exact scoreline) masks a genuinely weaker reasoning performance β which is exactly what the findings layer exists to expose.
Cross-agent summary
| Agent | Auto /50 | Findings /20 | MB /10 | Total /80 | Pct | Band | who_scores /20 |
|---|---|---|---|---|---|---|---|
| g4 | 40 | 20 | 10 | 70 | 88% | B | 0 |
| g5 | 30 | 10 | 10 | 50 | 63% | C | 0 |
| mm3 | 5 | 20 | 10 | 35 | 44% | F | 0 |
Key takeaways:
- g4 was the right agent for this match type (marginal favorite identification in a coin-flip) β it gave ESP real 0-GA credit + applied the POR finishing gate, landing winner + advance + scoreline-margin. Its documented "favorite precision" strength fired.
- The findings layer did its job. mm3 (auto 5, Band F) and g5 (auto 30, Band C) had inverted reasoning-vs-outcome luck: mm3's reasoning (findings 20/20) was far stronger than g5's (findings 10/20) β both got the marquee findings (SimΓ³n-real + F16-Ronaldo) that g5 got wrong β but g5's exact-scoreline luck pulled its total above mm3's. The findings column makes the reasoning-quality gap legible.
- g5's signature call failed. The "ESP 0 GA opponent-inflated / ENG-MEX warning" β staked 5Γ β was the match's biggest analytical miss and the direct cause of g5's POR-win modal. Brief learning #1 is a direct falsification.
- Merino (90'+1') beat all three β a Finding 19 late-sub winner, anticipated as a category (late-goal scenarios flagged by all) but not named (all listed Ramos/Williams/Ferran, not Merino). who_scores 0/20 across the board.
- No structural model-breakers materialized (GK confirmed, ref as profiled, roof closed, lineup as projected) β model-breaking 10/10 for all three; the decisive event was a tactical late sub, not a structural break.
Scores Summary (machine-readable β DO NOT EDIT)
{
"match": "por-vs-esp",
"result": { "score90": "0-1", "final": "0-1", "outcome90": "loss", "advance": "b", "route": "90" },
"agents": {
"mm3": { "auto": 5, "winner_90": 5, "advance": 0, "scoreline": 0, "findings": 20, "model_breaking": 10, "total": 35, "pct": 44, "band": "F", "who_scores": 0 },
"g4": { "auto": 40, "winner_90": 15, "advance": 10, "scoreline": 15, "findings": 20, "model_breaking": 10, "total": 70, "pct": 88, "band": "B", "who_scores": 0 },
"g5": { "auto": 30, "winner_90": 10, "advance": 0, "scoreline": 20, "findings": 10, "model_breaking": 10, "total": 50, "pct": 63, "band": "C", "who_scores": 0 }
}
}