Per-Agent Calibration Scores: England vs Mexico
Result: 3-2 at 90' β 3-2 final | ENG win in 90' (advance route: 90')
Scorers: ENG Bellingham 36', Quansah 54', Kane 60' / MEX QuiΓ±ones 42', JimΓ©nez 69'
Headline: ENG won a 5-goal thriller with 33% possession β the favorite played reactive at the Azteca and was clinical (5 SoT, 3 goals, 100% conversion from 1.55 xG). The 0-GA fortress was broken (3 conceded, more than the entire group stage + R32 combined). Modal across all three agents was 1-1 / 0-0 / ENG 1-0 β the low-scoring draw thesis was unanimous and wrong.
Auto-scored baseline: winner_90 + advance + scoreline (mm3 25/50 Β· g4 40/50 Β· g5 20/50) | who_scores separate (mm3 20/20 Β· g4 20/20 Β· g5 20/20)
Scoring note. All three writers staked the same four core findings (0-GA/hot-GK block, F12 conservation, Kane milestone, F11 altitude), so they are scored on a comparable set. Differentiation comes from (a) how each hedged the 0-GA β g4 was the only one to flag it as "untested vs Tier 1" with an explicit +2% ENG quality caveat, which is exactly what broke; (b) g5's isolated F13 sterile-domination bet (the only agent to stake it fires, modal 0-0) β the most emphatically wrong call of the three; (c) g5's possession-flip cap (58%) β the only writer to frame the favorite as reactive, closest to the 33% actual. The finishing-variance reversal (ENG 1.55 xG β 3 goals = clinical; MEX 1.94 xG β 2 = normal) is a genuine model-limitation (cf. BRA-NOR), not a writer failure, and is treated as unanticipatable here.
mm3
Findings: 13/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | F4/R32-5f hot-GK Rangel Γ0.5-0.6 + 0-GA = "model anchor in MEX's favorβ¦ a system product, NOT a fluke, like Vozinha/Room" (Β§6.5 4b/4c, Β§10 #2) | β | β | 0 | Rangel conceded 3 from 1.55 xG; the 0-GA streak "broke spectacularly." Brief learning #1: "0 GA vs weak opp (RSA/KOR/CZE/ECU) was NOT predictive vs a Tier-1 attack β opponent-quality-inflated." mm3's emphatic conviction ("NOT a fluke") was the wrong side of exactly this bet. (mm3 did flag "MEX's first Tier-1 testβ¦ the largest single unknown" as a hedge, but the modal stake was Γ0.5-0.6 applied + draw pushed to 26% β bet it holds.) |
| 2 | F12 = 0% both (conservation_discount=0; R16 must-win + Tuchel aggressive + Aguirre farewell + Kane milestone) | β | β | 5 | F12 = 0% both confirmed in brief. Cleanest call. |
| 3 | R32-5g Kane milestone = goals AND minutes (record-chaser, 6 in 4, tied Klose at 16, chasing Messi 18+) | β | β | 5 | Kane scored 60' β now 17 WC goals, surpassed Klose. Record-chaser β goals validated again (6/6 in knockouts). |
| 4 | F11 altitude = largest single deviation factor, +7-10pt MEX / β9pt ENG, "ENG press degrades after 60'" (Β§6.5 3, Β§10 #1) | β | β | 3 | Altitude confirmed; ENG gave up 67% possession and could not impose at altitude β the press-degradation mechanism was real and directionally right. But ENG WON in 90' via clinical countering, so the predicted MEX edge / draw did not materialize. Direction right, consequence/magnitude wrong. |
Model-breaking: 5/10
mm3 caught the structural model-breakers cleanly: no GK change (both Pickford + Rangel confirmed), lineup matched the injury-sweeper projection, pen band 30-40% sized as a band (0 pens confirmed β within band, "tournament-regressed"), F12 = 0 confirmed. It also flagged the RB crisis as "the model-breaking matchup" (Β§6.3) and flagged the 0-GA as "the largest single unknown." But: mm3 BET the 0-GA holds (Γ0.5-0.6, "not a fluke"), so it did not anticipate the 0-GA actually breaking (the #1 material model-breaker); it predicted ~50/50 possession and did not anticipate the possession flip (ENG 33%); the Quansah goal (the flagged "crisis" position scoring) was unanticipated β mm3 staked edge MEX on the RB. Clinical finishing unanticipated. Net: easy/structural + pivotal-position flags caught; the two dynamic model-breakers (0-GA breaking, possession flip) missed. Partial.
TOTAL: 43 / 80 β 54% β Band D
(auto 25 + findings 13 + model-breaking 5) Β· who_scores 20/20 (separate)
g4
Findings: 16/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | F4/R32-5f hot-GK Rangel Γ0.55 = "model anchor β BUT untested vs Tier 1" + explicit +2% ENG "0-GA quality caveat (no Tier-1 opponent)" (TL;DR, R32-5f, Β§6.5 row 5) | β | β | 3 | The 0-GA broke (3 conceded) exactly as g4's hedge warned. Uniquely among the three, g4 flagged the 0-GA as opponent-quality-inflated and added a +2% ENG caveat for it. Direction right on the caveat β the prescient call of the match. But g4 still applied Γ0.55 + "genuine draw compound" + draw band 27%, so the magnitude was wrong (block broke, not held). Direction right, magnitude wrong. |
| 2 | F12 = 0 (conservation_discount=0) | β | β | 5 | F12 = 0% both confirmed. |
| 3 | R32-5g Kane milestone β goals (+3% ENG) | β | β | 5 | Kane scored 60'. 6/6 record-chasers in knockouts. |
| 4 | F11 altitude + 48h acclim = swing factor, "ENG pressing effectiveness degrades after ~60', ENG more measured in possession after the hour" (TL;DR, Β§6.5) | β | β | 3 | ENG at 33% possession confirms the press-degradation / measured-possession mechanism. But ENG won in 90' via countering β the predicted MEX swing did not decide it. Direction right, consequence wrong. |
Model-breaking: 6/10
g4 was the most alert to the key model-breaker. Its "MEX 0 GAβ¦ BUT untested vs Tier 1" hedge β plus the explicit +2% ENG "0-GA quality caveat (no Tier-1 opponent)" β was the only writer-level flag that the model's anchor might not hold, and it didn't (3 conceded). g4 also correctly down-weighted the pen factor ("NOT load-bearing"; 0 pens confirmed), confirmed GK/lineup, and called F12 = 0. Missed: the possession flip (predicted ~50%, actual ENG 33%) and did not flag Quansah as a scoring threat (framed RB only as an injury-backup risk, line 49). Clinical finishing unanticipated. Net: anticipated the #1 model-breaker risk best + caught the structural ones; missed the possession flip + finishing variance. Best of the three on model-breaking anticipation.
TOTAL: 62 / 80 β 78% β Band B
(auto 40 + findings 16 + model-breaking 6) Β· who_scores 20/20 (separate)
g5
Findings: 10/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | F13 sterile-domination FIRES (finding_13_fires=true; "the sterile-domination screen is live, not theoretical"; draw-forcer + ENG profligacy) | β | β | 0 | ENG scored 3 from 1.55 xG = CLINICAL, not sterile. A 5-goal thriller. F13 fired the opposite way. g5's headline bet (modal 0-0 at 14%) was the most emphatically wrong of the three. |
| 2 | F4/R32-5f hot-GK Rangel Γ0.4 + 0-GA = "textbook draw-forcer compound" (Step 4, no Tier-1 hedge on the defense) | β | β | 0 | Rangel conceded 3; the 0-GA broke. g5 staked it as a draw-forcer (most bullish β modal 0-0, hot-GK Γ0.4) with no "untested vs Tier 1" caveat on the defense. (g5 did hedge the Azteca fortress record as built "vs CONCACAF minnows," but not the 0-GA defense itself.) Block was opponent-quality-inflated. |
| 3 | F12 = 0 (conservation_discount=0) | β | β | 5 | F12 = 0% both confirmed. |
| 4 | R32-5g Kane milestone β goals (5/5 R32) | β | β | 5 | Kane scored 60'. |
Model-breaking: 5/10
g5's standout catch: the only writer to put a specific possession cap (58%) and explicitly frame the favorite as reactive β the possession flip (ENG 33%) was the match's #2 headline surprise (brief learning #2), and g5 anticipated its direction (if not magnitude) best of the three. g5 also confirmed GK/lineup, sized the pen band (0 pens), called F12 = 0, and caught the Kane milestone. But: g5 bet the 0-GA holds as a draw-forcer (no Tier-1 hedge on the defense) β missing the #1 material model-breaker (0-GA broke); staked F13 sterile-domination fires (the opposite materialized β ENG clinical); the clinical finishing was unanticipated. Net: one big catch (possession-flip direction) + one big miss (0-GA breaking) + the structural ones. Partial.
TOTAL: 35 / 80 β 44% β Band F
(auto 20 + findings 10 + model-breaking 5) Β· who_scores 20/20 (separate)
Cross-agent notes
- The 0-GA blind spot was shared β but hedged unevenly. All three weighted MEX's 0-GA (4 CS vs RSA/KOR/CZE/ECU) as predictive vs a Tier-1 attack. Brief learning #1 makes the miss explicit. Differentiation is honest: g4 flagged it "untested vs Tier 1" with a +2% ENG caveat (3/5); mm3 staked "NOT a fluke, system product" but flagged it "the largest single unknown" (0/5 β the modal conviction was wrong); g5 staked "textbook draw-forcer" + F13 fires + modal 0-0 with no defense-level hedge (0/5). g4's reasoning was best on the match's defining factor.
- The possession flip (ENG 33%) was unanticipated in magnitude by all three (predicted 50-58%), but g5 was closest β it was the only writer to cap ENG possession explicitly (58%) and frame the favorite as reactive ("a favorite with a lead plays reactive"). mm3 and g4 predicted ~50/50. g5's best call of the match; it lives in the model-breaking score, not findings (g5's modal findings were the hot-GK/F13 bets).
- Shared wins, all clean: F12 = 0 (all 5/5), Kane milestone (all 5/5 β 6/6 record-chasers in knockouts now), and both #1 scorers (all 20/20 who_scores: Kane + JimΓ©nez). The scorer layer was the pipeline's strongest component here.
- Don't double-penalize the outcome. g5's winner_90 (modal loss, 10/15) + scoreline (0-0, 0/25) are already docked by the auto-script. The F13-fires finding (0/5) judges the reasoning (staked sterile-domination fires) β a distinct miss from the wrong modal, scored separately per the rules. Likewise mm3's winner_90 was right (modal win = actual win); its findings miss is the 0-GA conviction, not the outcome.
- The finishing-variance reversal is a model-limitation, not a writer failure. The xG projections were roughly right on ENG (mm3 1.6 / g4 1.5 / g5 1.4 vs actual ENG 1.55); the scoreline miss was ENG's 100% SoT conversion (5 shots, 3 goals) β a clinical over-performance no writer is expected to call, analogous to the BRA-NOR xG reversal. Treated as unanticipatable; not held against the model-breaking scores.
- Quansah scoring (the RB "crisis" position) was a genuine surprise (brief learning #3: "the crisis became the advantage"). mm3 flagged the RB as the model-breaking matchup (edge MEX); g4 flagged it as an injury-backup risk; none anticipated it as a scoring source. Not credited to any agent.
Scores Summary (machine-readable β DO NOT EDIT)
{
"match": "eng-vs-mex",
"result": { "score90": "3-2", "final": "3-2", "outcome90": "win", "advance": "a", "route": "90" },
"agents": {
"mm3": { "auto": 25, "winner_90": 15, "advance": 10, "scoreline": 0, "findings": 13, "model_breaking": 5, "total": 43, "pct": 54, "band": "D", "who_scores": 20 },
"g4": { "auto": 40, "winner_90": 15, "advance": 10, "scoreline": 15, "findings": 16, "model_breaking": 6, "total": 62, "pct": 78, "band": "B", "who_scores": 20 },
"g5": { "auto": 20, "winner_90": 10, "advance": 10, "scoreline": 0, "findings": 10, "model_breaking": 5, "total": 35, "pct": 44, "band": "F", "who_scores": 20 }
}
}