Per-Agent Calibration Scores: England vs France
Result: ENG 6-4 FRA (90') β ENG win bronze in 90' (no ET, no pens)
Auto-scored baseline: winner_90 + advance + scoreline = varies (see per-agent) | who_scores tracked separately
π¨ CALAMITOUS ENSEMBLE MISS. All three writers were UNANIMOUS: FRA 2-1, FRA advance ~55%. Actual: ENG 6-4 β wrong winner (all 3), wrong scoreline by a factor of ~3Γ (10 goals vs modal 3). The two theses that carried every prediction β "MbappΓ© retained = decisive edge" and "ENG attack gutted (F4a)" β both INVERTED. The 3rd-place context produced the opposite of the modeled dynamic: the deflated SF-loser (FRA) underperformed while ENG's hungry rotated backups (Saka Γ2, Rice, Konsa) + a benched-star sub winner (Bellingham 90'+8') ran riot. Heat (F11, 30-32Β°C) fired 3-4Γ harder than modeled.
β οΈ SCORELINE AUTO-SCORE IS A PARSER ARTIFACT β read this before interpreting the bands.
- g4 / g5 scoreline = 15/25 is spurious. Their unlabeled
"2-1"(intended = France 2-1, FRA win) is regex-parsed by the validator as teamA(England) 2-1 β "win" β matches the actual ENG win β 15. g4 even labels"2-1_ENG"separately, proving unlabeled"2-1"meant FRA. The validator cannot read that intent.- mm3 scoreline = 0/25 is spurious for the opposite reason. mm3 used a capitalized
"Other"key (18%), and the validator's scoreline filter only drops lowercase"other"(k !== "other"), so"Other"survived as the modal β no digit-match β 0. mm3's top specific scoreline was"FRA 2-1"at 12% β the identical prediction g4/g5 made.- Artifact-corrected "true" picture: all three modal-predicted FRA 2-1 with the WRONG winner (FRA) and WRONG margin (1 vs actual 2). Honest scoreline for all three = 0/25 β auto = 10/50 β all three β 27-28/80 β 34% β all Band F. The mm3(F)/g4(D)/g5(D) split below is purely a parser artifact, not a quality difference. The findings + model-breaking scores (which I control) are identical/symmetric across all three because the reasoning was genuinely uniform.
mm3
Findings: 11/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | F4a underdog-side INVERSE β ENG gutted (Kane+Bellingham+Pickford) β FRA gains ~3pt (Β§6.5 Step 6) | β | β | 0 | The most sophisticated gutting framing of the three (mm3 recognized F4a is usually favorite-side and tried to invert it) β but direction WRONG. ENG's backups scored 6 and WON. Brief learning #4: "Rotation gutting β F4aβ¦ backups out-performed stars." |
| 2 | F11 heat fully applies β both Tier-2 (rotated) β full penalty; open-game force, no draw compound (Β§6.5 Steps 3-4) | β | β | 3 | Direction clean: 30-32Β°C confirmed, no draw (ENG won in 90'), open game. But magnitude catastrophically under-modeled β mm3 projected a "2-3 goal scenario" / "3+ goal finale"; actual 10 goals (~3-4Γ). mm3 was the closest to anticipating a high-scoring finale, but still 3Γ short. |
| 3 | F12 = N/A, 0% discount β dead rubber, no conservation (Β§6.5 Step 2) | β | β | 5 | Confirmed: brief F12 = N/A (dead rubber, no conservation). Clean. |
| 4 | Henderson GK = SWE-negative / ENG defensive liability β untested (0 KO min), capped confidence (Β§6.5 Step 6 + F14) | β | β | 3 | Direction RIGHT β Henderson conceded 4 (liability confirmed). Magnitude wrong: the liability was symmetric (Maignan also conceded 6) and didn't flip the match to FRA; ENG outscored it. "ENG-only liability" thesis was half the story. |
Model-breaking: 6/10
Flagged the π¨ MODEL-BREAKING GK CHANGE (PickfordβHenderson) as the headline β exactly the model-breaker lineups.md flagged with π¨. Caught the 8/11 ENG + 6/11 FRA rotation fully. Staked F11 heat applies fully. Predicted 1-3 yellows vs actual 0 β direction right on low intensity, but the extreme "most lenient match of the tournament" (dead rubber = no tension) was under-modeled. Did NOT anticipate the 3rd-place motivation inversion (a reasoning-direction error, scored in finding #1) nor the 10-goal magnitude. The catastrophe was a direction/thesis failure, not a detection failure β all structural breakers (GK, lineup, weather) were caught.
TOTAL: 27 / 80 β 34% β Band F
(auto 10 = w90 10 + adv 0 + scor 0[artifact] + findings 11 + MB 6) Β· who_scores 12/20 (MbappΓ© #1 FRA β brace + Saka #2 ENG β; Rogers #1 ENG β). Artifact-corrected "true" band = F (uniform with g4/g5).
g4
Findings: 11/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | F4a fires for England β gutted creative spine (Kane main striker + Bellingham main creator + Pickford GK) + known-unknown in goal | β | β | 0 | Standard gutting framing β direction WRONG. ENG's "gutted" attack scored 6. Brief learning #4 explicit: the F4a framework does not apply to voluntary 3rd-place rotation. |
| 2 | F11 heat applies β Hard Rock open-air, heat-index 33-38Β°C; pressing degrades after 60β² | β | β | 3 | Direction right (heat fired, pressing collapsed). Magnitude badly wrong: projected_xg 1.4 / modal 2-1 (β3 goals) vs actual 10. ~3Γ under. |
| 3 | 3rd-place = open/goals β lower draw band (~20%), wider scoreline bands, no clean sheet, profligacy > clean sheet | β | β | 3 | Direction right (open game, no draw, no clean sheet for either). But g4 used this to lock a 2-1 modal β catastrophically under-weighting the goal TOTAL despite correctly calling "open." Left 19% on "other" but modal stayed low. |
| 4 | F12 = 0 β conservation_discount = 0 (dead rubber) | β | β | 5 | Confirmed. Clean. |
Model-breaking: 5/10
Caught the GK change but bundled it into F4a ("known-unknown in goal") rather than flagging it as a discrete model-breaker β weaker model-breaker emphasis than mm3/g5. Caught the rotation + heat (TL;DR "Heat applies"). Predicted 1-3 yellows vs actual 0. Same structural detection as the others, but the most generic framing β no explicit model-breaker callout, and the "goalkeeping gap = swing variable" thesis swung the wrong way (it was symmetric, not FRA-favoring).
TOTAL: 41 / 80 β 51% β Band D
(auto 25 = w90 10 + adv 0 + scor 15[artifact] + findings 11 + MB 5) Β· who_scores 10/20 (MbappΓ© #1 FRA β brace; Gordon #1 ENG β). β οΈ The D band is entirely the spurious scoreline=15 (parser read "2-1" as ENG 2-1). Artifact-corrected "true" total = 26/80 = 33% = Band F (uniform with mm3/g5).
g5
Findings: 11/20
| # | Finding staked | Direction | Magnitude | Score | Note |
|---|---|---|---|---|---|
| 1 | "Model-breakers STACK FOR FRANCE; motivation asymmetry IS the match" β Deschamps retained MbappΓ©/Olise/DouΓ©/Rabiot (wants to win); Tuchel benched stars (development fixture) | β | β | 0 | The most explicitly-committed central thesis of the three β and the most decisively INVERTED. Brief learning #1: "3rd-place context INVERTS the retained-stars thesis." The asymmetry cut the other way: ENG's hungry backups wanted it more; FRA's retained stars (fresh off losing a final spot) underperformed. |
| 2 | F11 heat applies (MEDIUM confidence β flagged missing NWS forecast) β pressing degrades, FRA amplifier | β | β | 3 | Direction right (heat fired). The honest "MEDIUM confidence / no archived NWS" flag was a transparency plus β but the magnitude was 3Γ under (actual 10 goals), and it amplified ENG as much as FRA, not FRA-only. |
| 3 | ENG GK change = SWE-negative (R32-5f), not draw-forcer β tilts xG-conversion toward FRA | β | β | 3 | Direction right (Henderson conceded 4 β SWE-negative confirmed) and the cleanest GK framing of the three (correctly rejected the draw-forcer read). Magnitude wrong: the tilt was symmetric (Maignan conceded 6); didn't produce a FRA win. |
| 4 | F12 = 0 β conservation_discount = 0 (dead rubber) | β | β | 5 | Confirmed. Clean. |
Note: g5 also explicitly staked F22 blowout EXCLUSION ("no pressing-side openerβ¦ blowout tail ~12% IF Henderson collapses"). Direction defensible on MARGIN β ENG 6-4 is a 2-goal margin, not a 4+ blowout β so F22-as-margin didn't fire. But g5 used the exclusion to cap scorelines, contributing to the 10-goal under-weight; and the Henderson-collapse conditional it down-weighted actually fired (conceded 4). Not separately scored (5-slot cap); noted as a near-miss that leaned the wrong way on goal-total.
Model-breaking: 6/10
Model-breaker detection lane β flagged the GK change prominently (SWE-negative, R32-5f citation), caught the rotation asymmetry, staked F11 heat, and honestly flagged the MEDIUM-confidence weather gap (missing NWS forecast) rather than bluffing. Predicted 2-3 yellows vs actual 0. Detection of structural breakers was strong (g5's specialty). BUT the model-breaker interpretation β "stack for FRA" β was inverted (scored in finding #1), and the F22 blowout exclusion under-weighted the high-total tail. The decisive material outcome (ENG rout) was unanticipated, but the headline breakers (GK, rotation, weather) were all caught.
TOTAL: 42 / 80 β 53% β Band D
(auto 25 = w90 10 + adv 0 + scor 15[artifact] + findings 11 + MB 6) Β· who_scores 10/20 (MbappΓ© #1 FRA β brace; Gordon #1 ENG β, Rogers #2 ENG β). β οΈ The D band is entirely the spurious scoreline=15 (parser read "2-1" as ENG 2-1). Artifact-corrected "true" total = 27/80 = 34% = Band F (uniform with mm3/g4).
Cross-agent notes
- All three are genuinely at the SAME analytical level. Identical modal winner (FRA), identical modal scoreline (FRA 2-1), identical advance (FRA ~55%), identical miss. The findings (11/20) and near-identical model-breaking (5-6/10) scores reflect this uniformity. The D/F band split in the JSON is 100% a scoreline-parser artifact (g4/g5 unlabeled "2-1" read as ENG 2-1; mm3 capitalized "Other" not filtered). Artifact-corrected, all three = Band F (~34%).
- The two load-bearing theses both inverted: (1) "MbappΓ© retained = decisive edge" β MbappΓ© scored a brace but FRA lost 4-6; ENG's distributed backup attack (Saka Γ2, Rice, Konsa, Bellingham sub) was MORE clinical. (2) "ENG attack gutted (F4a)" β the "gutted" XI scored 6; brief learning #4 is explicit that F4a does not apply to voluntary 3rd-place rotation.
- F11 heat: universal direction-right, universal magnitude-fail. All three staked heat applies (30-32Β°C confirmed) β open game, no draw. All three under-weighted the goal total ~3Γ (modal 2-1 / xG 1.4-1.7 vs actual 10 goals). The "3rd-place = more open, more goals" pattern fired 3-4Γ harder than any agent modeled β the draw band was correctly cut (no draw), but the goal TOTAL was never up-weighted to the 30-40% it warranted.
- GK liability was symmetric, not FRA-favoring. All three correctly read Henderson as a liability (SWE-negative) β and he DID concede 4 (direction right). But the same dead-rubber defensive collapse hit Maignan too (conceded 6). The "ENG-only liability tilts to FRA" framing was half the story; in a 3rd-place shootout BOTH GKs concede. Brief learning #3.
- MbappΓ© #1 FRA scorer: unanimous HIT (brace, 48'+66'). But the ENG modal scorers (Rogers/Gordon β neither started prominently / neither scored) all missed. Actual ENG scorers: Rice, Konsa (CB!), Saka Γ2, Bellingham (sub) β none was any agent's ENG #1. Only mm3 had Saka in its list (ENG #2, 12%) β the lone extra who_scores hit.
- Ref: universal partial miss. All three predicted 1-3 (mm3/g4) or 2-3 (g5) yellows β 0 cards. Valenzuela's leniency was extreme in a tension-free dead rubber (FRA 14 fouls, ENG 8). Direction right (low), magnitude off (zero).
- Process integrity held. This was flagged pre-match as "the lowest-confidence pre-match prediction of the tournament" (lineups.md); honest spreads were wide-ish (36-39% ENG). The framework correctly signaled uncertainty β but uncertainty is not prescience, and a 10-goal ENG rout lived in a tail none of the three weighted. This is the worst ensemble miss of the tournament and a clean demonstration that the 3rd-place fixture needs its own context rule (motivation inversion + goal-total up-weight) for future cycles.
Scores Summary (machine-readable β DO NOT EDIT)
{
"match": "eng-vs-fra",
"result": { "score90": "6-4", "final": "6-4", "outcome90": "win", "advance": "a", "route": "90" },
"agents": {
"mm3": { "auto": 10, "winner_90": 10, "advance": 0, "scoreline": 0, "findings": 11, "model_breaking": 6, "total": 27, "pct": 34, "band": "F", "who_scores": 12 },
"g4": { "auto": 25, "winner_90": 10, "advance": 0, "scoreline": 15, "findings": 11, "model_breaking": 5, "total": 41, "pct": 51, "band": "D", "who_scores": 10 },
"g5": { "auto": 25, "winner_90": 10, "advance": 0, "scoreline": 15, "findings": 11, "model_breaking": 6, "total": 42, "pct": 53, "band": "D", "who_scores": 10 }
}
}