Per-Agent Calibration Scores: France vs Morocco

Result: FRA 2-0 MAR (90') → FRA 2-0 (final) → FRA advance to SF (won in 90')
Scorers: Mbappé 60' (Doué assist), Dembélé 66' (Mbappé assist) | Maignan clean sheet
Referee: Facundo Tello — 1 yellow (Diop 63', = SF ban, moot as MAR eliminated), 0 pens
Auto-scored baseline: winner_90 15 + advance 10 + scoreline (mm3 0 / g5 15) | who_scores 10/20 each (Mbappé ✅, Ounahi ❌)

⚠️ g4 ABSENT. glm-5-turbo (g4) failed to rate-limit 3× and produced NO analysis output. No fra-vs-mar-2026-pure-football-analysis-g4.md file exists. g4 is scored as absent below and set to null in the machine-readable footer (validate-master renders it as ? ?).


mm3 (MiniMax-M3)

Findings: 20/20

# Finding staked Direction Magnitude Score Note
1 F17 4th flag fires — MAR go Díaz false-9, no specialist #9 in squad (En-Nesyri/Aguerd/Ezzalzouli/Ziyech all OUT) 5 Díaz scored 0; MAR scored 0 — the no-finisher thesis was the brief's headline validation (Learning #2). Staked as "limits MAR win conversion / MAR win −2"; actual MAR win = 0%. Fired as staked.
2 MAR R16 xG regression (0.79 xG → 3-goal over-performance reverts) 5 MAR 0 goals from ~0.5 xG vs FRA's elite 3-GA-in-5 defense — full regression, exactly the central MAR downgrade (Learning #3).
3 Bounou hot-GK ×0.25 in open 4-2-3-1 (R32-5f) — NOT a regular-time draw-forcer; edge lives at pens 5 Bounou conceded 2; FRA won in 90'; 25% draw band did NOT fire (Learning #4). Hot-GK edge correctly deferred to a pen phase that never arrived.
4 F8 possession-flip risk — R32-5c cap; "MAR's possession 4-2-3-1 may actually dominate the ball mid-game" (§7 weakness) 5 MAR 51.9% > FRA 48.1% — the flip fired (Learning #1, now 6/6 in KOs). mm3 explicitly raised it as a FRA weakness.

mm3 findings = 20/20. Also staked & validated: F11 neutralized (~26°C, T-60 verify flag → brief confirms net neutral); Diop SF-ban caution (materialized — Diop 63' yellow = SF ban, moot as MAR eliminated); Tello LOW-PEN 7-12% (0 pens ✅).

Honesty flag (not docked — captured in scoreline): mm3's modal scoreline was 1-1 (draw), which is internally inconsistent with its own findings #1–3 (MAR won't score / Bounou won't force a draw ⇒ FRA clean win). mm3 flagged every risk correctly but hedged its synthesis toward a cagey draw its own findings didn't support — it priced MAR at 29% win / 1 goal when the mechanism it identified pointed lower. That inconsistency is paid for in the auto scoreline (0/25, modal draw ≠ FRA win); the findings themselves held.

Model-breaking: 10/10

All 5 material model-breaker categories caught:

mm3 TOTAL: 25 (auto: w90 15 + adv 10 + scor 0) + 20 findings + 10 MB = 55/80 → 69% → Band C | who_scores 10/20 (Mbappé ✅ / Ounahi ❌)


g5 (zai/glm-5.2)

Findings: 20/20

# Finding staked Direction Magnitude Score Note
1 MAR false-9 pivot = model-breaker flag — no true finisher vs FRA's elite Saliba-Upamecano axis (3 GA/5) → lower MAR goal ceiling 5 Díaz 0; MAR 0 — g5 called it the headline model-breaker; "lower MAR goal ceiling" collapsed to a zero ceiling.
2 MAR R16 regression (0.79 xG → 3G "borrowed"; expect ~1 not 3) 5 MAR 0 goals — regression fired harder than the "~1" hedge; direction + thesis validated (Learning #3).
3 Bounou diluted ×0.25-0.3 open shape (R32-5f) — pen-phase edge only, not a 90' draw-forcer 5 Conceded 2; no 90' draw; FRA won in 90'. Hot-GK edge correctly confined to a pen phase that never came.
4 F8 possession: "MAR may split/eclipse possession" (cap ~52%, R32-5c) 5 MAR 51.9% > FRA 48.1% — g5 explicitly raised MAR eclipsing possession; the match's defining dynamic (Learning #1).

g5 findings = 20/20. Also staked & validated: F13 correctly assessed as NOT firing (FRA high-xG; MAR regression, not sterility); F11 net neutral (~26-29°C, Tier-1 reduction); Tello LOW-PEN 10-15% (0 pens ✅); Diop SF-ban risk flagged; R32-3 CONMEBOL-reactive correctly ruled out (MAR = CAF possession side).

g5's synthesis was internally consistent with its findings: modal 2-1 (FRA win) matches the false-9/regression/Bounou-dilution triad pointing at a FRA win with a low MAR total. The "~1 MAR goal" hedge overshot by one (got 0) — within precision noise for a regression thesis.

Model-breaking: 10/10

All 5 material model-breaker categories caught — g5's validated strength (model-breaker detection) on full display:

g5 TOTAL: 40 (auto: w90 15 + adv 10 + scor 15) + 20 findings + 10 MB = 70/80 → 88% → Band B | who_scores 10/20 (Mbappé ✅ / Ounahi ❌)


g4 (glm-5-turbo)

ABSENT — not scored

g4 failed to rate-limit 3× and produced no output. No fra-vs-mar-2026-pure-football-analysis-g4.md file exists in the match dir. Recorded as null in the machine-readable footer (validate-master renders the cell as ? ? and excludes it from the g4 running average). No findings / model-breaking / scoreline to assess.

Process note: g4's absence cost the ensemble its best-on-clear-favorites + scoreline-precision input for this match. g5 (B, 70/80) and mm3 (C, 55/80) both independently arrived at FRA-advance with sharp qualitative reasoning, so the consensus direction held — but g4's contrarian/scoreline-precision voice was missing.


Cross-agent honesty note

mm3 and g5 were near-identical in qualitative substance — both staked the same four core theses (false-9 failure, R16 regression, Bounou open-shape dilution, possession flip) plus the same secondary calls (Tello low-pen, F11-neutral, Diop SF-ban, Doué change). The brief credits both agents for the false-9 and regression calls (Learnings #2, #3, #4). The sole differentiator is the scoreline modal: g5's 2-1 (FRA win, margin within 1 of the actual 2-0 ⇒ auto 15/25) vs mm3's 1-1 (draw, wrong winner in the scoreline ⇒ auto 0/25). g5's synthesis was internally consistent with its own findings; mm3 hedged toward a cagey draw its findings didn't support — a reasoning-quality gap that the findings rubric (correctly) doesn't double-penalize but the scoreline component does.


Scores Summary (machine-readable — DO NOT EDIT)

{
  "match": "fra-vs-mar",
  "stage": "qf",
  "result": { "score90": "2-0", "final": "2-0", "outcome90": "win", "advance": "a", "route": "90" },
  "absent": ["g4"],
  "absent_reason": "g4 (glm-5-turbo) failed to rate-limit 3x and produced no output file",
  "agents": {
    "mm3": { "auto": 25, "winner_90": 15, "advance": 10, "scoreline": 0, "findings": 20, "model_breaking": 10, "total": 55, "pct": 69, "band": "C", "who_scores": 10 },
    "g4": null,
    "g5": { "auto": 40, "winner_90": 15, "advance": 10, "scoreline": 15, "findings": 20, "model_breaking": 10, "total": 70, "pct": 88, "band": "B", "who_scores": 10 }
  }
}