Per-Agent Accuracy Report β WC2026 Pipeline
42 ensemble matches tracked. Auto-extracted from TOML blocks + post-event reviews.
Generated: 2026-06-28. Post-TASK-99 independent resolution active since Jun 23.
Group stage COMPLETE (48/48). This is the FINAL group-stage report. Round of 32 starts Jun 28.
Executive Summary
The 42-match, group-stage-closed picture confirms a clear division of labor, and it sharpened β not blurred β over the MD3 wave. g5 (zai/glm-5.2) is the aggregate winner-accuracy leader at 68% and the cleanest scoreline (avg dist 2.0, exact 17%); mm3 (MiniMax-M3) is the "least-wrong" specialist, leading closest-on-scoreline at 22x and tying for the best scorer hit rate (44%); g4 (glm-5-turbo) is the volume backbone (N=41, never absent) with the tightest close-game calibration.
The MD3 wave (Jun 25-26, +12 matches) was the hardest stretch of the tournament and it produced a blunt verdict: all three agents share the same blindspots and none "broke out." The blowout-underestimate failure fired three more times (SEN 5-0, NED 5-1, FRA 4-1 β all modal 2-0/1-1), and three upsets landed in the tail with zero correct winner calls (GER 1-2 ECU, USA 2-3 TUR, KOR 0-1 RSA) β exactly the pattern the wide spreads (4/18/8pt) honestly flagged as "uncertain" rather than predicted. Findings 23 (GD-Chase Mode) and 24 (Bracket Optimization) were ADDED on Jun 27 because of these MD3 failures β they did not help during MD3 (they didn't exist yet), and with the group stage over they are retrospective lessons, not live tools for R32. TASK-99's verdict is unchanged: independence makes the pipeline honest (8.6pt avg spread), not prescient β it correctly signals where uncertainty lives, which is the most an ensemble of three LLM agents can do.
Per-Agent Summary Table
| Agent | Model | Matches | Winner % | Exact Score % | Score β€1 goal % | Avg Scoreline Dist | #1 Scorer Hit % | Closest on Score |
|---|---|---|---|---|---|---|---|---|
| mm3 | MiniMax-M3 | 36 | 64% (23) | 11% (4) | 33% (12) | 2.1 goals | 44% (16) | 22x |
| g4 | glm-5-turbo | 41 | 59% (24) | 12% (5) | 32% (13) | 2.0 goals | 44% (18) | 14x |
| g5 | zai/glm-5.2 | 41 | 68% (28) | 17% (7) | 37% (15) | 2.0 goals | 41% (17) | 6x |
g5 = best aggregate (winner + exact + avg dist). mm3 = closest-most-often + tied-best scorer. g4 = never absent, tightest on close games. N mismatch: mm3 has 36 vs 41 β five matches ran a 2-variant ensemble (mm3 absent); no evidence of systematic mm3 exclusion bias.
Key Findings
1. Winner Prediction
g5 leads winner accuracy at 68% (28/41), edging mm3 (64%, 23/36) and g4 (59%, 24/41). g5's edge held through MD3, where it correctly tabbed the routine favorites (ARG-JOR, NED-TUN, CRO-GHA, ESP-URU) and the one clean draw-band read that mattered (EGY-IRN 1-1, exact modal).
The headline failure mode is shared, not agent-specific: on the three MD3 upsets (GER-ECU, USA-TUR, KOR-RSA) all three agents picked the wrong winner β 0/3 on each. None of the three models has an upset-picking edge; they all anchor on team-strength priors. Where the ensemble did add value was in signaling uncertainty: those three matches carried the widest spreads of the wave (4pt / 18pt / 8pt), so a lead reading "this is genuinely uncertain, treat the favorite as soft" was the correct meta-call even though the specific pick missed. Conclusion: don't expect any single agent to "call upsets." Expect the spread to tell you which matches to hedge.
2. Scoreline Precision
g5 has the best scoreline profile: lowest avg distance (2.0), most exacts (7, 17%), most within-1 (15, 37%). g4 matches the avg distance (2.0) but with fewer exacts; mm3 is the loosest on aggregate (2.1 avg, only 4 exacts).
But the closest-on-score crown belongs to mm3 (22x), well ahead of g4 (14x) and g5 (6x). This is the "least-wrong specialist" pattern: mm3 tends to land near the actual scoreline more often even when its modal pick isn't exact. MD3 illustrates it cleanly β on FRA-NOR (4-1) mm3's 3-1 was the closest at dist 1; on KOR-RSA (0-1) and USA-TUR (2-3) mm3 was "closest" by being slightly less wrong than g4/g5, not by being right. g5 trades closest-frequency for outright exacts; mm3 trades exacts for closest-frequency. Both are real, complementary strengths.
The systemic blindspot is blowout margins (Finding 22), and MD3 re-confirmed it three more times:
- SEN 5-0 IRA β modal 2-0 across all three (dist 3-5). GD-chase mechanism (Finding 23, codified from this) never weighted.
- NED 5-1 SWE β modal 1-1 / 2-1 (dist 3-4). Same miss.
- FRA 4-1 NOR β modal 2-0/3-0 (dist 1-3). Same miss.
No agent has a blowout edge. The "g4 = blowout specialist" label (carried in earlier notes) is wrong β all three under-predict 4+ margins equally when an elite attack meets a non-veteran-GK block. The fix is a framework-level rule (weight 4+ margin at 15-20%, +10-15pt in GD-chase), not an agent re-weight.
3. Who Scores #1 Accuracy
mm3 and g4 tie for the best #1-scorer hit rate at 44%; g5 trails at 41%. mm3's 44% comes on a smaller base (16/36), g4's on volume (18/41). The "generational milestone = goal" rule (Finding 12 interaction) continued to pay in MD3 β IsmaΓ―la Sarr (#1, ~22%) hit SEN-IRA; the unanimous-with-mechanism pattern held.
The structural caveat (flagged in the PAR-AUS 0-0 review) bites harder in MD3's draw-heavy wave: a 0-0 result is unscorable (0/20 on the rubric) regardless of how good the macro reasoning was. PAR-AUS, POR-COL, CPV-KSA, EGY-IRN were all 0-0 or low-scoring β three of four capped the scorer ceiling. This is a rubric artifact, not an agent weakness, and it should be read as such: scorer-% in MD3 is depressed by the draw epidemic, not by a drop in agent picker quality.
Late-sub / bench-scorer risk (Finding 19, plus the dead-rubber amplifier added from USA-TUR) is the area all three agents systematically miss. Pape Gueye's brace (SEN-IRA, 59'/71', a sub), Ayhan's 97:50 winner (USA-TUR, an 88' sub), 3-of-5 SEN goals from subs β the modal Who Scores lists start-heavy and miss the bench explosion. This is a framework fix (add a "deep-arriving sub" to every list on dead rubbers), not an agent re-weight.
4. TASK-99 Impact (Independence vs Accuracy)
Average spread across the 41 multi-agent matches: 8.6 pts (up from the ~2pt artificial-anchoring era). The MD3 wave pushed this up: USA-TUR (18pt), KOR-RSA (8pt), ESP-URU (19pt), AUT-ALG (15pt), NED-TUN (16pt), ARG-JOR (15pt) all carried genuinely wide, honest spreads.
The core question β does independence correlate with accuracy? β has a nuanced answer:
- Independence correlates with HONESTY, not with PRESCIENCE. On the three MD3 upsets, the widest spreads (USA-TUR 18pt, KOR-RSA 8pt) were the matches the ensemble got wrong on the specific pick but right on the meta-call ("this is uncertain, don't trust the favorite"). The narrow-spread matches that missed (GER-ECU 4pt, NED-SWE 2pt) were failures of calibration, not independence β the agents anchored on a shared wrong prior and the narrow spread gave false confidence.
- Widest spread β closest agent. On the 4 highest-spread MD3 matches, the closest agent was mm3 twice, g4 once, g5 once β no consistent "the outlier agent is right" pattern. Independence surfaces disagreement, not truth.
- TASK-99's value is operational, not predictive: it stops the ensemble from manufacturing false consensus. A 1-2pt spread should now be read as a warning (anchoring on shared input), which is why the report flags the ~15 narrow-spread matches (β€2pt) as "β οΈ anchoring?" β several of those (ENG-GHA 0-0, JPN-SWE 1-1, GER-CIV 2-1) were quiet successes, but GER-ECU (4pt, narrow) was a silent miss.
Verdict: keep TASK-99. Independence made the pipeline's uncertainty legible. It did not β and was never going to β make a 3-LLM ensemble call upsets. The pre-TASK-99 era's 1pt spreads were worse than useless (they hid disagreement); the post-TASK-99 8.6pt average is the honest floor.
5. Per-Agent Strengths / Weaknesses
mm3 (MiniMax-M3) β 36 matches
- β Closest-on-scoreline leader (22x). Lands near the actual scoreline more often than any agent β the "least-wrong" hedge. Strong on MD3 routine favorites (FRA-NOR dist 1, ARG-JOR, CRO-GHA).
- β Tied-best scorer picker (44%). Best Who Scores lists in the ensemble.
- β Loosest aggregate scoreline (2.1 avg, 4 exacts). Wins closest-frequency by being slightly-off-a-lot, not by being exact. Lowest winner % of the active three would be the read if N were equal β but its smaller N (36) and draw-hedging posture bias both numbers.
- β Draw-hedging posture inflates closest-count on draw days (ENG-GHA, POR-COL, JPN-SWE) but produces wrong-favorite calls on upset days where the draw band isn't the answer.
g4 (glm-5-turbo) β 41 matches
- β Volume backbone β never absent (N=41). Most reliable coverage; the only agent on USA-AUS (2-0) and several single-variant runs.
- β Tightest on close games (exact 12%, within-1 32%). When a match is genuinely close, g4's modal is the most disciplined.
- β Lowest winner % of the three (59%) and lowest closest-count-for-its-volume (14x on N=41). g4 "shows up" but rarely "wins" a match outright.
- β Contrarian-but-often-wrong on dead rubbers. Flagged USA-TUR and POR-COL draws/upsets as live β the uncertainty read was right, the specific call was wrong (USA-TUR 2-3, POR-COL 0-0 vs g4's 2-1).
g5 (zai/glm-5.2) β 41 matches
- β Best aggregate accuracy β winner 68%, exact 17%, avg dist 2.0, within-1 37%. The cleanest Finding-11/decision-graph reasoning in the ensemble.
- β Best draw-band reader of MD3. Nailed PAR-AUS (0-0 modal runner-up), CPV-KSA (1-0 dist 1), EGY-IRN (1-1 exact), NED-TUN (3-0 dist 1). When the draw floor is real, g5 finds it.
- β Lowest closest-on-score (6x on N=41). When g5 is wrong, it's wrong decisively β fewer "near-miss" landings than mm3.
- β Scorer lists weakest (41%). Marginally behind mm3/g4; no milestone-mechanism edge in MD3.
6. Calibration Rubric Comparison
Per-match consensus calibration (100-pt rubric from doc-1, adopted TASK-97) for the MD3 wave (Jun 25-26):
| Match | Result | Consensus Cal | Closest Agent | Pattern |
|---|---|---|---|---|
| ARG-JOR | 3-1 | ~80% B | mm3 | clean favorite |
| NED-TUN | 3-1 | ~80% B | g5 | clean favorite (g5 exact-adjacent) |
| JPN-SWE | 1-1 | ~80% B | g4 | draw hit |
| CRO-GHA | 2-1 | ~70% C | mm3 | clean favorite |
| ESP-URU | 1-0 | ~70% C | g4 | clean favorite |
| PAR-AUS | 0-0 | ~70% C | g5 | draw-band hit (g5 modal) |
| CPV-KSA | 0-0 | ~68% C | g5 | draw-band hit |
| BEL-NZL | 1-5 | ~68% C | g4 | blowout miss (upset + margin) |
| SEN-IRA | 5-0 | ~68% C | g4 | blowout miss (GD-chase) |
| EGY-IRN | 1-1 | ~78% B | g5 | draw exact |
| FRA-NOR | 4-1 | ~73% B | mm3 | blowout miss |
| NED-SWE | 5-1 | ~55% F | g4 | blowout miss (biggest scoreline miss) |
| AUT-ALG | 3-3 | ~55% F | mm3 | wild draw/shootout |
| GER-ECU | 1-2 | ~52% F | mm3 | upset miss |
| USA-TUR | 2-3 | ~42% F | mm3 | upset miss (widest spread, 18pt) |
| POR-COL | 0-0 | ~41% F | mm3 | draw miss (all picked POR) |
| KOR-RSA | 0-1 | ~40% F | mm3 | upset miss |
MD3-wave average calibration β 65% (band C) β the weakest matchday of the tournament, dragged down by 4 F-grade misses (3 upsets + POR-COL draw). The clean-favorite and draw-band matches scored B/C as expected; the failures cluster entirely on upsets, blowouts, and the POR-COL dead-rubber draw β exactly the three categories no agent has an edge on.
Per-agent closest-share in the F-grade misses (4 matches): mm3 4/4, g4 0/4, g5 0/4. mm3's "least-wrong" posture is what surfaces in the hard matches β useful for not being embarrassingly wrong, but it did not produce a single correct winner call in any of them.
Recommendations for the Framework (R32 and beyond)
-
Give g5 a plurality weight on winner outcome in R32 consensus. g5's 68% winner + cleanest scoreline is the most defensible aggregate lead. Concretely: in reconciliation, treat g5's winner call as the default and require mm3/g4 to overturn it with explicit reasoning, rather than the current flat average. (Do NOT drop mm3/g4 β their disagreement IS the signal.)
-
Weight mm3's scorer lists at 1.1-1.2x for Who Scores. Tied-best hit rate (44%) on a smaller base means mm3's picker is genuinely sharp; give its #1/#2 picks slight priority in the consensus Who Scores list. Keep g5's draw-band read as the tiebreaker when mm3 and g4 split.
-
Do NOT re-weight for blowouts per-agent. All three agents under-predict 4+ margins identically. Keep the fix at the framework level (Finding 22: 15-20% on 4+ margin; Finding 23 GD-chase: +10-15pt to 25-30%) β applied by the lead in reconciliation, not delegated to a "blowout agent" that doesn't exist.
-
Retain the spread-as-uncertainty rule for R32, but lower the "wide spread = hedge" threshold for knockouts. R32 has fewer true upsets than the group stage and no draws in the final outcome (extra time). A 10pt+ spread in R32 is a stronger upset-warning signal than the same spread in MD3, because the prior should be tighter. Treat any R32 spread >12pt as "the underdog is genuinely live" and write that into the TL;DR explicitly.
-
Findings 23 & 24 are group-stage-specific and will NOT fire in R32. GD-Chase (F23) requires a standings table; Bracket-Optimization (F24) requires group-position incentive. Both are zero in pure knockouts. Archive them as validated group-stage lessons (3-4 strong validations each: SEN-IRA, FRA-NOR, NED-SWE, USA-TUR). Do not let agents invoke them in R32 β there is no draw-suffices, no best-3rd-place, no GD math in the Round of 32. (They may marginally re-fire in the final group-stage-style contexts of future tournaments, but not here.)
-
Do not retire any agent. g4's 59% winner is the lowest, but its N=41 volume and close-game discipline make it the ensemble's stability anchor; dropping it would cost the closest-frequency disagreement that makes TASK-99 spreads honest. The division of labor (g5 = aggregate, mm3 = least-wrong + scorer, g4 = volume + close games) is intact and each role is earning its place.
Win% Spread Analysis (independence indicator, TASK-99)
Average spread across agents: 8.6 pts (41 matches with 2+ agents)
| Match | Spread | Actual | Notes |
|---|---|---|---|
| alg-vs-jor | 1pt | 2-1 | β οΈ narrow (anchoring?) |
| arg-vs-aut | 5pt | 2-0 | β honest spread |
| arg-vs-jor | 15pt | 3-1 | β honest spread |
| aut-vs-alg | 15pt | 3-3 | β honest spread |
| bel-vs-irn | 4pt | 0-0 | |
| bel-vs-nzl | 5pt | 1-5 | β honest spread |
| bra-vs-hai | 2pt | 3-0 | β οΈ narrow (anchoring?) |
| col-vs-drc | 1pt | 1-0 | β οΈ narrow (anchoring?) |
| cpv-vs-ksa | 4pt | 0-0 | |
| cro-vs-gha | 5pt | 2-1 | β honest spread |
| cro-vs-pan | 0pt | 1-0 | β οΈ narrow (anchoring?) |
| cur-vs-civ | 55pt | 0-2 | β honest spread |
| drc-vs-uzb | 10pt | 3-1 | β honest spread |
| ecu-vs-cur | 1pt | 0-0 | β οΈ narrow (anchoring?) |
| egy-vs-irn | 6pt | 1-1 | β honest spread |
| eng-vs-gha | 2pt | 0-0 | β οΈ narrow (anchoring?) |
| eng-vs-pan | 1pt | 2-0 | β οΈ narrow (anchoring?) |
| esp-vs-ksa | 0pt | 4-0 | β οΈ narrow (anchoring?) |
| esp-vs-uru | 19pt | 1-0 | β honest spread |
| fra-vs-ira | 3pt | 3-0 | |
| fra-vs-nor | 2pt | 4-1 | β οΈ narrow (anchoring?) |
| ger-vs-civ | 2pt | 2-1 | β οΈ narrow (anchoring?) |
| ger-vs-ecu | 4pt | 1-2 | |
| jpn-vs-swe | 2pt | 1-1 | β οΈ narrow (anchoring?) |
| kor-vs-rsa | 8pt | 0-1 | β honest spread |
| mar-vs-hai | 6pt | 4-2 | β honest spread |
| mex-vs-cze | 18pt | 3-0 | β honest spread |
| ned-vs-swe | 2pt | 5-1 | β οΈ narrow (anchoring?) |
| ned-vs-tun | 16pt | 3-1 | β honest spread |
| nor-vs-sen | 1pt | 3-2 | β οΈ narrow (anchoring?) |
| nzl-vs-egy | 2pt | 1-3 | β οΈ narrow (anchoring?) |
| par-vs-aus | 11pt | 0-0 | β honest spread |
| por-vs-col | 21pt | 0-0 | β honest spread |
| por-vs-uzb | 1pt | 5-0 | β οΈ narrow (anchoring?) |
| sco-vs-bra | 56pt | 0-3 | β honest spread |
| sco-vs-mar | 6pt | 0-1 | β honest spread |
| sen-vs-ira | 14pt | 5-0 | β honest spread |
| tun-vs-jpn | 2pt | 0-4 | β οΈ narrow (anchoring?) |
| tur-vs-par | 4pt | 0-1 | |
| uru-vs-cpv | 1pt | 2-2 | β οΈ narrow (anchoring?) |
| usa-vs-tur | 18pt | 2-3 | β honest spread |
Per-Match Detail
| Match | Actual | mm3 modal (dist) | g4 modal (dist) | g5 modal (dist) | Closest | Consensus Cal |
|---|---|---|---|---|---|---|
| alg-vs-jor | 2-1 | 1-0 (2) | 2-1 (0) | 2-1 (0) | g4 | ~88% |
| arg-vs-aut | 2-0 | 2-0 (0) | 2-1 (1) | 2-0 (0) | mm3 | ~92% |
| arg-vs-jor | 3-1 | 2-0 (2) | 2-0 (2) | 2-0 (2) | mm3 | ~80% |
| aut-vs-alg | 3-3 | 1-1 (4) | 1-0 (5) | 1-1 (4) | mm3 | ~55% |
| bel-vs-irn | 0-0 | 1-1 (2) | 2-0 (2) | 2-0 (2) | mm3 | ~75% |
| bel-vs-nzl | 1-5 | 1-0 (5) | 1-1 (4) | 1-0 (5) | g4 | ~68% |
| bra-vs-hai | 3-0 | 2-0 (1) | 3-0 (0) | 3-0 (0) | g4 | ~88% |
| col-vs-drc | 1-0 | β | 1-0 (0) | 2-0 (1) | g4 | ~82% |
| cpv-vs-ksa | 0-0 | 1-1 (2) | 1-1 (2) | 1-0 (1) | g5 | ~68% |
| cro-vs-gha | 2-1 | 1-0 (2) | 1-0 (2) | 1-0 (2) | mm3 | ~70% |
| cro-vs-pan | 1-0 | 1-0 (0) | 2-0 (1) | 2-0 (1) | mm3 | ~68% |
| cur-vs-civ | 0-2 | 2-1 (3) | 1-0 (3) | 1-0 (3) | mm3 | ~80% |
| drc-vs-uzb | 3-1 | 2-1 (1) | 2-0 (2) | 2-0 (2) | mm3 | ~78% |
| ecu-vs-cur | 0-0 | β | 2-0 (2) | 2-0 (2) | g4 | ~45% |
| egy-vs-irn | 1-1 | 1-0 (1) | 1-0 (1) | 1-1 (0) | g5 | ~78% |
| eng-vs-gha | 0-0 | 2-0 (2) | 2-0 (2) | 2-0 (2) | mm3 | ~45% |
| eng-vs-pan | 2-0 | 2-0 (0) | β | 2-0 (0) | mm3 | ~93% |
| esp-vs-ksa | 4-0 | 3-0 (1) | 2-0 (2) | 2-0 (2) | mm3 | ~85% |
| esp-vs-uru | 1-0 | β | 2-0 (1) | 2-1 (2) | g4 | ~70% |
| fra-vs-ira | 3-0 | 2-0 (1) | 2-0 (1) | 2-0 (1) | mm3 | ~82% |
| fra-vs-nor | 4-1 | 3-1 (1) | 3-0 (2) | 2-0 (3) | mm3 | ~73% |
| ger-vs-civ | 2-1 | 1-1 (1) | 2-1 (0) | 2-1 (0) | g4 | ~82% |
| ger-vs-ecu | 1-2 | 2-0 (3) | 2-0 (3) | 2-0 (3) | mm3 | ~52% |
| jpn-vs-swe | 1-1 | β | 1-0 (1) | 2-1 (1) | g4 | ~80% |
| kor-vs-rsa | 0-1 | 1-0 (2) | 1-0 (2) | 2-0 (3) | mm3 | ~40% |
| mar-vs-hai | 4-2 | β | 2-0 (4) | 2-0 (4) | g4 | ~80% |
| mex-vs-cze | 3-0 | 2-0 (1) | 2-0 (1) | 2-0 (1) | mm3 | ~70% |
| ned-vs-swe | 5-1 | 1-1 (4) | 2-1 (3) | 2-2 (4) | g4 | ~55% |
| ned-vs-tun | 3-1 | 2-0 (2) | 2-0 (2) | 3-0 (1) | g5 | ~80% |
| nor-vs-sen | 3-2 | 2-1 (2) | 2-1 (2) | 2-1 (2) | mm3 | ~85% |
| nzl-vs-egy | 1-3 | 1-0 (3) | 1-2 (1) | 0-2 (2) | g4 | ~72% |
| par-vs-aus | 0-0 | 1-1 (2) | 1-1 (2) | 0-0 (0) | g5 | ~70% |
| por-vs-col | 0-0 | 2-1 (3) | 2-1 (3) | 2-1 (3) | mm3 | ~41% |
| por-vs-uzb | 5-0 | 2-0 (3) | 2-0 (3) | 2-0 (3) | mm3 | ~80% |
| sco-vs-bra | 0-3 | 2-0 (5) | 2-0 (5) | 1-2 (2) | g5 | ~83% |
| sco-vs-mar | 0-1 | 0-1 (0) | 0-1 (0) | 1-1 (1) | mm3 | ~82% |
| sen-vs-ira | 5-0 | 1-1 (5) | 2-0 (3) | 2-1 (4) | g4 | ~68% |
| tun-vs-jpn | 0-4 | 1-1 (4) | 1-1 (4) | 0-1 (3) | g5 | ~55% |
| tur-vs-par | 0-1 | 1-0 (2) | 2-1 (2) | 2-1 (2) | mm3 | ~55% |
| uru-vs-cpv | 2-2 | 1-0 (3) | 2-0 (2) | 1-0 (3) | g4 | ~80% |
| usa-vs-aus | 2-0 | β | 1-1 (2) | β | g4 | ~75% |
| usa-vs-tur | 2-3 | 2-1 (2) | 2-0 (3) | 1-0 (4) | mm3 | ~42% |
Final group-stage report. 42 ensemble matches. Group stage complete (48/48). Round of 32 begins Jun 28. Findings 23 (GD-Chase) + 24 (Bracket Optimization) added Jun 27, retrospective to MD3, non-applicable in knockouts. Not published β lead runs gister.