Per-Agent Accuracy Report β€” WC2026 Pipeline

42 ensemble matches tracked. Auto-extracted from TOML blocks + post-event reviews.
Generated: 2026-06-28. Post-TASK-99 independent resolution active since Jun 23.
Group stage COMPLETE (48/48). This is the FINAL group-stage report. Round of 32 starts Jun 28.

Executive Summary

The 42-match, group-stage-closed picture confirms a clear division of labor, and it sharpened β€” not blurred β€” over the MD3 wave. g5 (zai/glm-5.2) is the aggregate winner-accuracy leader at 68% and the cleanest scoreline (avg dist 2.0, exact 17%); mm3 (MiniMax-M3) is the "least-wrong" specialist, leading closest-on-scoreline at 22x and tying for the best scorer hit rate (44%); g4 (glm-5-turbo) is the volume backbone (N=41, never absent) with the tightest close-game calibration.

The MD3 wave (Jun 25-26, +12 matches) was the hardest stretch of the tournament and it produced a blunt verdict: all three agents share the same blindspots and none "broke out." The blowout-underestimate failure fired three more times (SEN 5-0, NED 5-1, FRA 4-1 β€” all modal 2-0/1-1), and three upsets landed in the tail with zero correct winner calls (GER 1-2 ECU, USA 2-3 TUR, KOR 0-1 RSA) β€” exactly the pattern the wide spreads (4/18/8pt) honestly flagged as "uncertain" rather than predicted. Findings 23 (GD-Chase Mode) and 24 (Bracket Optimization) were ADDED on Jun 27 because of these MD3 failures β€” they did not help during MD3 (they didn't exist yet), and with the group stage over they are retrospective lessons, not live tools for R32. TASK-99's verdict is unchanged: independence makes the pipeline honest (8.6pt avg spread), not prescient β€” it correctly signals where uncertainty lives, which is the most an ensemble of three LLM agents can do.

Per-Agent Summary Table

Agent Model Matches Winner % Exact Score % Score ≀1 goal % Avg Scoreline Dist #1 Scorer Hit % Closest on Score
mm3 MiniMax-M3 36 64% (23) 11% (4) 33% (12) 2.1 goals 44% (16) 22x
g4 glm-5-turbo 41 59% (24) 12% (5) 32% (13) 2.0 goals 44% (18) 14x
g5 zai/glm-5.2 41 68% (28) 17% (7) 37% (15) 2.0 goals 41% (17) 6x

g5 = best aggregate (winner + exact + avg dist). mm3 = closest-most-often + tied-best scorer. g4 = never absent, tightest on close games. N mismatch: mm3 has 36 vs 41 β€” five matches ran a 2-variant ensemble (mm3 absent); no evidence of systematic mm3 exclusion bias.

Key Findings

1. Winner Prediction

g5 leads winner accuracy at 68% (28/41), edging mm3 (64%, 23/36) and g4 (59%, 24/41). g5's edge held through MD3, where it correctly tabbed the routine favorites (ARG-JOR, NED-TUN, CRO-GHA, ESP-URU) and the one clean draw-band read that mattered (EGY-IRN 1-1, exact modal).

The headline failure mode is shared, not agent-specific: on the three MD3 upsets (GER-ECU, USA-TUR, KOR-RSA) all three agents picked the wrong winner β€” 0/3 on each. None of the three models has an upset-picking edge; they all anchor on team-strength priors. Where the ensemble did add value was in signaling uncertainty: those three matches carried the widest spreads of the wave (4pt / 18pt / 8pt), so a lead reading "this is genuinely uncertain, treat the favorite as soft" was the correct meta-call even though the specific pick missed. Conclusion: don't expect any single agent to "call upsets." Expect the spread to tell you which matches to hedge.

2. Scoreline Precision

g5 has the best scoreline profile: lowest avg distance (2.0), most exacts (7, 17%), most within-1 (15, 37%). g4 matches the avg distance (2.0) but with fewer exacts; mm3 is the loosest on aggregate (2.1 avg, only 4 exacts).

But the closest-on-score crown belongs to mm3 (22x), well ahead of g4 (14x) and g5 (6x). This is the "least-wrong specialist" pattern: mm3 tends to land near the actual scoreline more often even when its modal pick isn't exact. MD3 illustrates it cleanly β€” on FRA-NOR (4-1) mm3's 3-1 was the closest at dist 1; on KOR-RSA (0-1) and USA-TUR (2-3) mm3 was "closest" by being slightly less wrong than g4/g5, not by being right. g5 trades closest-frequency for outright exacts; mm3 trades exacts for closest-frequency. Both are real, complementary strengths.

The systemic blindspot is blowout margins (Finding 22), and MD3 re-confirmed it three more times:

No agent has a blowout edge. The "g4 = blowout specialist" label (carried in earlier notes) is wrong β€” all three under-predict 4+ margins equally when an elite attack meets a non-veteran-GK block. The fix is a framework-level rule (weight 4+ margin at 15-20%, +10-15pt in GD-chase), not an agent re-weight.

3. Who Scores #1 Accuracy

mm3 and g4 tie for the best #1-scorer hit rate at 44%; g5 trails at 41%. mm3's 44% comes on a smaller base (16/36), g4's on volume (18/41). The "generational milestone = goal" rule (Finding 12 interaction) continued to pay in MD3 β€” IsmaΓ―la Sarr (#1, ~22%) hit SEN-IRA; the unanimous-with-mechanism pattern held.

The structural caveat (flagged in the PAR-AUS 0-0 review) bites harder in MD3's draw-heavy wave: a 0-0 result is unscorable (0/20 on the rubric) regardless of how good the macro reasoning was. PAR-AUS, POR-COL, CPV-KSA, EGY-IRN were all 0-0 or low-scoring β€” three of four capped the scorer ceiling. This is a rubric artifact, not an agent weakness, and it should be read as such: scorer-% in MD3 is depressed by the draw epidemic, not by a drop in agent picker quality.

Late-sub / bench-scorer risk (Finding 19, plus the dead-rubber amplifier added from USA-TUR) is the area all three agents systematically miss. Pape Gueye's brace (SEN-IRA, 59'/71', a sub), Ayhan's 97:50 winner (USA-TUR, an 88' sub), 3-of-5 SEN goals from subs β€” the modal Who Scores lists start-heavy and miss the bench explosion. This is a framework fix (add a "deep-arriving sub" to every list on dead rubbers), not an agent re-weight.

4. TASK-99 Impact (Independence vs Accuracy)

Average spread across the 41 multi-agent matches: 8.6 pts (up from the ~2pt artificial-anchoring era). The MD3 wave pushed this up: USA-TUR (18pt), KOR-RSA (8pt), ESP-URU (19pt), AUT-ALG (15pt), NED-TUN (16pt), ARG-JOR (15pt) all carried genuinely wide, honest spreads.

The core question β€” does independence correlate with accuracy? β€” has a nuanced answer:

Verdict: keep TASK-99. Independence made the pipeline's uncertainty legible. It did not β€” and was never going to β€” make a 3-LLM ensemble call upsets. The pre-TASK-99 era's 1pt spreads were worse than useless (they hid disagreement); the post-TASK-99 8.6pt average is the honest floor.

5. Per-Agent Strengths / Weaknesses

mm3 (MiniMax-M3) β€” 36 matches

g4 (glm-5-turbo) β€” 41 matches

g5 (zai/glm-5.2) β€” 41 matches

6. Calibration Rubric Comparison

Per-match consensus calibration (100-pt rubric from doc-1, adopted TASK-97) for the MD3 wave (Jun 25-26):

Match Result Consensus Cal Closest Agent Pattern
ARG-JOR 3-1 ~80% B mm3 clean favorite
NED-TUN 3-1 ~80% B g5 clean favorite (g5 exact-adjacent)
JPN-SWE 1-1 ~80% B g4 draw hit
CRO-GHA 2-1 ~70% C mm3 clean favorite
ESP-URU 1-0 ~70% C g4 clean favorite
PAR-AUS 0-0 ~70% C g5 draw-band hit (g5 modal)
CPV-KSA 0-0 ~68% C g5 draw-band hit
BEL-NZL 1-5 ~68% C g4 blowout miss (upset + margin)
SEN-IRA 5-0 ~68% C g4 blowout miss (GD-chase)
EGY-IRN 1-1 ~78% B g5 draw exact
FRA-NOR 4-1 ~73% B mm3 blowout miss
NED-SWE 5-1 ~55% F g4 blowout miss (biggest scoreline miss)
AUT-ALG 3-3 ~55% F mm3 wild draw/shootout
GER-ECU 1-2 ~52% F mm3 upset miss
USA-TUR 2-3 ~42% F mm3 upset miss (widest spread, 18pt)
POR-COL 0-0 ~41% F mm3 draw miss (all picked POR)
KOR-RSA 0-1 ~40% F mm3 upset miss

MD3-wave average calibration β‰ˆ 65% (band C) β€” the weakest matchday of the tournament, dragged down by 4 F-grade misses (3 upsets + POR-COL draw). The clean-favorite and draw-band matches scored B/C as expected; the failures cluster entirely on upsets, blowouts, and the POR-COL dead-rubber draw β€” exactly the three categories no agent has an edge on.

Per-agent closest-share in the F-grade misses (4 matches): mm3 4/4, g4 0/4, g5 0/4. mm3's "least-wrong" posture is what surfaces in the hard matches β€” useful for not being embarrassingly wrong, but it did not produce a single correct winner call in any of them.

Recommendations for the Framework (R32 and beyond)

  1. Give g5 a plurality weight on winner outcome in R32 consensus. g5's 68% winner + cleanest scoreline is the most defensible aggregate lead. Concretely: in reconciliation, treat g5's winner call as the default and require mm3/g4 to overturn it with explicit reasoning, rather than the current flat average. (Do NOT drop mm3/g4 β€” their disagreement IS the signal.)

  2. Weight mm3's scorer lists at 1.1-1.2x for Who Scores. Tied-best hit rate (44%) on a smaller base means mm3's picker is genuinely sharp; give its #1/#2 picks slight priority in the consensus Who Scores list. Keep g5's draw-band read as the tiebreaker when mm3 and g4 split.

  3. Do NOT re-weight for blowouts per-agent. All three agents under-predict 4+ margins identically. Keep the fix at the framework level (Finding 22: 15-20% on 4+ margin; Finding 23 GD-chase: +10-15pt to 25-30%) β€” applied by the lead in reconciliation, not delegated to a "blowout agent" that doesn't exist.

  4. Retain the spread-as-uncertainty rule for R32, but lower the "wide spread = hedge" threshold for knockouts. R32 has fewer true upsets than the group stage and no draws in the final outcome (extra time). A 10pt+ spread in R32 is a stronger upset-warning signal than the same spread in MD3, because the prior should be tighter. Treat any R32 spread >12pt as "the underdog is genuinely live" and write that into the TL;DR explicitly.

  5. Findings 23 & 24 are group-stage-specific and will NOT fire in R32. GD-Chase (F23) requires a standings table; Bracket-Optimization (F24) requires group-position incentive. Both are zero in pure knockouts. Archive them as validated group-stage lessons (3-4 strong validations each: SEN-IRA, FRA-NOR, NED-SWE, USA-TUR). Do not let agents invoke them in R32 β€” there is no draw-suffices, no best-3rd-place, no GD math in the Round of 32. (They may marginally re-fire in the final group-stage-style contexts of future tournaments, but not here.)

  6. Do not retire any agent. g4's 59% winner is the lowest, but its N=41 volume and close-game discipline make it the ensemble's stability anchor; dropping it would cost the closest-frequency disagreement that makes TASK-99 spreads honest. The division of labor (g5 = aggregate, mm3 = least-wrong + scorer, g4 = volume + close games) is intact and each role is earning its place.

Win% Spread Analysis (independence indicator, TASK-99)

Average spread across agents: 8.6 pts (41 matches with 2+ agents)

Match Spread Actual Notes
alg-vs-jor 1pt 2-1 ⚠️ narrow (anchoring?)
arg-vs-aut 5pt 2-0 βœ… honest spread
arg-vs-jor 15pt 3-1 βœ… honest spread
aut-vs-alg 15pt 3-3 βœ… honest spread
bel-vs-irn 4pt 0-0
bel-vs-nzl 5pt 1-5 βœ… honest spread
bra-vs-hai 2pt 3-0 ⚠️ narrow (anchoring?)
col-vs-drc 1pt 1-0 ⚠️ narrow (anchoring?)
cpv-vs-ksa 4pt 0-0
cro-vs-gha 5pt 2-1 βœ… honest spread
cro-vs-pan 0pt 1-0 ⚠️ narrow (anchoring?)
cur-vs-civ 55pt 0-2 βœ… honest spread
drc-vs-uzb 10pt 3-1 βœ… honest spread
ecu-vs-cur 1pt 0-0 ⚠️ narrow (anchoring?)
egy-vs-irn 6pt 1-1 βœ… honest spread
eng-vs-gha 2pt 0-0 ⚠️ narrow (anchoring?)
eng-vs-pan 1pt 2-0 ⚠️ narrow (anchoring?)
esp-vs-ksa 0pt 4-0 ⚠️ narrow (anchoring?)
esp-vs-uru 19pt 1-0 βœ… honest spread
fra-vs-ira 3pt 3-0
fra-vs-nor 2pt 4-1 ⚠️ narrow (anchoring?)
ger-vs-civ 2pt 2-1 ⚠️ narrow (anchoring?)
ger-vs-ecu 4pt 1-2
jpn-vs-swe 2pt 1-1 ⚠️ narrow (anchoring?)
kor-vs-rsa 8pt 0-1 βœ… honest spread
mar-vs-hai 6pt 4-2 βœ… honest spread
mex-vs-cze 18pt 3-0 βœ… honest spread
ned-vs-swe 2pt 5-1 ⚠️ narrow (anchoring?)
ned-vs-tun 16pt 3-1 βœ… honest spread
nor-vs-sen 1pt 3-2 ⚠️ narrow (anchoring?)
nzl-vs-egy 2pt 1-3 ⚠️ narrow (anchoring?)
par-vs-aus 11pt 0-0 βœ… honest spread
por-vs-col 21pt 0-0 βœ… honest spread
por-vs-uzb 1pt 5-0 ⚠️ narrow (anchoring?)
sco-vs-bra 56pt 0-3 βœ… honest spread
sco-vs-mar 6pt 0-1 βœ… honest spread
sen-vs-ira 14pt 5-0 βœ… honest spread
tun-vs-jpn 2pt 0-4 ⚠️ narrow (anchoring?)
tur-vs-par 4pt 0-1
uru-vs-cpv 1pt 2-2 ⚠️ narrow (anchoring?)
usa-vs-tur 18pt 2-3 βœ… honest spread

Per-Match Detail

Match Actual mm3 modal (dist) g4 modal (dist) g5 modal (dist) Closest Consensus Cal
alg-vs-jor 2-1 1-0 (2) 2-1 (0) 2-1 (0) g4 ~88%
arg-vs-aut 2-0 2-0 (0) 2-1 (1) 2-0 (0) mm3 ~92%
arg-vs-jor 3-1 2-0 (2) 2-0 (2) 2-0 (2) mm3 ~80%
aut-vs-alg 3-3 1-1 (4) 1-0 (5) 1-1 (4) mm3 ~55%
bel-vs-irn 0-0 1-1 (2) 2-0 (2) 2-0 (2) mm3 ~75%
bel-vs-nzl 1-5 1-0 (5) 1-1 (4) 1-0 (5) g4 ~68%
bra-vs-hai 3-0 2-0 (1) 3-0 (0) 3-0 (0) g4 ~88%
col-vs-drc 1-0 β€” 1-0 (0) 2-0 (1) g4 ~82%
cpv-vs-ksa 0-0 1-1 (2) 1-1 (2) 1-0 (1) g5 ~68%
cro-vs-gha 2-1 1-0 (2) 1-0 (2) 1-0 (2) mm3 ~70%
cro-vs-pan 1-0 1-0 (0) 2-0 (1) 2-0 (1) mm3 ~68%
cur-vs-civ 0-2 2-1 (3) 1-0 (3) 1-0 (3) mm3 ~80%
drc-vs-uzb 3-1 2-1 (1) 2-0 (2) 2-0 (2) mm3 ~78%
ecu-vs-cur 0-0 β€” 2-0 (2) 2-0 (2) g4 ~45%
egy-vs-irn 1-1 1-0 (1) 1-0 (1) 1-1 (0) g5 ~78%
eng-vs-gha 0-0 2-0 (2) 2-0 (2) 2-0 (2) mm3 ~45%
eng-vs-pan 2-0 2-0 (0) β€” 2-0 (0) mm3 ~93%
esp-vs-ksa 4-0 3-0 (1) 2-0 (2) 2-0 (2) mm3 ~85%
esp-vs-uru 1-0 β€” 2-0 (1) 2-1 (2) g4 ~70%
fra-vs-ira 3-0 2-0 (1) 2-0 (1) 2-0 (1) mm3 ~82%
fra-vs-nor 4-1 3-1 (1) 3-0 (2) 2-0 (3) mm3 ~73%
ger-vs-civ 2-1 1-1 (1) 2-1 (0) 2-1 (0) g4 ~82%
ger-vs-ecu 1-2 2-0 (3) 2-0 (3) 2-0 (3) mm3 ~52%
jpn-vs-swe 1-1 β€” 1-0 (1) 2-1 (1) g4 ~80%
kor-vs-rsa 0-1 1-0 (2) 1-0 (2) 2-0 (3) mm3 ~40%
mar-vs-hai 4-2 β€” 2-0 (4) 2-0 (4) g4 ~80%
mex-vs-cze 3-0 2-0 (1) 2-0 (1) 2-0 (1) mm3 ~70%
ned-vs-swe 5-1 1-1 (4) 2-1 (3) 2-2 (4) g4 ~55%
ned-vs-tun 3-1 2-0 (2) 2-0 (2) 3-0 (1) g5 ~80%
nor-vs-sen 3-2 2-1 (2) 2-1 (2) 2-1 (2) mm3 ~85%
nzl-vs-egy 1-3 1-0 (3) 1-2 (1) 0-2 (2) g4 ~72%
par-vs-aus 0-0 1-1 (2) 1-1 (2) 0-0 (0) g5 ~70%
por-vs-col 0-0 2-1 (3) 2-1 (3) 2-1 (3) mm3 ~41%
por-vs-uzb 5-0 2-0 (3) 2-0 (3) 2-0 (3) mm3 ~80%
sco-vs-bra 0-3 2-0 (5) 2-0 (5) 1-2 (2) g5 ~83%
sco-vs-mar 0-1 0-1 (0) 0-1 (0) 1-1 (1) mm3 ~82%
sen-vs-ira 5-0 1-1 (5) 2-0 (3) 2-1 (4) g4 ~68%
tun-vs-jpn 0-4 1-1 (4) 1-1 (4) 0-1 (3) g5 ~55%
tur-vs-par 0-1 1-0 (2) 2-1 (2) 2-1 (2) mm3 ~55%
uru-vs-cpv 2-2 1-0 (3) 2-0 (2) 1-0 (3) g4 ~80%
usa-vs-aus 2-0 β€” 1-1 (2) β€” g4 ~75%
usa-vs-tur 2-3 2-1 (2) 2-0 (3) 1-0 (4) mm3 ~42%

Final group-stage report. 42 ensemble matches. Group stage complete (48/48). Round of 32 begins Jun 28. Findings 23 (GD-Chase) + 24 (Bracket Optimization) added Jun 27, retrospective to MD3, non-applicable in knockouts. Not published β€” lead runs gister.