World Cup 2026 Project — Glossary
Terms used throughout the pipeline. Football-specific, project-specific, and pipeline-specific jargon.
Pipeline & Architecture
MD1 / MD2 / MD3
Matchday 1/2/3 — the three rounds of group-stage fixtures. Each team plays 3 matches (one per matchday). MD3 matches within a group kick off simultaneously to prevent collusion.
Dead Rubber
A match where neither team has anything competitive at stake — both already eliminated, both already qualified, or the result can't change tournament position. Affects predictions: motivation is uncertain, lineups rotate, intensity drops, Finding 12 (conservation) behaves differently.
| Type | Example | Prediction impact |
|---|---|---|
| True dead rubber | BiH-QAT (both eliminated), USA-TUR (qualified vs eliminated) | Unreliable — rotation + motivation dominate |
| Soft dead rubber | GER-ECU (qualified vs barely-alive), MEX-CZE (group won vs must-win) | Semi-reliable — one team has pride/stakes |
| NOT a dead rubber | CAN-SUI (1st-place seeding mattered) | Should always be staged |
RAPID Analysis
A no-gathering ensemble — writers fire from MD1 form only, no live gathering (scout/injury/referee), no T-60 confirmed lineups, no match-day weather. Used when T-8 or less to kickoff (full pipeline can't fit). All XIs flagged PROJECTED-UNCONFIRMED. Validated: reliable for clear favorites (>65%) + generational scorers; unreliable for coin-flips.
T-60 Workflow
The standard pre-match timing: gather early (T-24h+), then pull confirmed XIs ~60 minutes before kickoff, then fire the ensemble. Ensures lineups are confirmed, not projected.
TASK-99 (Independent Resolution)
Pipeline change (Jun 23, 2026): each match-writer resolves the 6-step decision graph (doc-5) independently, instead of the lead pre-resolving it and passing a synthesis block. Produces honest spreads (3-7pt) instead of artificial 1pt convergence. Pre-TASK-99 per-bot stats were noise (all 3 anchored on lead's synthesis); post-TASK-99 they're genuine signal.
TASK-97 (Post-Event Writer + Rubric)
Pipeline change (Jun 23, 2026): a post-event-writer agent + 100-point calibration rubric (doc-1) replaces the lead's vibe-based ~X% scoring. Components: Winner (0-25), Scoreline (0-25), Who Scores (0-20), Findings (0-20), Model-breaking (0-10), with −5 deductions. Bands: A (90+), B (75-89), C (60-74), D (45-59), F (<45).
Agents (7 total)
Gathering Layer (run in PARALLEL)
| Agent | Model | Role |
|---|---|---|
match-scout |
MiniMax-M3 | Venue, roof, weather, H2H, group context, managers, tiers |
injury-sweeper |
MiniMax-M3 | Injuries, suspensions, GK confirmation, projected XIs (both languages) |
referee-profiler |
MiniMax-M3 | Referee card/pen rates, strictness, tournament discount |
Synthesis Layer (run in PARALLEL, 3 models for ensemble)
| Agent | Model | Notes |
|---|---|---|
match-writer |
MiniMax-M3 | Default. Upset/draw specialist (highest winner rate, widest draw band). |
match-writer-g4 |
glm-5-turbo | Scoreline/blowout specialist (lowest avg distance, fires Finding 22 most aggressively). |
match-writer-g5 |
zai/glm-5.2 | Best discipline (cleanest Finding reasoning). Weakest scoreline precision. |
match-writer-g51 |
zai/glm-5.1 | A/B control pin during glm-5.2 validation window. |
Post-Event Layer
| Agent | Model | Role |
|---|---|---|
post-event-scout |
MiniMax-M3 | Gathers 12-category factual brief (score, stats, lineups). Does NOT write the review. |
post-event-writer |
zai/glm-5.2 | Writes the rubric-scored calibration review from the scout brief. |
accuracy-reporter |
zai/glm-5.2 | Reads computed metrics JSON, writes per-agent narrative accuracy report. |
The 22 Calibration Findings (doc-1)
The empirical rules validated across the tournament. Key ones referenced frequently:
| # | Finding | Summary |
|---|---|---|
| 1 | Draw Epidemic | 34% draw rate; draw band 25-33% for low-block matchups |
| 4 / 4a | GK Changes / Attack-Gutted | GK changes are model-breaking (Finding 14). Losing 2+ of {creator, striker, GK} = gutting (-25-30pt). SINGLE absence = -3-5pt only. |
| 6 | Hot-GK | Veteran/GK heroics force draws (Vozinha, Room, Mpasi, Mendy, Beiranvand, Asare). 15-20% draw-floor band. |
| 11 | Physical/Environmental | Heat (>28°C) degrades pressing by 60'. Neutralized if roof closed. Altitude (>2000m) = pace-tax. Rain ≠ heat (different mechanism). |
| 12 | Tournament Conservation | Context-dependent. Defending champions REDUCE discount (make statements). Manager personality > tier. Milestone motivation = NO conservation. |
| 13 | Sterile-Domination | Tier 1 favorite + low block + hot-GK = draw. Can resolve via HT subs. Can fire on favorites too (ENG-GHA). |
| 14 | GK Changes Model-Breaking | Pre-match GK confirmation critical. Iraq (Hassan→Basil), SEN (Mendy injury), UZB (Yusupov→Nematov), TUR (Çakır→Günok). |
| 15 | Pen Rate as Band | Referee pen rate presented as 30-40% band, NOT near-certainty. High-pen ref ≠ pen likely. |
| 16 | Age-Decline | Applies to scorers 33+ (Modrić, Ronaldo, Son, Arnautović). Hits finishing, not playmaking for deep creators. Can be REFUTED if milestone + finishing mechanism exist. |
| 17 | Sterile-Domination (Direction-Agnostic) | Fires on ANY possession-dominant team (not just favorites). PAN 62% poss → 0 goals. Resolves via aggressive subs. |
| 19 | Sub-Scorer Rule | Fresh legs vs tiring defense. One of the most reliable scorer signals. (Benbouali, Budimir, Trézéguet, Leão.) |
| 20 | Double-Underestimate | Favorite both underpriced on win% AND margin. Fires when 4/4 triggers present (Tier 1 + milestone striker + aggressive mgr + ageing opponent). |
| 22 | Blowout Modes | 4+ goal margins weighted at 15-20% when: elite attack + non-veteran-GK block + (desperate underdog OR MD1-stutter→MD2-statement). Has LIMIT: deep block + veteran GK excludes. g4 systematically right on blowouts. |
Key Concepts
Modal Scoreline
The single most-likely exact score (highest probability in the distribution). E.g., "POR 2-0 at 17%." Hitting the exact modal is rare (~20% of matches) and scores 25/25 on the rubric.
Who Scores #1
The top-ranked scorer pick for each team, with a probability %. The pipeline's scorer predictions are one of its strongest signals — especially for generational milestone-chasing players (Messi, Mbappé, Haaland, Ronaldo all braced when #1 unanimous).
Finding Resolution
Each writer walks doc-5's 6-step decision graph and explicitly states whether each Finding fires (applies), doesn't fire, or fires partially — with reasoning. This is the analytical engine, not just number-crunching.
Spread (Win% Spread)
The range of favorite win% across the 3 models. Pre-TASK-99: 1-2pt (artificial anchoring). Post-TASK-99: 3-7pt (honest uncertainty). The widest-spreading (most defensive) agent tends to be closest on upset/draw matches.
Category 0 Gate
The no-fabrication hard rule: the final score MUST be confirmed from ≥2 independent sources (BBC/Guardian/FotMob/ESPN) before it enters any review or summary. A single stale live-blog is NOT sufficient (CZE-RSA was mis-reported as 1-0; actual was 1-1).
Project Infrastructure
Canonical Profiles (sources/reference/)
Stable, shared intelligence updated per matchday — NOT contamination:
teams/{code}.md— squad, manager, tactical signature, GK, formreferees/{surname}.md— career profile + assignments logvenues/{city-slug}.md— roof, capacity, altitude, climate-control
Per-Match Directory (sources/matches/{code}/)
Thin delta — match-specific only:
scout/,team-news-{t1}/,team-news-{t2}/,referee/,weather/,lineups/,post-event/,archive-superseded/
RAG (Retrieval-Augmented Generation)
Local semantic search database (rag/). 800+ chunks embedded via ollama nomic-embed-text (768d). Search: node .pi/skills/mundial-rag/search.mjs "query". Complements, not replaces, direct profile reading.
TOML Metrics Block
Machine-parseable block at the end of every analysis file ([match], [win_probability], [scorelines], [who_scores], [calibration], [predicted_stats]). Auto-extracted by rag/toml-extract.mjs and agent-accuracy.mjs.
Calibration Rubric (doc-1)
100-point scoring system for post-event reviews:
| Component | Points | Criteria |
|---|---|---|
| Winner | 0-25 | 25 = strict modal won; 20 = in top-3; 15 = co-modal; 0 = upset missed |
| Scoreline | 0-25 | 25 = exact modal; 20 = top-3; 15 = winner + margin ≤1; 10 = margin off 2+ |
| Who Scores | 0-20 | 20 = #1 brace+; 15 = #1 scored; 10 = alt pick; 5 = assist only |
| Findings | 0-20 | 5 each × up to 4 findings the analysis explicitly staked a position on |
| Model-breaking | 0-10 | 10 = all caught; 5 = partial; 0 = key factor missed |
| Deductions | −5 each | Unpredicted sub-scorer; unflagged material factor; misleading review |
Bands: A (90+, exceptional), B (75-89, good), C (60-74, mixed), D (45-59, weak), F (<45, miss).
File Naming Conventions
| Type | Pattern | Example |
|---|---|---|
| Pre-match (per model) | {code}-2026-pure-football-analysis-{mm3|g4|g5}.md |
por-vs-uzb-2026-pure-football-analysis-g4.md |
| Consensus | {code}-2026-consensus-ensemble.md |
por-vs-uzb-2026-consensus-ensemble.md |
| Post-event | {code}-2026-post-event-review.md |
por-vs-uzb-2026-post-event-review.md |
| Match codes | lowercase 3-letter, -vs- |
por-vs-uzb, cze-vs-rsa |
Last updated: June 25, 2026. Living document — updated as the pipeline evolves and new findings are validated.