World Cup 2026 Project — Glossary

Terms used throughout the pipeline. Football-specific, project-specific, and pipeline-specific jargon.


Pipeline & Architecture

MD1 / MD2 / MD3

Matchday 1/2/3 — the three rounds of group-stage fixtures. Each team plays 3 matches (one per matchday). MD3 matches within a group kick off simultaneously to prevent collusion.

Dead Rubber

A match where neither team has anything competitive at stake — both already eliminated, both already qualified, or the result can't change tournament position. Affects predictions: motivation is uncertain, lineups rotate, intensity drops, Finding 12 (conservation) behaves differently.

Type Example Prediction impact
True dead rubber BiH-QAT (both eliminated), USA-TUR (qualified vs eliminated) Unreliable — rotation + motivation dominate
Soft dead rubber GER-ECU (qualified vs barely-alive), MEX-CZE (group won vs must-win) Semi-reliable — one team has pride/stakes
NOT a dead rubber CAN-SUI (1st-place seeding mattered) Should always be staged

RAPID Analysis

A no-gathering ensemble — writers fire from MD1 form only, no live gathering (scout/injury/referee), no T-60 confirmed lineups, no match-day weather. Used when T-8 or less to kickoff (full pipeline can't fit). All XIs flagged PROJECTED-UNCONFIRMED. Validated: reliable for clear favorites (>65%) + generational scorers; unreliable for coin-flips.

T-60 Workflow

The standard pre-match timing: gather early (T-24h+), then pull confirmed XIs ~60 minutes before kickoff, then fire the ensemble. Ensures lineups are confirmed, not projected.

TASK-99 (Independent Resolution)

Pipeline change (Jun 23, 2026): each match-writer resolves the 6-step decision graph (doc-5) independently, instead of the lead pre-resolving it and passing a synthesis block. Produces honest spreads (3-7pt) instead of artificial 1pt convergence. Pre-TASK-99 per-bot stats were noise (all 3 anchored on lead's synthesis); post-TASK-99 they're genuine signal.

TASK-97 (Post-Event Writer + Rubric)

Pipeline change (Jun 23, 2026): a post-event-writer agent + 100-point calibration rubric (doc-1) replaces the lead's vibe-based ~X% scoring. Components: Winner (0-25), Scoreline (0-25), Who Scores (0-20), Findings (0-20), Model-breaking (0-10), with −5 deductions. Bands: A (90+), B (75-89), C (60-74), D (45-59), F (<45).


Agents (7 total)

Gathering Layer (run in PARALLEL)

Agent Model Role
match-scout MiniMax-M3 Venue, roof, weather, H2H, group context, managers, tiers
injury-sweeper MiniMax-M3 Injuries, suspensions, GK confirmation, projected XIs (both languages)
referee-profiler MiniMax-M3 Referee card/pen rates, strictness, tournament discount

Synthesis Layer (run in PARALLEL, 3 models for ensemble)

Agent Model Notes
match-writer MiniMax-M3 Default. Upset/draw specialist (highest winner rate, widest draw band).
match-writer-g4 glm-5-turbo Scoreline/blowout specialist (lowest avg distance, fires Finding 22 most aggressively).
match-writer-g5 zai/glm-5.2 Best discipline (cleanest Finding reasoning). Weakest scoreline precision.
match-writer-g51 zai/glm-5.1 A/B control pin during glm-5.2 validation window.

Post-Event Layer

Agent Model Role
post-event-scout MiniMax-M3 Gathers 12-category factual brief (score, stats, lineups). Does NOT write the review.
post-event-writer zai/glm-5.2 Writes the rubric-scored calibration review from the scout brief.
accuracy-reporter zai/glm-5.2 Reads computed metrics JSON, writes per-agent narrative accuracy report.

The 22 Calibration Findings (doc-1)

The empirical rules validated across the tournament. Key ones referenced frequently:

# Finding Summary
1 Draw Epidemic 34% draw rate; draw band 25-33% for low-block matchups
4 / 4a GK Changes / Attack-Gutted GK changes are model-breaking (Finding 14). Losing 2+ of {creator, striker, GK} = gutting (-25-30pt). SINGLE absence = -3-5pt only.
6 Hot-GK Veteran/GK heroics force draws (Vozinha, Room, Mpasi, Mendy, Beiranvand, Asare). 15-20% draw-floor band.
11 Physical/Environmental Heat (>28°C) degrades pressing by 60'. Neutralized if roof closed. Altitude (>2000m) = pace-tax. Rain ≠ heat (different mechanism).
12 Tournament Conservation Context-dependent. Defending champions REDUCE discount (make statements). Manager personality > tier. Milestone motivation = NO conservation.
13 Sterile-Domination Tier 1 favorite + low block + hot-GK = draw. Can resolve via HT subs. Can fire on favorites too (ENG-GHA).
14 GK Changes Model-Breaking Pre-match GK confirmation critical. Iraq (Hassan→Basil), SEN (Mendy injury), UZB (Yusupov→Nematov), TUR (Çakır→Günok).
15 Pen Rate as Band Referee pen rate presented as 30-40% band, NOT near-certainty. High-pen ref ≠ pen likely.
16 Age-Decline Applies to scorers 33+ (Modrić, Ronaldo, Son, Arnautović). Hits finishing, not playmaking for deep creators. Can be REFUTED if milestone + finishing mechanism exist.
17 Sterile-Domination (Direction-Agnostic) Fires on ANY possession-dominant team (not just favorites). PAN 62% poss → 0 goals. Resolves via aggressive subs.
19 Sub-Scorer Rule Fresh legs vs tiring defense. One of the most reliable scorer signals. (Benbouali, Budimir, Trézéguet, Leão.)
20 Double-Underestimate Favorite both underpriced on win% AND margin. Fires when 4/4 triggers present (Tier 1 + milestone striker + aggressive mgr + ageing opponent).
22 Blowout Modes 4+ goal margins weighted at 15-20% when: elite attack + non-veteran-GK block + (desperate underdog OR MD1-stutter→MD2-statement). Has LIMIT: deep block + veteran GK excludes. g4 systematically right on blowouts.

Key Concepts

The single most-likely exact score (highest probability in the distribution). E.g., "POR 2-0 at 17%." Hitting the exact modal is rare (~20% of matches) and scores 25/25 on the rubric.

Who Scores #1

The top-ranked scorer pick for each team, with a probability %. The pipeline's scorer predictions are one of its strongest signals — especially for generational milestone-chasing players (Messi, Mbappé, Haaland, Ronaldo all braced when #1 unanimous).

Finding Resolution

Each writer walks doc-5's 6-step decision graph and explicitly states whether each Finding fires (applies), doesn't fire, or fires partially — with reasoning. This is the analytical engine, not just number-crunching.

Spread (Win% Spread)

The range of favorite win% across the 3 models. Pre-TASK-99: 1-2pt (artificial anchoring). Post-TASK-99: 3-7pt (honest uncertainty). The widest-spreading (most defensive) agent tends to be closest on upset/draw matches.

Category 0 Gate

The no-fabrication hard rule: the final score MUST be confirmed from ≥2 independent sources (BBC/Guardian/FotMob/ESPN) before it enters any review or summary. A single stale live-blog is NOT sufficient (CZE-RSA was mis-reported as 1-0; actual was 1-1).


Project Infrastructure

Canonical Profiles (sources/reference/)

Stable, shared intelligence updated per matchday — NOT contamination:

Per-Match Directory (sources/matches/{code}/)

Thin delta — match-specific only:

RAG (Retrieval-Augmented Generation)

Local semantic search database (rag/). 800+ chunks embedded via ollama nomic-embed-text (768d). Search: node .pi/skills/mundial-rag/search.mjs "query". Complements, not replaces, direct profile reading.

TOML Metrics Block

Machine-parseable block at the end of every analysis file ([match], [win_probability], [scorelines], [who_scores], [calibration], [predicted_stats]). Auto-extracted by rag/toml-extract.mjs and agent-accuracy.mjs.

Calibration Rubric (doc-1)

100-point scoring system for post-event reviews:

Component Points Criteria
Winner 0-25 25 = strict modal won; 20 = in top-3; 15 = co-modal; 0 = upset missed
Scoreline 0-25 25 = exact modal; 20 = top-3; 15 = winner + margin ≤1; 10 = margin off 2+
Who Scores 0-20 20 = #1 brace+; 15 = #1 scored; 10 = alt pick; 5 = assist only
Findings 0-20 5 each × up to 4 findings the analysis explicitly staked a position on
Model-breaking 0-10 10 = all caught; 5 = partial; 0 = key factor missed
Deductions −5 each Unpredicted sub-scorer; unflagged material factor; misleading review

Bands: A (90+, exceptional), B (75-89, good), C (60-74, mixed), D (45-59, weak), F (<45, miss).


File Naming Conventions

Type Pattern Example
Pre-match (per model) {code}-2026-pure-football-analysis-{mm3|g4|g5}.md por-vs-uzb-2026-pure-football-analysis-g4.md
Consensus {code}-2026-consensus-ensemble.md por-vs-uzb-2026-consensus-ensemble.md
Post-event {code}-2026-post-event-review.md por-vs-uzb-2026-post-event-review.md
Match codes lowercase 3-letter, -vs- por-vs-uzb, cze-vs-rsa

Last updated: June 25, 2026. Living document — updated as the pipeline evolves and new findings are validated.