About - Methodology

TL;DR

No odds. No expert picks. Just verifiable football data. Every match prediction on this site is built from scratch using results, player profiles, injury reports, tactical systems, and referee statistics. We archive every source, review our accuracy after each match, and calibrate the model against what actually happened. After 11 matches, the model is running at ~71% overall accuracy β€” and getting better.


Why This Exists

Most football predictions are built on betting markets and media consensus. This project does the opposite. Every analysis starts from raw football data and builds its own probability model β€” no sportsbook odds, no pundit aggregators, no prediction markets.

The goal is simple: can pure football analysis outperform the market?


Core Principles

  1. No betting odds or expert picks β€” Sportsbook odds, prediction markets (Kalshi, Polymarket), and media pundit predictions are ignored entirely. They reflect market sentiment, not football reality.

  2. Verifiable data only β€” Every claim traces to a specific source: match results, squad announcements, injury reports, tactical descriptions, referee statistics.

  3. Own the prediction β€” Probability rankings are built from raw analysis, not aggregated from others' predictions.

  4. Archive everything β€” All source articles are saved for post-event review. Sources are rated on quality and tracked for accuracy over time.

  5. Calibrate continuously β€” After every match, empirical findings are fed back into the model. The methodology is a living document, updated with what the tournament actually teaches us.


What Goes Into Each Analysis

Team Assessment

⚠️ Injury & Lineup Check (added after ESP-CPV miss)

A dedicated injury sweep is run for every team, separate from tactical research:

Tactical Analysis

Referee Profiling

Context

Continental Pedigree Check

Tournament Energy Management (Context-Dependent Conservation)

A World Cup is a 7-match marathon β€” but not all contenders conserve energy in group matches. It depends on manager personality and motivation, not just tier.

Validated Jun 17 (ARG-ALG): Argentina, the defending champions, did NOT conserve β€” they scored at 17', Messi played 90' and scored a hat-trick. They wanted to make a STATEMENT (first opener win as defending champions since 1982/1990 losses). The original hypothesis (defending champions conserve MORE) was decisively wrong.

Calibration rule (corrected):


Report Types

Type What It Covers When
Pure Football Analysis Full pre-match breakdown with scenario probabilities Before kickoff
Live Match Report Real-time reassessment based on match events During the match
Post-Event Review Prediction accuracy, referee profile validation, lessons learned After full-time
Predictions Summary Cumulative accuracy tracker across all matches Rolling

Every report opens with a TL;DR in a tight bullet format: 1 bold headline line (win% / scoreline / who scores) β†’ 4-6 short bullets (key findings, no paragraphs) β†’ 1 "Watch:" line (the single swing variable). Keep it scannable, not a text wall.


How Predictions Work

Each match produces 5 ranked scoreline scenarios with estimated probabilities that sum to ~100%. Every scenario is justified by specific evidence β€” not gut feeling. A Who Scores section ranks the most likely scorers per team, each with a probability and the most likely scoring mechanism (open play, set piece, penalty, counter).

After each match, a post-event review compares:


Tournament Results & Calibration (11 matches, Jun 11-15, 2026)

Accuracy Snapshot

Component Hit Rate Grade
Winner prediction 4.5/10 (45%) 🟑 Moderate β€” draws underpriced
Exact scoreline (top 3) 3/10 (30%) πŸ”΄ Margins systematically wrong
Who Scores β‰₯1 hit 6/9 (67%) 🟒 Good
Who Scores #1 hit 5/9 (56%) 🟒 Good
Possession/flow 8/10 (80%) 🟒 Good
Referee red cards 9/9 (100%) 🟒 Perfect
Referee card count 5/9 (56%) 🟑 Tournament leniency
Overall average ~71% 🟒 Good β€” improving

Key Calibration Findings

These are empirical patterns observed in actual match data, now baked into every prediction:

  1. Draws are systematically underpriced β€” 4 of 10 matches ended draws (40-50% rate). For evenly-matched teams, draw probability is now set at 25-35%.

  2. Margins fail in two opposite directions β€” blowouts are underestimated (GER 7-1, SWE 5-1) AND margins are overestimated vs defensive teams (ESP 0-0 vs Cape Verde). Both failure modes now have scenarios.

  3. Tournament referee leniency β€” card counts come in ~15-20% below career averages. Red cards are almost always 0 in group stage.

  4. Lineup news can invalidate the whole model β€” the biggest misses came from lineup surprises (Yamal benched vs CPV, GK change in AUS-TUR). The injury sweep + lineup check now catches this.

  5. AFCON/continental pedigree β‰  minnow status β€” Cape Verde (FIFA #68, AFCON QF) held Spain (FIFA #2) to 0-0. Continental tournament results matter more than FIFA ranking.

  6. The low-block + hot goalkeeper problem β€” Spain generated 2.29 xG and scored 0 against Vozinha (40yo GK). A hot goalkeeper is impossible to price β€” acknowledged as irreducible uncertainty.

  7. CB set-piece goals are underpriced β€” Schlotterbeck and Van Dijk both scored from set pieces. CBs are now included as Who Scores options.

  8. Possession can flip β€” defensive teams can win the possession battle. Possession is a flow indicator, not a result predictor.

  9. Coin-flip framework β€” for evenly-matched teams, either side winning validates the model's balance. Winner prediction and exact scoreline are tracked as separate categories.

  10. Comeback probability β€” tactically flexible teams (Japan, South Korea) have higher 2nd-half comeback odds. Pre-tournament friendlies (Tunisia 5-0 vs Belgium) are real signals.

  11. Physical & environmental conditions are a first-class input β€” every tournament "shock" had a physical explanation hiding in plain sight. Uruguay played at "~60% energy" in 32Β°C Miami heat (flight delays + base camp commute); Iran conceded at 7' after a Tijuana base camp + same-day US travel; Spain and Belgium were sluggish. A physically depleted favorite plays at ~60% energy β€” functionally a ~15-20% weaker team. Rules: heat (>28Β°C + humidity) degrades pressing teams by ~60'; same-day travel costs ~5-8% off the favorite; intercontinental teams need 5-7 days acclimatization; high-energy pressing styles (Bielsa) are uniquely heat-vulnerable. This stacks with the attack-gutting rule for compounding downgrades.


Source Strategy

Sources are searched in both teams' local languages, not just English. Bosnian/Croatian sources provided lineup scoops for Canada vs BiH that were unavailable in any English outlet. Spanish-language sources uncovered the USA-Paraguay brawl narrative. Arabic, Korean, Persian, and Portuguese sources have all contributed unique data.

All sources are archived in a structured directory (sources/matches/{match}/) and rated for quality post-event.


The Calibration Adjustment Checklist

Before finalizing any prediction, the model runs through a 10-point checklist:

  1. Lineup check β€” confirmed starting XI? Any stars benched?
  2. Draw weight β€” is the draw probability high enough?
  3. Blowout scenario β€” is there a 4+ goal margin scenario for quality gaps?
  4. Low-block discount β€” is the opponent defensively organized with a capable GK?
  5. Referee leniency β€” discount career cards by 20-25% (was 15-20%, upgraded after 5 consecutive lenient performances); check personality.
  6. Continental pedigree β€” is the "minnow" actually a continental quarterfinalist?
  7. CB set-piece β€” include an aerial CB as a Who Scores option.
  8. GK heroics β€” acknowledge the ~15-20% "hot goalkeeper saves a draw" scenario (was 8-12%, upgraded after 4 validations: Vozinha, Beach, Shobeir, Al-Owais).
  9. Coin-flip framing β€” are the teams within 5-10 points? Increase draw weight.
  10. Final stats only β€” never use halftime figures.
  11. Physical & environmental β€” heat >28Β°C? Same-day travel? Disrupted base camp? High-energy pressing style? If yes, the favorite plays at ~60% energy β€” add 5-10% to draw/underdog. Weather MUST use the match-day NWS/official forecast with timestamp β€” NOT general climate guides (a stale guide caused a 3-5% error on FRA-SEN). Same-day travel here means INTERCONTINENTAL or major disruption, not a routine same-state bus ride.
  12. Star-as-creator β€” model the best player ASSISTING as well as scoring (Salah, Wood both assisted twice).
  13. Substitution risk β€” a wasteful #9 has >30% chance of being subbed by 60'. Don't assume your #1 scorer plays 90'.
  14. Goalkeeper confirmed β€” GK changes are model-breaking. Confirm before kickoff (FRA-SEN had an unresolved Mendy-vs-Diaw ambiguity).

Notable Matches

Spain 0-0 Cape Verde (Jun 15) β€” the model's biggest miss and most important lesson. FIFA #2 vs FIFA #68, predicted ESP 3-0, actual 0-0. Yamal + Williams benched (hamstrings) invalidated the attacking model. Vozinha (40yo GK) produced a 1-in-10 masterclass. Led directly to the injury-sweep process and the low-block finding.

Saudi Arabia vs Uruguay (Jun 15) β€” first real application of the new lineup check. An injury sweep hours before kickoff revealed Uruguay's entire CB pairing (AraΓΊjo + GimΓ©nez) and chief creator (De Arrascaeta) were OUT. Uruguay's win probability was revised from 65% to 49% before kickoff.

Sweden 5-1 Tunisia (Jun 14) β€” best Who Scores result: both #1 (GyΓΆkeres) AND #2 (Isak) predicted correctly, the first #1+#2 double hit of the tournament.


Who Runs This

This is an independent football analysis project for the 2026 FIFA World Cup. No affiliations with bookmakers, media outlets, or football governing bodies.

Analyses are produced using a multi-agent research pipeline: specialized AI agents gather team news, injury data, and referee profiles in parallel (searching in both English and local languages), then a synthesis agent applies this calibrated methodology to produce the pre-match analysis. High-stakes matches may run through multiple models and be averaged into a consensus. All raw sources are archived for post-event review.

Methodology last updated: June 16, 2026 (14 matches calibrated, 11 Calibration Findings, multi-agent pipeline validated on FRA-SEN)