Argentina vs. Algeria - Post-Event Review

2026 World Cup, Group J, Match 19 | June 17, 2026 | Kansas City Stadium

Predicted vs. Actual | Calibration: ~80% (raw consensus) / ~55% (Finding 12 overlay)

⚠️ THE DEFINITIVE FINDING 12 TEST CASE

TL;DR

Result: Argentina 3-0 Algeria (HT 1-0). Messi hat-trick (17', 60', 76'). Argentina won as predicted β€” but the Finding 12 Conservation Overlay was DECISIVELY WRONG on every single prediction it made beyond the base consensus. Argentina did NOT conserve. Messi played the full 90 and scored his 3rd at 76'. Lautaro was the one subbed, not Messi. The scoreline was a comfortable 3-0, not the tight 1-0/2-0/1-1 cluster the overlay predicted. The raw consensus (62/24/15) was better calibrated than the overlay (55/31/14) on every dimension.

Watch (from the overlay): "Messi's sub minute. Off before 70' = conservation confirmed." β†’ REALITY: Messi played 90', scored at 76'. Conservation REFUTED.


1. The Experiment β€” Raw Consensus vs Finding 12 Overlay

This match was set up as a controlled experiment. Two predictions were published:

Prediction ARG Draw ALG Top scoreline Who Scores #1
Raw consensus (4 models) 62% 24% 15% ARG 2-0 (21%) Messi β‰ˆ Lautaro (~22% each)
Finding 12 overlay 55% 31% 14% ARG 1-0 (18%) Lautaro (~24%) > Messi (~19%)
ACTUAL Win 0% 0% 3-0 Messi (hat-trick)

Verdict: Raw consensus wins on EVERY dimension

Dimension Raw consensus Finding 12 overlay Winner
Winner βœ… ARG (62%) βœ… ARG (55%) Tie (both right, raw more confident)
Scoreline ⚠️ 2-0 modal, 3-0 at 8% (#5) ❌ 1-0 modal β€” 3-0 nowhere near Raw
Draw probability 24% (too high, but closer) 31% (way too high) Raw
Who Scores #1 βœ… Messi co-fav (~22%) ❌ Demoted Messi to ~19% Raw
Sub prediction β€” ❌ "Messi subbed by 65-70'" Raw (Messi played 90')

The Finding 12 overlay made the prediction WORSE on every measurable dimension. This is the clearest possible refutation of over-applying conservation to a defending champion.


2. Why Finding 12 Failed Here β€” The Psychology Was Backwards

The overlay's core thesis: "Argentina are the defending champions with maximum marathon incentive β†’ maximum conservation β†’ lean to upper end (7-8%)."

The reality was the exact opposite. Argentina attacked aggressively from the opening whistle:

The psychology was backwards. A defending champion doesn't coast on MD1 β€” they want to make a STATEMENT. They have something to prove. "We are still the champions." Conservation is for teams protecting a lead in the marathon; statement-making is for teams asserting dominance at the start.

The France/Argentina contrast

This is the key insight. Both are Tier 1 contenders, but they behaved OPPOSITELY:

France (FRA-SEN) Argentina (ARG-ALG)
Manager Deschamps (notoriously conservative) Scaloni (aggressive, freedom-based)
MD1 behavior 0-0 at HT, sluggish, conserved Goal at 17', aggressive from the start
Star subbed? No (MbappΓ© played 90') No (Messi played 90')
Result Won 3-1 (late goals) Won 3-0 (early + late goals)
Conservation? βœ… YES ❌ NO

Finding 12 is not just about tier β€” it's about the TEAM'S PERSONALITY and the MANAGER'S APPROACH. Deschamps' France naturally conserves. Scaloni's Argentina naturally attacks. The rule needs a personality modifier.


3. The Messi Milestone Factor

Individual milestone motivation overrides team-level conservation.

Messi entered this match chasing:

He equalled the record with a hat-trick. The Guardian's Pablo Iglesias Maurer quote captures it perfectly:

"Messi, put simply, is in extra time at this point... Entirely unburdened, the Argentinian is playing his final World Cup free from the expectations... That sort of freedom can liberate and empower a player."

When a star player is chasing a record, they will NOT be conserved or subbed early. The milestone is bigger than the marathon. This is a hard rule: milestone motivation = no conservation for that player.

The overlay explicitly predicted "Messi more likely to be subbed by 65-70'" β€” this was the single worst call in the entire tournament. Messi played 90' and scored his 3rd at 76'.


4. Calibration Scorecard (Raw Consensus)

Category Prediction Actual Hit?
Winner Argentina (62%) Argentina 3-0 βœ…
Scoreline (modal) ARG 2-0 (21%) ARG 3-0 ⚠️ (3-0 at 8%, #5)
Who Scores #1 Messi (~22%, co-fav) Messi hat-trick βœ…βœ…βœ…
Penalty 30-40% chance No penalty ❌ (but 60-70% chance of no pen)
Algeria defense Low-block + clean-sheet streak Shattered at 17' ❌
Marciniak cards Tournament-discounted Minimal (Messi escaped a card) βœ… (lenient)
Tagliafico cover Facundo Medina at LB Medina started at LB βœ…

Overall (raw consensus): ~80% β€” Winner βœ…, Who Scores βœ…βœ…βœ…, cards βœ…, lineup βœ…. Misses: exact scoreline and the penalty prediction.


5. Version-by-Version Comparison

Version Win% (ARG/DRAW/ALG) Scoreline #1 3-0 in predictions? Messi as #1?
A (human) 66/20/14 ARG 2-0 (21%) 3-0 ~8% Messi or Lautaro
mm3 63/22/15 ARG 2-0 (18%) ❌ Messi
g4 62/24/14 ARG 2-0 (21%) ❌ Lautaro
g5 58/27/15 ARG 2-0 (15%) ❌ Messi
Consensus 62/24/15 ARG 2-0 (21%) ~8% (#5) Messi β‰ˆ Lautaro

No version predicted 3-0 as a top-3 scoreline. Every model picked ARG 2-0. The margin was slightly underestimated across the board β€” but not as badly as IRA-NOR. The 3-0 was in the human version's top 5 (8%), which is why the raw consensus outperformed.

The Messi/Lautaro split was resolved decisively in Messi's favor. mm3 and g5 picked Messi; g4 picked Lautaro. The human (A) hedged. Reality: Messi hat-trick, Lautaro subbed. mm3 and g5 were right; g4 was wrong again on scorer #1 (consistent with the g4 scorer-outlier pattern from the benchmark).


6. Finding 12 Calibration β€” What Must Change

This match provides the clearest calibration data for Finding 12 in the entire tournament. The rule needs THREE modifications:

Fix 1: Defending champions β†’ REDUCE or ELIMINATE the conservation discount

❌ Current rule (wrong): "Lean to the upper end (7-8%) for defending champions β€” maximum marathon incentive."

βœ… Corrected rule: "Defending champions want to make a STATEMENT in their opener, not coast. They have something to prove (Argentina lost openers in 1982/1990). REDUCE the conservation discount to 0-3% for defending champions, or eliminate it entirely. The 'maximum marathon incentive' psychology was backwards β€” defending champions assert dominance, they don't protect a lead."

Fix 2: Team personality matters more than tier

Conservation applies to: Conservative/defensive managers (Deschamps' France, Bielsa when fatigued)
Conservation does NOT apply to: Aggressive/statement-making managers (Scaloni's Argentina)

Add a personality modifier: aggressive managers (Scaloni, Nagelsmann, de la Fuente when motivated) discount less. Conservative managers (Deschamps, Southgate) discount more.

Fix 3: Individual milestone motivation = no conservation for that player

Hard rule: When a star player is chasing a record (Messi β†’ Klose's WC scoring record), they will NOT be conserved or subbed early. The milestone overrides the marathon. Do not apply substitution risk to milestone-chasing stars.


7. The Penalty Miss

Marciniak's 0.42/game penalty rate produced NO penalty. The prediction was 30-40% pen chance β€” the highest of any match analyzed. All 3 Messi goals were from open play (curler, tap-in, one-two finish).

This isn't a calibration error β€” 30-40% means 60-70% chance of NO penalty. The model correctly identified Marciniak as penalty-prone (the highest in the dataset), and a VAR goal was disallowed early (Marciniak WAS involved in box drama). The penalty just didn't happen. This is irreducible referee variance.

However, Marciniak WAS lenient on cards β€” Messi escaped a yellow for a reckless challenge on Mandi. Finding 3 (tournament leniency) validated for the 8th consecutive time.


8. Key Learnings

  1. Finding 12 OVER-APPLIED on defending champions. The overlay made the prediction worse on every dimension. The psychology was backwards: defending champions make STATEMENTS, they don't conserve. Fix: reduce to 0-3% or eliminate for defending champions.

  2. Team personality > tier for conservation. Deschamps' France conserves; Scaloni's Argentina attacks. The rule needs a manager-personality modifier.

  3. Milestone motivation overrides conservation. Messi chasing Klose's record played 90' and scored a hat-trick. Never apply substitution risk to milestone-chasing stars.

  4. The raw consensus (62/24/15) outperformed the overlay (55/31/14). This validates the multi-model ensemble approach β€” the grand average was closer than the human overlay. The lesson: trust the model consensus over single-analyst overlays, even when the overlay seems theoretically sound.

  5. Algeria's defensive streak was overvalued. 393 minutes without conceding ended at 17'. Finding 5 (continental pedigree) doesn't make a team draw-proof against a Tier 1 contender in statement mode.

  6. g4 wrong on scorer again (Lautaro over Messi). Consistent with the g4 scorer-outlier pattern. mm3 and g5 correctly picked Messi.

  7. Messi hat-trick = best Who Scores hit of the tournament (tied with MbappΓ©'s brace). When the consensus picks a generational scorer as #1, they deliver.


9. Updated Cumulative Stats (17 matches)

Metric Before (16) After ARG-ALG (17) Trend
Winner hit 50% 53% (9/17) ↑
Exact scoreline top-3 38% 35% (6/17) ↓ (3-0 was #5)
Who Scores β‰₯1 hit 69% 71% (12/17) ↑ (Messi hat-trick!)
Referee red cards 100% 100% (0 reds) β€”
Overall ~71% ~72% ↑

Post-event review for TASK-43. Sources archived: sources/matches/arg-vs-alg/post-event/01-guardian-bbc-report.md. This is the definitive Finding 12 calibration data point. No betting odds used. FINAL stats only.