France vs. Senegal - Consensus / Ensemble Analysis

2026 World Cup, Group I | June 16, 2026 | MetLife Stadium, East Rutherford, NJ

Kickoff: 3:00 PM ET (21:00 CEST) | Referee: Alireza Faghani (Australia/Iran)

What this is: A meta-analysis averaging five independent predictions for the same match — two human-authored (A = original, B = scout-corrected) and three model-authored (MiniMax-M3, grok-4.3, glm-5.1), all run on identical, clean source inputs (sources/matches/fra-vs-sen-B/). The purpose is to (1) find the consensus call, (2) measure where human and machine judgment diverge, and (3) validate the multi-agent research pipeline. All five source analyses are linked at the bottom.


TL;DR

France clear-but-not-dominant favorite: ~56% / ~25% / ~19%. Machine ensemble more cautious (~53/27/20). Top scoreline: FRA 2-1 (~17%).

Watch: Senegal's confirmed GK at kickoff. If it's Diaw not Mendy, the defensive projection shifts.


1. The Experiment

Five analyses of France vs Senegal were produced and compared:

ID Author Model Notes
A Human (lead analyst) Original. Used contaminated weather source (30-32°C general guide). Superseded by B.
B Human (lead analyst) Scout-corrected. Uses real NWS match-day forecast (26°C). The clean human baseline.
mm3 match-writer agent MiniMax-M3 Caught the GK ambiguity (Mendy vs Diaw). Most cautious on France.
g4 match-writer-g4 agent grok-4.3 Vivid prose; highest Mbappé scorer weight.
g5 match-writer-g5 agent glm-5.1 Cleanest Finding 11 reasoning; best finding-discipline.

All model runs used the same self-contained source directory (fra-vs-sen-B/) with the contaminated A-version weather archived away, so divergence reflects genuine judgment, not input noise.


2. Win Probability Consensus

All five versions side-by-side

Version Author France Draw Senegal
A (superseded) Human 60% 22% 18%
B (clean human) Human 63% 21% 16%
mm3 MiniMax-M3 48% 28% 24%
g4 grok-4.3 55% 28% 17%
g5 glm-5.1 55% 25% 20%

Consensus calculations

Ensemble France Draw Senegal Spread on France
Machine only (mm3+g4+g5) 53% 27% 20% 48–55%
Human only (A+B) 61% 22% 17% 60–63%
All five averaged 56% 25% 19% 48–63%

Headline consensus: France ~56% / Draw ~25% / Senegal ~19%

(The machine-only ensemble, 53/27/20, is the more defensible number — it excludes the contaminated A and is the true multi-model signal.)


3. Scoreline Consensus

Averaging the top scorelines across all versions:

Score Consensus % How it's read
France 2-1 ~17% Modal pick — 3 of 5 versions rank it #1 or near-top
1-1 Draw ~17% The draw is a near co-favorite scoreline (models heavier than human)
France 1-0 ~15% Tight Mbappé-decides-it win
France 2-0 ~13% Clean control
Senegal 1-0 ~8% 2002 rerun via counter + GK heroics
France 3-1 ~8% Quality gap opens late
2-2 Draw ~5% Open game

Top-3 consensus scorelines: France 2-1, 1-1 Draw, France 1-0. Notably, the draw is effectively tied with the favorite win as the top single scoreline — a direct consequence of the tournament's draw epidemic (Finding 1) that every model weighted.


4. Who Scores Consensus

🇫🇷 France — strong agreement

Rank Player Consensus Range across versions
#1 Kylian Mbappé ~25% 22% (A/B/mm3) · 32% (g4) · 26% (g5)
#2 Michael Olise ~12% Promoted by B & g5 after his Jun 8 hat-trick
#3 Ousmane Dembélé ~12% Consistent across versions

All five agree Mbappé is #1. g4 is the outlier at 32% (over-aggressive); the consensus ~25% is sounder.

🇸🇳 Senegal — real disagreement (this is the interesting signal)

Player Human models (A/B) Machine models (mm3/g4/g5) Consensus
Nicolas Jackson (striker, pace) #1 (~15%) lower / #2 ~13%
Sadio Mané (winger, big-game) #2 (~12%) #1 (~14-16%) ~13%

The humans favor Jackson (the pace striker); the machines favor Mané (the experienced leader). This is a genuine analytical split: Jackson is the higher-upside, higher-variance pick (pace in behind, but disciplinary/substitution risk); Mané is the floor pick (100+ caps, last WC, big-game temperament). The consensus rates them roughly co-favorites at ~13% each. Ismaïla Sarr sits at #3 (~8%) across the board.


5. Where Everyone Agrees

  1. France are favorites, but not dominantly. No version put France above 63%; all land in the 48–63% band — a clear edge, not a stroll.
  2. Mbappé is the #1 scorer. Unanimous. The only disagreement is magnitude (22–32%).
  3. Conditions are comfortable (26°C). The contamination fix held — every clean version correctly drops the heat thesis. No version applies a heat penalty.
  4. The draw is live (22–28%). Every version rates the draw ≥22%, reflecting the tournament's draw epidemic (Finding 1, 50% draw rate).
  5. Mendy / hot-GK is the key draw driver (Finding 6, 15-20%). Senegal's CL-winning keeper is the mechanism by which a draw or upset happens.
  6. Mbappé vs Koulibaly is the decisive matchup, with Faghani's 0.29 pen/game rate making a penalty a ~15-20% live scenario.

6. Where They Disagree (the analytical signal)

6.1 Human vs machine: the 9-point France gap

Humans (A+B) average 61% France; machines (mm3+g4+g5) average 53% — a 9-point spread. The gap comes almost entirely from one input: the travel asymmetry. All three models weighted France's 3.5-4 hour matchday bus ride from Waltham, MA as a meaningful fatigue/ freshness factor (−2 to −3%). The human models judged this over-weighted — a same-state bus the day before is routine for elite athletes, and Finding 11's same-day-travel rule is really about intercontinental disruption, not regional ground transport.

Verdict: the humans are likely right to discount the bus. But the machines' caution is not unreasonable — it's a defensible reading of Finding 11. The grand average (56%) is a fair hedge.

6.2 The Senegal #1 scorer (Jackson vs Mané)

Covered in §4. The models lean Mané; humans lean Jackson. The unresolved variable that would resolve this: the confirmed starting XI. If Jackson leads the line, his pace-vs-high-line upside rises; if Mané plays centrally, his big-game floor dominates.

6.3 g4's Mbappé outlier (32%)

g4 rates Mbappé at 32% — far above the ~22-26% cluster. g4 was also the model that, in the contaminated run, re-imported the heat thesis. This suggests g4 has a tendency toward confident-sounding extremes — useful for vivid narrative, riskier for calibrated probabilities. Its other numbers are sound; the Mbappé figure is the one to discount.


7. The One Variable That Could Break the Model

⚠️ Senegal's goalkeeper is unconfirmed.


8. Final Consensus Call

Outcome Probability Confidence
France win ~56% (machine ensemble: 53%) High agreement on direction
Draw ~25% (machine ensemble: 27%) High agreement
Senegal win ~19% (machine ensemble: 20%) Moderate agreement

Top scoreline: France 2-1 (~17%), with 1-1 Draw (~17%) as an equal co-favorite.

Who Scores #1: Mbappé (~25%); Senegal #1: Jackson or Mané (~13% each).

The single thing to watch: Senegal's confirmed goalkeeper. If Mendy plays, the consensus is firm. If Diaw plays, nudge France up ~4 points.


9. Source Analyses (linked)


10. Pipeline & Methodology Notes

This consensus is itself a pipeline validation artifact. Key lessons confirmed:

  1. Source hygiene is non-negotiable. The contaminated A-version weather (30-32°C) produced unpredictable, model-specific noise — mm3 and g4 swung in opposite directions (−12 and +11 points) when it was removed. Conflicting source files → nondeterministic results. Fix: each match directory must be self-contained; superseded files archived to archive-superseded/, never left active.
  2. Ensemble averaging reduces single-model error. The grand average (56%) hedges both the human optimism and the machine caution. No single version is as defensible as the average.
  3. g5 (glm-5.1) is the strongest single synthesis model — cleanest Finding 11 reasoning, best finding-discipline, math-clean. It is now the project's default match-writer model.
  4. mm3's GK-ambiguity catch is the kind of detail that justifies running multiple models: a single model might miss it, but the ensemble surfaced it and it's the match's biggest unresolved variable.
  5. The human-machine travel-asymmetry gap (9 points) is a calibration finding in itself: Finding 11's travel rule may need a "regional ground transport ≠ intercontinental disruption" clarifier so models don't over-apply it.

Consensus meta-analysis. No betting odds used. Aggregates five independent predictions on identical clean source inputs. Source archive: sources/matches/fra-vs-sen-B/