Omaha 2.1.0
Released 29 September 2026
The engine can now tell a tied game from a three-point lead. None of the eleven figures below moved: the average is +9.7% before and after, and what improved is confined to close games late.
How to read these numbers
Every figure below compares the simulation against a player’s own season average so far, scaled to how much of the game is left, roughly what a reader could work out without us. Zero means we add nothing over that. Positive means we beat it; negative means we are worse than it.
Measured across the 2024 season, weeks 1–18, rebuilding the model each week using only what was known at the time. 400 simulations per situation.
The change column is simply this version’s accuracy minus 2.0.0’s, each measured across everything that version covers, so the two columns you can see always account for it. We also measure the change on shared situations alone, which is a more sensitive test; where the two disagree we publish this one.
Accuracy by stat
| Stat | vs. season average | Change from 2.0.0 |
|---|---|---|
| Receiving touchdownsrare event | +20.3% | +0.05−0.05 to +0.26 |
| Rushing touchdownsrare event | +16.8% | −0.06−0.37 to +0.12 |
| Receiving yards | +14.1% | −0.04−0.05 to +0.02 |
| Pass attempts | +12.8% | +0.10−0.05 to +0.24 |
| Passing yards | +11.1% | +0.04−0.12 to +0.21 |
| Rushing yards | +10.9% | −0.04−0.10 to +0.01 |
| Completions | +9.5% | +0.05−0.10 to +0.21 |
| Passing touchdownsrare event | +7.0% | −0.07−0.73 to −0.05 |
| Targets | +1.9% | +0.03−0.02 to +0.11 |
| Receptions | +1.7% | −0.01−0.05 to +0.08 |
| Rush attempts | +0.8% | −0.03−0.08 to +0.07 |
| Average across all figures | +9.7% |
The average is an unweighted mean across the figures above. It is a summary of how we are doing, not a single score for the model: yards and counts are not measured in the same units, so combining them any more cleverly than this would just let passing yards decide the answer.
Each change carries the range the measurement can actually support. A change whose range crosses zero is shown in neutral rather than as a gain: we cannot tell it apart from no change, and we are not claiming it as one.
What changed
Play calling reads the score through a band, and that band put a tied game and a three-point lead together. Those are opposite intentions: tied, a team is trying to win, and up three inside two minutes it is running the clock out. The engine read one play-calling rate for both. There are now seven bands instead of four, cut where the game actually changes: one to three points is field-goal range, four to eight is one score needing a touchdown, nine or more is two scores, and a tied game is its own state rather than the bottom of a lead.
Inside two minutes, real teams throw on 67% of snaps when tied and 16% when up three. The engine read 42% for both, and now reads 60% and 25%. Across the down-and-score grid its play-calling error falls by 62%.
Getting the play calling right in those states changes how often the ball is thrown in them, which is what a passing-yards line is made of. This is the only place the change reaches, and it is the situation the engine was weakest in.
Passing yards score 0.85 points better in a one-score fourth quarter, measured against a control that changes nothing and confirmed at two different simulation counts. No other stat in that situation moved beyond what the control moved.
The average is +9.7% before and after, and no individual stat moved by more than a tenth of a point. That is not a disappointment being buried: a one-score fourth quarter is about a tenth of what these figures are averaged over, so an improvement confined to it is arithmetically invisible here. We measured it where it happens instead, and we are telling you the headline number did not move.
Mean skill +9.72% at 2.0.0 and +9.72% at 2.1.0, on the same 272 games.
What this version still gets wrong
- Nothing outside a close, late game changed. Blowouts, first halves and pre-game projections are exactly as they were.
- Rushing touchdowns read 0.4 points worse on one of the two measures we score them by, and 0.1 points worse on the other. A control that changes nothing moved the same figures by similar amounts, so we do not believe either is real, but we would rather show you the number than decide for you.
- The improvement was found by auditing which situations the engine is worst in, and the worst one is still a close fourth quarter. This narrows that gap; it does not close it.