HOW IT WORKS

Was this game unusual?

We compare what happened with what similar games, teams and officiating histories would lead us to expect.

Five levels, one overall rating

Game rating scale
RatingMeaning
Fair · 1/5The outcome and penalty pattern fit historical expectations.
Debatable · 2/5A statistical result stands out after accounting for the comparisons we run.
Hmm · 3/5A more unusual statistical result, or two drive-extending penalties on one drive.
Sus · 4/5A strong outcome or penalty anomaly, or three-plus drive-extending penalties.
RIGGED? · 5/5A strong statistical anomaly and a three-plus penalty sequence both favored the winner.

The rating uses three statistical comparisons: the penalty pattern, the final margin given the box score, and the final margin against the spread. The first two use predictions made from earlier seasons, followed by comparison with errors in later historical games. A rating does not increase simply because the play scan found a lot of notable plays.

For each available comparison, we multiply its historical tail frequency by three, capped at 100%, to account conservatively for running three comparisons. Adjusted frequencies at or below 20%, 10% and 3% qualify for Debatable, Hmm and Sus respectively. We take the highest qualifying level, not a sum. Spread alone is capped at Hmm. These are published screening thresholds, not an independently calibrated overall probability.

Two distinct defensive penalties awarding a first down on third or fourth down during the same drive establish a minimum rating of Hmm; three establish Sus. RIGGED? requires both an unusually favorable box-score outcome for the winner (adjusted frequency at most 3%) and a three-plus penalty sequence for that same winner. Duplicate plays, separate drives, declined or offsetting penalties, and independent rushing or passing conversions cannot manufacture a sequence.

Fair means the available checks fit the usual range. RIGGED? expresses the strongest statistical concern; it does not establish intent or show that a call was incorrect. Unrated means the required comparison is missing, inconsistent or too small. Ties can be rated when the new comparisons are supported. Current rules: game-suspicion-v3. Older report versions retain their original rating rules.

Expected versus actual

Penalties

We compare accepted penalties and yards with the league average, the team’s usual committed penalties, and the opponent’s usual penalties drawn. The baseline uses the preceding five completed regular seasons. Each team and opponent average is pulled toward the league average with a 20-game prior; their deviations receive equal weight. Available referee effects are then applied as described below.

The penalty comparison evaluates both total penalties across the game and the imbalance between teams, using counts and yards. These correlated measures form one family: its statistic is the largest absolute standardized residual. They never become four separate votes for a higher rating.

Final score

A linear historical model estimates the final home margin from net offensive yard differential and turnover margin. This is a retrospective box-score expectation, not a pregame forecast. Return touchdowns, field position, kicking and other football events can explain a large residual; an outlier is a prompt to understand the game, not an attribution of misconduct.

Historical checks

For a 2026 game, the baseline uses 2021–2025 and the calibration games come from 2023–2025. Each calibration game is predicted by a model fit only on its own earlier five seasons. Current-season and future results are excluded. At least 500 calibration games are required. The tail estimate is (games at least as unusual + 1) / (comparison games + 1), preventing unsupported zero-probability claims.

The new reference supports regular-season games. Postseason comparisons remain unavailable until separately supported. Penalty comparisons are per game, so overtime exposure and pace can affect them. The older fixed win-profile comparisons remain available as additional context, but do not drive the new rating.

What referee history tells us

Reports identify the head official and show earlier crew penalty totals, home/away splits, and each team’s record and penalties in games involving that referee. These are not the individual referee’s personal flag totals, and a head official’s supporting crew can change.

Assignments come from nflverse schedules and the officials release. We prefer GSIS game keys for matching, retain explicitly documented name aliases, and exclude conflicting assignments from referee comparisons. Official IDs changed namespaces in 2023, so they are not treated as stable cross-era identities. When the full-crew release is missing, an unambiguous schedule assignment is identified as schedule-only.

The referee adjustment uses the historical penalty residual after accounting for the teams and opponents. It requires at least 10 games, is shrunk with a 40-game prior, and is split across both sides. When unavailable, we use a separately calibrated team/opponent baseline. Referee–team win records are descriptive context only; small samples and differences in team strength prevent a causal interpretation.

The spread versus the score

The recorded closing spread supplies the expected winning margin. We show the actual margin, the size of the miss and which side covered. Its historical reference uses earlier seasons, requires at least 500 games, and preserves the source line with its snapshot checksum. We do not invent a bookmaker or quote time.

The spread enters the same three-comparison adjustment described above and can reach at most Hmm. An upset or cover alone does not imply anything unusual. Older ratings used the earlier spread thresholds preserved with their report version.

Your take on the rating

A thumbs-up or thumbs-down saves one public response. The slider and explanation are optional; adding or changing them updates that same response. The average uses only ratings people actually chose on the slider. The public count, agreement breakdown and comments show the latest response from each browser for that game, across report versions.

Feedback records which version and rating the visitor saw. Earlier private comments stay private unless their author explicitly shares them. This is an informal browser-based poll, not verified one-person-one-vote voting. Fan feedback does not alter the calculated game rating.

Notable plays and limitations

The automatic scan identifies recorded drive-extending penalties, reversals, nullified scores and selected late-game events. These details are optional and collapsed on the game page. A flagged play is not a confirmed mistake and its count does not set the rating.

Play-by-play does not reliably record missed fouls. Statistics cannot determine whether an uncalled hold happened. The public report runs automatically and does not wait for a person to review it. Any separate call-correctness finding requires evidence outside this statistical rating.

Additional calculations

The expandable analysis retains supported estimates for ruling impact, fourth-down choices, fumble recoveries, kicking, execution charting and called penalties. Each includes its evidence and assumptions. These calculations do not all enter the overall rating. Observed win-probability movement includes everything on a play and is not automatically the effect of a flag.

Rarity is not intent

Family comparisons describe how often historical games produced an error at least as large. They are not the odds that a game was rigged. The five-level scale uses transparent editorial thresholds and has not been independently validated as a measure of misconduct. Correlated signals are not added together.

Missing data stays missing

Reports run when complete final-game data becomes available and are checked again for provider updates. Missing totals or insufficient historical samples can leave the overall rating unavailable; missing referee coverage alone does not. Unavailable values never silently become zero.

Game momentum

The chart shows before-play win probabilities in recorded play order, not elapsed time. Available values are connected for readability; intermediate estimates are not invented. Each recorded entry remains in the data table. An explicitly recorded overtime final result is shown separately as an observation.

Overtime

A separate experimental model compares prior games with matching possession rules and similar game states. It requires at least 20 distinct earlier games, with one state per game. Win, loss and tie are separate outcomes. Unsupported rule eras, possession sequences and postseason states stay unavailable. These estimates describe momentum and do not supply officiating costs.

A report you can trace

Source URLs, retrieval times, checksums and model versions are retained. Updated sources or rating models create a new report version; older versions and published-post records stay intact. Historical data is provided by nflverse/nflfastR and Lee Sharpe under CC BY 4.0.

Data sources ↗ · Report changes ↗