Methodology

What we measure, and why. Current forecast accuracy: Brier 0.098 against a 0.250 coin-flip baseline. See /calibration for the scoring record.

Three legs of the forecast

The Datum Signals is not a mood-only model. "Mood" is one of three legs, and the race forecast rests on all three:

  1. Structure (fundamentals). The partisan lean of the state, incumbency and the incumbent's prior personal vote, and candidate-level campaign money (the Democratic share of two-party receipts through June 30 of the cycle: FEC filings for Senate races, state disclosure board filings for Governor races where one is on file). These are the slow-moving facts of a race.
  2. Behavioral mood. The cross-sectional state-mood index built from publicly reported behavioral data. This is the layer the name points to, and it moves the structural baseline, it does not replace it.
  3. The thermostat (long-run swing). A deep-estimated midterm correction: the president's party tends to lose ground in midterms, applied as a calibrated national swing.

Structure anchors the forecast, mood adjusts it, and the thermostat applies the long-run midterm regularity. Where a large mood signal runs against a decisive structural and incumbency lean, the race page says so plainly and treats the call with caution.

Core Premise

Polling measures stated opinion. We measure revealed behavior. People can lie to a stranger on the phone; they cannot lie to their unemployment filing, their voter registration record, their vehicle miles, or their gasoline receipt. The composite index and the forecast model are built entirely on publicly reported behavioral data, calibrated against historical election outcomes.

Composite Index (descriptive layer)

The state composite mood score for each state is built from 14 behavioral dimensions normalized cross-sectionally against the other fifty states. The score is descriptive: it tells you the current mood reading, not a prediction. The "National Weather Service for civic emotions" voice belongs to this layer.

The dimensions

DimensionWhat it measuresRepresentative source
MobilityVehicle miles, transit ridership, household movementFHWA TVT, DOL
SecurityFear, threat preparationFBI UCR, FBI NICS
EconomyUnemployment, gas prices, initial claimsBLS LAUS, EIA, DOL
SpendingRetail and household economic commitmentBLS CES, Census ACS
CivicTrustVoter registration, civic participationCensus CPS Voting Supplement, MIT EL turnout
DomesticStabilityHousehold formation, migration persistenceIRS SOI migration, Census ACS
IsolationWithdrawal, alienationCDC mortality (suicide), Census ACS
RiskSpeculation, bet-on-future behaviorMultiple (composite)
EscapeStress and fatigue retreatCDC overdose, Google Trends
InformationSearch interest, civic-anxiety queriesGoogle Trends
Information_candidatesCandidate-quality signal from polling/coverageVoteHub aggregates
Pulse (reserved)Dimension reserved for high-cadence stress signals(currently empty, see caveats)

Normalization

Each indicator reading is normalized against the cross-section of the other fifty states, not against its own time series. Election outcomes are state-versus-state contests, so the relevant baseline is "how anxious is Wisconsin relative to Michigan right now," not "how anxious is Wisconsin relative to its own two-year average." Slow-moving dimensions favor the most recent reading over a stale average, so a dimension that updates quarterly does not dilute one that updates weekly.

Forecast Model (experimental layer)

The race-level forecast maps a state's structural position and behavioral mood reading to P(Democrat wins). It is labeled experimental everywhere it appears and always carries its current Brier score. The forecast layer is the only place predictive language is permitted.

What the forecast reads

The forecast draws on the three legs described above: the structural position of the race (state partisan lean, incumbency, candidate money), the behavioral mood dimensions listed in the table above, and the long-run midterm correction. Mood adjusts the structural baseline rather than replacing it, and where the two disagree sharply the model deliberately holds back rather than resolving the conflict in favor of either.

Display dimensions vs forecast inputs

All 14 behavioral dimensions are computed daily and surfaced on the dashboard, state detail panels, and the public API. A subset feeds the forecast. Two dimensions are intentionally display only:

Dimensions are added to the forecast only after enough history exists to score them out-of-sample. We would rather show a dimension and exclude it than include it on thin evidence.

Calibration

The forecast is fit on recent completed senate and governor cycles and scored against a wider historical corpus. Accuracy is reported out-of-sample: every scored cycle is predicted by a model that did not train on that cycle, so the headline number reflects performance on races the model had not seen. It is not a single favorable split. The current reported Brier is 0.098 against a 0.250 coin-flip baseline. The full scoring record, including per-cycle results, is published at /calibration.

What our 0.098 Brier score means

0.098 is our out of sample score on 2024, the most recent completed cycle, which the model never trained on. 0.124 is the pooled record across every cycle from 2014 to 2024. Both are below.

A Brier score measures how far a forecast lands from what actually happened, so lower is better. The two ends of the scale, then ours.

Learn more about our Brier score

LOCO means leave one cycle out. To test the model on 2018, we train it on every other cycle and make it predict 2018 cold, then score the misses. Repeat for each cycle. The model never grades its own homework: every score comes from elections it had not seen, the same position it is in for 2026.

The headline 0.098 is one such test: an identically specified model trained on the two prior cycles and scored on the most recent completed cycle, which it had never seen. Run the same test on every cycle from 2014 to 2024 and the pooled out of sample score is 0.124 across 1,404 scored race cutoffs, ranging from 0.086 in 2022 to 0.172 in 2018.

Across the 360 out of sample calls where the model gave a side 80% or better on average, that side won 91% of the time. The model is if anything under confident at that end: it wins more often than it says it will, which costs it Brier points in the safe direction.

Cycle held outOut of sample BrierRaces scored
20140.136288
20160.137184
20180.172284
20200.111184
20220.086280
20240.087184

Scores depend on the slate. A cycle full of safe seats scores better than a cycle of toss ups, for any forecaster. That is why we publish the per cycle numbers above rather than one average, and why a single number should never be read as a ranking. See the full record.

How 0.098 compares

ForecasterBrierUses survey inputReference
Always guessing 50/50 0.250 No source reviewed 2026-08-29
Perfect oracle 0.000 No source reviewed 2026-08-29
FiveThirtyEight congressional and governor forecasts 2018 to 2022 0.030 to 0.100 Yes source reviewed 2026-08-29
Academic fundamentals-only election models 0.120 to 0.200 No source reviewed 2026-08-29
NWS 24-hour precipitation forecasts 0.100 to 0.150 No source reviewed 2026-08-29
NFL pregame win-probability models 0.200 to 0.220 No source reviewed 2026-08-29
DatumSignals, held-out cycle 0.098 No per cycle record
DatumSignals, pooled 2014 to 2024 0.124 No per cycle record

Scores depend on the slate. A cycle full of safe seats scores better than a cycle of toss ups, for any forecaster, so the same model can post very different numbers on two different sets of races. Other models here were scored on different races in different years, often with survey inputs ours does not use. Treat these comparisons as directional, not as a ranking.

Accuracy over time

The forecast has improved from roughly 0.247 at first release in May 2026 to 0.098 today, across eight documented revisions. That trajectory is not monotonic: one intermediate version posted a better-looking score that we retired after tracing it to a data error rather than a genuine gain, and the corrected version scored worse and shipped anyway. Revisions are logged in /changelog as they happen, including the ones that moved the number the wrong way.

Confidence interval. Each Senate/Governor race's displayed ±N is a per-race, per-rating-tier figure, not a global constant. It floors at a rating tier's own out-of-sample calibration error (leave-one-cycle-out, 2014 to 2024: Safe R 5.6pp, Safe D 8.6pp, Likely R 11.1pp, Toss-up 18.7pp, Lean R 19.0pp, Likely D 19.1pp, Lean D 37.9pp -- the last of these on a thin n=17 holdout sample, flagged rather than trusted at face value), then adds a penalty for missing model dimensions and for how close the race sits to a 50/50 toss-up, capped at 17pp. This changed 2026-08-27: the prior version used one shared base term (a function of the model's overall test Brier score) for every race regardless of rating tier, which understated uncertainty specifically in the Lean and Likely tiers -- see the full derivation in the repository's docs/MODEL_CALIBRATION.md.

House Forecast

The 435 House districts run on a separate, simpler model from the Senate and Governor forecast above: the district's cross-cycle Partisan Voting Index (PVI), incumbency, open-seat status, the incumbent's prior personal vote, candidate fundraising split where available, and a midterm-correction thermostat term, calibrated to a win probability. This is a deliberate scope decision, not an oversight -- Senate and Governor races carry a state-level behavioral mood reading; no district-level analogue of that mood reading exists for all 435 House seats. Rather than force the three-leg model onto a feature set it does not have, the House forecast is honest about running on structure alone.

This means the House forecast has no mood component and no per-feature decomposition to show on a race page's WHY panel -- only the district lean itself, and, where known, whether the seat is open. As of 2026-08-26 the House model (model_version house-2026-panel-v1) is trained on five pooled cycles (2016-2024, both midterm and presidential years) and validated out-of-sample on every held-out cycle from 2018 through 2024, rather than fit on a single prior cycle. The prior version was fit on 2022 alone and carried that cycle's specific national environment into 2026 unchanged; see /changelog.

Coverage is capped at districts whose current lines have been measured: 400 of 435 districts are in scope as of this writing; the rest sit on a map superseded by a mid-cycle redistricting and are held back from the public forecast until their PVI is re-measured on the current boundaries, rather than shown on stale lines. See the House table on /forecast for the full, filterable list and each district's last-verified candidate check.

Lean basis, per cycle and per office. The House model's lean is 2024 presidential results measured on the 2026 district lines, as a deviation from the certified 2024 national two-party Democratic share of 0.49250. It was 2020 results on the previous lines until 2026-08-30, when ten states had redrawn their maps and no 2020-by-2026-lines measurement existed for them; see /changelog. The Senate and Governor model's state_lean remains on the 2020 baseline of 0.5227, since state boundaries do not change. Because the model standardizes lean within each cycle's cross-section, changing the basis does not change any coefficient.

Lean scale. PVI is computed as round((d_share/(d_share+r_share) - national_baseline) * 100, 2) -- a margin-style deviation from that cycle's national baseline, in percentage points -- and is stored divided by 50 in both the House model's pvi column and the Senate/Governor model's state_lean feature. Every "favors [party] by about N points" line on a race page and every "D+N"/"R+N" lean label on the House table multiplies that stored value back out by 50 to recover the real margin figure, consistently across all three offices.

2028 Candidate Fit

The 2028 POTUS page ranks potential candidates by fit, not win odds: how well a candidate's behavioral profile matches the current state-mood snapshot. Each candidate has a hand-set profile vector (8 of the same dimensions used elsewhere on the site, on a -1 to +1 scale) and each state has a live mood vector from the daily composites. Fit is the cosine similarity between the two, rescaled from [-1, +1] to a 0-100 score where 50 is neutral. This is a judgment-driven fit exercise, not a validated forecast, and it is labeled experimental everywhere it appears.

The headline number is an electoral-vote-weighted mean across the other 49 states -- a candidate's home state is excluded so a favorite-son effect in their own state cannot inflate the national figure. The page also reports the single strongest and weakest state outside home, and a separate battleground fit averaged over a configurable set of swing states (AZ, GA, MI, NV, NC, PA, WI by default), shown alongside the national number with a toggle to sort by either.

Party is stored per candidate as a single current-affiliation field. Where that affiliation has changed recently or is otherwise contested, the candidate is flagged on the page with a note explaining the history -- Tulsi Gabbard, who ran as a Democrat through 2020, registered independent in 2022, then joined the Republican Party and now serves as Director of National Intelligence under a Republican administration, is the current example.

Honest Caveats

Backtest coverage

Every dimension actually in the served forecast has been leave one cycle out backtested across the model's full training window; none of the model's coefficients rest on unbacktested history. The behavioral block in sengov-2026-govmoney-v1's trained feature set is Mobility, Security, Economy, Spending, CivicTrust, DomesticStability, Isolation, Risk, Escape and Information_candidates, each scored out of sample on every held out cycle the model was trained across. Money, the bare Information dimension and Pulse are computed and shown for context on the dashboard, state panels and race pages, but are not model inputs: they are excluded from P(Democrat wins) entirely, not partially weighted, pending a deep enough historical record to score them honestly. See Display dimensions vs forecast inputs above for why each is held out.

Forecasting Accuracy (Brier Backtest)

Brier Score: 0.098
PUBLISHABLE
EXPERIMENTAL v1.0
on 2024 held-out races, never seen during training. 52% better than naive baseline (0.202).

The Brier score measures forecast calibration. A perfect predictor scores 0. A random coin flip scores 0.25. Lower is better. Cited Brier values throughout this site refer to the out-of-sample score, last calibrated 2026-09-11.

How accuracy is measured

Accuracy is scored on races the forecast did not train on. For each race we reconstruct the state's behavioral readings at four points in the run-up to election day, using only data that existed on or before each of those dates. Nothing that happened after a cutoff is visible to the forecast being scored at that cutoff, so the reported number reflects what the forecast would actually have said at the time. No leakage.

Brier by Cycle

CycleBriern cells

Brier by Office

OfficeBriern cells
v10 calibrates 11 dim features: Mobility, Security, Economy, Spending, CivicTrust, DomesticStability, Isolation, Risk, Escape, Information (six fixed political-engagement terms), and Information_candidates (per-active-roster aggregate). The two Information signals were backfilled separately for 2018-2024 via Google Trends interestByRegion, then trained as independent features so the model learns each weight independently. Information_candidates z-scores are capped symmetrically at +/-1.5 at both training and inference time, after v9's 2018 OOS regression diagnosis showed extreme candidate-fame density in FL 2018 (z=+2.43) was driving wrong-direction P(Dem) shifts. The cap recovers half of the 2018 Brier loss while preserving the 2024 holdout improvement. Pulse and state-level Money remain live in current scoring but not in the backtest, pending historical backfill of their underlying sources. The model also uses incumbent_party_d and Economy x incumbent as derived features keyed on a hand-curated state_incumbents table (2018-2026 Senate and Governor).

Last calibrated: 2026-09-11. See /methodology for what we measure and /calibration for the scoring record.

National Mood Index

The VoteROI National Mood Index aggregates the 50 state composite scores (plus DC) into a single number on a -50 R / +50 D scale, where 0 is a dead tie. The headline figure shown on /forecast is the population-weighted reading; an electoral-vote-weighted lens is published as a secondary view.

WeightingRoleWeight basisWhat it captures
Population-weightedHeadline2020 Census apportionmentWhere Americans live. The default "how does the country feel" reading.
EV-weightedSecondary lens2024 Electoral College allocationWhere presidential elections are actually decided. Useful for handicapping 2028, less so for everyday mood.
Competitive-weightedInternal diagnostic2026 competitive race count per stateWhere current-cycle action concentrates. Surfaced in the API but not in the hero readout.
EnsembleInternal diagnosticEqual-weight mean of the threeHeld in the API for backward compatibility with v1.0 consumers.

Each state contributes its composite score minus 50 (the temporal baseline), weighted by the method. The sign for the state is +1 if the state's relevant 2026 incumbent is Democratic, -1 if Republican, 0 if independent or unknown. The per-state pull is then (composite - 50) * sign * weight. A high mood reading in a Democratic-incumbent state reads as Democratic reward (positive on the index); the same high mood reading in a Republican-incumbent state reads as Republican reward (negative). This captures the textbook incumbent-reward and incumbent-blame dynamic at state granularity rather than averaging it out.

The headline convention (population-weighted) reflects everyday civic mood. The electoral-vote lens is reported alongside because the same mood translates differently into a presidential outcome depending on which states feel it most. When the two diverge, that gap is itself information.

vs. Kalshi KPOW

Both VoteROI National Mood Index and the Kalshi American Power Index (KPOW) aim to be a single-number readout of where US political power sits. They use different signals and answer different questions:

VoteROI National Mood IndexKalshi KPOW
InputBehavioral data (FEC, BLS, Census, CDC, FHWA, etc.)Prediction-market prices (Kalshi)
RefreshDaily, after composites recomputeReal-time, tick by tick
ReadsState-level fundamentals (the "why")Market consensus (the "what")
Scale-50R / +50D, signed by federal incumbent-50R / +50D, market consensus mapping
StrengthCaptures structural shifts before markets price them; works without an active liquid marketCaptures news shocks in real time; price discovery in liquid contracts
WeaknessMood-to-partisan translation needs an incumbent-sign convention; lagged by composite cadenceReflexive (reads its own market); requires liquid contracts to exist

The two indices are complementary, not competing. KPOW tells you where the market thinks power is. VoteROI tells you where the behavioral fundamentals say it ought to be heading. When they disagree, that gap is itself information.

What We Are Not

Versioning and Audit

Every published forecast is timestamped and tied to a dated revision of the model. Methodology changes are documented in /changelog before going live, and material changes to a published probability are logged in /corrections. Historical forecasts are preserved as issued, so a call we made in the past can be checked against what actually happened rather than quietly restated.

Questions: press@parallaxadvisory.llc.