How Vugol Me aggregates polls
The model is intentionally simple and fully documented. Every published number is reproducible from a versioned function and a stored input snapshot.
What is validated, and when
Four things on this site are easy to mix up. They are separate.
- Vector is an experiment. It is a direction rating: which way a race points and how firmly. It is not a forecast, not a win probability and not a vote share. It has not yet been checked against real results. How Vector is built.
National environment (since vector-1.1.0). A seat's last result and its presidential vote were cast in 2024's national mood (or its own year's). Vector moves both by how far the country has swung since:swing = generic-ballot average − national House vote of that year; component = stored margin + swing. Today our generic-ballot average is D+8.0 (28 pollsters, as of 2026-09-30) against the 2024 House vote of R+2.6, a swing of D+10.6: a seat that voted R+5 for president now counts as D+5.6 on that input. - Attention (Pulse) measures noise, not support. It counts how much news and online activity a race gets. A high score means people are paying attention; it says nothing about who is ahead. How Attention is built.
- Verified polls are separate. Poll averages use only polls from pollsters staff have verified, and they are shown as polls, apart from Vector and Attention.
- Community polls never feed any number. Answers from site visitors are shown for interest only. They do not enter a poll average, Vector or Attention.
Milestones
- Roster accuracy — ongoing, nightly. Integrity checks run every night and send roster mismatches (retirements, primary losers still listed, missing incumbent flags) to staff. See race data.
- Calibration check — by October 20, 2026. Where a race has both a Vector rating and a verified poll average, we publish how often they point the same way, and where they disagree.
- Backtest — published by December 15, 2026. After the November 3 election we compare every Vector with the official results and publish the full list, misses included.
Four words, used precisely
- Verified — staff have confirmed the pollster’s identity (the organization exists and answers for its work) and that it publishes its methodology. It is a statement about the source, not an accuracy score. Only verified pollsters enter aggregates.
- Checksummed — a release carries a SHA-256 over its inputs, model version and outputs; pages recompute it at render time, and a changed byte anywhere shows up as a mismatch.
- Reproducible — the release manifest stores every input and the model version, so anyone can rerun the published function and obtain the same numbers.
- Audited — we do not use this word for our own work. It applies only once an independent third party has examined the ledger and published its findings; none has yet.
What enters an aggregate
- Only polls from verified pollsters. New pollsters are quarantined until reviewed.
- Each poll is weighted by sample precision (effective sample size), recency (older polls decay), and a reduced weight for partisan-sponsored polls — which are down-weighted, never silently dropped.
- Syndicated polls that share an underlying sample are collapsed to a single measurement, so one poll re-reported by several outlets is not counted several times.
- Anonymous or self-selected “community” responses never enter a scientific aggregate, regardless of volume.
What triggers human review
Rather than a single fixed rule, a newly-arrived poll is held for a human when it does any of the following:
- changes which option is leading;
- moves a statewide average by more than 2 points;
- moves a national average by more than 1.5 points;
- accounts for more than 25% of the aggregate's weight; or
- lands more than 2.5 standard deviations from the model's expectation.
These thresholds are starting points, to be calibrated against past elections before any confidence labels are published. A held poll is decided within 24 hours on business days; see the editorial process.
National trackers and issue polls
Presidential approval, the generic ballot, approval of Congress, the Supreme Court and cabinet members, favorability, and issue polls load automatically every night.
- Sources, licence first. The VoteHub polls API (CC BY 4.0), which links every poll to the pollster’s own release, and Wikipedia’s poll tables (CC BY-SA 4.0), linked to the exact revision read. We never copy another site’s average, and we never read sites whose terms forbid it.
- Provenance. Every fetch is hashed (SHA-256) and stored with its URL, retrieval time and licence; every poll records the parser version that read it and links to its source. Parsing is deterministic code, with no language model involved.
- Automatic publishing. A poll publishes without a person only when it passes every check: field dates in order, not in the future and at most 45 days long; a stated sample of 200 to 60,000; a stated population (adults, registered or likely voters); toplines between 0 and 100 that sum to about 100; not a duplicate of a poll already on file (from either source); and its lead number within 12 points of the median of other pollsters’ polls of the same question in the surrounding month. Anything else waits for staff, with the reason shown. The operator can pause loading, turn automatic publishing off, or change the outlier distance at any time.
- Our averages. Each pollster counts once (its latest poll in the last 60 days), weighted by sample size, recency (14-day half-life) and half weight for partisan sponsors: the same weights as our race averages. We show the weighted share as polled, so undecided respondents are not reallocated. At least three pollsters, or no number.
- Question wording. The sources give toplines, not always the verbatim question, so each tracker stores a standard wording marked as such; the source link has the pollster’s exact text.
- Not verified releases. Pollsters not yet verified by staff are included in these trackers and labelled; national trackers never enter a checksummed release.
Corrections
Published releases are immutable. When a pollster corrects a number, we publish a new release that supersedes the old one and records what changed; the earlier release stays accessible, marked superseded. How to report an error is on the corrections page.
Attention, Vector, Bench
Polls cover a few dozen races. In-house signals give every race a brief. None of them is a poll, and every page that shows one says so.
- Attention score (0–100) measures attention: Wikipedia pageviews, GDELT news volume and tone, and Bluesky posts for each candidate, log-scaled per source and summed. Attention (50%) is the race's percentile among all rated races that day; momentum (30%) compares the trailing 7 days to the 21 days before; breadth (20%) is the share of tracked sources with activity. News tone is reported separately and never folded into intensity. Model
pulse-1.0.0(the product was called Pulse when the model was named; the id is kept so stored rows stay traceable). - Vector — Experimental. Says which way a race points and how firmly: a direction (R or D), a magnitude 0–100, and a band — Anchored (70+), Held (45–69), Tilting (20–44), Open (under 20), or Unrated. Inputs, each optional with weights renormalized over what is stored: the seat's last result (0.35), its presidential lean (0.30), incumbency (0.15, ±6 points), FEC fundraising (0.10, capped ±8), attention share (0.10, capped ±3), and a verified poll release (0.60 when present). Confidence is high with a poll or both result inputs, medium with one, low with incumbency or money only. Missing inputs lower confidence — they are never filled in. Vector has not been backtested against past elections yet; treat it as a labeled experiment, not a forecast. Model
vector-1.1.0. - Vector's national environment. The last result and presidential lean describe the year they were cast, not 2026. Standard forecasters move them by the national swing, and so does Vector: swing = elasticity × (generic-ballot margin − national House popular vote in the input's year), added to each of those two inputs. The generic ballot is our own average of every national generic-ballot poll (the number on /generic-ballot, with its date and pollster count stored on every rating). Baselines are the national House vote from the Clerk of the House: 2024 R+2.6, 2022 R+2.7, 2020 D+3.1; a result without a stored year is taken from the seat's last regular election (House 2024, Senate 2020, governors 2022). Elasticity is 1 for House and Senate races, which have tracked the national vote almost point for point, and 0.8 for governors and other statewide offices, where candidates matter more. Today: D+8.0 against R+2.6, a D+10.6 swing for House races. Incumbency, money and attention are not moved. A verified poll already measures today, so it is never moved either, and where a race has a verified polling average that average, not Vector, colours the race. With no generic-ballot average the swing is left out and the rating lists the national environment as a missing input. The calibration check (October 20) and the backtest (December 15) score the ratings with this term; earlier
vector-1.0.0rows stay stored as they were. - Bench is an index of disclosed campaign expenditures, built from FEC Schedule B operating expenditures: payees normalized, purposes categorized by keyword, operational spend (payroll, rent, travel, refunds, taxes) excluded, ranked by clients then dollars, every row linked to the filings. Only organizations are indexed: payments to individuals (FEC entity type individual or candidate, or a payee named as a person) are never stored or shown. Bench is paused until after November 3 while the index is rebuilt with amendment and memo-entry handling. Model
bench-1.1.0. - Race briefs are written by a language model from the stored numbers only; the generator may not add facts, and any number in the text that is absent from its inputs rejects the brief. The model id, prompt version and input hash are stored with each one. Brief generation is paused until after November 3.
- After the November 2026 election we publish a backtest of every Vector, misses included, by December 15, 2026. Dates for every check are under what is validated, and when.
Geography
ZIP codes are used only for coarse discovery. A ZIP does not reliably determine a voter's ballot — some cross county and district lines — so district-level claims require address-level resolution, not a ZIP-to-ballot guess.
Public polls API
GET /api/public/polls — no auth, read-only, the same data /latest-polls shows. Query params, all optional:
office—senate|house|governor|statewidestate— two-letter state codepollster— substring match against pollster name or poll titlecontestKey— exact race key, e.g.US-NC(the same key in a race's URL)format—json(default) orcsvlimit— max polls returned, default 100, hard cap 500
Rate limit: 30 requests per 60 seconds per caller IP, fixed window; a caller over the limit gets 429 with a Retry-After header. Responses are cached 60 seconds at the edge. Every field mirrors v_poll_feed: pollster name and tier (verified/quarantined — see pollsters), field dates, sample size, computed margin of error, and topline shares per candidate.
Race polls, results, money and dates
A nightly job reads the Wikipedia article for every Senate, Governor and House delegation race (licensed CC BY-SA 4.0; we link the exact revision) and the Federal Election Commission's public API (public domain). Parsing is fixed code, not a language model. Every stored row keeps its source URL, retrieval time, parser version and the SHA-256 of the exact bytes we read.
- Polls. Only individual polls are read; poll averages published by other sites are never copied. Each poll links to the pollster's own release. General-election polls feed the race page; primary, runoff and hypothetical-matchup polls are kept on file and never averaged. Pollsters new to us start quarantined (shown, not averaged) until staff verify them. We do not copy proprietary race ratings; the rating shown is our own.
- Automatic publication. A poll publishes without a reviewer only if all of these hold: sample size known and between 150 and 30,000; field period at most 60 days and ended after the 2024 election and not in the future; candidate shares at least 30% in total and all shares at most 103%; a link to the pollster's release; and, when the race has 3 or more recent polls, a margin within 15 points of their median. Anything else, and any poll whose numbers changed after we first read it, waits for staff review. Staff can tune every threshold (admin setting
race_polls_rules). - Results. Primary and runoff vote counts are transcribed from the article's result tables, which cite the state's official results. A table publishes when its percentages add to 100 (±1.5) and its votes add to the stated total (±2%); otherwise it is held (
race_results_rules). The general-election table lists who is on the November ballot; it drives “Incumbent running” / “Open seat” and the “Lost primary” tags. A nightly check sends roster mismatches (retiring incumbents, missing incumbent flags, primary losers still listed as running) to staff. - Money. Raised, spent and cash on hand per candidate committee from the FEC candidate totals, as of the committee's latest report.
- Dates. Primary, runoff, special and general election dates by state from the FEC's official list. Registration deadlines and early-voting windows are not yet carried: no official national source allows automated reuse.