Skip to content
PickAuditor
Methodology

How the audit works

Public-tier picks captured before kickoff, graded against official results. This page is the complete set of rules behind every number on the site; changes to it are dated in the changelog.

Revised · changelog

What we capture

PickAuditor reads the public pages of forecasting services and records the picks shown there. Every record on the site is built from picks that were publicly visible, captured before the event started, and graded against the official result.

  • Public pages only. A source is captured only where its content is visible without an account, a payment or a code. Items shown as paywalled, blurred or “subscribe to view” are not extracted; they are counted so the capture log records how much of a page we could not read. Premium content is never stored, in any form.
  • robots.txt is respected. A service that disallows crawling is not captured. It can send a feed instead — see For services.
  • Every 4 hours per source, six times a day. Each request carries a User-Agent that names PickAuditor and includes contact@pickauditor.com, so a service can see who is reading its pages and write to us.
  • Timestamped by us. The moment we fetch a page is the pick's captured_at. It is our timestamp, not the service's, and it is the one used to decide whether a pick was recorded before kickoff.
  • Evidence kept. The raw HTML of every capture is stored in private storage for at least 24 months. It is used to re-check a record when a dispute is filed and is never republished.
  • Quotation limits. At most 300 characters of a service's reasoning per pick and at most 200 characters of a marketing claim are stored, verbatim. Full text is never copied.

What counts as a pick

A pick is one explicit forecast on one event and one market: “Arsenal to win”, “Over 2.5 goals”, “Lakers −3.5”, “LeBron James over 24.5 points”.

  • Parlays and accumulators are split into their legs. Each leg becomes its own pick and carries a parlay_split warning.
  • Odds are stored as decimal. American and fractional prices are converted (+150 → 2.50, 5/2 → 3.50). A pick with no visible price has no quoted odds and is priced from the market at capture — see Odds source.
  • A probability is recorded only when the page frames it as a win probability. “92% confidence” is kept as a raw label, not as a probability, and does not feed the Brier score.
  • Nothing is inferred. If a team, line or market cannot be read from the page, the field stays empty and the pick ends up unmatched or ungradable rather than guessed.
  • The same service, event, market, selection and line is stored once. The earliest capture is the one that counts; later repeats of the same pick create no new record.
  • A service's own “published at” time is stored when shown, but the before-kickoff test always uses our capture time.

Matching to canonical events

Services write team and player names in many ways. Each captured pick is matched to one canonical event in our database, or to none.

  1. Candidates are generated in code: events in the same sport starting within ±36 hours of the time shown on the page, whose competitor names or known aliases have a trigram similarity of at least 0.3 with the names on the page (or one side above 0.6). At most 8 candidates are considered, and home and away are checked in both orders.
  2. A language model chooses one candidate or none. It is instructed to be conservative: a wrong match corrupts a record, a missed match only delays it. When two candidates remain plausible — a doubleheader, or the same teams in league and cup within 36 hours — the pick is left unmatched.
  3. Match confidence of 0.85 or higher is accepted automatically. Between 0.5 and 0.85 the pick goes to a human review queue and stays out of every record until a person confirms it. Below 0.5 it is unmatched.
  4. Each pick stores how it was matched: model, exact alias or human. A wrong match found later is corrected, the affected picks are regraded, and the correction is written to the changelog.

Odds source

Two prices matter for each pick: the price it is graded at, and the closing price it is compared with.

  • Reference bookmaker. Pinnacle whenever it prices the market; otherwise the bookmaker that has a snapshot for that selection and line.
  • Odds used. The decimal price the service itself quoted, if it showed one above 1.00. Otherwise the market price at capture: the snapshot taken at capture time (Pinnacle preferred), else the nearest snapshot taken no later than 30 minutes after capture. Each pick shows which of the two was used.
  • Closing odds. The reference bookmaker's closing snapshot for the same selection and the same line. When no closing snapshot was recorded, the last snapshot before kickoff is promoted to closing. A pick whose line matches no closing line has no closing price and therefore no CLV.
  • No price, no ROI. A graded pick with neither a quoted nor a market price keeps its win or loss for hit rate but has no profit figure, so it is not in the ROI denominator.

All odds on the site are decimal. All times are UTC.

Grading rules by market

One function grades every pick against the final score recorded for its event. Results come from a results API and are cross-checked against a second source when one is available. The same function grades every service and every consensus model.

Grading rules by market
MarketSelectionWonLostPush · void · ungradable
Match winnerhome / awayThe selected side has the higher final score.The other side has the higher final score. In soccer a tie is also a loss for a home or away pick.Push on a tie in every sport other than soccer, where the market has no draw option.
Match winnerdrawThe final score is a tie.Either side wins.—
Draw no bethome / awayThe selected side wins.The other side wins.Push on a tie.
Double chance1X / X2 / 121X: home wins or tie. X2: away wins or tie. 12: either side wins.The single outcome the selection does not cover.The label is read from the service's own text (“1X”, “home or draw”). If it cannot be read, the selection field decides: home → 1X, away → X2, otherwise 12.
Totalover / under, lineOver: the combined final score is above the line. Under: below the line.The opposite.Push when the total equals the line. Ungradable without a line.
Spreadhome / away, linemargin + line > 0, where margin is the selected side's score minus the other side's and line is the selected side's handicap (for example −3.5).margin + line < 0.Push when margin + line = 0. Ungradable without a line.
Both teams to scoreyes / noYes: both sides scored. No: at least one side did not.The opposite.—
Correct scoreexact scoreThe final score matches exactly (2-1 is not 1-2).Any other score.Ungradable when the score text cannot be parsed.
Player propover / under, lineThe player's official box-score figure for the named statistic is above (over) or below (under) the line.The opposite.Push when the figure equals the line. Ungradable without a line, or when no box-score figure is available 72 hours after the start.
Any market———Void when the event is postponed or cancelled, and when the pick was captured after kickoff. Markets we do not grade are stored as “other” and marked ungradable.

Ungradable picks appear in the pick counts on record pages and never in a hit rate, ROI or any ranked figure.

Metric definitions

Every figure is computed from graded picks — won, lost or push — that were matched and captured before kickoff, over the picks whose event started inside the period. Periods are 7, 30, 90 and 365 days and all time. Statistics are recomputed daily, and each page states the date they are as of.

Metric definitions
MetricDefinitionNotes
Hit ratewon ÷ (won + lost) × 100Pushes and voids are excluded from numerator and denominator. Computed for every graded pick.
Odds usedquoted price if shown and above 1.00, else market price at captureMarket price at capture: the snapshot taken at capture time (Pinnacle preferred), else the nearest snapshot taken no later than 30 minutes after capture. Each pick shows which source was used.
Flat-stake profitwon: odds used − 1 · lost: −1 · push: 0One unit on every pick, no staking plan. A pick with no usable odds has no profit figure.
ROIΣ flat-stake profit ÷ picks with odds × 100The denominator is the number of graded picks that have an odds value; it can be lower than the n shown for hit rate.
UnitsΣ flat-stake profitThe same sum, not divided.
CLV(odds used ÷ closing odds − 1) × 100Closing odds: the reference bookmaker's closing snapshot for the same selection and line; when none was recorded, the last snapshot before kickoff is promoted to closing. Averaged over picks that have both prices.
Brier score(stated probability − outcome)², outcome 1 for won, 0 for lostOnly when the service states a win probability for that pick. Pushes excluded. Averaged; lower is better; 0.25 is the score of always saying 50%.
Log scoreln(probability assigned to what happened), floored at 10⁻⁶Same condition as Brier. Stored per pick; profile pages show Brier.
Max drawdownmax(peak − running total), in units, in event orderRunning total of flat-stake profit ordered by event start time. The peak is floored at zero, so a record that never goes positive has a drawdown equal to its lowest point.
Longest losing streaklongest run of consecutive lost picks in event orderPushes and voids are removed from the sequence before counting, so they neither break nor extend a run.
p-value vs break-eventwo-sided normal approximation of a binomial test: won against (won + lost) × mean implied probabilityMean implied probability is the average of 1 ÷ odds used, i.e. the hit rate needed to break even at those prices. Computed when won + lost ≥ 5. A small value means the hit rate is unlikely under break-even; it says nothing about the future.
Sample flagtiny < 10 · small < 30 · ok < 200 · large ≥ 200Derived from n. Only ok and large are ranked; tiny and small are marked amber.

Longer explanations with worked examples: closing line value, Brier score, why win rate lies, how to read a track record.

Sample policy

A rate without a sample size is not a rate. These rules decide what is shown, what is ranked and what search engines are asked to index.

Sample and index policy
SituationRule
Any rate or returnShown with its n, always.
Service with fewer than 30 graded picks in totalThe profile renders, is marked noindex for search engines, and appears on the leaderboard as an “insufficient data” row without a rank.
Sport or market segment with fewer than 30 graded picksNot shown as a ranked row on the profile or on a segment leaderboard.
Leaderboard rankRequires n ≥ 30. Rankable services are ordered by the chosen metric; smaller samples are listed below them, unranked.
365-day ranking pagesn ≥ 200 and the 365-day period, ordered by ROI. No editorial ordering.
Event pageFewer than 3 service opinions → noindex.
Sample flagtiny < 10 · small < 30 · ok < 200 · large ≥ 200. Tiny and small are shown in amber.

“Insufficient data” and “no data yet” are different states. The first means graded picks exist but too few to rank; it is shown as an amber stamp or row. The second means no graded pick exists yet — a new source, or one whose events have not finished. Captures run every 4 hours; a service is ranked once it has 30 graded public picks.

What we do not do

  • We do not read, store or display premium, paid or member-only content. The database refuses a pick that is not marked public-tier.
  • We do not take or place bets and hold no funds.
  • We do not give advice. No page tells the reader what to bet, and no text on the site calls a pick safe, likely or recommended.
  • We do not sell picks. The consensus meta-models are shown to everyone and graded in public.
  • We do not rank by opinion. The order of every list comes from audited numbers and the sample rules above.

The number guard

Two checks run on every generated text before it is published.

  • Every number must exist in the data. Profile and event texts are generated from the same rows the tables are drawn from. A checker then extracts every numeric token in the text — percentages, counts, decimals, dates — and verifies that it appears in the source data, allowing only a percent-to-fraction conversion, rounding to one decimal and thousand separators. One violation fails the text; it is regenerated at most twice and then queued for a person. A failed text is never published, and the page shows its tables without it.
  • Forbidden vocabulary. Words that assert certainty, promise outcomes or judge character are never used in generated or hand-written copy; the exact list is enforced in code and a text containing one of them fails the guard. Gaps between a claim and an audited figure are described with numbers instead: the claim, the audited value, the sample and the scope.

Hand-written pages, including this one, follow the same two rules.

Right of reply

Every service has a right of reply, and every dispute is public.

  • A service, a reader or a member of staff can dispute a capture (the pick was not there, or was not public), a match (wrong event), a grade (wrong result or wrong rule) or a quoted claim.
  • Disputes are filed by email; the disputes page has the template.
  • We respond within 7 days with a decision — accepted or rejected — and a written resolution note.
  • An accepted dispute corrects the record: the pick is regraded, rematched or removed, statistics are recomputed on the next daily run, and an entry is written to the changelog. The dispute and its resolution stay visible on the disputes page and on the service's profile.
  • Rejected disputes stay visible too, with the reason.

Affiliate policy

Some services pay a commission when a reader subscribes through a link on this site. These rules keep that separate from the audit:

  • Affiliate status never affects ranking, order, colour, badges or any number. Ranking is computed by code from graded picks, and that code has no affiliate input.
  • The step that writes a service's profile text is not given the service's affiliate status.
  • Each service profile shows exactly one disclosure line stating whether an affiliate relationship exists.
  • Affiliate links appear only where a reader would look for the service's own site, never inside a record or a table cell that carries a figure.
  • The full policy, including what is recorded when a link is clicked, is at Affiliate disclosure.

Consensus meta-models

PickAuditor derives four meta-models from the picks it captures. They are the site's own forecasters, and they are graded like everyone else: the same grading function, the same metrics, the same 30-pick threshold, the same public record. If they do badly, they look bad on the same leaderboard. The market price for a consensus pick is the snapshot at the time it was computed.

Consensus meta-model rules
Meta-modelRuleProduces a pick
MajoritymajorityThe same market and selection is backed by at least 5 services and the agree share is at least 0.65.Daily, at least 2 hours before the event starts.
Record-weightedrecord_weightedEach service's vote is weighted by its 90-day ROI (negative ROI counts as zero); the weighted vote must reach 0.6.Same.
CLV-weightedclv_weightedWeight equals the service's 90-day average CLV (zero or below counts as zero); only services with positive CLV vote.Same.
ContrariancontrarianTakes the opposite side of the majority, only when the agree share is at least 0.8 and the market price is at least 2.20.Rarely, by construction.

Consensus picks are shown to everyone on the event page and graded after the event. They are not advice, and a high agree share is not a probability.

Limitations

  • Capture cadence. A pick that appears and disappears between two four-hour captures is missed. A pick first seen shortly before kickoff leaves the market little time to move, so its CLV is close to zero by construction.
  • Odds coverage. Not every market has a reference price. Picks without a price drop out of ROI and CLV but not out of hit rate, which is why n can differ between columns of the same row.
  • Quoted odds. A service may quote a price from a bookmaker that is higher than the reference bookmaker's price at the same moment. CLV measured against Pinnacle's close then looks stronger than it would against that bookmaker's own close. The odds source is shown per pick.
  • Public tier only. A service's public picks may differ from what it sells. We measure what anyone can see, and every page that shows a record says so.
  • Matching and results errors happen. Each is corrected when found, regraded and logged. The raw capture is kept so the correction can be verified.
  • Small samples. 30 picks is a floor for display, not proof of anything. Read the p-value and the period, and prefer 365-day figures where they exist.
  • Survivorship. A service that stops publishing keeps its record here, marked paused. A service that removes past picks from its own pages does not remove them from ours.
  • Time. All times are UTC. A pick captured at 23:30 UTC on a Monday belongs to Monday even where the local calendar says otherwise.

Rule changes are dated in the changelog. Questions about a specific record go to contact@pickauditor.com.