Track record

How accurate is our outlook, really?

Every projection W/ SEC makes gets graded against what actually happened, and nothing is quietly dropped from the count. This page is that record — updated from the same data the app itself reads, not a curated highlight reel.

Direction accuracy right direction, in or out of range
Hit rate landed inside the projected range
Companies covered distinct filers graded
Calls graded across four metrics

These four numbers come from our historical backtest — see "Backtest vs. live" below before treating them as a live forward-looking claim.

Definitions

Every projection is a directional call (up / down / flat) plus a range. Once the next filing lands, we grade it one of three ways:

This grades our forecast, not the company — a "wrong call" can still be great news for the business (e.g. we projected a decline and the actual number grew).

Inputs & exclusions

A projection is not the company's own guidance, and it isn't freely made up: the range comes from a deterministic statistical model — the empirical spread of what actually happened historically to companies whose metric was trending the same way (up / down / flat) as this one's most recent year. We do not use analyst estimates, guidance, or anything the company itself forecasts.

Net income's range runs noticeably wider than the other three metrics below. That's structural, not a modeling gap: net income is the GAAP line most exposed to one-time items (impairments, tax adjustments, asset sales) and to companies sitting near breakeven, both of which swing the percentage sharply. Treat the direction as the signal; the width is inherent to this one metric.

Backtest vs. live

The stats above are an out-of-sample historical backtest: each historical filing is graded against a range computed only from data available before that filing, the same way a live call would be — but run retrospectively across SEC history, so the record has real sample size from day one instead of starting empty. It is not the same claim as "these were live predictions we published in advance."

We also keep a separate, much smaller count of genuinely live calls — graded only after they were made and a real filing settled them going forward. It's shown below on its own, not folded into the backtest number above, specifically because it's too small yet to draw conclusions from.

Backtest accuracy, by metric

Hit Partial Wrong call

Loading…

Since public launch

Loading…

Limitations

Does a red flag predict an SEC enforcement action?

Our forensic score isn't only about fiscal projections — it also flags financial-statement red flags (leverage, accruals, liquidity, and three more deterministic checks). We wanted a harder test than grading our own forecasts: does that score actually run higher for companies the SEC later charged with accounting fraud than for the market at large?

We cross-referenced every SEC Accounting and Auditing Enforcement Release (AAER) since 2010 — the earliest year XBRL data exists to check against — matched each one to a specific public company by name, and reconstructed what our forensic score would have said the last time that company reported before the SEC acted, using only the data that was actually knowable at the time. The market-wide baseline runs the identical five checks across every US filer, same years, so "elevated" and "high" mean exactly the same thing on both sides of the comparison.

Enforcement companies flagged elevated/high, before the SEC acted
Market-wide baseline same flag rate, every US filer, same years
Lift enforcement flag rate ÷ market flag rate
Companies checked matched enforcement actions, 2010–present

What we found

Loading…

Inputs & exclusions

An AAER is matched to a CIK only by an exact, deterministic normalized-name match — the same matching primitive the app uses elsewhere for 13F holdings. A name that collides across more than one real company (a parent and its own subsidiary, or two unrelated companies sharing a dual-listed name) is left unmatched rather than guessed. Audit firms, individuals, and private advisers named alongside a company in the same release are excluded entirely — only the company itself is scored.

A company with too little XBRL data to run any of the five checks is counted as its own bucket, not silently dropped and not counted as a miss — both here and in the market-wide baseline, so neither side is inflated or deflated by how much data happened to exist.

Limitations