Why we publish our errors
Anyone can show you their wins — a highlight reel is the cheapest thing in AI. Publishing the misses, next to the hits, on a record that can’t be quietly edited, is the only version of a track record worth believing.
Every AI company can show you a highlight reel. Curated demos, hand-picked benchmark wins, the three anecdotes where the model was brilliant. Highlight reels are the cheapest artifact in this industry — which is exactly why they carry no information. If every vendor, good and bad, can produce one, then seeing one tells you nothing about which kind of vendor you’re looking at.
The misses are the signal
What a bad-faith vendor cannot cheaply produce is a public record of being wrong. Publishing errors is costly in precisely the way that makes it credible: it invites criticism, it arms competitors with quotes, and it forecloses the option of quietly improving the story later. That is the economics of the signal. When someone shows you their misses next to their hits, they are demonstrating — not asserting — that the record was not built by selection. A track record with no visible errors isn’t a clean record. It’s an edited one.
What we actually publish
The AI Trust Index is the standing example. It scores 18 frontier models over 13,171 scored pairs — 7,447 forecasts plus 5,724 knowledge questions — against ground truth, and the out-of-sample result, a Brier score of 0.1877 against a 0.25 uninformed baseline, is published with the errors on the record next to the hits. The methodology is hashed, so the scoring rules cannot quietly shift between runs to flatter a result. And the underlying record is hash-locked and recomputable: change a value and the arithmetic no longer checks out. You are not asked to trust our summary of how we did. You can re-derive it.
Why a proper score makes hiding pointless
There is also a mathematical reason we can commit to this. The Brier score (Brier, 1950) is a proper scoring rule — the only way to improve it is to make stated confidence match reality. You cannot hedge your way to a good number, and neither can we. Publishing under a proper scoring rule is a form of pre-commitment: we chose an instrument that punishes exactly the behavior — confident overstatement — that error-hiding is designed to conceal.
The standard this sets — for us
Once your public track record includes the misses, every other claim you make gets read against it, and that is the point. It disciplines the product side of the company too: our judgment engines compute cross-lab agreement from real referee positions rather than asserting consensus, and flag one-sided runs rather than presenting them as balanced, because the alternative — smoothing over inconvenient detail — is the same sin as deleting a miss from the index.
The standard it sets — for everyone you buy from
Here is the portable version of this essay: when any AI vendor shows you evidence, ask where their errors are. Not whether errors exist — they always exist — but whether the vendor’s own materials will show them to you, and whether the record they live in could be edited without a trace. A vendor who publishes misses on a tamper-evident record has made lying expensive for themselves. A vendor who shows you only wins has kept it free. Choose accordingly — including about us.