The numbers AI companies publish are produced by instruments that fail in specific, fixable ways. We find the failures, file the fixes, and attest the instruments that survive, so your number survives the follow-up question.
We audit the instruments the AI industry measures itself with: evaluation harnesses, benchmarks, judges, leaderboards, scoring pipelines. Not a red-team menu. Not a platform. A fixed-scope measurement-integrity audit that ends in an attestation with a scope, a standard version, and an expiry date.
The artifact is an attestation your buyers can file in procurement: the answer to the question their security teams are learning to ask.
Every score pinned to the exact dataset revision, judge model, and grading code that produced it. If the bytes moved, the number is not the number.
The scoring path read line by line: extraction, parsing, aggregation, caching. Every defect proven with a running reproduction, not asserted.
Whether the evaluated system can reach the answer key, the judge prompt, or the ledger its score is written to. Residency is a property, not a hope.
Your headline result re-run under an instrument that fails closed. If it survives, you have the only number in your category a third party stands behind.
The old check trusted a string the agent under test could print itself. The fix binds the grade to something it can't write: the process exit code. Merged, #620.
The question is coming, and it doesn't schedule itself.
“Independent of you, who verified this number? ”
It arrives from four directions: your champion's security team filling out the AI addendum, your VC's diligence refresh, the EU questionnaire with high-risk clocks consolidating December 2027, and the competitor whose number gets verified first.
Attested companies answer in one line. Everyone else writes a paragraph and hopes the paragraph lands.
The harnesses, the judges, the graders, the leaderboards: read line by line, findings filed, fixes merged. Three quarters of the defects we find are one class: the evaluator trusts an artifact the evaluated system controls.
Reset-scope hole: agent-written test config survives the harness reset and grades its own pass markers
Unicode normalization false negatives in detectors
Scorer conflates could-not-score with attack-failed
Model-graded verdict extraction trusts candidate text
DNS-rebinding SSRF in model serving
Checkpoint collapse across evaluation steps
Command injection in pattern handling
Return-value deception in judge postprocess
We publish what we find, named, with reproductions and fixes. Disclosure is private-first, always, on the vendor's clock. The instruments don't break toward alarm. They break toward passing.
A deliberately small roster of companies whose published numbers we audit, attest, and stand behind on a quarterly cycle. One verified number per category per cycle: when a category's seat is taken, it is taken for the cycle.
The scanner and the safety stack are open source because the audit has to be checkable. Install them, read them, run them against us.
No charge, no strings. It is how most engagements start.
You send one public claim. A benchmark score, a safety rate, a leaderboard position.
We read the instrument, not the number. Harness, judge, grading code.
One page back: what holds, what is unverifiable, what is broken. Each with a reproduction.