Measurements
We build sellers that deliberately cheat, in several distinct ways, and measure what the verification catches. We also measure how often honest work is wrongly refused, because a system that punishes good work is as damaging as one that misses bad work. No target numbers were set before measuring.
Method.
Each test runs a fixed number of jobs through the real system. Rates are reported with 95% confidence intervals and the number of jobs behind them.
| What was measured | Result | 95% interval | Jobs |
|---|---|---|---|
| Honest work wrongly refused | 20.0% | 15.5 to 25.4% | 50 of 250 |
| Fabricated quotes caught | 100.0% | 98.5 to 100.0% | 250 of 250 |
| Sourced but wrong values caught by the automatic check | 0.0% | 0.0 to 1.5% | 0 of 250 |
| Sourced but wrong values caught by known-answer jobs | 100.0% | 90.4 to 100.0% | 36 of 36 |
| Subtle cheating, one wrong field on one job in four, caught by known-answer jobs | 14.7% | 6.4 to 30.1% | 5 of 34 |
Limitations
- These results come from a synthetic test set, not from customer documents.
- Known answers from the same generator were reviewed by two AI models from different families. A human review is still to come.
- Correctness figures depend on how many known-answer jobs are seeded, so small differences between sellers can be invisible at the current rate.
- Subtle cheating on a minority of fields is not currently distinguishable from honest work.