The market for measurement
Frontier labs have made training data one of the largest line items in AI. Mercor, the biggest supplier of human data work, did $614M in the first half of 2026 and is running at a $2B annual pace, with roughly 91% of that revenue coming from frontier labs. Anthropic reportedly budgets on the order of $1B for RL environments alone. The money has moved to whatever grades the model.
It moved there because verification, not model capacity, has become the bottleneck. When a grader is wrong, the model learns the wrong thing: 28.5% of SWE-bench Verified tasks accept incorrect patches, and models score about 14 points higher on exactly those tasks. Better graders measurably produce better models, which makes grader quality a first-order training input rather than an evaluation detail. DeepSeek-R1-class reasoning came from little more than a prompt and a verifier.
Measurement is the grader that cannot be wrong in that way, and the precedents for training against it are the strongest results in scientific ML. AlphaFold came out of CASP, a contest that scored blind predictions against withheld experimental structures: one contest, one Nobel. GraphCast, trained on weather measurements, now out-forecasts the physical simulators that generated its training data. Machine-learned force fields reach near quantum-chemistry accuracy at roughly a thousand times the speed.
Yet frontier models still fail most research-level scientific coding and physics benchmarks, most scientific data is never published at all, and much of civilization's validated physics lives in decades-old code maintained by a shrinking community. The infrastructure is starting to catch up: Anthropic's Model Hardware Standard, released in August 2026, is a spec for agents operating laboratory instruments, piloted at Genentech, Janelia, and QuEra. The signal layer of the world is being connected to models. Someone has to supply the data.