← The frontier
Technology & AIAug 20, 2026

Phantom Gains: Auditing Self-Improvement Against a Measured Null

Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses.

Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts. Auditing three…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.