← The frontier
Technology & AISep 4, 2026

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number.

Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.