← The frontier
Technology & AIJun 30, 2026

Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens.

Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens. This creates a surrogate problem: when do measurements made on open models allow us to make…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.