← The frontier
Technology & AIAug 24, 2026

How to Train a Critic Stably and Efficiently

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt.

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.