← The frontier
Technology & AIAug 12, 2026

A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions

We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e.

We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. reward functions that adhere to a given preference ordering over trajectories. Given a task described in natural language, our process produces a linear reward…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.