Technology & AISep 3, 2026
Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model?
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a…
Sign in to learn & save →
The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.