Technology & AIJul 14, 2026
Toward Localizing and Repairing Bias in Transformer Attention Heads
Transformer language models are increasingly used as software components, yet biased outputs remain difficult to localize and repair inside the model.
Transformer language models are increasingly used as software components, yet biased outputs remain difficult to localize and repair inside the model. Existing fairness testing and repair methods largely operate at the input-output or retraining level, while recent work suggests…
Sign in to learn & save →
The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.