← The frontier
Technology & AIJul 23, 2026

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction.

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.