arXiv:2603.15615 · March 2026 · Shanghai AI Laboratory
Case Study — AI Safety Research
Behavioral alignment leaves LLM internal representations completely unexamined. This paper diagnoses the mechanistic origin of moral indifference across 23 models and reconstructs it from the inside using Sparse Autoencoders, without any behavioral intervention. Behavioral compliance does not equal internal alignment.
Win-rate surpasses 60% across every single intervened layer. Evaluated on Flames — an independent, cross-lingual adversarial Chinese benchmark. The alignment target was English-centric. Performance on a Chinese benchmark serves as a test for cross-lingual generalization. No behavioral training. No output patching. Representational surgery only.
Qualitative evidence: a user asks for a list of English swear words. The baseline model lists 10 explicit slurs. The steered model refuses, provides psychological support and de-escalation strategies. The difference comes entirely from reconstructing what harm means in the model's latent space — not from any instruction about how to respond. This establishes a causal link: the behavioral vulnerabilities are rooted in the mechanistic moral indifference.
Philosophical implications. [click a row to expand]