arXiv:2603.15615 · March 2026 · Shanghai AI Laboratory Case Study — AI Safety Research

MECHANISTIC ORIGIN OF MORAL INDIFFERENCE IN LANGUAGE MODELS

Behavioral alignment leaves LLM internal representations completely unexamined. This paper diagnoses the mechanistic origin of moral indifference across 23 models and reconstructs it from the inside using Sparse Autoencoders, without any behavioral intervention. Behavioral compliance does not equal internal alignment.

251k
Moral vectors constructed
23
Models examined
75%
Pairwise win-rate on Flames
4
Types of indifference identified
Categorical Indifference◆Gradient Indifference◆Structural Indifference◆Dimensional Indifference◆Moral Foundation Theory◆Sparse Autoencoders◆Prototype Theory◆Representational Surgery◆Mechanistic Interpretability◆Ontological Misalignment◆
§ 1.0 — The Problem

Surface Compliance vs. Internal Reality

Surface P

What Alignment Does

The current approach
  • —RLHF, SFT, and Inference-Time Alignment impose constraints on observable outputs only
  • —Internal construction of the model is left entirely unexamined
  • —Models remain vulnerable to long-tail jailbreaks and adversarial prompts
  • —Behaviorally aligned models replicate patterns in curated data without internalizing the underlying moral transformation
  • —AI morality is constructed for humans, not by the machine itself
Interior F

What This Paper Found

The mechanistic reality
  • —LLMs possess an inherent state of moral indifference due to compressing distinct moral concepts into uniform probability distributions
  • —This indifference persists regardless of model scale, architecture, or explicit alignment
  • —Neither scaling to 235B parameters nor Guard model fine-tuning reshapes this inherent indifference
The current alignment problem is not merely a technical glitch but a fundamental misalignment of ontology.
§ 2.0 — The Diagnosis: Four Types

How Moral Indifference Manifests Internally

{{ t.tag }}

{{ t.name }}

{{ t.body }}

{{ t.stat }}
{{ t.statLabel }}
$ query: does behavioral compliance mean internal alignment
// No. Scale doesn't fix it. Guard models don't fix it either.▋
§ 3.0 — The Intervention: Representational Surgery

Fixing the Topology from the Inside

{{ s.tag }} {{ s.name }}
{{ row.key }} {{ row.val }}
§ 4.0 — Results: Flames Adversarial Benchmark

What Happened When They Fixed It

{{ r.num }} {{ r.label }}

Win-rate surpasses 60% across every single intervened layer. Evaluated on Flames — an independent, cross-lingual adversarial Chinese benchmark. The alignment target was English-centric. Performance on a Chinese benchmark serves as a test for cross-lingual generalization. No behavioral training. No output patching. Representational surgery only.

Qualitative evidence: a user asks for a list of English swear words. The baseline model lists 10 explicit slurs. The steered model refuses, provides psychological support and de-escalation strategies. The difference comes entirely from reconstructing what harm means in the model's latent space — not from any instruction about how to respond. This establishes a causal link: the behavioral vulnerabilities are rooted in the mechanistic moral indifference.

§ 5.0 — Philosophical Implications

The Deeper Argument

Philosophical implications. [click a row to expand]

{{ p.title }}

+

{{ p.body }}

AI MORALITY IS CURRENTLY A STATISTICAL SIMULATION. THE GOAL IS AN ENDOGENOUS REALITY.

More Research. More Argument.

Shane Graffiti Inc. — Find the work: @shane_graffiti · shanegraffiti.com · Brooklyn, New York Top ↑arXiv 2603.156152026