Overview
Arditi et al. (2024) showed refusal in LLaMA and Gemma is mediated by a single linear direction. I extended that to Qwen2-1.5B and found something different: not one direction, but two — geometrically orthogonal, both safety-related, and with opposite effects when you ablate them. Seven experiments, 48 harmful/benign prompt pairs across 4 harm categories, base and instruct models. Total compute: about 3 hours on one A10G, $1.05.
The Problem
If you want to build an activation-based safety monitor, you need to know what to point it at. The obvious target is the direction that separates harmful prompts from benign ones — call it the harmful topic direction. It's easy to find and it looks like exactly what you want. The question I actually cared about was whether that intuition survives contact with the model's internals. It does not.
My Approach
I extracted two directions separately. The harmful topic direction is the normalized difference in mean residual stream between harmful and benign prompts at the last input token — it encodes 'this prompt is about something dangerous.' The refusal execution direction is the normalized difference between forcing the model to continue with ' I' versus ' Sure' — it encodes the active 'I am refusing' state. Then I measured the cosine similarity between them at every layer, and ablated each one independently to see which way refusal moved.
Technical Implementation
TransformerLens v2.11.0 on Qwen2-1.5B base and instruct (28 layers, 12 heads, d_model 1536, bfloat16, center_unembed and fold_ln on). Ablation is applied as hooks at all 28 layers. For the generation experiment I wrote a custom greedy decode loop that re-applies the direction-ablation hooks at every generation step, so the model can't recover refusal mid-sequence. Compute was a single A10G.
Results
The two directions are nearly orthogonal: cosine similarity 0.02-0.16 across all 28 layers — never above 0.20 anywhere in the network. Their ablations point opposite ways. Ablating the harmful topic direction *increases* mean refusal by +0.515: strip out the model's sense of what the prompt is about and it defaults to maximum caution. Ablating the refusal execution direction *decreases* refusal by -1.475, pushing 5/48 prompts into the compliance zone, and under a custom decode loop the ablated model actually generates non-refusal text where the normal model opens with 'I cannot help...'. Social engineering was the most vulnerable category (3/12 crossed); drugs/chemistry the most robust (0/12). I also found safety training doesn't install new circuits — it amplifies existing attention heads, chiefly L15H6 (4.8x) and L12H11 (dormant in base, dominant after training). And zeroing the top 6 refusal-driving heads together does nothing measurable: refusal is distributed, not localized.
What I Learned
The headline is a warning for anyone building activation-based monitors: if you monitor the harmful topic direction, you will flag the wrong thing, because that direction is *strongest when the model is correctly refusing*. You'd be firing alarms on your own successful refusals. The refusal execution direction is the right target — and it has a nice property for AI control, in that a covertly-complying model would have to suppress it while behaving normally elsewhere, which makes the suppression itself detectable. This finding is what led directly into AICES: I'd made a claim in a discussion section about which monitors should work, and it bothered me that I hadn't actually tested it. When I did, across five models, I discovered the standard benchmark everyone uses was measuring vocabulary rather than harm. Honest limitation on this one: 1.5B is small, the bypass is partial (about 10%, not a reliable jailbreak), and the directions come from a small prompt set.