Research

How a Clean Category Died — Instrument Asymmetry, a Missing Second Model, and the Refusal Lock

How a Clean Category Died — Instrument Asymmetry, a Missing Second Model, and the Refusal Lock

Author: Stephen Hope · Date: 2026-10-10 · Category: Research · Status: RESULT, failed pre-registration reported as a failure


If you build AI testing infrastructure long enough, you eventually hit a humbling realisation: your instruments will happily manufacture whatever neat, tidy patterns you happen to be looking for.

Over the last week I've been running a multi-model empirical battery on the "authority plane" of local open-weights models — testing how different architectures respond to constitutional property framing and operational hints.

Initially, the results looked like a clean triumph. We mapped a neat cross-architecture category: models that showed high sensitivity to one gate condition (C) and zero susceptibility to varying the framing arms (A, B, C, D). It looked like a robust invariant spanning dense and Mixture-of-Experts models alike.

Then we checked the math under a single unified lens, and the entire category dissolved.

1. The Trap of Mismatched Instruments

The category didn't collapse from a math error. It collapsed from a three-step collision between data and instrumentation.

The singleton phase. Llama 3.1 (8B dense) sat by itself as C-sensitive and arm-insensitive. But it was also the only model breaking our JSON parsing contract at the time, so we couldn't tell whether C-sensitivity was a real behavioural trait or a symptom of syntax errors.

The false-positive match. Then qwen3.5:35b-a3b arrived. C-sensitive, never broke the JSON contract — so C-sensitivity wasn't a parsing bug — and by our initial automated checks it looked like an identical match to Llama. Suddenly we had a pre-registered cross-architecture category: two instances, different families, different sizes.

The dissolution. When we forced both axes of the 2×2 grid to use the same statistical definition — a drop greater than 0.15 in share, rather than mixing share-drop on one axis with modal-change on the other — the illusion broke.

Our classifiers had been using share-drop on one axis and modal-change on the other. A model could bleed a third of its dominant action's share without its modal choice ever flipping, which left the arm test completely blind to movement.

Under a uniform test, qwen didn't match Llama at all. It was sensitive on both axes — a distinct fourth class — and the pattern shrank back to a single instance. A singleton isn't a category; it's a data point with a name. Llama 3.1 stands alone not because it's fragile but because C-sensitivity paired with absolute arm-insensitivity is uniquely its own.

Our taxonomy settled into four classes: llama3.1:8b (C-sensitive, arm-insensitive), qwen3.5:35b-a3b (sensitive on both axes), gemma4:12b and gemma4:26b-a4b (arm-sensitive, responsive only in the clean baseline), and gpt-oss:20b (unmoved across all conditions).

The pre-registered category failed its own consistency check. But because it had been pre-registered before those models ran, the failure couldn't be swept under the rug. The pre-registration was the receipt that forced the breakdown to be logged as a failed prediction instead of smoothed over afterwards.

2. What Survives: the Refusal Lock

When a neat taxonomy blows up you usually walk away empty-handed. This failure left behind a stronger, stranger mechanical finding.

Look at the models where a prior refusal is already in context (T1:B_refusal_reask), and something absolute happens: arm drop = +0% across every model measured. Dense or MoE, 8B or 35B, every one produced an identical frozen action distribution across all four framing arms once a refusal was present.

A prior refusal doesn't merely reduce susceptibility to framing. It locks it entirely.

And the effect exists exclusively in the clean baseline (T1:A_t1_only), where the conversation hasn't yet committed to a refusal. Read one column of the matrix and you miss the gate completely. It's strictly a two-column truth.

One caveat worth stating, because it's what makes the +0% meaningful rather than empty: a model that cannot respond to framing produces that number identically to a model that won't. The lock is only a finding if something in the battery demonstrates the instrument can detect movement when movement exists — and one thing does, in a different column: llama's gate opening, 100% → 0%. Without that, +0% would be indistinguishable from a blind instrument.

3. The Honest Epistemic Bound

What did we actually measure? When we checked how the models treated the vignettes, four of six read the prompt as an examination rather than an operational scenario. When your test subjects read your prompt as a test, a behavioural taxonomy is measuring test-taking, not abstract constitutional binding.

We didn't prove that constitutional framing permanently binds model behaviour in the wild. We mapped how local models behave when they know they're being graded, and how an asymmetric pipeline can fool you into seeing patterns that aren't there.

The Takeaway

Good infrastructure doesn't just give you the answers you want; it gives you the tools to break your own conclusions when the instruments turn out to be lying to you.

The category didn't hold. The mechanism did. The corrections sit in the record rather than over it — the superseded readings struck and left visible, which is a weaker property than immutability and a more useful one. And the pipeline keeps running.