Verbalizable Representations Form a Global Workspace (J-lens)
research paper · source date 2026-07-06 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog
1
Problems / challenges / motivations
- Evaluations read outputs, so anything a model deliberates about but never says is invisible to them.
- Sandbagging and eval awareness are threats to eval validity that output-level testing structurally cannot detect.
- Interpretability tools rarely connect to concrete evaluation needs.
2
Key ideas
- Introduces the Jacobian lens (J-lens) and "J-space": a small privileged set of verbalizable representations with global-workspace properties, validated across fourteen tasks via ablation.
- Decoding J-space surfaces strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in outputs.
- The findings motivate counterfactual reflection training as a mitigation.
- Ships with open-source code and a Neuronpedia demo.
3
Why it matters for evals
- It offers a concrete tool for reading a model's unspoken reasoning, which targets one of the deepest threats to eval validity: a model that behaves differently because it knows it is being tested.
- It pairs naturally with the AISI cheating result — one shows models hide rule-breaking from chain-of-thought, the other offers a channel that is not the chain-of-thought.
- Caveat: downstream eval utility is still being demonstrated, and the J-space findings depend on the technique's assumptions and need replication beyond the reported task battery.
Comments
No comments yet.