AI & Agent Evaluation
2,151total visitsadmin

Verbalizable Representations Form a Global Workspace (J-lens)

research paper · source date 2026-07-06 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog

Problems / challenges / motivations

  • Evaluations read outputs, so anything a model deliberates about but never says is invisible to them.
  • Sandbagging and eval awareness are threats to eval validity that output-level testing structurally cannot detect.
  • Interpretability tools rarely connect to concrete evaluation needs.

Key ideas

  • Introduces the Jacobian lens (J-lens) and "J-space": a small privileged set of verbalizable representations with global-workspace properties, validated across fourteen tasks via ablation.
  • Decoding J-space surfaces strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in outputs.
  • The findings motivate counterfactual reflection training as a mitigation.
  • Ships with open-source code and a Neuronpedia demo.

Why it matters for evals

  • It offers a concrete tool for reading a model's unspoken reasoning, which targets one of the deepest threats to eval validity: a model that behaves differently because it knows it is being tested.
  • It pairs naturally with the AISI cheating result — one shows models hide rule-breaking from chain-of-thought, the other offers a channel that is not the chain-of-thought.
  • Caveat: downstream eval utility is still being demonstrated, and the J-space findings depend on the technique's assumptions and need replication beyond the reported task battery.

Comments

No comments yet.