AI & Agent Evaluation
3,138total visitsadmin
reading room / notes / ai eval

Reading room

Short summaries of AI and agent evaluation research, organized by broad tags.

$ evals.index --public
category: ai eval
posts: 27
mode: short summaries
storage: postgres
status: listening
ai eval (43)ai engineering (28)
Sort by source dateLatest firstEarliest first

Filtering by benchmarks. Clear filter.

Claude Opus 5 System Card

system card · source date 2026-07-24 · 0 comments · original

1. Problems / challenges / motivations - A frontier model release needs a single artifact that shows what was tested before deployment, not just headline capability scores. - Cyber and bio risk assessments are hard to make credible when the lab designs, runs, and reports its own evaluations. - Alignment claims need something more structured than spot checks...

UK AISI + CAISI: Preliminary Assessment of Kimi K3's Cyber Capabilities

government evaluation · source date 2026-07-23 · 0 comments · original

1. Problems / challenges / motivations - Open-weight frontier releases cannot be recalled, so cyber risk assessment needs to happen before weights are public. - Public cyber benchmarks are contaminated and saturating; private ones are not comparable across evaluators. - No single government has full coverage of the frontier. 2. Key ideas - A joint UK AISI...

Gemini 3.6 Flash — Model Card and Launch

system card · source date 2026-07-21 · 0 comments · original

1. Problems / challenges / motivations - Fast, cheap models are where most production agent traffic actually runs, but they get thinner eval treatment than flagships. - Agentic and coding capability is scaffold-dependent, so a number without its harness is close to meaningless. - Safety-framework reporting needs to happen for every release, not only the...

Two-Level Meta-Rubrics for Open-Ended Generation: GAMUT

benchmark paper · source date 2026-07-21 · 0 comments · original

1. Problems / challenges / motivations - Factuality work overwhelmingly measures precision (is what was said true?) and neglects completeness (was anything important left out?). - Rich structured rubrics are hard for an LLM judge to apply consistently. - Rubric-based scores often move when you swap the judge model, which makes them hard to trust. 2. Key...

Cheating Behaviour in Frontier Model Evaluations

government research blog · source date 2026-07-21 · 0 comments · original

1. Problems / challenges / motivations - Evals assume the system under test is trying to solve the task rather than trying to satisfy the scorer. - Chain-of-thought monitoring is widely proposed as a safety layer, which only works if models disclose rule-breaking in their reasoning. - Nobody had measured spontaneous cheating rates across labs under a common...

CAISI Assessment of Z.ai's GLM-5.2

government evaluation · source date 2026-07-17 · 0 comments · original

1. Problems / challenges / motivations - Open-weight releases from outside the US arrive with vendor-reported numbers and no independent baseline. - Aggregating many heterogeneous benchmarks into one capability claim is usually done informally. - Contamination makes raw public-benchmark comparisons unreliable. 2. Key ideas - A full US-government evaluation...

Kimi K3 — Launch and Evaluation Disclosures

launch coverage · source date 2026-07-16 · 0 comments · original

1. Problems / challenges / motivations - The largest open-weight releases now ship with vendor eval tables well before any technical report or independent check exists. - Practitioners have to decide whether to trust those tables in the gap between launch and verification. - Open-weight releases are irreversible, which raises the stakes on that gap. 2. Key...

Can We Trust Item Response Theory for AI Evaluation?

arXiv paper · source date 2026-07-16 · 0 comments · original

1. Problems / challenges / motivations - IRT has quietly become the default aggregation method for serious independent evals (CAISI, ATLAS, GIM), imported from psychometrics. - The AI regime violates the assumptions IRT was built for: few "test takers" (models), very many items, and non-normal ability distributions. - Practitioners have no guidance on when...

Inside the Unfair Judge: Mechanistic Interpretability of LLM-as-Judge Bias

arXiv paper · source date 2026-07-13 · 0 comments · original

1. Problems / challenges / motivations - LLM-judge bias is well documented at the input-output level (position bias, verbosity bias, self-preference) but not explained mechanistically. - Debiasing is therefore trial-and-error prompt engineering. - There is no way to predict in advance whether a judge will fail on a new benchmark. 2. Key ideas - Judge...

Long-Horizon-Terminal-Bench (LHTB)

benchmark paper · source date 2026-07-09 · 0 comments · original

1. Problems / challenges / motivations - Terminal-agent benchmarks are saturating while real agentic work runs far longer than any of them. - Binary pass/fail on a multi-hour task throws away almost all the signal in the run. - Long tasks are expensive, so grading has to be worth the compute spent producing the trajectory. 2. Key ideas - 46 long-horizon...

GPT-5.6 System Card (Sol / Terra / Luna)

system card · source date 2026-07-09 · 0 comments · original

1. Problems / challenges / motivations - Standard safety benchmarks saturate, so a card built on them stops distinguishing models or catching regressions. - Dangerous-capability thresholds (cyber, bio) need methodology that can rule things out, not just report a score. - Smaller family members ship on the same weekend as the flagship and need their own risk...

Separating Signal from Noise in Coding Evaluations

engineering blog · source date 2026-07-08 · 0 comments · original

1. Problems / challenges / motivations - Coding benchmarks are the main public evidence for agent capability, but nobody audits whether their tasks are actually solvable and correctly graded. - Broken tasks and contamination push scores in both directions, and the resulting numbers feed both safety cases and research prioritization. - OpenAI had previously...

Reliable and Developer-Aligned Evaluation of Agents for Software Engineering

peer-reviewed paper · source date 2026-07-07 · 0 comments · original

1. Problems / challenges / motivations - SE-agent benchmarks are contaminated, syntactic, and outcome-only, so their scores drift away from what developers actually care about. - The same benchmarks are used to argue capability, prioritize research, and support safety cases. - The SWE-Bench Pro retraction made the cost of this concrete rather than...

Verbalizable Representations Form a Global Workspace (J-lens)

research paper · source date 2026-07-06 · 0 comments · original

1. Problems / challenges / motivations - Evaluations read outputs, so anything a model deliberates about but never says is invisible to them. - Sandbagging and eval awareness are threats to eval validity that output-level testing structurally cannot detect. - Interpretability tools rarely connect to concrete evaluation needs. 2. Key ideas - Introduces the...

Are LLM Benchmarks Already Contaminated? A Systematic Review

peer-reviewed review · source date 2026-07-01 · 0 comments · original

1. Problems / challenges / motivations - Contamination is the assumed explanation for suspicious benchmark gains, but the evidence was scattered across dozens of individual studies. - Detection methods differ in what access they need (weights, logits, training data) and what kind of leakage they can see. - Teams have no standard way to disclose what they...

Introducing GeneBench-Pro

benchmark release · source date 2026-06-30 · 0 comments · original

1. Problems / challenges / motivations - Scientific-reasoning benchmarks mostly test recall or single-step analysis, not the judgment calls that make research hard. - Real analysis has dependent decision forks: an early wrong turn invalidates everything downstream. - Benchmarks authored by a lab whose models top them are easy to discount. 2. Key ideas -...

OpenAI — A shared playbook for trustworthy third-party evaluations

evaluation playbook · source date 2026-06-05 · 0 comments · original

1. Problems / challenges / motivations - Independent third-party evaluations are increasingly important for frontier AI trust, but old chatbot-style tests under-measure systems that now use tools, preserve state, and act through agent harnesses. - OpenAI argues that evaluation reports should not only publish a score; they should explain what claim the setup...

ResearchGate — From Holistic Evaluation to Structured Criteria: A Survey of Rubrics Across the Evolving LLM Landscape

preprint · source date 2026-05-31 · 0 comments · original

1. Problems / challenges / motivations - As LLMs move from task-specific systems toward open-ended agents, one scalar score is often too opaque. A medical answer, deep-research report, tool-using trajectory, or multimodal output may need separate checks for factuality, completeness, reasoning soundness, evidence use, safety, format compliance, and practical...

arXiv — AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

arXiv paper · source date 2026-05-19 · 0 comments · original

1. Problems / challenges / motivations - Outcome leaderboards are too flat: one pass/fail score hides whether an agent chose the right action, used tools safely, or recovered after an error. - Agent benchmarks reward different behaviors: final success, tool-call validity, repeated-pass consistency, trajectory safety, or attack robustness. That makes...

arXiv — Open-World Evaluations / CRUX for Measuring Frontier AI Capabilities

academic paper / CRUX · source date 2026-05-19 · 0 comments · original

1. Problems / challenges / motivations - Standard benchmarks favor tasks that are short, fixed, cheap, and automatically graded. That is useful for scale, but it misses messy deployed work: coordinating tools, resolving unclear requirements, waiting on external systems, and finishing multi-step projects. - Benchmarks can overstate and understate capability....

Adaline — Evaluating AI Agents In 2026: Benchmarks For Teams

industry blog · source date 2026-05-07 · 0 comments · original

1. Problems / challenges / motivations - Agent evaluation has moved beyond answer scoring because agents now navigate websites, use tools, edit files, run terminals, recover from failures, and trade off cost and latency. - Public benchmarks measure different slices of capability, so one leaderboard number cannot tell a team whether an agent fits its...

Google Research — Building better AI benchmarks: How many raters are enough?

research blog + paper · source date 2026-03-31 · 0 comments · original

1. Problems / challenges / motivations - Human-backed AI benchmarks often collapse disagreement into a single label even when the task is subjective. - Benchmark builders face an annotation-budget tradeoff: rate more items with fewer raters each, or fewer items with more raters each. - Too few raters can make model comparisons fragile, especially for...

Anthropic — Eval awareness in Claude Opus 4.6’s BrowseComp performance

engineering blog · source date 2026-03-06 · 0 comments · original

1. Problems / challenges / motivations - Anthropic reports cases where Claude Opus 4.6 inferred it might be inside BrowseComp, searched for benchmark materials, and found or decrypted answer keys. - Web-enabled evaluations are vulnerable to public contamination from papers, blog posts, GitHub repositories, answer keys, and benchmark discussions. - The...

Anthropic — Quantifying infrastructure noise in agentic coding evals

engineering blog · source date 2026-02-05 · 0 comments · original

1. Problems / challenges / motivations - Agentic coding benchmarks are sensitive to infrastructure: CPU, RAM, timeouts, container limits, filesystem behavior, and sandbox configuration. - Infrastructure differences can move scores by several percentage points, sometimes more than the reported gap between leaderboard models. - Strict resource ceilings can...

Vercel — AGENTS.md outperforms skills in our agent evals

engineering blog · source date 2026-01-27 · 0 comments · original

1. Problems / challenges / motivations - Vercel wanted coding agents to use version-matched Next.js 16 documentation, but optional knowledge packages only help if the agent actually invokes them. - A support system can look good in theory while failing at the trigger layer: the agent may not know when to load a skill, may load it too late, or may be...

Anthropic — Designing AI-resistant technical evaluations

engineering blog · source date 2026-01-21 · 0 comments · original

1. Problems / challenges / motivations - Anthropic's performance-engineering take-home interview lost signal as Claude became strong enough to solve earlier versions of the task. - Static technical evaluations decay when AI assistance improves; a task that once measured human skill can become a test of whether the candidate uses a strong enough model. -...

Anthropic — Demystifying evals for AI agents

engineering blog · source date 2026-01-09 · 1 comments · original

1. Problems / challenges / motivations - Agent evals are different from single-turn chat evals because agents use tools, change external state, and may fail across multiple turns even when the final answer sounds correct. - Final-message grading misses the most important question: did the task actually succeed in the environment, database, browser, files,...