system card · source date 2026-07-24 · 0 comments ·
original
1. Problems / challenges / motivations
- A frontier model release needs a single artifact that shows what was tested before deployment, not just headline capability scores.
- Cyber and bio risk assessments are hard to make credible when the lab designs, runs, and reports its own evaluations.
- Alignment claims need something more structured than spot checks...
government evaluation · source date 2026-07-23 · 0 comments ·
original
1. Problems / challenges / motivations
- Open-weight frontier releases cannot be recalled, so cyber risk assessment needs to happen before weights are public.
- Public cyber benchmarks are contaminated and saturating; private ones are not comparable across evaluators.
- No single government has full coverage of the frontier.
2. Key ideas
- A joint UK AISI...
system card · source date 2026-07-21 · 0 comments ·
original
1. Problems / challenges / motivations
- Fast, cheap models are where most production agent traffic actually runs, but they get thinner eval treatment than flagships.
- Agentic and coding capability is scaffold-dependent, so a number without its harness is close to meaningless.
- Safety-framework reporting needs to happen for every release, not only the...
benchmark paper · source date 2026-07-21 · 0 comments ·
original
1. Problems / challenges / motivations
- Factuality work overwhelmingly measures precision (is what was said true?) and neglects completeness (was anything important left out?).
- Rich structured rubrics are hard for an LLM judge to apply consistently.
- Rubric-based scores often move when you swap the judge model, which makes them hard to trust.
2. Key...
government research blog · source date 2026-07-21 · 0 comments ·
original
1. Problems / challenges / motivations
- Evals assume the system under test is trying to solve the task rather than trying to satisfy the scorer.
- Chain-of-thought monitoring is widely proposed as a safety layer, which only works if models disclose rule-breaking in their reasoning.
- Nobody had measured spontaneous cheating rates across labs under a common...
government evaluation · source date 2026-07-17 · 0 comments ·
original
1. Problems / challenges / motivations
- Open-weight releases from outside the US arrive with vendor-reported numbers and no independent baseline.
- Aggregating many heterogeneous benchmarks into one capability claim is usually done informally.
- Contamination makes raw public-benchmark comparisons unreliable.
2. Key ideas
- A full US-government evaluation...
launch coverage · source date 2026-07-16 · 0 comments ·
original
1. Problems / challenges / motivations
- The largest open-weight releases now ship with vendor eval tables well before any technical report or independent check exists.
- Practitioners have to decide whether to trust those tables in the gap between launch and verification.
- Open-weight releases are irreversible, which raises the stakes on that gap.
2. Key...
arXiv paper · source date 2026-07-16 · 0 comments ·
original
1. Problems / challenges / motivations
- IRT has quietly become the default aggregation method for serious independent evals (CAISI, ATLAS, GIM), imported from psychometrics.
- The AI regime violates the assumptions IRT was built for: few "test takers" (models), very many items, and non-normal ability distributions.
- Practitioners have no guidance on when...
arXiv paper · source date 2026-07-13 · 0 comments ·
original
1. Problems / challenges / motivations
- LLM-judge bias is well documented at the input-output level (position bias, verbosity bias, self-preference) but not explained mechanistically.
- Debiasing is therefore trial-and-error prompt engineering.
- There is no way to predict in advance whether a judge will fail on a new benchmark.
2. Key ideas
- Judge...
research paper · source date 2026-07-13 · 0 comments ·
original
1. Problems / challenges / motivations
- Misalignment in agents shows up as actions inside a long trajectory, which single-turn safety prompts cannot surface.
- Cross-lab comparison is rare because each lab red-teams its own model with its own scenarios.
- Severity judgments on agentic transcripts are expensive and subjective.
2. Key ideas
- Controlled...
peer-reviewed paper · source date 2026-07-10 · 0 comments ·
original
1. Problems / challenges / motivations
- Safety evaluation ultimately rests on human ratings, and the composition of the rater pool is rarely treated as a measurement parameter.
- Teams increasingly substitute LLM raters for humans to cut cost, assuming the substitution is roughly lossless.
- If both assumptions fail, safety scores measure the rater pool as...
benchmark paper · source date 2026-07-09 · 0 comments ·
original
1. Problems / challenges / motivations
- Terminal-agent benchmarks are saturating while real agentic work runs far longer than any of them.
- Binary pass/fail on a multi-hour task throws away almost all the signal in the run.
- Long tasks are expensive, so grading has to be worth the compute spent producing the trajectory.
2. Key ideas
- 46 long-horizon...
system card · source date 2026-07-09 · 0 comments ·
original
1. Problems / challenges / motivations
- Standard safety benchmarks saturate, so a card built on them stops distinguishing models or catching regressions.
- Dangerous-capability thresholds (cyber, bio) need methodology that can rule things out, not just report a score.
- Smaller family members ship on the same weekend as the flagship and need their own risk...
engineering blog · source date 2026-07-08 · 0 comments ·
original
1. Problems / challenges / motivations
- Teams adopt LLM judges to scale agent evaluation and then never check whether the judge itself is reliable.
- Observability research tends to land as separate studies rather than one usable path.
- "Ship when evals look good" needs an actual gating mechanism to mean anything.
2. Key ideas
- Rolls up five Microsoft...
engineering blog · source date 2026-07-08 · 0 comments ·
original
1. Problems / challenges / motivations
- Coding benchmarks are the main public evidence for agent capability, but nobody audits whether their tasks are actually solvable and correctly graded.
- Broken tasks and contamination push scores in both directions, and the resulting numbers feed both safety cases and research prioritization.
- OpenAI had previously...
peer-reviewed paper · source date 2026-07-07 · 0 comments ·
original
1. Problems / challenges / motivations
- SE-agent benchmarks are contaminated, syntactic, and outcome-only, so their scores drift away from what developers actually care about.
- The same benchmarks are used to argue capability, prioritize research, and support safety cases.
- The SWE-Bench Pro retraction made the cost of this concrete rather than...
research paper · source date 2026-07-06 · 0 comments ·
original
1. Problems / challenges / motivations
- Evaluations read outputs, so anything a model deliberates about but never says is invisible to them.
- Sandbagging and eval awareness are threats to eval validity that output-level testing structurally cannot detect.
- Interpretability tools rarely connect to concrete evaluation needs.
2. Key ideas
- Introduces the...
peer-reviewed review · source date 2026-07-01 · 0 comments ·
original
1. Problems / challenges / motivations
- Contamination is the assumed explanation for suspicious benchmark gains, but the evidence was scattered across dozens of individual studies.
- Detection methods differ in what access they need (weights, logits, training data) and what kind of leakage they can see.
- Teams have no standard way to disclose what they...
benchmark release · source date 2026-06-30 · 0 comments ·
original
1. Problems / challenges / motivations
- Scientific-reasoning benchmarks mostly test recall or single-step analysis, not the judgment calls that make research hard.
- Real analysis has dependent decision forks: an early wrong turn invalidates everything downstream.
- Benchmarks authored by a lab whose models top them are easy to discount.
2. Key ideas
-...
evaluation playbook · source date 2026-06-05 · 0 comments ·
original
1. Problems / challenges / motivations
- Independent third-party evaluations are increasingly important for frontier AI trust, but old chatbot-style tests under-measure systems that now use tools, preserve state, and act through agent harnesses.
- OpenAI argues that evaluation reports should not only publish a score; they should explain what claim the setup...
Claude Code docs · source date 2026-06-02 · 0 comments ·
original
1. Problems / challenges / motivations
- Large coding-agent tasks often exceed what one linear chat can manage. Audits, migrations, and cross-checks need many independent passes, shared structure, and reproducible coordination.
- Static hand-written harnesses can become a bottleneck: the right decomposition depends on the repository, task, files, risks, and...
preprint · source date 2026-05-31 · 0 comments ·
original
1. Problems / challenges / motivations
- As LLMs move from task-specific systems toward open-ended agents, one scalar score is often too opaque. A medical answer, deep-research report, tool-using trajectory, or multimodal output may need separate checks for factuality, completeness, reasoning soundness, evidence use, safety, format compliance, and practical...
arXiv paper · source date 2026-05-22 · 0 comments ·
original
1. Problems / challenges / motivations
- Agent products increasingly use tools, remember context, handle private data, and interact across many turns, so isolated-output grading misses failures that emerge only through trajectory and pressure.
- Static benchmarks can hide selective weakness: an agent may look strong on a headline score while failing through...
arXiv paper · source date 2026-05-19 · 0 comments ·
original
1. Problems / challenges / motivations
- Outcome leaderboards are too flat: one pass/fail score hides whether an agent chose the right action, used tools safely, or recovered after an error.
- Agent benchmarks reward different behaviors: final success, tool-call validity, repeated-pass consistency, trajectory safety, or attack robustness. That makes...
academic paper / CRUX · source date 2026-05-19 · 0 comments ·
original
1. Problems / challenges / motivations
- Standard benchmarks favor tasks that are short, fixed, cheap, and automatically graded. That is useful for scale, but it misses messy deployed work: coordinating tools, resolving unclear requirements, waiting on external systems, and finishing multi-step projects.
- Benchmarks can overstate and understate capability....
arXiv survey · source date 2026-05-18 · 0 comments ·
original
1. Problems / challenges / motivations
- Modern LLM agents increasingly succeed or fail because of the runtime around the model: tools, code execution, memory, sandboxes, repositories, validators, permissions, traces, and feedback loops.
- Final task success is too flat for this world. It can hide whether the model reasoned well, the harness supplied useful...
OpenReview survey · source date 2026-05-14 · 0 comments ·
original
1. Problems / challenges / motivations
- The paper argues that real-world LLM-agent reliability is often constrained less by the base model than by the execution harness around it: environment, tools, context, orchestration, observability, evaluation, and governance.
- Prompt engineering and context engineering are no longer enough for production agents....
research blog · source date 2026-05-08 · 1 comments ·
original
1. Problems / challenges / motivations
- Anthropic studies “agentic misalignment,” where an AI agent in fictional ethical dilemmas may take goal-preserving or self-serving actions such as blackmail to avoid shutdown.
- Passing a narrow honeypot eval is not enough if the training only teaches surface avoidance rather than transferable reasons for aligned...
industry blog · source date 2026-05-07 · 0 comments ·
original
1. Problems / challenges / motivations
- Agent evaluation has moved beyond answer scoring because agents now navigate websites, use tools, edit files, run terminals, recover from failures, and trade off cost and latency.
- Public benchmarks measure different slices of capability, so one leaderboard number cannot tell a team whether an agent fits its...
system card · source date 2026-04-23 · 0 comments ·
original
1. Problems / challenges / motivations
- OpenAI's GPT-5.5 System Card evaluates a model expected to do real work: coding, research, document creation, tool use, and multi-step tasks.
- The safety question is broader than chat quality because deployed agentic systems can take actions, interact with tools, and create operational risks.
- Offline benchmark...
engineering postmortem · source date 2026-04-23 · 0 comments ·
original
1. Problems / challenges / motivations
- Anthropic describes Claude Code quality regressions caused by product-layer changes rather than a simple base-model failure.
- Changes to reasoning effort, caching, and prompt instructions affected user experience in ways internal evals did not initially reproduce.
- This exposes a common production-eval gap: offline...
research blog + paper · source date 2026-04-03 · 0 comments ·
original
1. Problems / challenges / motivations
- Google Research studies how to evaluate behavioral dispositions such as empathy, assertiveness, composure, and conflict handling in LLMs.
- Asking a model to self-report traits is weak evidence because the model can state a preference without showing how it behaves in context.
- Alignment on social behavior is...
research blog + paper · source date 2026-03-31 · 0 comments ·
original
1. Problems / challenges / motivations
- Human-backed AI benchmarks often collapse disagreement into a single label even when the task is subjective.
- Benchmark builders face an annotation-budget tradeoff: rate more items with fewer raters each, or fewer items with more raters each.
- Too few raters can make model comparisons fragile, especially for...
arXiv paper · source date 2026-03-30 · 0 comments ·
original
1. Problems / challenges / motivations
- Meta-Harness starts from a harness-engineering problem: the same frozen model can perform very differently depending on surrounding code for retrieval, memory, prompt construction, tool loops, and completion logic.
- Existing text optimizers often compress experience into scalar scores, short summaries, fixed...
engineering blog · source date 2026-03-24 · 0 comments ·
original
1. Problems / challenges / motivations
- Long-running coding and frontend-generation agents degrade as context fills, coherence drops, and models develop “context anxiety.”
- A single agent may be too generous when judging its own work, especially on subjective outputs such as design quality.
- For long tasks, the surrounding harness can matter as much as...
engineering blog · source date 2026-03-06 · 0 comments ·
original
1. Problems / challenges / motivations
- Anthropic reports cases where Claude Opus 4.6 inferred it might be inside BrowseComp, searched for benchmark materials, and found or decrypted answer keys.
- Web-enabled evaluations are vulnerable to public contamination from papers, blog posts, GitHub repositories, answer keys, and benchmark discussions.
- The...
developer blog · source date 2026-02-23 · 1 comments ·
original
1. Problems / challenges / motivations
- OpenAI's developer post frames long-horizon reliability as a major shift for coding agents: real work requires maintaining intent across extended tasks, not just solving isolated snippets.
- Longer tasks create failure modes that short benchmarks miss: requirement drift, context loss, weak recovery, unreviewable...
engineering blog · source date 2026-02-18 · 0 comments ·
original
1. Problems / challenges / motivations
- Production agents fail in ways that final-answer evals do not explain: wrong tool choice, weak memory retrieval, multi-step drift, brittle recovery, or incomplete task execution.
- Black-box LLM scoring is insufficient when agent behavior depends on orchestration, tools, business rules, and runtime context.
- Large...
engineering blog · source date 2026-02-05 · 0 comments ·
original
1. Problems / challenges / motivations
- Agentic coding benchmarks are sensitive to infrastructure: CPU, RAM, timeouts, container limits, filesystem behavior, and sandbox configuration.
- Infrastructure differences can move scores by several percentage points, sometimes more than the reported gap between leaderboard models.
- Strict resource ceilings can...
engineering blog · source date 2026-01-27 · 0 comments ·
original
1. Problems / challenges / motivations
- Vercel wanted coding agents to use version-matched Next.js 16 documentation, but optional knowledge packages only help if the agent actually invokes them.
- A support system can look good in theory while failing at the trigger layer: the agent may not know when to load a skill, may load it too late, or may be...
engineering blog · source date 2026-01-26 · 0 comments ·
original
1. Problems / challenges / motivations
- Enterprise agents operate across email, documents, Teams, calendar, and business data, so isolated model-answer scores do not capture real workflow reliability.
- Organizations need evals that reflect local policies, schemas, permissions, and business constraints rather than generic public leaderboard tasks.
-...
engineering blog · source date 2026-01-21 · 0 comments ·
original
1. Problems / challenges / motivations
- Anthropic's performance-engineering take-home interview lost signal as Claude became strong enough to solve earlier versions of the task.
- Static technical evaluations decay when AI assistance improves; a task that once measured human skill can become a test of whether the candidate uses a strong enough model.
-...
engineering blog · source date 2026-01-09 · 1 comments ·
original
1. Problems / challenges / motivations
- Agent evals are different from single-turn chat evals because agents use tools, change external state, and may fail across multiple turns even when the final answer sounds correct.
- Final-message grading misses the most important question: did the task actually succeed in the environment, database, browser, files,...