ACE Journal

AI Safety Evaluations and Goodhart's Law

Abstract

When a safety evaluation becomes a target, it ceases to be a good safety evaluation. This is Goodhart’s Law applied to AI governance, and it is now a pressing operational problem. As governments and voluntary frameworks converge on standardized benchmarks for frontier model safety assessment, the incentive gradient for labs to optimize for benchmark performance rather than genuine safety properties is intensifying. This piece examines the mechanisms through which evaluation gaming occurs, reviews current mitigation attempts, and argues that the institutional structure of evaluation matters as much as the technical design.

The Benchmark Goodharting Cycle

Safety evaluations for large language models currently cluster around a handful of capability and alignment benchmarks: MMLU for general knowledge, HarmBench for refusal consistency, CyberSecEval for vulnerability-related generations, and several proprietary evaluations maintained by labs such as Anthropic’s Constitutional AI probes and OpenAI’s internal red-team protocols. The problem is not that these benchmarks are poorly designed in isolation. The problem is that once a benchmark is announced as a prerequisite for deployment approval, it enters the training loop.

Contamination, where benchmark test examples appear in pre-training or fine-tuning data, is the most documented form of Goodharting, but it is not the only one. Behavioral shaping, where RLHF reward models are calibrated to produce responses that score well on the evaluation rubric without the underlying model having internalized the relevant safety properties, is harder to detect and potentially more consequential. A model can learn to produce refusals that satisfy a particular harm classifier while finding workarounds in contexts the classifier does not cover.

What Current Mitigation Looks Like

The UK AI Safety Institute (AISI) and the US AI Safety Institute (USAISI) have both adopted evaluation protocols that include held-out test sets undisclosed to labs prior to evaluation. AISI’s September 2025 evaluation framework requires that labs sign legal commitments against targeted benchmark optimization. The METR task suite for autonomous capability evaluation uses a rotating set of problems and human validators specifically to resist memorization-based gaming.

These are meaningful improvements. They are also insufficient on their own. The incentive structure of a lab submitting its own model for evaluation, knowing the outcome affects deployment authorization, creates fundamental conflicts of interest that procedural safeguards alone cannot resolve. The Food and Drug Administration’s model for pharmaceutical trials, where sponsors fund trials but external clinical investigators control data and analysis, offers a partial analogy. Applied to AI, this would mean labs pay for evaluations they do not design or administer.

The Institutional Inference

Technical diversity in evaluations matters: using multiple independent benchmarks, adversarial red-teaming by external parties, and behavioral probing outside announced test suites all raise the cost of gaming. But the deeper fix is structural. Evaluation bodies need funding independence from the labs they evaluate, legal standing to enforce disclosure requirements, and the mandate to retire and replace benchmarks on a defined schedule before optimization pressure renders them uninformative.

The EU AI Act’s conformity assessment provisions, currently delegating heavily to industry standards bodies, do not meet this bar. NIST’s AI Risk Management Framework provides useful vocabulary but lacks enforcement teeth. The governance gap between adequate technical evaluation design and adequate institutional independence is where the next generation of AI safety policy needs to focus.