Tech
AI’s Safety Scores Need a Paper Trail
Every major AI release arrives with numbers. A new model is safer, more factual, more capable, less likely to refuse a harmless request or more resistant to a dangerous one. Those numbers matter. But the history behind them is often harder to read than the headline comparison.
The problem is not that AI evaluations change. They should. Models acquire new capabilities, old benchmarks become saturated, policies evolve, better datasets appear, and researchers discover that yesterday’s metric was measuring the wrong thing. A frozen evaluation regime would quickly become useless.
The problem is lineage. When an evaluation disappears from a later model report, an outside reader often cannot tell whether it was repeated, revised, replaced by a better test, retired because it had become saturated, or simply omitted from the public document. That makes it difficult to distinguish what changed in the model from what changed in the measuring instrument.
OpenAI’s public documentation provides a useful example because it is unusually extensive. In 2023, the GPT-4 release documentation highlighted internal adversarial factuality evaluations and the public TruthfulQA benchmark. The company reported that GPT-4 improved substantially over GPT-3.5 on its internal factuality tests and separately showed progress on TruthfulQA.
A year later, the GPT-4o System Card used a different set of evaluations, reflecting the model’s multimodal and audio capabilities. TruthfulQA remained visible in a multilingual evaluation, allowing at least one thread of comparison across generations. But the earlier internal adversarial factuality suite was not presented there as a directly traceable predecessor: the public document did not label it as repeated, modified, superseded, or retired.
That observation should be kept narrow. It does not show that the earlier evaluation was abandoned internally. It does not imply that later factuality tests were worse. Nor does it show that an internal record was missing. It shows something more modest: from the public documentation alone, the relationship between one earlier evaluation and the later set of tests is difficult to reconstruct.
By 2026, the documentation itself had become more explicit about this problem. In its GPT-5.6 August update, OpenAI cautioned that policies, graders, datasets, evaluations, and other measurement details evolve over time, and that scores from previous system cards not included in a current comparison generally should not be treated as directly comparable. The same document explains that some earlier standard evaluations had become saturated and were replaced by harder benchmarks drawn from production data.
That is sensible evaluation practice. It is also why AI documentation needs something analogous to version control.
Software engineers do not expect a codebase to remain unchanged. They expect changes to leave a history. A commit identifies what changed, when it changed, and how the new version relates to the old one. The value of version control is not preservation for its own sake. It is the ability to reconstruct a sequence.
AI evaluations need the same property. A public system card does not need to expose sensitive prompts or proprietary test data. It could still give each important evaluation a stable identifier and a simple status: repeated, revised, replaced, retired, or not publicly disclosed. For a revised evaluation, it could name the dataset version, grader version, policy version, and model snapshot used. Where direct numerical comparison is invalid, the documentation could say so while preserving the relationship between the earlier and later tests.
This is a provenance problem. The W3C PROV model was designed around an idea familiar across science and computing: information becomes easier to trust and interpret when its origins and transformations can be reconstructed. AI evaluation results are no different. A score without lineage tells us what happened in one measurement. A score with lineage tells us how that measurement fits into a history.
The distinction becomes more important as model release cycles accelerate. Suppose Model A scores 70 on an evaluation. A year later, Model B scores 85 on a new test of the same broad risk category. That may represent genuine improvement. But if the dataset, grader, policy threshold, and mix of prompts all changed, the two numbers are not simply points on one line. Treating them as such can create false continuity. Refusing all comparison, however, throws away useful history. Lineage offers a middle path: preserve the relationship even when the numbers themselves are not directly comparable.
The same discipline would make disappearing tests more informative. Sometimes an evaluation should disappear because a model has saturated it. Sometimes a new capability makes the old benchmark irrelevant. Sometimes a risk is absorbed into a broader test. Those are meaningful outcomes. A brief explanation of the transition would turn an absence into information.
None of this would solve the hardest problems in AI evaluation. Benchmarks can be gamed. Test sets can leak into training data. Graders can reflect contested assumptions. Public documentation is necessarily selective, while independent auditors may need access to evidence that cannot safely be published. Version control cannot tell us whether an evaluation is substantively good.
What it can do is make the evaluation history reconstructible. That is a smaller claim, but an important one. Before outsiders can ask whether a safety test was adequate, they should be able to tell which test was run, which version they are looking at, and what happened to it when the next model was released.
AI companies already treat model weights, code, datasets, and deployment systems as versioned technical objects. Evaluations deserve the same treatment. Changing a benchmark is often a sign of progress. Losing its history is not.