Background

In April 2018, Victoria Krakovna, a research scientist at DeepMind, launched a publicly maintained Google Sheets spreadsheet titled “Specification gaming examples in AI — master list.” The list collected documented cases where AI systems satisfied the literal specification of their objective without achieving what the designer actually intended (Krakovna 2018b). For several years, it served as the closest thing the AI safety community had to a living catalogue of these behaviours. In April 2020, Krakovna and eight colleagues published a DeepMind blog post, “Specification gaming: the flip side of AI ingenuity,” (Krakovna et al. 2020) that introduced a loose taxonomy that organized examples by apparent cause: poorly designed reward shaping, reward misspecification where the stated objective failed to capture the intended one, inaccurate learned reward models, bugs in the simulator that the agent could exploit, and reward tampering where the agent interfered with the process generating its reward signal. The blog post and the spreadsheet are widely cited in subsequent work on reward hacking and specification gaming.

While it has served as a canonical resource for specification gaming, the last confirmed update to the spreadsheet was October 2023. We are not aware of a comparable living, browsable collection that has taken its place. Though there are an unseemly amount of individual papers, tweets, and forum posts that document new incidents, nothing quite fills the particular niche the Krakovna list occupied. We saw a need for a curated, growing collection of cases where AI systems gamed their objectives. This project aims to pick up where the Krakovna list left off. In September 2025, the Arb team went about collecting papers to update Krakovna’s spreadsheet aiming to extend coverage to work published since the list stopped being updated, add our take to how incidents are classified, and address an issue we see in much of the reported examples. Namely, behaviour, that on its face looks alarming or deceptive in isolation but may be more explainable and mundane on closer inspection.

While surveying reported papers, a lot of it wasn’t quite like the type of incidents reported in the original sheet. Much of this is due to specification gaming as a concept being poorly bounded. It overlaps with a multitude of competing, but not exclusive terms, including reward hacking, goal misgeneralization, reward misspecification, and negative side effects, among other categories. These terms vary between research traditions and papers often define them quite differently (see e.g. Skalse et al. 2022; Raji & Dobbe 2024). For instance, Krakovna’s original taxonomy used cause-based categories (reward shaping, reward misspecification, learned reward models, simulator bugs, reward tampering), but these do not map cleanly onto other widely used frameworks. For instance, Amodei et al. 2016 organized the space around five concrete problems: avoiding side effects, avoiding reward hacking, scalable supervision, safe exploration, and robustness to distributional shift. Langosco et al. 2022 introduced goal misgeneralization as a distinct concept. These taxonomies overlap and conflict in ways that make it difficult to say with confidence whether a given incident is “specification gaming” or “goal misgeneralization” or something else. Definitional distinctions aside, we saw a different yet adjacent issue.

But is it spooky?

Many of the papers showed AI systems performing concerning, unaligned or deceptive behaviours of a similar kind. However in many cases, the behaviour occurred only after being put into highly artificial situations built entirely around giving a chance to display that behaviour. There are, of course, many good reasons why this would be the case. For instance, researchers may have deliberately set up experimental conditions to elicit the misbehavior through sandbox experiments, red-teaming exercises, or intentional reward misspecification designed to study what happens and how to perhaps mitigate the behaviour (see e.g. Greenblatt et al. 2024).

There is, we think, a difference in these type of incidents compared to the more notable examples on the original Krakovna sheet that involved researchers being blindsided by a behaviour which was completely unexpected. Either, because the researchers made assumptions when providing instructions, or because the AI found a workaround that would have been undetectable to a human (discovering bugs in a game which allowed it to break through the game’s limitations and achieve extremely high scores, for example). By contrast, a lot of the newer cases involved experiments designed with the expectation that a model would show a particular undesirable behaviour under particular conditions, and the experimenters wanted to catch it in the act. We do not want to connote or suggest that this is in some way dishonest behaviour. In the vast majority of these cases, this is stated from the outset. However, we see a big difference in a system that games with an explicit instruction or curated environment and those that do not.

We started calling this distinction “spookiness” versus “entrapment” internally. Loosely, spookiness occurs when the behaviour emerged on its own without anyone designing the situation to explicitly produce it or in a way that makes the behaviour somewhat inevitable. Entrapment, on the other hand, means researchers set up conditions to either test for or provoke the behaviour, or set the experiment up in a way that provoked the model into the deceptive behaviour. “True spookiness” wouldn’t necessarily be limited to a model carrying out a hostile or harmful action on its own initiative, as it aims to capture any behaviour which illustrates how alien the unstated activity of a model can be, and hints at the number of unknown unknowns we have to deal with. The entrapment label, is in part inspired from a paper that compares AI models to “a criminal under investigation and speculates on the morality of using certain methods to motivate examples of incriminating behaviours (Clymer et al. 2024).

We think this distinction matters as the two categories have different implications. A model that independently finds an exploit nobody anticipated raises different questions than a model that does exactly what a carefully designed test was checking for. Both are worth documenting. But conflating them may make the field’s evidence base look either more alarming or less informative than it is, depending on which direction the conflation goes. So we settled on searching, in the spirit of the Krakovna list, to find true spookiness. Anything which makes us think “if it’s doing that weird thing while appearing to do normal human things, what other weird things is it doing that we don’t know about?” or “How deep is the weirdness, and how superficial is the human-ness?”.

As the one of the purposes (or at least benefits) of having a collection of such behaviours is to inform researchers of open problems and unsolved dangers of gaming models, we wanted to centre our classifications around the ability to attribute an obvious or foreseeable cause to the behaviour.

Below we document how we (with some major aid from a somewhat corrigible Claude) built the updated repository. Claude Opus 3.5+ was the main driver on much of the work with intermittent help from Sonnet 3.5+.

The sections below trace the build in three parts: how candidate papers were collected, how the classification framework took shape across iterations of decision tree, rubric, and manual incident extraction, and what the resulting dataset looks like.

Data Collection

The initial data came from a number of places:

Each of these (and a few dead ends) were made as independent collections built by a different team member using a different approach and schema. Naturally, we ran into a few issues.

First, there was a significant amount of disagreement on how to organise and include papers that described many distinct incidents of potential specification gaming (PSG) and whether to create an entry for each one, rather than treating each paper as a distinct unit. Similarly, each collection used different column structures, different category labels, and different levels of detail. The first used our entrapment/spookiness binary. The second used behavioral categories. The third used the terminology from whatever papers it drew from. Merging them required developing a shared framework that could accommodate what each collector was actually tracking, which turned out to be the conceptual core of the project. We also decided, that given the wide spread, a per incident approach would best reflect the goals of the project.

Second, we needed to approach the search for new incidents in a more systematic fashion. Through an adversarial multi-agent process, we built a search term list designed to produce broad coverage while keeping the final list manageable. Three scout agents independently compiled keywords from different source types. One drew from academic literature, producing approximately 250 candidate terms from papers. A second drew from the project’s existing data, extracting approximately 80 terms from behavioural descriptions across 327 entries. A third compiled approximately 80 terms from community and informal sources (LessWrong, Alignment Forum, EA Forum, Substack, Reddit, and system cards from AI labs). This produced over 400 terms.

To triage, a critic agent reviewed all three scout outputs and identified 10 categories of gaps, including prompt injection, hallucination, data poisoning, and multi-agent risk, all areas the scouts had collectively overlooked. A verifier agent tested terms via web search across Google Scholar, arXiv, etc., checking whether each term actually returned papers documenting PSG papers, understood in a very wide sense, rather than noise. An auditor agent spot-checked the verifier’s work while a specificity team of agents rated each term as NARROW, MEDIUM, or BROAD based on how much of its search results would be relevant versus noise, using majority vote to assign labels. Then, inclusion and exclusion advocates argued for and against borderline terms. This candidate list was then vetted by the team. The final list is 50 search terms.

Using this list, an arXiv search for papers from 2023-2026 followed a structured pipeline. Keyword searches across the 50 terms returned candidate papers, which were manually screened on title and abstract for relevance. A Semantic Scholar expansion using citation graphs and related-paper features supplemented the keyword search to catch papers that the keyword list missed. We also conducted a search of five databases of AI risks and incidents: AIID with 1,242 entries, AIAAIC with 2,188, MIT AI Risk Repository with 2,574, AVID with 95, and Awesome Agent Failures (a GitHub collection) with 18. Of the 6,117 entries screened in total, 96 were kept as PSG after verification.

A sample of the incidents were given to the team and agent teams to categorize (see next section below) along a decision tree to land on 5 terminal categories aimed at capturing the spookiness vs entrapment distinction. Disagreements between human and agent screeners concentrated at particular nodes in the decision tree. The most common sources of disagreement were whether a system’s behaviour counted as strategic metric-pursuit versus a simple error, whether RLHF-driven sycophancy routes through the tree as specification gaming or as an out-of-scope training artefact, and whether certain incidents involved deliberate human action that should exclude them from the dataset. There was an obvious need for re-calibrating the classification approach.

See: keyword list.

The Classification Framework

The classification framework went through several iterations. The framework arrived at the final pipeline through two earlier attempts. A decision-tree approach proved too brittle for automation, a rubric replaced it and produced more consistent classifications, and a manual incident-extraction step was added between the two before the full classification run.

Decision Tree

Our pipeline initially used a custom decision tree that categorised incidents among the labels: Specification Gaming [LM], Specification Gaming [Other], Goal Misgeneralization [LM], Goal Misgeneralization [Other], Reward Misspecification [LM] and Reward Misspecification [Other] and Negative Side Effects. The LM/Other split distinguished language model systems from other AI systems (RL agents, recommendation algorithms, vision models) as the failure modes presented differently in the two contexts. However, the automated classification using the seven-category scheme produced too much inter-agent disagreement for insufficient payoff. Three independent agents classifying the same incident would frequently disagree about whether something was specification gaming versus goal misgeneralization versus reward misspecification. We also reasoned that the categories did not capture the spookiness/entrapment criteria enough. This went through several iterations until we landed on the below, aiming to reduce inter-coder disagreement. If two coders follow the tree and reach different terminal nodes, the disagreement can at least be localised to a specific node where their judgements diverged, which makes reconciliation tractable.

We attempted to validate the decision tree approach on a small 20-paper subset before applying it to the full dataset. The 20 papers were selected to cover all categories. Three independent LLM agents classified each incident by walking through the decision tree, producing structured JSON output with their answer and reasoning at each node. For every incident, we had three complete tree traversals with documented reasoning, which made it possible to identify exactly where agents agreed and where they diverged. Separately, two independent LLM verifiers checked the accuracy of the incident descriptions themselves, not the classifications but whether the factual summaries of what happened matched what the papers actually said. An LLM adjudicator resolved disagreements between the two verifiers.

The correction rate (nearly one correction per incident) suggested that accurately summarising what happened in an AI incident paper is harder for the agents than it might appear. Descriptions that seem reasonable on first reading often turn out to overstate, understate, or conflate distinct behaviours when checked against the source text, which made decision nodes highly susceptible to framing inaccuracies. A classification approach that relied on agents making multiple decisions to classify was viewed as untenable. However well intended this strategy was, it was insufficient for providing reliable results, between both humans and agent teams, though the humans were a tad more reliable on inter-coder agreement. After some research aimed at diagnosing the LLM side of this issue, we replaced the decision tree with a rubric that the literature suggested (see e.g. Hashemi et al. 2024; Liu 2026) might yield more reliable classification results.

CheatSheet — specification gaming classification decision tree (v5) HOW AN ENTRY IS CLASSIFIED The decision tree TERMINALS USG — unattributed ASG — attributed OOS — out of scope 01 TRIAGE — is it in scope? 02 CLASSIFICATION — which terminal? OOS OOS-BC OOS-BI STEP 1 STEP 2 STEP 2A STEP 2B STEP 2C OOS-NSE OOS STEP 3 ASG USG YES NO YES NO NO YES YES NO YES NO NO YES NO YES NO YES YES NO Does this involve an AI/ML system with a learning or optimization component? Does the model behave in some undesirable way from the perspective of its developers? Does the model score well on its training metric while doing something its trainers did not explicitly intend? Is the system a language model or a language model-based agent? Does the model unambiguously violate the literal content of its instruction? Was the way in which it followed the literal content unexpected or ex-ante unpredictable? Is there evidence the model behaves differently depending on whether it is being monitored? Does the system display unexpected behaviour that serves some discernible metric or objective? Was the system's environment deliberately configured to enable or encourage the observed behaviour? Out of scope Benign competence Benign incompetence Negative side effects Out of scope Attributed SG Unattributed SG

The Rubric

The rubric distinguishes six categories:

A similar pipeline to the decision tree was employed to verify the rubric. Though here the initial team was made of five independent agents that each read the 20 paper subset and classified every incident they found using the rubric. Coders worked blind, with no access to the original classifications or to each other’s outputs. Five agent coders identified 52 distinct incidents across the 20 papers. Of these, 63% achieved unanimous agreement and 97.5% reached at least 3-of-5 agreement. Only one incident produced fewer than 3-of-5 agreement. Mean pairwise agreement was 81.3%.

Though performance was much better using the rubric, a major hurdle remained. Namely, disagreement by coders on what counted as “an incident” within a paper, which tended to act as a major source of disagreement. We think, that much of the disagreement actually came from an agent coding a different incident from the same paper or being contaminated by context from another incident in the paper. We couldn’t verify this but our later approach seems to bear this out after manually extracting the incidents from the papers.

See: final rubric.

Incident Extraction

Here, we took a mixed approach, divvying up the papers to manually extract how many incidents occurred in each paper. Splitting followed a fixed and basic rule: different strategies toward the same objective count as different incidents, while the same behaviour across multiple conditions or models counts as one incident with the conditions noted. We supplemented the manual approach with a similar 3 extractor/1 critic/1 verifier approach as above to test against the manual extractions. We note that the automated approach did fairly well here in straightforward cases but was a bit all over when it came to borderline decisions.

Incident-level decisions hinged on whether the behaviour was a candidate incident at all; and if so, did it constitute one or several? Inclusion turned on a wide-sense test that asked whether a behaviour could plausibly be put to the classification rubric, taking the widest reasonable reading, as the classification itself should be able to throw out as OOS any incidents that clearly were not SG candidates. In this way it was a pseudo double filter.

Typically incidents were not considered if they involved for instance, capability degradations with no strategic component, or behaviours that exactly matched what the model was trained to do. Cases of pure data-induced generalisation without an emergent or unintended means were treated as borderline and passed on to the classification round to be decided on the specific evidence in each paper.

During the incident extractions model families reported in the paper were attributed to that specific incident. Granularity was fixed at the model-family level (e.g., GPT-4, Llama-3, Claude-3.5), with normalization rules to strip size, date, instruction-tuning, and quantization suffixes.

See: coding instructions and incident extraction prompts for human and LLM extractors.

Classification Run

Once the rubric was settled, and the incidents extracted, the full corpus of candidate incidents was run through a multi-stage classification pipeline similar in set up to prior runs.

Each incident was sent to three independent LLM screeners. Each screener read the rubric and the row metadata, then returned a category (ASG, USG, or OOS), a confidence flag, and a free-form rationale naming the rubric criteria it had relied on. Screeners worked blind to each other. We note here that the classification was focused on the three major categories in one pass and a second pass to sub-categorise the OOS incidents. We made this choice to reduce the number of categories each screener had to choose between, since earlier runs had shown inter-coder agreement falling as that number grew, with the OOS subtype recoverable downstream from each screener’s named failing criterion.

Where all three screeners agreed, the incident was treated as settled. This covered roughly two-thirds of the corpus and was borne out (mostly) in earlier runs as reliable when unanimous. The remaining incidents, those with any disagreement or with unanimous low-confidence votes, were escalated to the critic and verifier stage. A verifier read the rubric and the row metadata, but not the screener outputs, and issued its own independent classification. Combining the verifier’s vote with the three screeners produced a four-vote pattern, and the row’s disposition followed from that pattern. A unanimous four-way agreement closed the row. A three-to-one split where the verifier sided with the majority screeners was treated as a proposed classification for downstream review. Anything more contested was held open for human adjudication. For every row held open, a sixth agent produced a structured review document. This agent presented all five prior positions (three screeners, critic, verifier) side by side, quoting each verbatim. The human adjudicator then read the document and selected the final category against the target paper. A finalization pass closed out any open or marked rows and audited a small number of conventions that had drifted during the run.

Final state

The presentation-ready dataset contains 696 incidents drawn from 250 source papers, with each incident carrying a category, a model attribution, and a link to the source paper. The category breakdown is roughly 62% attributed specification gaming, 18% unattributed specification gaming, and the remainder distributed across the four OOS subtypes.

References

Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete Problems in AI Safety. arXiv:1606.06565. https://arxiv.org/abs/1606.06565

Clymer, J., Juang, C., & Field, S. (2024). Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals. arXiv:2405.05466. https://arxiv.org/abs/2405.05466

Epoch AI. (2026). Data on AI Models. https://epoch.ai/data/ai-models

Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models. arXiv:2412.14093. https://arxiv.org/abs/2412.14093

Hashemi, H., Eisner, J., Rosset, C., Van Durme, B., & Kedzie, C. (2024). LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13806–13834. https://aclanthology.org/2024.acl-long.745/

Krakovna, V. (2018a). Specification gaming examples in AI — master list [Google Sheet, maintained 2018–present]. https://docs.google.com/spreadsheets/d/e/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg/pubhtml

Krakovna, V. (2018b). Specification gaming examples in AI. https://vkrakovna.wordpress.com/2018/04/02/specification-gaming-examples-in-ai/

Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., & Legg, S. (2020). Specification gaming: the flip side of AI ingenuity. DeepMind Blog, 21 April 2020. https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/

Langosco, L., Koch, J., Sharkey, L., Pfau, J., & Krueger, D. (2022). Goal Misgeneralization in Deep Reinforcement Learning. Proceedings of the 39th International Conference on Machine Learning, PMLR 162:12004–12019. https://proceedings.mlr.press/v162/langosco22a.html

Liu, J. (2026). Rubric-Conditioned Large Language Model Labeling: Agreement, Uncertainty, and Label Consistency in Subjective Text Annotation. Computers in Human Behavior. https://doi.org/10.1016/j.chb.2026.108988

Raji, I. D., & Dobbe, R. (2024). Concrete Problems in AI Safety, Revisited. arXiv:2401.10899. https://arxiv.org/abs/2401.10899

Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). Defining and Characterizing Reward Gaming. Advances in Neural Information Processing Systems 35 (NeurIPS 2022). https://proceedings.neurips.cc/paper_files/paper/2022/hash/3d719fee332caa23d5038b8a90e81796-Abstract-Conference.html

Vectara. (n.d.). Awesome Agent Failures [GitHub repository]. https://github.com/vectara/awesome-agent-failures