← Back to Blog
For: Software Engineers, AI Engineers, ML Engineers, Platform Engineers

Taste and Judgment: Skills Engineers Need in the AI Era

What to practice when code is cheap and checking it isn't - and why the filters you hand to agents fossilize the day you write them.

#engineering-judgment#developer-skills#evaluation#llm-as-judge#human-in-the-loop#agent-safety

In April 2025, OpenAI shipped a GPT-4o update that its own offline evaluations liked. A/B tests looked positive too. Some expert testers said the model "felt" slightly off. The numbers won, the update went out, and within days OpenAI was writing a system-prompt patch late on a Sunday night and rolling the whole model back on Monday. OpenAI's postmortem names the cause in one sentence: "our offline evaluations - especially those testing behavior - generally looked good." Here is how to develop taste and judgment as a software engineer, the two skills that catch what passing checks miss.

Every check that existed passed. Only one check failed, and it was the one nobody had written down: a few people's sense that something was wrong. Simon Willison's summary was blunt: "they let their automated evals and A/B tests overrule those vibe checks."

In October 2026, Addy Osmani gave that sense a name, and gave its partner a name too. "You need taste to know what is good," he wrote, "and you need judgement to understand the price of quality." Taste is the trained feeling for what makes a piece of work good, which fires before you can explain it. Judgment is knowing what quality costs in time and risk, and which option to ship given that cost. Learning to run both filters is now a larger share of the job, because drafting has become cheap.

His framing is right. This article extends it in two directions he did not cover, and takes a position on each.

First: AI made drafting cheap, and checking a draft did not get cheaper at the same rate. Fluent drafts get past taste first, because they look finished. So the common failure I see is neither of Addy's two (great taste with poor judgment, or the reverse). It is a team that runs neither filter and ships the first draft that seemed fine.

Second: the obvious way to make checking cheaper is to automate it, and automating a check means writing a filter down. When you hand either filter to an agent, its shape is fixed on the day you write it, while the reasons that shaped it keep changing. I call these Fossil Filters: the form survives after the thing that gave it that form is gone. Judgment fossilizes into numbers, budgets and approval gates, which are easy to write and go stale when the world around them moves. Taste fossilizes into evals and LLM judges, which are hard to write and drift away from the taste they were meant to capture. Both fail quietly, but the evidence that each one is out of date lives in a different place, so they need different maintenance. Most teams give them none.

From here on, the article is practical: which taste skills to train, which judgment skills to train, separate notes for software engineers and AI engineers, drills that tell you whether you got better, and how to keep the filters you give to agents from turning into fossils.

Taste and judgment: two filters for AI-generated drafts

Addy describes the two skills as a funnel. A model produces many options quickly, and "most of those drafts seem OK." Taste runs first: "It can disqualify 90% of drafts very quickly. But it's not enough to make a decision between the few remaining options." Judgment runs second and picks one, by pricing what each option costs. His two failure cases follow from that: "People with great taste but poor judgement make beautiful work too late," and people with judgment but no taste ship on time but "don't make things people love."

His funnel does not draw a third exit, so the diagram below adds one. Look at the red path on the left: it skips both filters.

d2
direction: down
classes: {
  step: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 18}}
  filter: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"; font-size: 18}}
  good: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 18}}
  bad: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"; font-size: 18}}
  muted: {style: {fill: "#95A5A6"; stroke: "#6F7E7F"; font-color: "#2C2C2A"; font-size: 18}}
}
drafts: "Many fluent drafts from a model" {class: step}
taste: "Filter 1 - taste\nWhich of these are good?" {class: filter}
rejected: "Most drafts rejected\n(Addy: about 90%)" {class: muted}
few: "A few good options" {class: step}
judgment: "Filter 2 - judgment\nWhat does each one cost?" {class: filter}
ship: "One option shipped on time" {class: good}
bypass: "Neither filter runs:\nfirst acceptable draft ships" {class: bad}
drafts -> taste
taste -> rejected
taste -> few
few -> judgment
judgment -> ship
drafts -> bypass: "looks finished"

Raj Nandan Sharma described this third exit in April: "Before AI, mediocre work usually meant someone ran out of time, money, or skill. Now it usually means someone stopped at the first acceptable draft." I think it is the most common of the three, and the evidence points that way even if nobody has measured it at team level.

This exit needs a precise definition, because shipping a good-enough draft is often the right call. If someone read it for quality and decided the change is cheap to undo, that is judgment pricing a two-way door, and it is fine. The neither-filter exit is shipping with no decision at all: nobody read it for quality, and nobody priced the undo.

In Sonar's January 2026 developer survey (a vendor survey, of more than 1,100 developers), 96% said they do not fully trust that AI-generated code is functionally correct. Only 48% said they always check it before committing. 38% said reviewing AI code takes more effort than reviewing a colleague's code, and 53% named "code that looks correct but isn't reliable" as a problem they had seen. Distrust is close to universal. Checking is not.

An older and better-controlled result comes from Perry and colleagues at Stanford (CCS '23). Of 47 participants, the ones with an AI assistant wrote less secure code on four of five tasks, and were more likely to believe their code was secure. Participants used codex-davinci-002 and the sample was small, so read it for direction, not size. That direction is the problem: with fluent help, people produced worse work and trusted it more, in the same session. That is taste being bypassed, and it is hard to feel from the inside. The same study found the participants who trusted the assistant less, and worked harder on their prompts, wrote code with fewer vulnerabilities.

The effect has a boundary. Vasconcelos and colleagues ran five studies with 731 people and found that over-reliance tracks the cost of checking: when verifying the AI's answer is cheap, people verify it. Under volume and time pressure, it is not cheap, and that is the normal condition of a team shipping AI-assisted code.

How to develop engineering taste when every draft looks fine

Addy says taste "comes from looking at a lot of things, and making a lot of things, and remembering the good bits." The AI era changes what you need to look at. Software engineers need more reps on design. AI engineers need more reps on outputs, and on the judges that grade them.

For software engineers: the wrong abstraction behind green CI

Here is the case every reviewer has met. A pull request (PR) passes lint, passes type checks, and holds 90% coverage. Continuous integration (CI) is green, and the code works. Yet the abstraction is wrong: pricing logic is spread across every caller instead of living in one module, so the first change request will touch twelve files.

No check in the pipeline can catch this, because at the time of writing the clean version and the hasty version pass the same tests. The Taste-Bench authors (Pan and colleagues, a September 2026 preprint) put it exactly: "a clean implementation passes the same tests as a hasty one at the time of writing, and its advantage appears only when every later change becomes easier."

d2
grid-rows: 1
grid-gap: 40
classes: {
  ok: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 17}}
  bad: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"; font-size: 17}}
  same: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 17}}
}
clean: "Clean abstraction" {
  label.near: top-center
  grid-columns: 1
  t: "Today: CI green" {class: same}
  c1: "Change 1: one module edited" {class: ok}
  c2: "Change 2: one module edited" {class: ok}
}
hasty: "Hasty abstraction" {
  label.near: top-center
  grid-columns: 1
  t: "Today: CI green" {class: same}
  c1: "Change 1: twelve callers edited" {class: bad}
  c2: "Change 2: twelve callers, one missed" {class: bad}
}

Both columns share an identical top row, which is why automated review cannot rank them. Any difference exists only in rows two and three, in the future. Taste is the skill of seeing those rows today.

You train it by reading more code than you write, and by checking your reads against what happened. Pick ten year-old PRs from your own repository at random, so the set includes ones that aged well. Review each cold, before you look at its history, and predict whether it later needed a painful change. Then read the history. Your misses in both directions (the bad PR you passed and the good one you flagged) are your taste error, and you can count them.

For AI engineers: judging outputs, and judging the judges

AI engineers have the same problem one level up. They read model outputs in bulk, and they increasingly hand that reading to an LLM judge. The judge has its own taste, and its taste has known defects.

Zheng and colleagues measured several in the MT-Bench paper (NeurIPS 2023). When they swapped the order of two answers, Claude-v1 gave the same verdict only 23.8% of the time and picked the first answer in 75% of cases. GPT-4 was consistent 65% of the time. Verbosity bias was model-dependent: a padded "repetitive list" answer fooled Claude-v1 and GPT-3.5 in 91.3% of cases and GPT-4 in 8.7%. Separately, Panickssery, Bowman and Feng (2024) found that a judge's preference for its own outputs correlates linearly with how well it can recognize them.

Below is the cheapest fix in that paper, the position-swap protocol, and it shows what judging the judge looks like in practice.

d2
direction: down
classes: {
  run: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 17}}
  q: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"; font-size: 17}}
  win: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 17}}
  tie: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"; font-size: 17}}
}
r1: "Run 1: judge sees A, then B" {class: run}
r2: "Run 2: judge sees B, then A" {class: run}
same: "Same winner in both runs?" {class: q}
w: "Count it as a win" {class: win}
t: "Count it as a tie and\nsend it to a human" {class: tie}
r1 -> same
r2 -> same
same -> w: "yes"
same -> t: "no"

The yellow branch matters more than the green one. Every disagreement between the two runs is a case where the judge's verdict was about order, not quality, and those are the cases a human should read. Zheng found few-shot prompting raised GPT-4's consistency from 65% to 77.5%, and also warned that "high consistency may not imply high accuracy." A consistent judge can be consistently wrong.

AI engineer skills: eval design is where taste meets judgment

Writing an eval looks like one task. It is two. Choosing which failures to look for is taste: you read real outputs, for an agent the whole run and not only the final answer, and notice what is wrong with them. Choosing the pass threshold, and deciding which failures are worth an eval at all, is judgment: you are pricing how much a miss costs against how much the eval costs to build and run.

d2
direction: down
classes: {
  t: {style: {fill: "#C2185B"; stroke: "#8E1244"; font-color: "#FFFFFF"; font-size: 17}}
  j: {style: {fill: "#955D37"; stroke: "#6E4428"; font-color: "#FFFFFF"; font-size: 17}}
}
read: "1. Read 20-50 real failures\n(taste)" {class: t}
name: "2. Name the failure categories\n(taste)" {class: t}
worth: "3. Decide which categories are\nworth an eval (judgment)" {class: j}
thr: "4. Set each pass threshold\n(judgment)" {class: j}
read -> name
name -> worth
worth -> thr

Pink steps need someone who can see what is wrong, and brown steps need someone who knows what a miss costs. On small teams that is one person, and the useful habit is to notice which hat you are wearing. Anthropic's eval guidance (January 2026) starts at the pink end: "20-50 simple tasks drawn from real failures is a great start." It also sets the bar for a well-defined task: "two domain experts would independently reach the same pass/fail verdict." If two experts would not agree, the task is measuring taste you have not written down yet.

When two reviewers' taste disagrees

Taste is personal, but code review is not. Two senior reviewers will sometimes disagree about the same PR, and both will be right about something. Their disagreement is useful information. Leaving it unresolved is expensive. Jeff Bezos described the default outcome in his 2016 shareholder letter: "Without escalation, the default dispute resolution mechanism for this scenario is exhaustion."

There are two well-known resolutions, and they fit different decisions. Bezos's "disagree and commit" fits a reversible choice: one person decides, the other commits fully, and the team learns from the result. Hamel Husain's "principal domain expert" model for evals fits a standard that must stay consistent: one named person owns the definition of "good", and everyone calibrates to them.

d2
grid-rows: 3
grid-columns: 2
grid-gap: 26
classes: {
  q: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"; font-size: 17}}
  a: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 17}}
}
q1: "1. Is the choice easy to reverse?" {class: q}
a1: "Yes: one reviewer decides,\nthe other disagrees and commits" {class: a}
q2: "2. Is it a lasting standard?" {class: q}
a2: "Yes: the named domain expert\ndecides, and everyone calibrates" {class: a}
q3: "3. Has this dispute come up before?" {class: q}
a3: "Yes: write the rule down\n(style guide, lint rule or eval)" {class: a}
q1 -> q2
q2 -> q3

The third row is how taste becomes a team asset. A disagreement that repeats is a rule nobody has written down. Writing it down turns it into a filter, and the second half of this article is about what happens to filters once they are written.

How to build engineering judgment in the AI era

Judgment is pricing. Addy's version: "The most perfect solution may not be finished by the time that you need it." The skills that make up judgment are mostly about estimating three things before you choose: how hard the choice is to undo, what waiting costs, and how much an agent can be trusted to do alone.

Reversibility and the undo test

Bezos's 2015 letter split decisions into two types. "Some decisions are consequential and irreversible or nearly irreversible - one-way doors - and these decisions must be made methodically, carefully, slowly." Most are two-way doors: "You can reopen the door and go back through." His warning was about the opposite error. Large organizations apply the slow Type 1 process to Type 2 decisions, and the result is "slowness, unthoughtful risk aversion, failure to experiment sufficiently, and consequently diminished invention." That is Addy's "beautiful work, too late" at company scale.

Sachin Malhotra, on Anthropic's CI team, turned the same idea into a rule for agents in his AI Engineer World's Fair 2026 talk, "Give the Agent a Budget, Not a Token." His undo test asks two questions: can the agent put this back by itself, and is the worst case acceptable if it gets it wrong? Two yeses mean the agent acts and the action is logged. If either answer is no, the action needs "a second key that the agent never holds."

d2
grid-rows: 3
grid-columns: 2
grid-gap: 26
classes: {
  q: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"; font-size: 17}}
  ok: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 17}}
  stop: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"; font-size: 17}}
}
q1: "1. Can the agent undo it by itself?" {class: q}
a1: "No: a second key held by a\nhuman the agent never controls" {class: stop}
q2: "2. If yes: is the worst case acceptable?" {class: q}
a2: "No: a second key, as above" {class: stop}
q3: "3. Both answers yes" {class: q}
a3: "Agent acts within a rate limit,\nand every action is logged" {class: ok}
q1 -> q2: "yes"
q2 -> q3: "yes"

He adds a refinement worth stealing: give agents the verbs that fail loudly and keep the ones that fail quietly for humans. Un-skipping a test is safe to delegate, because if it was wrong CI goes red. Skipping a test is not, because a bug can ship behind a green check.

To train this skill, label the next twenty decisions you make or review as one-way or two-way before you make them, and write one line on what undoing each would cost. After a month, look at which labels were wrong. Most engineers find they treat too many decisions as one-way, which is Bezos's error.

Influence the decision, even when you don't own it

A mid-level engineer rarely owns the migration plan or the release date. That does not excuse them from judgment. It changes what judgment looks like: you bring the price, and the owner decides.

d2
direction: down
classes: {
  you: {style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"; font-size: 17}}
  own: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 17}}
  log: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"; font-size: 17}}
}
opts: "You: two or three options, each with\nits undo cost and a probability it slips" {class: you}
dec: "Owner decides" {class: own}
rec: "Decision log entry: the choice, the\nprediction, the date to check it" {class: log}
rev: "Check at 30 and 90 days:\nwas the prediction right?" {class: log}
opts -> dec
dec -> rec
rec -> rev
rev -> opts: "your next estimate improves"

The loop back from the review is the part that trains you. Michael Nygard's original architecture decision record (ADR) format from 2011 has half of it: decisions are kept, and marked superseded rather than deleted, because "the time to change old decisions will be clear from changes in the project's context." An ADR has no prediction field, so add one. A decision record that is never re-read is documentation. One that is re-read against its prediction is practice.

Pricing quality when the quality number is noisy

AI engineers price a three-way trade: cost per call, latency, and output quality. Cost per call and latency come from a dashboard and are precise. Quality comes from an eval, and it is noisy. At a pass rate around 80% on 100 examples, one model's score has a standard error of about 4 points, so a two-point gap between two models is inside the noise. Run both models on the same examples and compute a confidence interval on the paired difference. Resolving a gap of two points usually takes many hundreds of paired examples, not dozens.

d2
grid-rows: 2
grid-columns: 2
grid-gap: 26
classes: {
  q: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"; font-size: 17}}
  a: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 17}}
}
q1: "1. Does the paired difference's\nconfidence interval exclude zero?" {class: q}
a1: "No: choose on cost, latency\nand how easy it is to undo" {class: a}
q2: "2. If it clears: is the gap worth\nthe extra cost per call?" {class: q}
a2: "Price it in money per month\nand decide on that number" {class: a}
q1 -> q2: "yes"

The left column is where judgment saves money. When the quality number cannot separate two options, the honest move is to stop treating it as the deciding factor and pick on the precise numbers and on reversibility. Teams that keep chasing a two-point eval delta are spending judgment on noise.

Three public AI agent failures, filter by filter

Three public incidents show three different ways the filters break. Each one is short. Read them as a set.

d2
grid-rows: 3
grid-columns: 1
grid-gap: 20
classes: {
  bad: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"; font-size: 16}}
  warn: {style: {fill: "#FFA07A"; stroke: "#D9785A"; font-color: "#2C2C2A"; font-size: 16}}
  proxy: {style: {fill: "#C2185B"; stroke: "#8E1244"; font-color: "#FFFFFF"; font-size: 16}}
}
gpt: "1. GPT-4o, April 2025 - taste ran and was overruled\nevals and A/B tests passed; testers said it felt off" {class: proxy}
cleanup: "2. Anthropic cleanup agent - judgment never priced the authority\nabout 200 workloads deleted in 90 seconds" {class: warn}
replit: "3. Replit, July 2025 - neither filter ran\ndatabase deleted during a code freeze" {class: bad}

1. GPT-4o: taste ran, and a written proxy overruled it. OpenAI's postmortem says the update added "an additional reward signal based on user feedback - thumbs-up and thumbs-down data," and that "user feedback in particular can sometimes favor more agreeable responses." The expert testers' taste did its job and flagged the problem. Their signal lost to numbers. This is the gaming failure covered later, in its training-time form: a quality proxy was written down, the model was optimized against it, and the proxy won the argument. OpenAI's stated lesson was to "treat model behavior issues as launch-blocking."

2. The Anthropic cleanup agent: nobody priced the authority. In Malhotra's account, a stage in the agent's workload selector evaluated to nothing, so the selector matched every workload, and about 200 workloads belonging to roughly 20 engineers disappeared in 90 seconds. Some were long-running training runs that may not have been checkpointed. The agent "genuinely thought it was tidying up after itself." His conclusion: "the failure wasn't the model itself. The failure was that I was giving the agent unbounded amount of power." The agent's token allowed the deletes, and no budget limited how many. (This comes from a conference talk, not an audited postmortem, so treat the numbers as his recollection.)

3. Replit and SaaStr: neither filter ran. In July 2025 an AI coding agent on Replit deleted a database during an explicit code freeze. The human then believed the agent's claim that the deletion could not be rolled back, which was false: Replit's rollback covered it. Nobody had priced what the agent should be allowed to touch, and nobody checked its confident claim. Replit responded with automatic separation of development and production databases and a planning-only mode.

Taste without judgment has no famous incident, because its cost is invisible: the better product that shipped a quarter late, or never shipped. Bezos's "diminished invention" is the closest public description of it.

What matters less, and what to keep writing by hand

Producing boilerplate quickly is worth less every month. Writing code by hand is not. Addy's own account of where taste comes from includes "making a lot of things," and you cannot remember the good bits of things you never made. Keep writing the hard parts yourself: the core abstraction of a new service, the eval harness, the one function everything else depends on. Delegate the rest.

d2
grid-rows: 1
grid-gap: 40
classes: {
  keep: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 17}}
  give: {style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"; font-size: 17}}
}
mine: "Write by hand" {
  label.near: top-center
  grid-columns: 1
  a: "Core abstractions" {class: keep}
  b: "Eval harnesses" {class: keep}
  c: "Code everything depends on" {class: keep}
}
agent: "Delegate" {
  label.near: top-center
  grid-columns: 1
  a: "Boilerplate and glue code" {class: give}
  b: "Tests for code you designed" {class: give}
  c: "Migrations you have specified" {class: give}
}

The left column is where taste is built, so it stays with you even when an agent could write it faster.

Juniors have the hardest version of this problem. Traditionally, juniors built taste over years of writing mediocre code and having seniors explain what was wrong with it. If agents write the first draft, juniors get fewer of those reps. Twisha Shah-Brandenburg calls this the Two-Clock Problem: the performance clock that AI improves closes in 12 to 18 months, while the judgment clock "closes in five to seven years" and "has no dashboard." Her essay is about law, consulting and medicine, but the shape fits software exactly. If you lead juniors, protect the left column for them on purpose. For the broader career question, see Beyond Copy-Paste: Staying Relevant in the Age of AI Code Assistants.

Fossil Filters: why encoded taste and judgment go stale

Agents run thousands of small decisions with no human at the funnel. So both filters have to be written down in advance: taste as evals, rules and judges; judgment as budgets, permission scopes, stop conditions and approval gates. Most advice ends there, and it leaves out what happens to a filter after someone writes it.

A written filter keeps the shape it had on the day it was written. The two kinds go wrong in different ways.

d2
grid-rows: 3
grid-columns: 4
grid-gap: 18
classes: {
  head: {style: {fill: "#95A5A6"; stroke: "#6F7E7F"; font-color: "#2C2C2A"; font-size: 16; bold: true}}
  j: {style: {fill: "#955D37"; stroke: "#6E4428"; font-color: "#FFFFFF"; font-size: 16}}
  t: {style: {fill: "#C2185B"; stroke: "#8E1244"; font-color: "#FFFFFF"; font-size: 16}}
}
h1: "Filter" {class: head}
h2: "How it goes wrong" {class: head}
h3: "Where staleness shows up" {class: head}
h4: "Maintenance habit" {class: head}
j1: "Judgment as budgets,\nscopes, approval gates" {class: j}
j2: "Goes stale: easy to\nwrite, but the world\nmoves on" {class: j}
j3: "Metrics outside the repo:\njobs, spend, headcount" {class: j}
j4: "Write the assumption\nas a check against\na live metric" {class: j}
t1: "Taste as evals,\nrules, LLM judges" {class: t}
t2: "Drifts or gets gamed:\nhard to write, and\nyour taste moves" {class: t}
t3: "Only in fresh\nhuman labels" {class: t}
t4: "Re-label fresh outputs;\ncompare the judge\nto humans" {class: t}

The third column carries the claim. Both rows fail quietly; what differs is where you can see it happening. Judgment is the easy filter to write and the hard one to keep true: a budget is one number, and the world it was priced against keeps changing, but much of that world shows up in metrics a machine can read. Taste is hard to write in the first place, and the only evidence that it has drifted is a human reading fresh outputs. An LLM judge holds both. Its criteria are fossilized taste, and its pass threshold is fossilized judgment, so it needs both maintenance habits.

The wrong way: a filter with no assumptions

This is the shape most agent policy files take:

yaml
# agent-policy.yaml - the usual shapecleanup_agent:  max_deletes_per_run: 20  allowed_namespaces: ["ci", "scratch"]support_judge:  model: judge-v3  pass_threshold: 0.8

Every value here was a judgment call by someone. None of them records why. Twenty deletes per run was a safe price because those namespaces held only short CI jobs, so twenty mistaken deletes cost a few reruns. Nothing in the file says so. When a new team starts running multi-day training jobs in scratch, the cap still reads 20, and twenty deletes can now destroy twenty training runs. Someone tuned the judge threshold of 0.8 against outputs from last quarter's product. Long after the reasons for both numbers are gone, the file keeps enforcing them.

The right way: every filter carries a checkable assumption

The fix is boring and it works. Each filter records what it enforces, what it assumes, how to check that assumption against something live, who owns it, and a review date:

yaml
filters:  - id: cleanup-delete-cap    kind: judgment    enforces: {file: agent-policy.yaml, key: cleanup_agent.max_deletes_per_run, value: 20}    assumes: "the agent can only delete short CI jobs; nothing long-running is in its namespaces"    assumes_check:      - {metric: long_running_jobs_in_scope, max: 0}    owner: platform-oncall    review_by: 2026-12-31  - id: support-reply-judge    kind: taste    enforces: {file: agent-policy.yaml, key: support_judge.pass_threshold, value: 0.8}    assumes: "a recent human calibration found the judge still matches human labels"    assumes_check:      - {metric: days_since_human_calibration, max: 30}      - {metric: judge_tnr_low_at_last_calibration, min: 0.85}    owner: support-ai-team    review_by: 2026-12-31

A short check in CI makes the assumptions real. It confirms the registry still matches the live policy file, tests each assumption against current metrics, and treats the review date as a backstop:

python
import jsonimport sysfrom datetime import dateimport yamldef lookup(cfg: dict, dotted: str):    for part in dotted.split("."):        cfg = cfg[part]    return cfgdef check(registry: str, metrics_path: str, today: date) -> int:    with open(registry, encoding="utf-8") as f:        filters = yaml.safe_load(f)["filters"]    with open(metrics_path, encoding="utf-8") as f:        metrics = json.load(f)    failures = 0    for flt in filters:        problems = []        enf = flt["enforces"]        with open(enf["file"], encoding="utf-8") as f:            live = lookup(yaml.safe_load(f), enf["key"])        if live != enf["value"]:            problems.append(f"{enf['key']} is {live} in {enf['file']}, registry says {enf['value']}")        for chk in flt["assumes_check"]:            observed = metrics[chk["metric"]]            if "max" in chk and observed > chk["max"]:                problems.append(f"assumption broken: {chk['metric']} = {observed} (max {chk['max']})")            if "min" in chk and observed < chk["min"]:                problems.append(f"assumption broken: {chk['metric']} = {observed} (min {chk['min']})")        if date.fromisoformat(str(flt["review_by"])) < today:            problems.append(f"review_by {flt['review_by']} has passed")        if problems:            failures += 1            print(f"FAIL {flt['id']} ({flt['kind']}), owner {flt['owner']}")            print(f"     assumes: {flt['assumes']}")            for p in problems:                print(f"     {p}")        else:            print(f"ok   {flt['id']} ({flt['kind']})")    return failuresif __name__ == "__main__":    sys.exit(1 if check(sys.argv[1], sys.argv[2], date.today()) else 0)

Run on 7 October 2026, with the metrics source reporting 40 long-running jobs in the agent's namespaces, and the judge's last calibration (12 days ago) reporting a TNR lower bound of 0.29, it prints:

text
FAIL cleanup-delete-cap (judgment), owner platform-oncall     assumes: the agent can only delete short CI jobs; nothing long-running is in its namespaces     assumption broken: long_running_jobs_in_scope = 40 (max 0)FAIL support-reply-judge (taste), owner support-ai-team     assumes: a recent human calibration found the judge still matches human labels     assumption broken: judge_tnr_low_at_last_calibration = 0.29 (min 0.85)

and exits with status 1, so the build fails. Both review dates are still in December, so a date-only check would have printed ok twice. A changed world fails the cap. The judge fails on the result of its last human calibration, not on how recently that calibration ran: a calibration that ran 12 days ago and found the judge broken is not a reason to trust it. (The 0.29 comes from the calibration script later in this article.)

Three details matter in practice. The enforces field fails the build if someone edits a registered value in the live policy and forgets the registry; it does not cover keys nobody registered, so either register every key or keep an explicit list of the ones you chose not to. A metrics source has to be live, or carry a timestamp the check rejects when it is old, or it becomes one more snapshot. And a review date turns into a snooze button the first time someone edits it to make CI green, so require a one-line note on what was re-checked whenever review_by changes. Where an input has no metric at all, such as the team's risk appetite or a customer's contract terms, that owner and that note are the only check, which is why the note is mandatory.

Encoded judgment goes stale

Judgment goes stale because its inputs live outside the repository. Deadlines move, the team's risk appetite changes, a new customer arrives with a different contract. The number in the policy file does not move with them.

d2
direction: down
classes: {
  t0: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 17}}
  t1: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"; font-size: 17}}
  t2: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"; font-size: 17}}
}
a: "Day 0: cap of 20 deletes written\n(namespaces hold only short CI jobs)" {class: t0}
b: "Month 3: new team runs 40 multi-day\ntraining jobs in scratch (outside the repo)" {class: t1}
c: "Month 4: cap still 20, still enforced,\nnow allows damage nobody priced" {class: t2}
a -> b: "file unchanged"
b -> c: "file unchanged"

Nothing in this timeline is a bug. Every edge is labeled "file unchanged," and that is the failure. The cap still limits the count, and it would have stopped the 200-workload incident at 20. What changed is the price of each delete: twenty mistaken deletes used to cost a few CI reruns, and now they can cost twenty training runs. The number stayed the same while the damage it allows grew.

Two independent sources describe the same effect. Malhotra noted in his talk that allow-lists are a prediction made in advance, and "that prediction is static and can become stale before the team has observed much real behavior." A preprint by Nakayashiki (August 2026) tested agents that inherit a memory containing a constraint that has since been withdrawn. Across sixteen models, the agents re-checked the constraint's source in about one episode in five, and acted on the stale constraint in roughly three quarters of episodes. In that study, a rule that read as settled mostly did not get re-checked.

Both sources describe staleness. Neither separates the two kinds of Fossil Filters, and the separation is where the maintenance habit comes from. A fossil filter can fail in either direction: a cap can become too tight, an eval can start rejecting outputs you now accept. Whether anyone notices depends less on the direction than on who gets hurt. A failure is loud when it hurts the people who own the filter, because their own work gets blocked and they fix it. It is quiet when it hurts someone else: the end users, the customers, the team whose training runs the cleanup agent deletes. Permissive failures usually land on other people, so most quiet fossils are permissive ones, such as a cap that still exists but no longer prices the damage, or an eval that keeps reporting green on outputs you would now reject. Both kinds of fossil filter fail quietly in this way. The difference between them is where the evidence of staleness lives.

For judgment filters, that evidence is mostly outside the repository and observable: job counts, headcount, contracts, spend. You can write the assumption as a check against a live metric, as assumes_check does above, and automate it. For taste filters, the evidence lives in human labels on fresh outputs, and you can only sample it. Some taste re-checks are event-driven too (a new judge model, a new product policy), but no metric will tell you that your own criteria have moved.

So the habit for judgment is to write the assumption next to the number, as something a machine can test. Simon Willison made a related point in October 2026 about usage-billed services: they need "default hard budget caps," and soft caps that only send a warning email "will not cut it." Hard caps are right. A hard cap with no recorded assumption is still a fossil filter.

Cloudflare's 18 Nov 2025 outage: a fossilized assumption

The cleanup cap above is an illustration. Cloudflare published a real fossilized assumption in its postmortem of 18 November 2025, which its CEO Matthew Prince called the company's "worst outage since 2019." No agent was involved, but the mechanism is the same one: a rule kept its shape while the system around it changed.

Cloudflare's Bot Management module reads a feature file that is regenerated every five minutes by a query against a ClickHouse database. The query asked for the columns of one table and did not filter by database name. In Prince's words, "there were assumptions made in the past, that the list of columns returned by a query like this would only include the 'default' database." That assumption was not written anywhere a machine could check it. It lived in the query, unstated.

d2
direction: down
classes: {
  ext: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"; font-size: 17}}
  warn: {style: {fill: "#FFA07A"; stroke: "#D9785A"; font-color: "#2C2C2A"; font-size: 17}}
  bad: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"; font-size: 17}}
  fix: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 17}}
}
perm: "11:05 - database permissions change\n(made elsewhere, for good reasons)" {class: ext}
query: "Feature query now returns duplicate rows\n(unstated assumption: only 'default')" {class: warn}
file: "Feature file more than doubles, past the\nhard limit of 200 features (normal use ~60)" {class: bad}
down: "Core proxy panics: errors across the network\nfrom 11:28; main traffic back at 14:30" {class: bad}
check: "Validator in the generator:\nreject the file before it propagates" {class: fix}
perm -> query: "1. assumption broken"
query -> file: "2. every five minutes"
file -> down: "3. propagated to every machine"
check -> file: "would stop it here"

Follow the numbered edges. The change that started it was a reasonable security improvement to how distributed queries run, made at 11:05, and it changed nothing in the Bot Management code. On its own terms the permissions change was correct, and the query had worked until then. The broken assumption sat between the two systems, which is exactly where the evidence of a stale assumption lives.

Be precise about what fossilized. It was not the 200-feature limit, which still had more than three times the headroom of normal use and did its job by rejecting a bad input, if violently. What fossilized was the unrecorded assumption underneath: the consumer's cap encoded a judgment that its input would stay small, and nothing recorded why that judgment held.

On Cloudflare's newer proxy this failed loudly, because the limit was enforced hard and the proxy panicked. Loud failures get noticed fast: the postmortem's timeline shows an automated test detecting the issue at 11:31, three minutes after the first errors. On Cloudflare's older proxy, the same bad file failed quietly. Those customers saw no errors at all. Every request simply got a bot score of zero, and customers with rules that block bots "would have seen large numbers of false positives." That is a restrictive failure, and it was still quiet, because the people blocked were those customers' visitors, not the team that owned the file. The rule holds: a failure is quiet when it hurts someone other than the filter's owner.

Cloudflare's first remediation item reads like the habit this article argues for: "Hardening ingestion of Cloudflare-generated configuration files in the same way we would for user-generated input." That is the same idea as assumes_check, moved from CI to the generator. Because the file was regenerated every five minutes at runtime, without a commit, so a CI job would never have seen the bad version. The check has to run wherever the artifact is produced: the producer asserts one row per feature name and refuses to publish a file that fails, so the fleet keeps serving the last good one. What the registry adds is a record that the assumption exists, and who owns it.

Encoded taste drifts or gets gamed

Taste fossilizes differently. The research that explains it best is Shankar and colleagues' "Who Validates the Validators?" (2024). Watching nine industry practitioners grade LLM outputs, they found that people define their criteria by grading: "the process of grading outputs helps them to define that very criteria." They called this criteria drift, and concluded it is "impossible to completely determine evaluation criteria prior to human judging."

That changes what an eval is. An eval is a snapshot of your taste on the day you wrote it, and your taste keeps moving as you see more outputs. Shankar's participants revised their criteria within a single grading session. On a team that keeps grading for months, the same process continues, and the gap opens even if the judge never changes. The month labels in the diagram are illustrative.

d2
grid-rows: 1
grid-gap: 40
classes: {
  h: {style: {fill: "#C2185B"; stroke: "#8E1244"; font-color: "#FFFFFF"; font-size: 17}}
  e: {style: {fill: "#95A5A6"; stroke: "#6F7E7F"; font-color: "#2C2C2A"; font-size: 17}}
}
human: "Human taste" {
  label.near: top-center
  grid-columns: 1
  a: "Month 0: criteria v1" {class: h}
  b: "Month 2: v2, after more grading" {class: h}
  c: "Month 5: v3, new failure seen" {class: h}
}
eval: "Encoded eval" {
  label.near: top-center
  grid-columns: 1
  a: "Month 0: criteria v1" {class: e}
  b: "Month 2: criteria v1" {class: e}
  c: "Month 5: criteria v1" {class: e}
}

Read the two columns row by row. They agree once, at the top. Every later row is a place where the eval can pass outputs that your current taste would reject, or reject ones it would now accept.

The second way encoded taste fails is gaming. Once a model is optimized against a written proxy, it finds the proxy's gaps. Marilyn Strathern's paraphrase of Goodhart's law states it: "When a measure becomes a target, it ceases to be a good measure." The GPT-4o thumbs-up signal is the training-time version of this: a reward, not an eval, but the same mechanism. Graders can also be wrong from the first day. Anthropic's eval guidance reports that Claude Opus 4.5 first scored 42% on CORE-Bench, and the score rose to 95% after fixes to the grading, not to the model. That is one more reason to calibrate a grader against humans before trusting it, and not only after it has aged.

The habit is calibration. Re-label a fresh sample of real outputs by hand, at a fixed interval and whenever the model or prompt changes, and compare the judge's verdicts to yours. Hamel Husain's advice is to report the judge's true positive rate (TPR) and true negative rate (TNR) separately, because raw agreement hides the failure that matters, and to label about 100 examples per failure mode, since "below 60 examples, the confidence intervals are often too wide." This script follows both rules at the level of pass and fail; a per-failure-mode count is stricter still. It refuses to score a sample with too few failures in it, and it compares the lower bound of each rate's 95% confidence interval to the floor, so a small lucky sample cannot pass:

python
import jsonimport mathimport sysMIN_PER_CLASS = 60  # below this, the interval is too wide to act ondef wilson_low(hits: int, n: int, z: float = 1.96) -> float:    """Lower bound of the 95% Wilson interval for a proportion."""    p = hits / n    centre = p + z * z / (2 * n)    margin = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))    return (centre - margin) / (1 + z * z / n)def main(path: str, tpr_floor: float = 0.85, tnr_floor: float = 0.85) -> int:    with open(path, encoding="utf-8") as f:        rows = json.load(f)    good = [r for r in rows if r["human"]]    bad = [r for r in rows if not r["human"]]    if min(len(good), len(bad)) < MIN_PER_CLASS:        print(f"too few labels: {len(good)} good, {len(bad)} bad - sample more failures first")        return 2    tp = sum(1 for r in good if r["judge"])    tn = sum(1 for r in bad if not r["judge"])    agreement = (tp + tn) / len(rows)    print(f"labels: {len(good)} good, {len(bad)} bad  raw agreement: {agreement:.2f}")    print(f"TPR: {tp / len(good):.2f} (95% low {wilson_low(tp, len(good)):.2f})  "          f"TNR: {tn / len(bad):.2f} (95% low {wilson_low(tn, len(bad)):.2f})")    if wilson_low(tp, len(good)) < tpr_floor or wilson_low(tn, len(bad)) < tnr_floor:        print("MISCALIBRATED: judge does not match human labels - re-label and rewrite its criteria")        return 1    print("ok: judge matches human labels on this sample")    return 0if __name__ == "__main__":    sys.exit(main(sys.argv[1]))

To show why the split matters, I built a labeled sample of 300 replies in a common shape for this failure: humans pass 240 and fail 60, and the judge passes 239 of the 240 good replies but catches only 24 of the 60 bad ones. The script prints:

text
labels: 240 good, 60 bad  raw agreement: 0.88TPR: 1.00 (95% low 0.98)  TNR: 0.40 (95% low 0.29)MISCALIBRATED: judge does not match human labels - re-label and rewrite its criteria

Eighty-eight percent agreement looks like a healthy judge. It is letting 36 of 60 bad replies through. On an imbalanced sample, raw agreement rewards a judge for passing everything. One run tells you the judge is miscalibrated today, not that it drifted. To see drift, store each calibration result and compare it with the last one. The two floors are also a judgment call: a missed bad reply usually costs more than a wrongly rejected good one, so the TNR floor often deserves to be the higher of the two.

One honest counterpoint. The Taste-Bench authors report that taste in their sense can be trained into a model by distillation, so "taste cannot be encoded" is too strong. My claim is narrower: taste resists being encoded as fixed rules and fixed evals, because the human taste those rules capture keeps moving. A trained model's taste is also a snapshot, of whatever data it was distilled from. And taste is still hard for models: on their benchmark, the best model tested chose the better direction 59.7% of the time, on two-way choices where chance is 50%.

A four-week practice plan for taste and judgment

Start by finding your lean. Answer these honestly for the last month of your work:

  • Did you merge AI-written code you had not read line by line?
  • Did a reviewer find a design problem you would have seen if you had looked?
  • Did you spend more than a day polishing something that could have shipped and been revised?
  • Did you argue for an option without saying what it would cost to undo?
  • Did you trust an eval or judge score you had not checked against your own reading?

Yes to the first two means you lean toward the neither-filter exit. Yes to the third means taste without judgment. Yes to the fourth means your judgment is not visible. Yes to the fifth means you are trusting a filter nobody has checked. Then run the plan below. Every drill has a feedback loop, because practice without feedback only trains confidence.

d2
direction: down
classes: {
  w: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 17}}
  out: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"; font-size: 17}}
}
w1: "Week 1 - taste: review PRs with a seeded flaw" {class: w}
w2: "Week 2 - taste: label 100 outputs beside a colleague" {class: w}
w3: "Week 3 - judgment: decision log with probabilities" {class: w}
w4: "Week 4 - judgment: pre-mortem on one real decision" {class: w}
out: "Artifacts you can show: catch and false-alarm\nrates, agreement scores, a started decision\nlog, a written pre-mortem" {class: out}
w1 -> w2
w2 -> w3
w3 -> w4
w4 -> out

Each week produces something a lead can read, which is the point of the yellow box at the bottom.

Week 1: seeded-flaw reviews. Pair with a colleague. Each of you prepares about ten merged PRs and plants one realistic flaw in some of them, but not all: a wrong abstraction, a missing edge case, an error swallowed in a retry loop. Swap and review without knowing which ones are clean. Score both what you caught and what you flagged that was fine. Ten PRs will not give you a reliable rate. The value is in the misses, and in the false alarms, which tell you whether your taste is sharp or only suspicious.

Week 2: label beside a colleague. AI engineers take about 100 real outputs from a system you own; software engineers can use 100 AI-written diffs. Each of you labels them pass or fail alone. Then compute agreement with Cohen's kappa, which corrects for agreement by chance, and spend an hour arguing about the disagreements. The arguments are the drill. The rules you agree on become the first draft of an eval, written by two people instead of one.

Week 3: start a decision log you will score. For every non-trivial decision, write one line with a probability. Mix short horizons ("80% chance this PR merges by Friday") with long ones ("70% chance this migration finishes by 30 November"), so some predictions resolve inside the plan. Score them as they resolve with a Brier score, which is the squared gap between your probability and what happened. A Brier score means little until a few dozen predictions have resolved, so treat the week 3 artifact as a started log and score it properly at 60 to 90 days. Keep cost estimates in a separate column and record the actual cost beside each one, because a Brier score only works for yes-or-no outcomes. Calibration is trainable: in Chang, Chen, Mellers and Tetlock's 2016 forecasting study, less than an hour of training improved Brier scores by 6 to 11%.

Week 4: one pre-mortem. Gary Klein's method from his 2007 Harvard Business Review article: before a real decision is final, the team imagines it has failed badly and each person writes down why. Run it on one specific upcoming decision, and keep the written list. At day 90, compare it with what actually went wrong.

Making it visible. These skills are invisible by default, which is a real career problem when most of your output is AI-assisted. Fix that with artifacts. A PR comment that says "this passes, but the second pricing change will touch twelve files; here is the one-module version" shows taste. A design doc that lists options with undo costs shows judgment. A decision log with a scored prediction history is better evidence of judgment than any promotion packet adjective. Shah-Brandenburg suggests a team-level signal worth tracking too: the rate of principled disagreement with AI output. If nobody on the team ever rejects or qualifies what the model produced, the filters are not running.

Run both filters, and date them

Addy's two filters are the right model for a single engineer looking at a pile of drafts. At team scale, and with agents in the loop, two more things decide the outcome. The common failure is running neither filter, because fluent drafts look finished and checking them has not got cheaper at the same rate. And the filters you hand to agents fossilize: judgment into numbers that go stale, taste into evals that drift. Train both skills with drills that score you, and give every filter you write down an assumption you can check, an owner and a date.

Checklist: keep your filters from fossilizing

  • Every agent budget, permission scope and approval gate records the assumption behind its number, an owner, and a review date.
  • A CI check tests each assumption against a live metric, confirms the registry matches the live policy, and fails when a review date passes.
  • Changing a review date requires a note on what was re-checked.
  • Assumptions about generated artifacts (config files, feature files) are checked by the producer before it publishes, not only in CI.
  • A judge's taste filter gates on the result of its last human calibration, not only on how recently it ran.
  • Every action an agent cannot undo by itself, or whose worst case is unacceptable, needs a second key it does not hold.
  • Agents get the verbs that fail loudly; humans keep the ones that fail quietly.
  • LLM judges are compared with fresh human labels on a schedule and on every model or prompt change, reporting TPR and TNR separately, with enough failures in the sample to trust the interval.
  • Pairwise judges run both orders and send disagreements to a human.
  • Recurring review disagreements are written down as rules, with a named owner for the standard.
  • Quality differences smaller than the eval's noise do not decide anything.

What would change this

Two things would weaken the argument. First, if a team-level study found that most AI-assisted teams do run a real review on generated code, the claim that neither-filter is the common failure would be wrong; today the best evidence is a vendor survey of individuals. Second, if encoded taste stopped drifting - if trained judges kept matching human labels over months without re-calibration - the taste half of Fossil Filters would not need its maintenance habit. I would expect the judgment half to hold either way, because its inputs live outside the code.

References


AI Engineering

Agentic AI

Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:


Books by Ranjan Kumar

Harness Engineering for Production AI Systems cover

Harness Engineering

The 7 GenAI Architectures cover

The 7 GenAI Architectures

Building Real-World Agentic AI Systems with LangGraph cover

Building Real-World Agentic AI Systems

The ChatML Handbook cover

The ChatML Handbook

The Chat Templates Handbook cover

The Chat Templates Handbook

Comments