When you ask a model for a component or a migration plan, you get a lot of options, very quickly. Most of them look fine. That is the new problem: not too few ideas, but too many that pass the basic bar.
Addy Osmani describes this with a simple pair of filters: taste and judgment. As he puts it, "taste knows what's good, judgment knows the price of quality." This guide turns that framing into a working method you can use when reviewing AI drafts: what each filter does, how to run them in order, what to ask a model to help with, and what to keep for yourself. We are crediting the idea to Osmani and writing our own practical version; read his original for his full argument.
TL;DR: the two filters at a glance
| Filter 1: Taste | Filter 2: Judgment | |
|---|---|---|
| Question it answers | Is this good? | Can we afford this good thing, now? |
| What it is | A feel for quality, often felt before you can explain it | Understanding the price of quality in time and risk |
| Built from | Looking at and making many things | Domain knowledge and past decisions |
| What it removes | The bulk of the drafts, fast | All but one of the remaining options |
| Fails by | Letting weak work through | Choosing the ideal option that cannot ship |
| Missing it gives you | Work people do not love | Beautiful work, too late |
| Output | A few good options | One option, delivered |
Why people mix them up
Both feel like "knowing what's right," so we treat them as one skill. They answer different questions and fail differently.
- A reviewer with strong taste says "this is clumsy" or "that error message is aggravating." They can disqualify a draft in seconds. They cannot tell you whether the clean rewrite is worth missing the release.
- A reviewer with strong judgment says "we cannot migrate users during the final upgrade." They understand constraints. They may still approve a draft that works but nobody enjoys using.
When one person carries both, the two blur. When a team splits them, say designers and engineering leads, the handoff matters: each side needs to know which filter it is applying.
Filter 1: use taste to cut to a few
Taste is the first filter because it is fast and removes most of the pile. Run it before you think about schedules.
How to run it
- Collect the drafts. Generate several options on purpose rather than accepting the first. Three to eight is plenty.
- Read each for 60 seconds. Do not fix anything yet. Note your gut reaction and one reason, even a rough one.
- Cut hard. Drop drafts that feel clumsy, over-engineered, confusing or wrong for the product. Expect to cut most of them.
- Keep two to four. If more than four survive, your bar is too low. If none survive, regenerate with a sharper prompt.
Make taste explicit with a rubric
Taste is a feeling first, but you can write down what your feeling usually reacts to. A short rubric lets you apply it consistently and lets a model pre-screen for you.
| Check | Question |
|---|---|
| Clarity | Can a new teammate read this in one pass? |
| Fit | Does it match how our product and codebase already work? |
| Simplicity | Is anything here there "just in case"? |
| Care | Do the error states, empty states and names feel considered? |
| Surprise | Would a user or maintainer be surprised, in a bad way? |
For code, taste often shows up as the sense that something is over-built. Our coverage of Opus 5 over-engineering is a good example of readers reacting to exactly that smell.
A prompt that helps the first filter
You can use a model to apply your rubric. Paste your own criteria; do not let it invent them.
Here are four drafts of the same change and my review rubric.
For each draft, give a score of 1 to 5 on each rubric line and one sentence of evidence
from the draft itself. Do not recommend a winner. List the two drafts that
violate the rubric most and say why.
Asking it not to pick a winner keeps the second filter with you.
Filter 2: use judgment to choose one
Judgment starts where taste stops. By now you have a few good options, and "which is better?" no longer has a clean answer. The question becomes: what does each option cost, and what can we pay?
The three prices
| Price | Questions to ask |
|---|---|
| Time | Will this be done when we need it? What does the schedule really allow? |
| Risk | What breaks if this is wrong? Can we roll it back? Who is affected? |
| People | Who has to learn it, review it, migrate to it and maintain it? |
Osmani's examples are concrete: the most perfect solution may not be finished by the time you need it, and the sanest architecture may force users to migrate in the middle of a final upgrade. Judgment is knowing what you give up to get good things "for the sake of the users, for the sake of the team."
How to run it
- Write the constraints first. Deadline, who depends on this, what cannot break, who will maintain it.
- Cost each survivor. For each, estimate time to ship, risk, and the work it creates for others. Rough numbers are fine; explicit beats vague.
- Cross out what you cannot afford. An option that misses the deadline or the risk limit is out, however good it is.
- Pick the best of what is left. If nothing is left, that is information: cut scope, move the date, or choose to accept more risk on purpose and say so.
- Write down the decision and why. This is how judgment compounds.
Try it in the lab below. Move the taste bar and the two judgment sliders and watch which draft survives.
Notice the pattern: the highest-quality draft often loses to a slightly lower one that fits the deadline and the risk you can accept. That is judgment working, not taste failing.
A prompt that helps the second filter
A model can surface tradeoffs, but only from the facts you give it.
Context: ship by Friday, 3 engineers, no downtime allowed, 40% of users
on the old API. Below are three drafts that passed our quality review.
For each, list: time to ship, what could fail, how to roll back, and what
work it creates for other teams. Flag any assumption you had to make.
Do not choose for me.
Check its assumptions. A model that does not know your deadline or your users will happily invent plausible ones.
Worked example: a migration plan
You ask a model for a database migration plan and get six drafts.
Taste pass. Two are clearly over-built, with extra abstraction layers nobody asked for. One has confusing step names. One skips verification. That leaves two: a staged migration with a feature flag and a big-bang cutover with a clean final state.
Judgment pass. The deadline is a quarter-end release in two weeks. A big-bang cutover is simpler to reason about, but it requires downtime and a hard rollback. The staged plan takes four extra days and leaves two code paths alive for a while. The team is small and on call. You choose the staged plan, because the cost of a failed cutover during the release window is high and the extra days fit. You record the decision and a cleanup date for the second code path.
Neither filter alone gets you there. Taste removed the bad drafts. Judgment chose between two good ones.
Worked example: a UI component
A model gives you eight versions of a settings panel. Taste cuts five: one feels cluttered, one hides the key action, two use patterns that do not match the rest of the app, one has harsh error text. Three remain.
Judgment: one needs a new dependency, one reuses existing components, one needs a design review you cannot get this week. You ship the version that reuses existing components, file a follow-up for the nicer interaction, and note the tradeoff.
Failure modes to watch for
| Pattern | What is missing | Typical symptom |
|---|---|---|
| Endless polishing | Judgment | Beautiful work, too late |
| Choosing the "best" draft every time | Judgment | Missed releases, risky cutovers |
| Accepting whatever runs | Taste | Clumsy products, noisy codebases |
| Letting the model pick the winner | Both | Decisions nobody owns |
| Vague criteria | Taste | Review feels arbitrary |
| No record of past calls | Judgment | The same debates every sprint |
| Approving drafts you did not read | Both | See vibe coding mistakes |
Fatigue plays a part too. Reviewing many plausible drafts is tiring, and tired reviewers lower their bar without noticing. We covered that in the agentic fatigue analysis.
How to build each skill
Building taste
Taste is built by experience and practice: looking at a lot of things, making a lot of things, and remembering the good parts. Osmani notes the idea in Kant's aesthetics, appreciating form without a stake in the outcome. Practical habits:
- Read excellent code and use excellent products on purpose, then write one sentence on why they are good.
- Keep a swipe file of error messages, interfaces and pull requests you admire.
- Give yourself a small "name what is wrong" exercise: when something feels off, do not move on until you can say why.
- Compare your reactions with a respected colleague's, and note where you differ.
For the broader argument that taste is the scarce skill when making is cheap, see Jason Liu's taste essay.
Building judgment
Judgment is made from domain knowledge and experience from earlier decisions. Practical habits:
- Keep a decision log. Date, options considered, what you chose, what you expected.
- Review it after releases. Compare expectations with outcomes and note what you misjudged.
- Learn the operational side. Deployment, rollback, on-call load and migration costs are where time and risk hide.
- Ask who pays. For every option, name the person who absorbs its cost.
This is also where human-in-the-loop design comes in: decide in advance which choices an agent may make alone and which need a person to apply the second filter.
Using the filters with a team
- Say which filter you are applying in review comments. "Taste: this error text is aggravating." "Judgment: this misses the freeze date."
- Separate the two meetings or the two passes. First agree on what is good enough, then choose.
- Let the owner of the outcome apply judgment. Taste can be shared widely; judgment needs the person who carries the risk.
- Give your agent the first filter, not the second, unless you have written the constraints it needs. For tooling that reduces the cost of review, see token-efficient AI code review.
What this means for what you build
If you build with AI daily, the bottleneck has moved from producing options to choosing among them. Treat selection as a skill with two parts, practice each deliberately, and write down your criteria so both you and your tools can use them. A good habit is a one-line template at the top of any review of AI output: "Cut by taste: X. Chosen by judgment: Y because of Z."
For the wider skills picture, see Andrew Ng's AI engineering skills map, and for how over-reliance on generated output goes wrong, how to stop vibe coding and AI mania and decision-making.
Summary
Filter to a few with taste. Filter to one with judgment. Taste tells you what is good and removes most drafts quickly. Judgment tells you what you can afford and picks the one to ship. You need both, and the most common mistake is treating them as the same thing.
Related reading on explainx.ai
- Jason Liu on taste and AI creation
- Vibe coding nightmares and how to avoid them
- Agentic fatigue and the developer productivity paradox
- Human-in-the-loop AI: when to let an agent run
- Opus 5 over-engineering: the Reddit reaction
- How do we stop vibe coding?
- Andrew Ng's AI engineering skills map
- What is vibe coding?
The two-filter framing is Addy Osmani's; the method, prompts, rubric, examples and lab are explainx.ai's own. We did not locate the original post's URL, so quotations are taken from the shared text and graphic. Examples are illustrative, not drawn from a real project.
