A benchmark with a striking sentence landed in October 2026: Claude Opus 5.5 exceeds the human expert baseline on experimental research taste by a factor of 2.3, at roughly one thirtieth of the cost. The paper, TasteVal (arXiv 2610.06824), tries to put a number on something researchers usually describe vaguely: the instinct for which experiment to run next.
This explainer covers what TasteVal measures, how the setup works, why the confidence interval matters more than the headline, and what the result does and does not say about AI doing research. We read the paper's abstract and summary, not the full text, so details in the body may add nuance we have not seen.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A benchmark of experimental research taste, as compute efficiency. |
| Headline result? | Opus 5.5: 2.3x multiplier over the best expert attempts. |
| Confidence interval? | 95% CI 1.15 to 4.37, which is wide. |
| Cost? | About 1/30 of the human baseliners' average per-run cost. |
| Trend? | Frontier taste doubling about every 3.0 months since Dec 2025. |
| Tasks public? | No, held private to avoid contamination. |
| Does AI now out-research humans? | No. It is a narrow, bounded measure. |
What "research taste" means here
In AI research circles, "taste" is the judgment that separates a good researcher from a merely competent one: picking problems worth solving, designing experiments that actually discriminate between hypotheses, and interpreting messy results correctly. The authors operationalize it narrowly: experimental research taste as compute efficiency. If one researcher reaches the same score as an expert while using half the serial experimental compute, they have twice the taste.
That definition is practical and also limited. It rewards getting to a good result with few experiments. It does not directly measure picking an important problem in the first place, although the paper defines taste broadly as including that. The benchmark focuses on the experimental part: design experiments, run them, interpret results and iterate.
How the benchmark is set up
Based on the abstract and summary:
- Eight tasks. Novel, challenging, open-ended problems representative of frontier AI R&D.
- Two agents. A Researcher agent designs experiments, while a fixed Coder agent implements them. Separating the roles is meant to isolate research judgment from coding ability, so that a model does not score well merely because it writes better code.
- A fixed budget. Each attempt has 40 H100 hours or 120 wall-clock hours to work.
- Twenty models. Systems released between 2023 and 2026 were evaluated, and Opus 5.5 achieved the best result.
- Human baseline. Twenty-four expert researchers, at least two per task, with the best expert attempt per task setting the baseline.
- Private tasks. "To keep TasteVal uncontaminated, we do not release the tasks."
The best-of-human baseline is demanding. Comparing a model with the best of several experts is tougher than comparing it with an average researcher, which makes a result above 1 more notable.
Reading the number: 2.3x, with a wide interval
The headline is a compute multiplier of 2.3 for Opus 5.5 relative to the expert baseline, at roughly one thirtieth of the baseliners' average per-run cost. The 95 percent confidence interval runs from 1.15 to 4.37.
That interval is the important part. It says the data are consistent with Opus 5.5 being barely better than the best human attempts, or nearly four and a half times better. With only eight tasks and a modest number of expert attempts, the estimate is noisy. A fair summary is "likely above the expert baseline, magnitude uncertain," not "exactly 2.3 times better." Treat any social-media version that drops the interval as overclaiming.
The cost comparison is also worth a note. A model run costing a thirtieth of an expert's run is largely a statement about how cheap model compute is compared with expert time. It does not account for the cost of building the surrounding scaffold, or the model-training cost.
The trend line: doubling every three months
The paper also reports how taste has changed over time. For frontier models, the compute multiplier has doubled about every 3.0 months since December 2025 (95 percent CI 1.7 to 5.0), compared with doubling about every 14 months from 2023 to December 2025. If that holds, it would be a sharp acceleration.
Again, caution. A trend fitted to a small number of frontier models over a handful of months is fragile, and the interval from 1.7 to 5.0 months spans nearly a factor of three. A single new model can move the line. Trends like this are useful for noticing direction, and poor for forecasting. For other benchmark readings of Opus 5.5, see our coverage of its Epoch Capabilities Index score, and for how to read these numbers generally, our AI benchmarks guide.
What TasteVal does and does not show
It supports: on a set of bounded, compute-limited AI R&D experiments, a frontier model can match or exceed the best expert attempts on an efficiency measure, cheaply. That is relevant evidence for the idea that AI systems are becoming useful in the experimental loop of research.
It does not show:
- Problem selection. Choosing which question matters is a different skill from running experiments well.
- Breadth. Eight tasks in AI R&D do not cover chemistry, biology or mathematics, though other work points to progress there, as in the Opus 5.5 agents that found magnetic semiconductor candidates.
- Long horizons. Budgets of 40 H100 hours or 120 hours of wall-clock time are short compared with real research programs that run for months.
- Scaffold dependence. Results depend on the Researcher and Coder setup. Different scaffolds might change rankings.
- Reproducibility. Because the tasks are private, outsiders cannot rerun them, and must trust the authors' implementation. That is a deliberate trade-off against contamination, and a real limitation.
- Independent replication. We saw no replication by other groups.
Why this matters for builders
If you use AI in R&D workflows, the practical takeaway is not that models replace researchers. It is that the experimental loop, propose, implement, run, interpret, is increasingly something you can delegate to an agent, with a human choosing the questions and checking the work. Teams that set up that loop well, with good evaluation harnesses and compute budgets, may get more experiments per researcher-hour. Our look at how Claude is shaping scientific workflows discusses similar patterns.
It also raises governance questions. If model taste doubles every few months, capabilities relevant to AI development itself may advance faster than oversight. That is a policy debate beyond this post, but it is the reason benchmarks like this attract attention.
How to read the next benchmark headline
A few habits apply to TasteVal and to similar claims.
- Find the confidence interval, and ask whether the headline ignores it.
- Check the baseline. Best expert, median expert and novice are very different comparisons.
- Check task count and diversity. Eight tasks is small.
- Ask who can reproduce it. Private tasks protect integrity and limit verification.
- Separate the measured quantity from the claimed meaning. "Compute efficiency on eight tasks" is not "better researcher."
- Look for replications and critiques over the following weeks.
Questions the full paper should answer
If you read the full text, a few details will decide how far to trust the result. How many independent attempts did each model get per task, and how was variance handled? How were the human experts recruited and incentivized, and did they use AI assistance during their runs? Were the compute budgets the same for humans and models, and how was serial compute defined? What does the Coder agent do, and how much of a model's score depends on that fixed component? How sensitive are the rankings to dropping any single task? And how was the doubling-time trend fitted, with how many models in the recent period? Answers to these would tell you whether the 2.3 is robust or an artifact of a few tasks.
What this means for what you build or pay
For most people the immediate effect is none. For teams building research agents, TasteVal is a signal to invest in the scaffolding, evaluation and compute-budgeting that make an agent loop reliable, and to keep human judgment on problem selection. And if you cite the result, cite it with its interval and its limits, which is the most useful thing you can do for the quality of public discussion about AI progress.
Related reading
- Claude Opus 5.5 tops the Epoch Capabilities Index
- Opus 5.5 agents and magnetic semiconductor candidates
- How Claude is shaping science: bootloops
- AI benchmarks: a complete guide
- Did Claude break the 3SUM conjecture?
- Hugging Face RL environments and OpenEnv
- OpenAI's math manuscripts: what to check
Primary: "TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts," arXiv 2610.06824 (October 2026)
Details are accurate as of October 7, 2026 and are based on the paper's abstract and summary, not the full text. The tasks are not public, we saw no independent replication, and the confidence intervals are wide.
