Vals AI says a team of Claude Opus 5.5 agents, working with researcher Geby Jaff, has flagged two candidate room-temperature magnetic semiconductors: one newly designed compound and one that chemists first made in 1999. The post, published October 4, 2026 on vals.ai, drew attention on Hacker News. The honest one-line summary is that this is a computational prediction, not a lab result, and the most useful thing in it for builders is how the agents were kept honest.
TL;DR: the claim, the numbers, the catch
| Question | Answer |
|---|---|
| Who did it? | Claude Opus 5.5 agents with researcher Geby Jaff, published by Vals AI on October 4, 2026 |
| What was the target? | A semiconductor that is antiferromagnetic (no net magnetization) but still sorts electrons by spin, called a Luttinger-compensated magnet |
| What are the candidates? | YBaMnFeO5 (new design) and KV[Cr(CN)6] (first made in 1999) |
| How was it computed? | Density functional theory in Quantum ESPRESSO 7.5 on Modal cloud CPUs, October 1 to 4, 2026 |
| Is it confirmed? | No. Neither material's spin sorting has been directly measured |
| Can I verify it? | Yes, inputs, raw outputs and a checker are in the public ledger repo |
What is a Luttinger-compensated magnet, in plain language?
Most magnets you know are ferromagnets: the atomic moments line up, the material has a net magnetization, and it leaks a magnetic field. That leaked field is useful for a fridge door and a nuisance inside a dense chip, because neighboring bits disturb each other.
An antiferromagnet has the opposite arrangement. Neighboring moments point in opposite directions and cancel, so there is no stray field and the material is robust against outside fields. The usual downside is that, electronically, it looks nonmagnetic: the electrons are not sorted by spin, so you cannot easily use it to build spin-based logic or memory.
A Luttinger-compensated material is the sought-after in-between case. The atomic moments cancel exactly, but the electronic bands near the gap still split by spin, the way they would in a ferromagnet. Vals AI frames the goal as a semiconductor with this property that also works at room temperature. If it exists and can be made, you get spin control without the stray field, which is the pitch behind much of spintronics research.
The authors give one useful yardstick for what counts as a meaningful spin split: at room temperature, thermal agitation only causes around 26 meV of fluctuation. A spin window much larger than that survives the heat.
The two candidates and their predicted numbers
The figures below come from the more accurate HSE06 calculations, which Vals AI says it used for its headline results, with the faster PBE+U method as a cross-check.
| Property | YBaMnFeO5 | KV[Cr(CN)6] |
|---|---|---|
| Origin | Newly designed, never synthesized | First made in 1999, existing experimental data |
| Predicted band gap | 2.35 eV | 2.1 eV |
| Spin window, holes | 1.0 eV | 2.6 eV |
| Spin window, electrons | 1.4 eV | 1.6 eV |
| Magnetic ordering temperature | about 420 K raw, about 490 K calibrated | 376 K measured (103 degrees Celsius) |
Every spin window in the table is far above the 26 meV thermal scale, which is why the authors consider the candidates interesting. The ordering temperatures are above room temperature, about 293 K, which is the other half of the "room-temperature" claim.
The cyanide compound has an advantage the new design lacks: someone already made it, so its magnetic ordering is a measured number rather than a calculated one. The new oxide has the opposite profile, a cleaner design on paper but no sample anywhere.
What the agents actually did
The public ledger describes a campaign, not a single prompt. According to the compensated-magnet-ledger repository, agents ran roughly 750 jobs across several search lanes, and the repository documents only the lane whose results cleared internal thresholds. The agents designed some compounds from scratch and screened existing ones against fixed criteria.
Two process choices stand out:
- Pre-registered pass or fail rules. The criteria for a candidate to count were written down before the decisive calculations ran, which limits the temptation to move the goalposts after seeing a good number.
- Adversarial referee review. Separate agents were set to challenge a result before it was accepted, a pattern similar to the review loops described in Claude-shaped science.
The repository is large for what it is. It holds 876 inputs (868 plane-wave runs plus 8 post-processing steps) and a ledger of 61 traceable claims, about 130 MB gzipped.
How to check the work: three verification levels
This is the part most AI-for-science announcements skip. The repo offers a ladder of effort:
- Level 1, seconds. Run
verify.pyto recompute arithmetic from the raw outputs. The reported result is 58 checks passing, 0 failing, and 3 that are not computable. - Level 2, minutes. Rerun the exchange-constant and Monte Carlo models that turn DFT energies into ordering temperatures.
- Level 3, CPU hours. Reproduce the quantum calculations from scratch with the original pseudopotentials.
Levels 1 and 2 test whether the bookkeeping is right. Only Level 3 tests whether the physics reproduces, and even then it tests the model, not the material. Nothing here replaces a measurement.
Caveats the authors state themselves
Vals AI is unusually direct about the limits, and they are worth repeating because the headline "AI finds magnetic semiconductors" invites overreading.
- No measured spin sorting. Neither material has a measured band gap or spin polarization in the sense the claim needs.
- Zero-kelvin ideal crystals. The predictions describe perfect crystals at absolute zero, not real powders or films at 300 K.
- DFT error bars. Energy levels can be misplaced by 0.1 eV or more, and the spin windows depend on those levels.
- Water in the 1999 sample. The old cyanide sample contained water in its pores, which leaves a small residual magnetism of 0.125 Bohr magnetons. PBE+U and HSE06 disagreed on how much the water changes the spin sorting.
- Synthesis difficulty. YBaMnFeO5's ordered structure is predicted to become unstable above about 950 K, while making it would need temperatures of roughly 900 to 1300 degrees Celsius. In plain terms, the atoms may scramble during synthesis and destroy the ordering the design relies on.
A reasonable reading, then, is: a short list of two worth a chemist's time, produced faster than a human team would have, with the failure modes written down.
Why this is more than a materials story
The materials result will take months or years to resolve in a lab. The agent result is already usable as a template.
Long-horizon agentic work is getting measured. This materials campaign is long-running agent capability applied outside a benchmark. For how the model performs elsewhere, see the Opus 5.5 launch coverage and the Fable 5.1 vs Opus 5.5 comparison.
It is a pattern of agents as verifiers, not just generators. A related example is the formal verification of the Claude Agent SDK, where the model's output was checked by a proof assistant. Another is the 370-year-old cipher solved by Fable 5.1, where the answer could be checked against the text. The common thread is that the claim lives or dies on an external check.
The cost is not disclosed. The write-up gives no dollar figure. If you plan something similar, the per-task budgeting advice in what a Claude Code task costs on Opus 5.5 is the closest published reference on this site, though DFT compute on cloud CPUs is a separate line item from tokens.
What people are asking
Is this the AI scientist moment?
Not yet, by the authors' own account. The agents did real work: designing, screening, running about 750 jobs, and self-reviewing. But the output is two hypotheses, and science treats a hypothesis as the start of a process. Expect the story to move when someone synthesizes YBaMnFeO5 or measures the spin polarization of the 1999 cyanide.
Why are the 1999 compound and the new design both included?
They hedge different risks. The old compound removes the synthesis question because it exists, leaving the electronic-structure question. The new design is the reverse. A positive result on either would be informative, and the pairing makes the post more than a single lucky guess.
Could an outside lab test it quickly?
The cyanide material is the faster path, since a synthesis route is known. Measuring spin-resolved electronic structure is the hard part and needs specialized equipment, which is why the authors flag the missing measurement rather than implying it is routine.
What should I take from it if I build agents?
Three habits from this campaign transfer directly: write the acceptance criteria before the run, assign a skeptical second agent to attack each result, and keep a ledger where every claim is traceable to a file. None of this needs a physics background, and all of it improves any agent loop that produces claims.
How to try a similar loop yourself
You do not need a DFT cluster to borrow the structure. A minimal version for your own domain:
1. Write the pass/fail rule for a "good" result before any run.
2. Let one agent propose candidates and run the computation.
3. Let a second agent try to break each result using only raw outputs.
4. Record every accepted claim in a ledger with a link to the file that supports it.
5. Ship a checker script a stranger can run in seconds.
For the agent-orchestration side, the Opus 5.5 prompting guide covers how to structure long tasks, and the Opus 5.5 vs Sonnet 5.5 comparison helps decide which model to put in the proposer and referee seats.
Bottom line
Vals AI's post is a strong example of how to publish an AI-assisted science claim: specific numbers, a public ledger, a ladder of verification, and caveats stated up front. The materials themselves remain unconfirmed. If a lab confirms either candidate, that will be the real headline. Until then, treat it as an unusually well-documented hypothesis generated by Opus 5.5 agents, and watch the repository for updates.
Related reading
- Claude Opus 5.5 launch: benchmarks, pricing, reactions
- Fable 5.1 vs Claude Opus 5.5
- Boris Cherny's formal verification of the Agent SDK with Opus 5.5
- Claude-shaped science: pick the model's problems
- Claude Fable 5.1 solves the Cyphral Distich cipher
- What a Claude Code task costs on Opus 5.5
- Primary sources: Vals AI write-up, compensated-magnet-ledger on GitHub
Details reflect the Vals AI post and repository as of October 5, 2026; the candidates are unconfirmed predictions and the repository may change.
