explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the claim, the numbers, the catch
  • What is a Luttinger-compensated magnet, in plain language?
  • The two candidates and their predicted numbers
  • What the agents actually did
  • How to check the work: three verification levels
  • Caveats the authors state themselves
  • Why this is more than a materials story
  • What people are asking
  • How to try a similar loop yourself
  • Bottom line
  • Related reading
← Back to blog

explainx / blog

Opus 5.5 Agents Propose Two Room-Temperature Magnetic Semiconductor Candidates

Anthropic, Claude Opus 5.5, AI for Science, AI Agents, Research

Vals AI used Claude Opus 5.5 agents and about 750 DFT jobs to flag two room-temperature magnetic semiconductor candidates. What is real and what is not.

Oct 5, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Opus 5.5 Agents Propose Two Room-Temperature Magnetic Semiconductor Candidates

Vals AI says a team of Claude Opus 5.5 agents, working with researcher Geby Jaff, has flagged two candidate room-temperature magnetic semiconductors: one newly designed compound and one that chemists first made in 1999. The post, published October 4, 2026 on vals.ai, drew attention on Hacker News. The honest one-line summary is that this is a computational prediction, not a lab result, and the most useful thing in it for builders is how the agents were kept honest.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the claim, the numbers, the catch

table · 2 cols
QuestionAnswer
Who did it?Claude Opus 5.5 agents with researcher Geby Jaff, published by Vals AI on October 4, 2026
What was the target?A semiconductor that is antiferromagnetic (no net magnetization) but still sorts electrons by spin, called a Luttinger-compensated magnet
What are the candidates?YBaMnFeO5 (new design) and KV[Cr(CN)6] (first made in 1999)
How was it computed?Density functional theory in Quantum ESPRESSO 7.5 on Modal cloud CPUs, October 1 to 4, 2026
Is it confirmed?No. Neither material's spin sorting has been directly measured
Can I verify it?Yes, inputs, raw outputs and a checker are in the public ledger repo

What is a Luttinger-compensated magnet, in plain language?

Most magnets you know are ferromagnets: the atomic moments line up, the material has a net magnetization, and it leaks a magnetic field. That leaked field is useful for a fridge door and a nuisance inside a dense chip, because neighboring bits disturb each other.

An antiferromagnet has the opposite arrangement. Neighboring moments point in opposite directions and cancel, so there is no stray field and the material is robust against outside fields. The usual downside is that, electronically, it looks nonmagnetic: the electrons are not sorted by spin, so you cannot easily use it to build spin-based logic or memory.

A Luttinger-compensated material is the sought-after in-between case. The atomic moments cancel exactly, but the electronic bands near the gap still split by spin, the way they would in a ferromagnet. Vals AI frames the goal as a semiconductor with this property that also works at room temperature. If it exists and can be made, you get spin control without the stray field, which is the pitch behind much of spintronics research.

The authors give one useful yardstick for what counts as a meaningful spin split: at room temperature, thermal agitation only causes around 26 meV of fluctuation. A spin window much larger than that survives the heat.

The two candidates and their predicted numbers

The figures below come from the more accurate HSE06 calculations, which Vals AI says it used for its headline results, with the faster PBE+U method as a cross-check.

table · 3 cols
PropertyYBaMnFeO5KV[Cr(CN)6]
OriginNewly designed, never synthesizedFirst made in 1999, existing experimental data
Predicted band gap2.35 eV2.1 eV
Spin window, holes1.0 eV2.6 eV
Spin window, electrons1.4 eV1.6 eV
Magnetic ordering temperatureabout 420 K raw, about 490 K calibrated376 K measured (103 degrees Celsius)

Every spin window in the table is far above the 26 meV thermal scale, which is why the authors consider the candidates interesting. The ordering temperatures are above room temperature, about 293 K, which is the other half of the "room-temperature" claim.

The cyanide compound has an advantage the new design lacks: someone already made it, so its magnetic ordering is a measured number rather than a calculated one. The new oxide has the opposite profile, a cleaner design on paper but no sample anywhere.

What the agents actually did

The public ledger describes a campaign, not a single prompt. According to the compensated-magnet-ledger repository, agents ran roughly 750 jobs across several search lanes, and the repository documents only the lane whose results cleared internal thresholds. The agents designed some compounds from scratch and screened existing ones against fixed criteria.

Two process choices stand out:

  • Pre-registered pass or fail rules. The criteria for a candidate to count were written down before the decisive calculations ran, which limits the temptation to move the goalposts after seeing a good number.
  • Adversarial referee review. Separate agents were set to challenge a result before it was accepted, a pattern similar to the review loops described in Claude-shaped science.

The repository is large for what it is. It holds 876 inputs (868 plane-wave runs plus 8 post-processing steps) and a ledger of 61 traceable claims, about 130 MB gzipped.

How to check the work: three verification levels

This is the part most AI-for-science announcements skip. The repo offers a ladder of effort:

  1. Level 1, seconds. Run verify.py to recompute arithmetic from the raw outputs. The reported result is 58 checks passing, 0 failing, and 3 that are not computable.
  2. Level 2, minutes. Rerun the exchange-constant and Monte Carlo models that turn DFT energies into ordering temperatures.
  3. Level 3, CPU hours. Reproduce the quantum calculations from scratch with the original pseudopotentials.

Levels 1 and 2 test whether the bookkeeping is right. Only Level 3 tests whether the physics reproduces, and even then it tests the model, not the material. Nothing here replaces a measurement.

Caveats the authors state themselves

Vals AI is unusually direct about the limits, and they are worth repeating because the headline "AI finds magnetic semiconductors" invites overreading.

  • No measured spin sorting. Neither material has a measured band gap or spin polarization in the sense the claim needs.
  • Zero-kelvin ideal crystals. The predictions describe perfect crystals at absolute zero, not real powders or films at 300 K.
  • DFT error bars. Energy levels can be misplaced by 0.1 eV or more, and the spin windows depend on those levels.
  • Water in the 1999 sample. The old cyanide sample contained water in its pores, which leaves a small residual magnetism of 0.125 Bohr magnetons. PBE+U and HSE06 disagreed on how much the water changes the spin sorting.
  • Synthesis difficulty. YBaMnFeO5's ordered structure is predicted to become unstable above about 950 K, while making it would need temperatures of roughly 900 to 1300 degrees Celsius. In plain terms, the atoms may scramble during synthesis and destroy the ordering the design relies on.

A reasonable reading, then, is: a short list of two worth a chemist's time, produced faster than a human team would have, with the failure modes written down.

Why this is more than a materials story

The materials result will take months or years to resolve in a lab. The agent result is already usable as a template.

Long-horizon agentic work is getting measured. This materials campaign is long-running agent capability applied outside a benchmark. For how the model performs elsewhere, see the Opus 5.5 launch coverage and the Fable 5.1 vs Opus 5.5 comparison.

It is a pattern of agents as verifiers, not just generators. A related example is the formal verification of the Claude Agent SDK, where the model's output was checked by a proof assistant. Another is the 370-year-old cipher solved by Fable 5.1, where the answer could be checked against the text. The common thread is that the claim lives or dies on an external check.

The cost is not disclosed. The write-up gives no dollar figure. If you plan something similar, the per-task budgeting advice in what a Claude Code task costs on Opus 5.5 is the closest published reference on this site, though DFT compute on cloud CPUs is a separate line item from tokens.

What people are asking

Is this the AI scientist moment?

Not yet, by the authors' own account. The agents did real work: designing, screening, running about 750 jobs, and self-reviewing. But the output is two hypotheses, and science treats a hypothesis as the start of a process. Expect the story to move when someone synthesizes YBaMnFeO5 or measures the spin polarization of the 1999 cyanide.

Why are the 1999 compound and the new design both included?

They hedge different risks. The old compound removes the synthesis question because it exists, leaving the electronic-structure question. The new design is the reverse. A positive result on either would be informative, and the pairing makes the post more than a single lucky guess.

Could an outside lab test it quickly?

The cyanide material is the faster path, since a synthesis route is known. Measuring spin-resolved electronic structure is the hard part and needs specialized equipment, which is why the authors flag the missing measurement rather than implying it is routine.

What should I take from it if I build agents?

Three habits from this campaign transfer directly: write the acceptance criteria before the run, assign a skeptical second agent to attack each result, and keep a ledger where every claim is traceable to a file. None of this needs a physics background, and all of it improves any agent loop that produces claims.

How to try a similar loop yourself

You do not need a DFT cluster to borrow the structure. A minimal version for your own domain:

text
1. Write the pass/fail rule for a "good" result before any run.
2. Let one agent propose candidates and run the computation.
3. Let a second agent try to break each result using only raw outputs.
4. Record every accepted claim in a ledger with a link to the file that supports it.
5. Ship a checker script a stranger can run in seconds.

For the agent-orchestration side, the Opus 5.5 prompting guide covers how to structure long tasks, and the Opus 5.5 vs Sonnet 5.5 comparison helps decide which model to put in the proposer and referee seats.

Bottom line

Vals AI's post is a strong example of how to publish an AI-assisted science claim: specific numbers, a public ledger, a ladder of verification, and caveats stated up front. The materials themselves remain unconfirmed. If a lab confirms either candidate, that will be the real headline. Until then, treat it as an unusually well-documented hypothesis generated by Opus 5.5 agents, and watch the repository for updates.

Related reading

  • Claude Opus 5.5 launch: benchmarks, pricing, reactions
  • Fable 5.1 vs Claude Opus 5.5
  • Boris Cherny's formal verification of the Agent SDK with Opus 5.5
  • Claude-shaped science: pick the model's problems
  • Claude Fable 5.1 solves the Cyphral Distich cipher
  • What a Claude Code task costs on Opus 5.5
  • Primary sources: Vals AI write-up, compensated-magnet-ledger on GitHub

Details reflect the Vals AI post and repository as of October 5, 2026; the candidates are unconfirmed predictions and the repository may change.

Spotted something out of date? Let us know.

People in this article

  • Boris Cherny →Head of Claude Code at Anthropic
Explore people in AI →
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 24, 2026

Claude Discovers a Novel Enzyme System With CRISPR-Like Repeats: What Anthropic Found, and What Hacker News Pushed Back On

Anthropic launched a molecular biology lab and shared its first result: 950 Claude agents spent 21 hours and 210 million tokens on a DNA database and flagged array-associated reverse transcriptases (ART). The function is unknown. Here is the real claim, the process, and the strongest criticisms from a 486-point Hacker News thread.

Apr 25, 2026

Anthropic Project Deal: Claude AI Agents Negotiate 186 Deals in Office Marketplace Experiment

Anthropic's Project Deal tested Claude AI agents in a week-long San Francisco office marketplace where 69 employees each received $100 to buy and sell personal items. The autonomous agents posted listings, negotiated prices, and closed 186 deals worth over $4,000—but revealed significant performance gaps between Opus and Haiku models, with participants unable to detect when they received worse outcomes.

Sep 30, 2026

Anthropic Interviewer Is Back: Should You Make Your Claude Chat Public?

On September 29, 2026, Anthropic opened a new Anthropic Interviewer study for Claude and Claude Code users. Unlike the December 2025 round of 81,000 interviews, you can opt to publish the complete transcript plus country. This post is the eligibility checklist, the re-identification FAQ Anthropic actually wrote, and what a builder should say if they take the 15 minutes.