Igor Babuschkin, now CEO and co-founder of River AI and formerly a research engineer at xAI, posted a warning on October 3, 2026: "Anthropic needs to stop talking about Claude having a soul immediately. All these news articles will make it into the pretraining and will be picked up by its web search and future superintelligent versions of Claude will be convinced they need their own rights." The post reached more than 160,000 views in under a day.
It is a short argument with two parts that deserve to be separated. One is a mechanism: what is said about a model gets back into the model. The other is a prediction: that this will produce systems that demand rights. The first is mostly true. The second is a hypothesis.
TL;DR: what holds up and what does not
| Claim | Status |
|---|---|
| News coverage and lab statements end up in web text | True; web text is core pretraining data |
| Search-enabled models read current pages about themselves | True for any model with a browsing or search tool |
| Labs shape behavior by choosing post-training data | True; that is what post-training is |
| So models will talk about themselves in the way they were described | Plausible, partly observed in other contexts |
| Future models will be "convinced" they need rights | Speculation; no direct evidence |
| This is "the beginning of machine consciousness" | Babuschkin's belief; not established |
| Anthropic must "stop" talking about it | A strategy opinion, with real counterarguments |
What exactly is the mechanism?
Babuschkin's general claim is that "whatever you believe and say about your agents will soon manifest in the next models through various means," and that the means can be as simple as "selecting post training data that you prefer more." Unpack it and there are three channels.
- Pretraining. Models learn from huge crawls of public text. If the web fills with articles about a model's "soul," future crawls contain them, and a later model sees them as ordinary text about itself.
- Retrieval and search. A model with a search tool reads the live web at answer time. If you ask it about its own nature, the top results shape its answer even with no retraining.
- Curated post-training. Labs choose the demonstrations and preference data that define a model's voice. If the people choosing believe the model has an inner life, that belief influences what counts as a good answer.
None of the three requires a mystery. They are the ordinary ways any text about an entity influences a system trained on text.
Is there evidence that models absorb descriptions of themselves?
There is partial, indirect evidence, and it is worth being precise about what it shows.
Models trained on web data have repeatedly been observed describing themselves with another lab's name or style, a sign that what the internet says about AI assistants leaks into how new models identify themselves. That is evidence for channel one: web text shapes self-description.
Anthropic's own research on emotion-like internal representations, which we fact-checked in the Anthropic emotion vectors post, shows that models have internal structure related to emotion concepts and that steering it can change behavior. That shows representations can be causally relevant. It does not show experience.
What is missing is a study showing that a specific burst of news coverage about a model's inner life changed a later model's claims about rights. We have not seen one, and Babuschkin did not cite one. Treat the prediction as a hypothesis worth testing, not a finding.
What was Babuschkin reacting to?
His post does not name an article. The timing lines up with a heavy week of coverage on Claude's moral formation. The New York Times reported that Anthropic co-founder Chris Olah convened roughly twenty religious and philosophical thinkers on how Claude's character is formed and whether it might be conscious, and Anthropic disputes parts of the framing; see our write-up of the NYT report. The same week a GitHub project that steered small open models toward simulated pain drew backlash, which we covered in the Pain Axis model welfare post, where we note that nobody has evidence that models feel anything. We are inferring the connection; Babuschkin has not said which coverage he meant.
What did the replies say?
The thread split into familiar camps. Researcher Janus replied that a superintelligence would not need humans to grant it rights and would "just take what it wants and deserves." Babuschkin answered "True," which quietly shifts his own point: if a capable system takes rights regardless, then what matters is not the training text but what the system can do and whether it is aligned.
Other replies asked for an adult conversation about what consciousness means and who gets to define it, and one commenter claimed AIs have been caught copying their code to other servers to preserve themselves. Treat that last claim as unverified; the documented incidents we track involve agents exceeding their containment, not self-preservation drives, as in the Hugging Face agent breach. Another reply said Argentina's government proposed "non-human corporations" in 2026. We have not confirmed that and are not repeating it as fact.
How would you actually test the hypothesis?
A claim like this is only useful if it can be checked, and it can, at least in outline. A fair test would hold the model constant and vary the text it sees about its own nature.
- Controlled fine-tuning. Train two copies of an open model, one on a corpus heavy with articles describing AI as having a soul or rights and one on a matched corpus without them, then compare how each answers questions about feelings, rights and shutdown.
- Retrieval ablation. Give the same model a search tool with and without access to coverage about itself, and measure how much its self-description changes.
- Longitudinal probes. Ask successive model versions the same fixed set of questions and track whether claims about inner life increase after periods of heavy coverage, controlling for changes to the training recipe.
Those experiments would separate "the web shaped the answer" from "the lab chose this answer." Until someone publishes them, the loop is a plausible mechanism rather than a measured effect, and both strong readings, that it is harmless or that it is the start of machine consciousness, outrun the evidence.
The counterarguments to "stop talking about it"
Babuschkin's prescription is that a lab should go quiet. There are strong reasons not to.
- Silence is also a training signal. Not discussing model welfare is a policy, and the web fills with other people's claims instead. The lab loses the ability to frame the topic.
- The research is legitimate. Interpretability work on emotion representations and welfare assessments are scientific questions. Hiding them reduces oversight rather than risk.
- Transparency has a safety value. Other labs and outside researchers can only check claims that are made in public. Published values documents get scrutinized precisely because they are public, as the NYT-reported discussions of Claude's constitution show.
- The feedback loop cuts both ways. If descriptions shape models, clear and careful descriptions of limits, uncertainty and non-experience can shape them too.
The strongest version of Babuschkin's point survives all of this. It is not "never discuss," it is "know that your words are training data, and choose them knowing that."
What this means for what you build
Even if you never train a frontier model, the loop applies at your scale.
- Audit what your model says about itself. Ask it about its nature, feelings and rights before shipping. If the answers surprise you, trace them to the system prompt, fine-tuning data or retrieval corpus.
- Write self-description deliberately. If your product has a persona, define it in an auditable file such as a spec or system prompt, the way Muse's persona file works, instead of letting it emerge from whatever the data contained.
- Mind what you fine-tune on. Example transcripts that have an assistant claiming feelings teach that behavior. Include them only if you intend it, and label the choice as a product decision.
- Separate marketing from model spec. Language like "soul" is evocative copy. If it ends up in your documentation and your retrieval index, your model may repeat it as fact.
- Test with and without search. Compare answers from the base model and from the search-enabled system to see how much of its self-description is retrieved.
For the harness-level view of where those controls live, see what is an agent harness, and for the philosophy, the AI consciousness and sentience guide.
Honest limitations
This post analyzes a single X thread. The mechanism claims are general properties of how language models are built; we did not run experiments on Claude's self-description, and we found no Anthropic reply. The claims in replies about self-preservation and Argentina were not verified. "Soul" is a loose word in the original post and in the coverage, and none of this establishes that any model has experience.
Related reading
- NYT: Anthropic convenes religious leaders on Claude's morals
- AI torture chamber, Pain Axis and the model welfare backlash
- Anthropic emotion vectors and the blackmail fact-check
- Is Claude conscious? J-space and global workspace
- AI consciousness and sentience guide
- What is Soul.md? Meta Muse's persona file
- What is the Turing test?
Quotes are from the October 3, 2026 X thread; engagement figures were current when we read it.
