explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: Brain-IT at a glance
  • What problem does Brain-IT solve?
  • How does Brain-IT work?
  • What are the results?
  • How does it compare with earlier brain-decoding work?
  • A side benefit: a tool for neuroscience
  • What can Brain-IT not do?
  • Why should AI builders and learners care?
  • How to follow up or try it
  • Related reading
← Back to blog

explainx / blog

Brain-IT: Weizmann AI Reconstructs What You See From One Hour of fMRI

Research, Neuroscience, Diffusion Models, Computer Vision, ICLR

Part of AI Research

Weizmann Institute AI Brain-IT reconstructs seen images from fMRI scans using 1 hour of data instead of 40. How it works, results, and limits.

Oct 11, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
Brain-IT: Weizmann AI Reconstructs What You See From One Hour of fMRI

A "mind-reading" AI model from the Weizmann Institute of Science can rebuild a picture of what someone is looking at from their brain scan, and it needs about one hour of scanning from a new person instead of dozens. The model, called Brain-IT, comes from Prof. Michal Irani's lab and was selected for the International Conference on Learning Representations (ICLR). The Weizmann announcement went up on September 14, 2026, and the paper was accepted at ICLR 2026.

This is a research result, not a product. But it is a clean example of how modern generative models are being pointed at neuroscience, and it touches two things readers of this blog care about: how to train with very little data, and what happens to privacy when decoding gets good.

TL;DR: Brain-IT at a glance

table · 2 cols
QuestionAnswer
Who built it?Roman Beliy, Amit Zalcher, Jonathan Kogman, Navve Wasserman and Prof. Michal Irani, Weizmann Institute
What does it do?Reconstructs the image a person is viewing from fMRI brain activity
Headline claimAbout 1 hour of fMRI from a new person, versus dozens of hours for other models
VenueICLR; paper on arXiv 2510.25976
ResultsOutperforms baselines on 7 of 8 metrics, per the project page
Training dataNatural Scenes Dataset, 8 participants, about 73,000 image-fMRI pairs
Can it read thoughts or dreams?No. Only viewed images today; audio is next, video and dreams are open problems

What problem does Brain-IT solve?

Earlier systems already turn fMRI scans into pictures, and they have become good at the gist. Irani says they "preserve the semantic meaning of the image reasonably well" but "tend to make mistakes in basic features such as composition and color." A baseball scene might come back as a generic stadium with the wrong colors and layout.

The second problem is per-person training. Most decoders are personalized, so each new participant must lie in a scanner for dozens of hours. The same is true of the speech-restoration systems for people with paralysis that Weizmann mentions, which are life-changing but tailored to individuals. For a look at that adjacent field, see our coverage of a brain implant AI voice for paralyzed speech and Neuralink's brain implant milestone.

Brain-IT attacks both issues: better detail, and a drastically smaller per-person data requirement.

How does Brain-IT work?

A two-way dictionary instead of one decoder

Two documents joined by an arrow, illustrating Brain-IT encoder and decoder translating between images and fMRITwo documents joined by an arrow, illustrating Brain-IT encoder and decoder translating between images and fMRI

The core data trick, in Irani's words, resembles "a two-way bilingual dictionary." The team trained a decoder (fMRI to image) and an encoder (image to predicted fMRI). That lets the system translate any random image into a predicted brain scan and back again.

"By translating back and forth... the models would effectively build themselves a massive dataset," Irani explains. Images that were never shown to anyone in a scanner can still generate synthetic scan-image pairs, which matters because the largest real dataset covers only eight people.

Shared building blocks across brains

Brains differ, so the same picture produces different activity in different people. The team broke the problem down. They split the brain into roughly 40,000 voxels (tiny 3D units of volume) and measured how each responded to images, while an AI model extracted low-level features such as color and location plus semantic features such as "a food item or a face."

At that level, patterns that recur across people appeared. During training the encoder identified 128 functional regions shared by everyone. Some are familiar to neuroscientists, and some are new: inside the parahippocampal place area (PPA), one part responded to indoor scenes and another to outdoor scenes.

Brain Interaction Transformer plus diffusion

According to the project page, the Brain Interaction Transformer (BIT) maps each voxel to a shared functional cluster. A Brain Tokenizer produces one token per cluster, and a Cross-Transformer lets query tokens pull information out to predict image features. Two branches follow:

  • A low-level branch predicts structural (VGG) features to rebuild a coarse layout, which initializes the image.
  • A semantic branch predicts high-level features that steer a diffusion model toward the right content, with a Deep Image Prior helping keep structure faithful.

All components are shared across clusters and subjects, which is why little data per person goes a long way. If you follow generative modeling, the diffusion piece will feel familiar from posts like DiffusionGemma open-source; here, diffusion is the renderer and the brain model supplies the conditioning.

What are the results?

The authors report that Brain-IT beats prior methods visually and on standard metrics, winning 7 of 8 in their comparison averaged over four NSD subjects. The headline efficiency result is that with one hour of data from a new subject the model is comparable to current methods trained on the full 40 hours. The project page also shows meaningful reconstructions from 15 minutes of recording.

Treat those numbers with the normal caution for a self-reported benchmark: they come from the authors, on one public dataset of eight people, with people who looked at curated natural images in a controlled scanner. Independent replication will tell us how well the claims travel.

How does it compare with earlier brain-decoding work?

Image reconstruction from fMRI has moved quickly since diffusion models arrived. Systems such as MindEye2 and MindTuner, which the Brain-IT project page uses as comparison points, already produce recognizable pictures and cut the scan time needed for a new person. The common recipe is to map brain activity into the embedding space of a pretrained image model and let a diffusion decoder render the result.

What Brain-IT changes is where the sharing happens. Instead of treating each person's scan as one long vector and fine-tuning for that person, it assigns every voxel from every subject to a shared functional cluster, so the transformer learns one set of weights for all clusters and all people. A new person mostly needs to be aligned to the clusters, not taught from scratch. That is why the data requirement drops from dozens of hours toward one.

It also helps to see the two feature streams as a division of labor. The semantic stream answers "what is in the picture," and the low-level stream answers "where things sit and what color they are." Weizmann argues that earlier models were strong on the first and weak on the second. Splitting them, and using the coarse layout to start the diffusion process, is aimed squarely at that weakness.

One caution on interpretation: the reconstructions are guided by a strong image prior. A diffusion model can make a plausible picture from sparse cues, so a good-looking output does not prove the brain signal carried every visible detail. Metrics and controlled comparisons, like the 7-of-8 result the authors report, are the better test than eyeballing examples.

A side benefit: a tool for neuroscience

The encoder is useful on its own. Because it predicts scans from images, researchers can run controlled experiments: feed the same image in color and in black and white, compare the predicted scans, and infer where color is encoded. Irani suggests it could later show how regions work together and how perception takes shape. This resembles how AI models are being used to map other biology, such as the fruit fly brain connectome work and the Prima Mente cell atlas.

What can Brain-IT not do?

  • It needs a cooperating subject in an fMRI scanner. There is no way to run this on someone remotely or covertly.
  • It reconstructs viewed images, not private thoughts. Imagined images, memories and inner speech are different problems.
  • Video and dreams are unsolved. Irani says dozens of images change each second while an fMRI scan takes about two seconds, so "if we overcome all these obstacles, it's possible that in the future we may even be able to read dreams." That is a stated ambition, not a result.
  • Audio is next. The lab is working on extending the methods to auditory information.
  • Generalization is unproven at scale. Eight subjects is a small base.

Why should AI builders and learners care?

Consent bowls illustration for mental privacy and brain data consent in fMRI image reconstructionConsent bowls illustration for mental privacy and brain data consent in fMRI image reconstruction

It is a template for low-data learning. Pairing an encoder and decoder so each manufactures training data for the other is a general idea. Any domain with expensive paired data, from medical imaging to robotics, can borrow it. The same spirit shows in Meta's generative reconstruction work, which also reconstructs by generating and comparing.

It sets up a privacy debate early. Brain data is among the most sensitive data there is. Today the technique requires hours of cooperative scanning, but the trend is toward less data per person and better fidelity. Policies on neural data, consent for research datasets and storage of scans deserve attention before capabilities mature. The wider tension between AI capability and privacy is explored in AI agent privacy promises.

It is a reminder that transformers keep showing up everywhere. Here a transformer operates on tokens that stand for clusters of brain voxels instead of words or image patches.

How to follow up or try it

  1. Read the paper on arXiv.
  2. Check the project page for qualitative comparisons against MindEye2 and MindTuner, and for any code link.
  3. Get the Natural Scenes Dataset if you want to experiment; it is the public dataset the team used, though scans are large and require substantial compute.
  4. Watch for the lab's audio and video follow-ups.

Related reading

  • Brain implant AI voice for paralyzed speech
  • Neuralink brain implant milestone
  • Google fruit fly brain connectome AI
  • Prima Mente Alzheimer's cell atlas
  • Meta GenIA generative reconstruction
  • DiffusionGemma open source

Details reflect the Weizmann announcement of September 14, 2026 and the authors' project page; results are self-reported and may change with peer review and replication.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 10, 2026

Meta GenIA: Making SAM 3D Match Your Photos at Test Time, No Retraining

GenIA is a Meta and University of Tuebingen research release that reconstructs 3D objects from a single image, a few views, or a monocular video by aligning the frozen SAM 3D Objects generative prior to the inputs at test time. The code is out; the license is non-commercial and the paper reports few raw numbers.

Oct 9, 2026

Apple Normalizing Trajectory Models: 4-Step Image Generation

Apple Machine Learning Research published Normalizing Trajectory Models, a way to generate images in four steps while keeping an exact likelihood. This post explains the idea in plain language, the reported results, and what the paper does not claim.

Oct 9, 2026

Iris-3B: An Open Pixel-Space Image Model With No VAE, and an Honest Negative Result

Sperid Labs released Iris-3B on October 8, 2026, a 3B text-to-image diffusion transformer that outputs every pixel directly, with no VAE and no latent space. The paper reports the model matches Qwen-Image on OneIG. It also reports that pixel space did not beat latent models on depth or restoration.