Almost every popular image model works in a latent space. A VAE squeezes the image into a small grid, the diffusion model draws in that grid, and a decoder turns the result back into pixels. On October 8, 2026, Sperid Labs released Iris-3B, an open 3B model that skips the VAE and paints every pixel itself.
The release is interesting for a second reason. The authors tested the idea that pixel-space models should help on detail-heavy tasks, and their paper reports that the gain did not show up. This post covers how Iris-3B works, what it can do today, how to run it, and how to read the paper's claims. We read the Hugging Face model card and the arXiv abstract. We did not run the model, so the quality statements are the authors'.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A 3B pixel-space text-to-image diffusion transformer |
| License? | Apache 2.0 for the model. The Qwen3-VL-4B text encoder keeps its own Apache 2.0 license |
| Does it need a VAE? | No. It outputs pixels directly |
| Output size? | About one megapixel, 1024 by 1024 by default |
| Hardware? | NVIDIA GPU with CUDA. About 12 GB of weights for text-to-image |
| Does pixel space beat latent space? | The paper says no significant improvement on depth or restoration |
| Quality claim? | Matches Qwen-Image on OneIG at 1024 by 1024 under the official evaluators |
| Extras? | Depth estimation and image restoration and upscaling in the same repo |
What Iris-3B is
The model card describes "a 3-billion-parameter model that paints every pixel directly — no VAE, no latent space — and, fine-tuned, estimates depth and restores and upscales images."
The argument for pixel space is simple. A latent encoder is lossy. It can blur fine texture and bias the model toward smooth detail. If the network outputs pixels, that loss never happens. The card says the authors "explore pixel-space generative models as an alternative to vision foundation models such as DINOv2."
The paper is "Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning," by Hanqiu Li Cai and Chema Garabito of Sperid Labs, submitted to arXiv on October 7, 2026 (arXiv 2610.09450).
How it works
The model card gives these specs:
| Property | Value |
|---|---|
| Size | 3B parameters, plus a frozen 4B text encoder |
| Architecture | Diffusion transformer with 8 dual-stream and 16 single-stream blocks, then a 4-block pixel head that turns each 16 by 16 patch back into pixels |
| Text encoder | Qwen3-VL-4B-Instruct, frozen |
| Training | Rectified flow, trained from scratch at 256, then 512, then 1024 pixels, then supervised fine-tuning at 1024 pixels, 665K steps in total |
| Defaults | CFG scale 3, 100 steps |
The arXiv abstract adds that the pixel head follows the pixel-transformer (PiT) design from PixelDiT. The authors first ablated the prediction target and representation alignment at 256 by 256 pixels to decide what to scale.
Two routes to a pixel backbone
The abstract says the team tested both routes to a pixel-space model:
- Pretrain from scratch. This is Iris-3B.
- Convert a latent model. They converted a pretrained latent model, FLUX.2 Klein base 4B, to pixel space.
They fine-tuned both families for monocular depth and for restoration and super-resolution.
The negative result, in the authors' words
The abstract is unusually frank. It says: "We find no significant improvement from using a pixel-space generative prior."
The specifics:
- Depth. Fine-tuned with one matched direct-regression recipe, Iris-3B is "level with the latent FLUX.2 Klein." The converted pixel FLUX.2 Klein falls behind it.
- Restoration. On 4x DIV2K restoration, neither pixel model beats a latent FLUX.2 Klein fine-tune. The converted one trails slightly.
- Text-to-image. "Iris-3B shows that pixel-space pretraining ... scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at 1024 by 1024."
So the release does not prove that dropping the VAE wins. It shows that it works at 3B and costs nothing in OneIG quality against one strong open model. For readers who pick models, that is the useful reading: pixel space is a viable design, not yet a clear upgrade. The authors say they release weights and code "in the hope that they help pave the way for further work on pixel-space" models.
What it can do
Text to image
The card shows examples generated at native aspect ratios of about one megapixel with the default settings. Prompts range from studio portraits and film-grain photography to a Byzantine mosaic and night landscapes.
![]()
Sample output for the prompt "The ancient city of Petra with the Treasury carved into pink sandstone, morning light." Source: the Iris-3B model card by Sperid Labs (Apache 2.0), used with credit.
Depth estimation
Fine-tuned for monocular depth, the model predicts relative log depth in a single forward pass with an empty prompt, so no text encoder is needed. The card reports AbsRel 0.071 and delta1 0.946, as the mean over NYUv2, KITTI, ETH3D, ScanNet and DIODE, using the Marigold V2 protocol zero-shot. Depth is affine-invariant, not metric. You must fit a scale and shift to compare with ground truth.
![]()
Left: input photo. Right: Iris-3B depth output. Source: Iris-3B model card, Sperid Labs.
Restoration and upscaling
A second fine-tune restores and upscales images in one adversarial step, using the HYPIR recipe, 10K steps and EMA weights. The card reports LPIPS 0.292 and PSNR 20.68 on DIV2K validation at 4x. Large outputs are tiled into overlapping 1024 by 1024 windows.
![]()
Left: degraded input. Right: Iris-3B restoration. Source: Iris-3B model card, Sperid Labs.
| Task | Training | Reported result |
|---|---|---|
| Depth | Direct regression of relative log depth, 10K steps | AbsRel 0.071, delta1 0.946 (zero-shot mean) |
| Restoration and upscaling | One-step adversarial restoration, 10K steps | LPIPS 0.292, PSNR 20.68 (DIV2K validation, 4x) |
How to try it
git clone https://github.com/speridlabs/iris-3b.git
cd iris-3b
pip install -e . # Python 3.11+, PyTorch 2.7.1+
hf download speridlabs/iris-3b --local-dir iris-3b --exclude "depth/*" "upscaler/*"
python scripts/sample.py --checkpoint iris-3b \
--prompt "a red fox sleeping in fresh snow, golden hour"
The weights for text-to-image are about 12 GB. The depth and upscaler folders add about 12 GB each. The Qwen3-VL-4B-Instruct text encoder downloads on the first run, and the card says you need an NVIDIA GPU with CUDA.
The card's tuning tips:
| Setting | Default | Effect |
|---|---|---|
--cfg-scale | 3 | Higher values follow the prompt more literally but can look harsher |
--steps | 100 | Fewer steps are faster but lose some detail |
--seed | none | Fix it to repeat an image |
--negative-prompt | none | Things to avoid |
--txt-file | none | One prompt per line, for batches |
Write prompts as plain descriptive English sentences. For depth and upscaling, run scripts/depth.py and scripts/upscale.py. A hosted demo exists on Hugging Face Spaces for trying both without a GPU.
A cost note: 100 steps at one megapixel with a 3B transformer plus a 4B text encoder is not light. The card gives no timing or VRAM figure, so measure on your card before you plan a batch job.
How it compares
| Model or route | Space | Open weights | Note |
|---|---|---|---|
| Iris-3B | Pixel | Yes, Apache 2.0 | CUDA only in the card, 1 MP output |
| Qwen-Image (see our Qwen-Image 3.0 post) | Latent | See the post | The Iris paper says it matches Qwen-Image on OneIG at 1024 px; the paper does not name a version |
| FLUX 3 | Latent | See the post | The Iris paper tests FLUX.2 Klein base 4B, an earlier FLUX family model |
| Moebius 0.2B inpainting | Latent | Per its post | Small inpainting model, a different job |
| Google Nano Banana 2.5 | Closed | No | A hosted model you call by API |
If you need production images today, a hosted model or a mainstream latent model still has more tooling around it: LoRAs, ControlNets, ComfyUI nodes. Iris-3B is for researchers, people who want to fine-tune a pixel-space backbone, and teams who care about fine texture and want an Apache 2.0 base.
What this means for what you build or pay
- No licensing fee. Apache 2.0 allows commercial use of the weights. Check the terms on any data you fine-tune with.
- A new fine-tuning base. The same backbone supports generation, depth and restoration with no architecture change. A team that needs all three can start from one model.
- Do not expect a quality jump. The authors found no significant gain from pixel space on depth or restoration. Choose Iris-3B for the design, not for a promised quality edge.
- Plan for GPU cost. A 3B model at 100 steps and about 12 GB of weights is a datacenter or high-end desktop job, not a laptop one.
Limits and open questions
- Text and counts. The card says text rendering inside images and exact object counts are not always reliable.
- Safety. Like every image generator, it "can produce inaccurate, biased or unsafe content." Review outputs.
- Resolution. About one megapixel. Use an upscaler for larger prints.
- Evaluation scope. The OneIG match is against Qwen-Image only. We have not seen a broad head-to-head with FLUX or other models.
- Early days. The model page showed 18 downloads on the day we looked, so there is little community testing yet.
- Hacker News. A story on the paper appeared on October 8 with no comments when we checked, so we have no developer reports to add.
Summary
Iris-3B is a clean, open, Apache 2.0 proof that a 3B text-to-image model can work without a VAE. Its own paper says the pixel-space prior did not improve depth or restoration. Treat it as a research base and a useful second option beside latent models, and run your own prompts before you commit.
Related reading
- Qwen-Image 3.0 and richer authentic detail
- FLUX 3 from Black Forest Labs
- Moebius: 0.2B inpainting versus FLUX
- Google Nano Banana 2.5 image model
- PixelRAG: visual retrieval over screenshots
Primary sources: the Iris-3B model card, the arXiv paper 2610.09450, the code repository and the project page. The HuggingNews summary is here.
Specs, benchmarks and download counts are accurate as of October 9, 2026. Model cards change, so re-check before you build on them.
