explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • Quick answers
  • What exactly did Kandinsky Lab release?
  • How does it make audio and video together?
  • How good is it, really?
  • What do you need to run it?
  • Why the MIT license is the real story
  • Where does it sit among open and closed video models?
  • What people will ask next
  • A practical first-hour plan
  • Bottom line
  • Related reading
← Back to blog

explainx / blog

Kandinsky 6.0 Video: A 29B Open Model That Makes Sound and Picture Together

Kandinsky, Video Generation, Open Source AI, Open Weights, Audio Generation

Kandinsky 6.0 Video ships a 29B Pro and 3B Lite model with synced audio under an MIT license. Here is what it does, what it needs, and how to try it.

Oct 6, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
Kandinsky 6.0 Video: A 29B Open Model That Makes Sound and Picture Together

Kandinsky 6.0 Video is an open, MIT-licensed pair of video models that generate picture and sound in one pass. Kandinsky Lab published the code, checkpoints, and an arXiv technical report on October 6, 2026. There are two sizes: a 29-billion-parameter Pro for cinematic quality and a 3-billion-parameter Lite for quick iteration. Each makes five-second clips with synchronized 44 kHz audio, and a separate super-resolution stage raises the picture to 1920 by 1080.

That combination matters because open video models have mostly been silent. Adding speech, effects, and lip-sync usually means a second model and a manual alignment step. Kandinsky 6.0 trains both streams together. The headline for builders is the license: MIT, with no territory carve-outs. That is a sharp contrast to MiniMax H3, whose open weights exclude the US, EU, UK, and South Korea.

Three illustrated storyboard frames leading to a play button, representing short clips generated by Kandinsky 6.0 Video

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

Quick answers

table · 2 cols
QuestionShort answer
What was released?Kandinsky 6.0 Video Pro (29B) and Lite (3B), plus a video super-resolution model
License?MIT for code, checkpoints, and diffusers integration, per the repository and report
What does it generate?Five-second clips with synchronized 44 kHz audio, text-to-audio-video and image-to-audio-video
Output resolution?Up to Full-HD (1920x1080) via built-in super-resolution
What hardware?An NVIDIA GPU and Python 3.13 or 3.14; Pro HD runs 292 seconds on an H100 (non-distilled)
Can I try it hosted?Yes: a Hugging Face Space demo; also vLLM-Omni and ComfyUI nodes
Biggest caveat?The quality comparison is the authors' own human study; clips are only five seconds

What exactly did Kandinsky Lab release?

According to the project repository, Kandinsky 6.0 Video consists of two diffusion models. Both support text-to-audio-video (T2AV) and image-to-audio-video (I2AV). In the second mode you hand the model a still image and a prompt, and it animates the image and generates a matching soundtrack.

The Hugging Face paper page for the report lists the team as a large collective, with 88 contributors on the arXiv entry. The repository also links a diffusers integration, so the models can be loaded through the same Python library many readers already use for image models. For the technical background on the earlier generation, see the Kandinsky 5.0 report and the diffusers pipeline documentation.

There is also a dedicated video super-resolution repository. That split is practical: the base models generate at a lower resolution, and the upscaler is a separate plugin you can swap or skip.

How does it make audio and video together?

The report describes a dual-stream CrossDiT architecture. One stream is a pretrained video diffusion transformer. The second is a newly trained audio stream. The two are connected by bidirectional cross-attention, meaning video tokens can look at audio tokens and the reverse at every layer where they are linked. That is how the model is supposed to keep a closing door and its thud on the same frame, or lip movement aligned with speech.

The training recipe, as the authors summarize it, has several stages:

  1. Pretrain the audio stream on large-scale audio corpora.
  2. Train the video and audio streams jointly on paired audio-video data.
  3. Supervised fine-tuning on curated examples.
  4. Reinforcement-learning post-training.
  5. Distillation down to 10 sampling steps for a faster variant.

Two numbers stand out. The RL stage reportedly cut speech word error rate by 47 percent, which is a direct measure of how intelligible generated speech is when transcribed. The report also lists a 10-step distillation stage, and the quick-start downloads the pro-distill checkpoint by default.

For a plain-language refresher on how diffusion transformers differ from autoregressive models, our guide to video generation tools across Sora, Runway, and Kling covers the basics of the category.

How good is it, really?

The authors report a human evaluation in which Kandinsky 6.0 Pro clearly beats Kandinsky 5.0 Pro and is preferred over LTX 2.5 on visual quality, motion realism, visual prompt following, artifact reduction, and overall task solving. On speech quality the report claims a statistically significant advantage over LTX 2.5.

Read that carefully. This is a first-party study run by the team that built the model. It does not compare against the strongest closed systems, and the report describes the Pro model as "competitive" with leading models on speech rather than ahead of all of them. We have not independently reproduced the evaluation, and independent leaderboards typically take days to weeks to list a new open release. Wait for those before treating the ranking as settled.

What the claim does tell you: for an openly licensed model you can run yourself, the team is aiming at the same tier as current open competitors like LTX 2.5, with audio treated as a first-class output instead of an add-on.

What do you need to run it?

The repository is explicit about the basics. You need an NVIDIA GPU and Python 3.13 or 3.14, plus uv and just. The install flow uses the just task runner:

bash
git clone https://github.com/kandinskylab/kandinsky-6
cd kandinsky-6
just setup
just download pro-distill
just generate "A street musician plays violin in a rainy alley, close-up, ambient rain and reverb"

Weights are cached under ~/.cache/kandinsky, and each run writes a dated output folder containing the video, the expanded prompt, logs, and the configuration, which makes runs easy to reproduce and compare. Prompts are automatically expanded before generation, so keep the original short and descriptive.

The per-GPU timings in the README are worth reading before you rent hardware. For the non-distilled Pro model at HD, the README lists 292 seconds on an H100 and 2,716 seconds on an RTX 5060 Ti for a 5-second clip, excluding weight loading and encoding. The attention backend also varies by card: Hopper GPUs use FlashAttention 3, while Ampere, Ada, and consumer Blackwell cards need SageAttention 2.2.0, which you compile yourself.

table · 3 cols
ChoiceBest forTrade-off
Pro (29B), distilledHighest quality, final rendersSlow on consumer cards, heavy memory use
Lite (3B)Fast iteration, prompt testingLower fidelity than Pro
Hugging Face SpaceNo GPU, quick testsShared demo, Pro distill only, data leaves your machine
ComfyUI node packageExisting node-based workflowsSeparate node installs via ComfyUI Manager
vLLM-OmniServing at scaleRequires an inference engineering setup

For a first experiment, we would start with Lite or the hosted demo to learn how prompts behave, then move to Pro only for clips you intend to keep.

Why the MIT license is the real story

Open video models come with very different terms. Some restrict regions, some forbid using outputs to train other models, some require a commercial agreement above a revenue threshold. Our explainer on choosing open-weight versus closed models walks through why license text matters as much as benchmark scores when you plan a product.

MIT is about as permissive as it gets: use, modify, and redistribute, including commercially, with attribution and the license notice preserved. Practically, that means a studio, a startup, or a hobbyist can fine-tune Kandinsky on their own footage and ship the result. That is a different proposition from the gated terms we flagged in MiniMax H3 and from the hosted-only access of Seedance 2.5.

One honest caution: an MIT license on the weights does not remove other obligations. If your outputs feature real people's voices or likenesses, disclosure and consent rules still apply, and some jurisdictions are adding synthetic-media labeling laws, as we covered in the New York AI video disclosure law post.

Where does it sit among open and closed video models?

Video generation has moved quickly in 2026. Alibaba's Wan 3.0 and HappyHorse 1.1 pushed Chinese video models forward. MiniMax's fast H3 variant showed faster-than-playback generation on Blackwell hardware. On the closed side, Google's Gemini Omni Flash and ByteDance's Seedance 2.5 offer longer clips and higher peak quality through paid APIs.

Against that field, Kandinsky 6.0 has a clear niche:

  • Audio and video in one model, rather than a silent clip plus a separate audio tool.
  • Permissive license with no territory restrictions.
  • A usable small model, the 3B Lite, which is rare in a field where open releases are often 30B and above.
  • Broad tooling on day one: diffusers, vLLM-Omni, ComfyUI, and a public demo.

It also has clear limits. Five seconds is short next to closed models advertising clips up to 30 seconds. Full-HD comes through an upscaling stage, not native generation. And the Pro model is expensive in GPU time unless you use the distilled checkpoint or an H100-class card.

What people will ask next

Can it make a whole scene or a short film? Not directly. Five-second clips are building blocks. Teams stitch them with an editing workflow, use image-to-video to keep a character consistent between shots, and rely on an editor for pacing. Our guides to agentic video production show how multi-shot pipelines are assembled around short-clip models.

Is the audio good enough for dialogue? The report highlights lip-sync and a speech word error rate reduction from RL post-training, and claims a significant speech-quality edge over LTX 2.5. Quality in your language, accent, and noise conditions is something to test yourself. Start with short lines of dialogue and transcribe the output with a speech-to-text tool to measure intelligibility.

Will it run on my gaming GPU? Probably, slowly. The README lists benchmarks for seven GPU models across SD, HD, and Full-HD, and consumer cards appear in the table. A card with limited memory will push you toward Lite.

Is it safe to deploy? Open video-and-voice models raise obvious misuse risks, including impersonation. If you build on it, add disclosure labels, avoid cloning real voices without consent, and log generations. See our report on deepfake video-call fraud for what is at stake.

A practical first-hour plan

  1. Open the Hugging Face Space demo linked from the repository and try three prompts: a silent-looking scene, a scene with obvious sound, and a short spoken line.
  2. If the results fit your use case, run Lite locally with just generate and compare it against the demo output.
  3. Run the same prompt on the distilled Pro checkpoint and note the render time on your hardware.
  4. Use image-to-video with one of your own reference images to test consistency.
  5. Only then decide whether the super-resolution stage is worth the extra seconds for your delivery format.

Keep a log of prompts and seeds. The repository writes the expanded prompt and configuration for each run, which is enough to reproduce a clip later.

Bottom line

Kandinsky 6.0 Video is a credible, permissively licensed entry in the open video race, and its main differentiator is native audio generation in a model you can actually download. The quality ranking vs LTX 2.5 is the authors' own and should be confirmed by independent leaderboards. The five-second clip length and GPU cost are real constraints. If you build with video AI and care about a clean license, it belongs on your test list this week.

Related reading

  • MiniMax H3: open video model with territory restrictions
  • MiniMax fast H3 and real-time open video on Blackwell
  • Alibaba Wan 3.0 video model
  • Video generation AI complete guide: Sora, Runway, Kling
  • How to choose open-weight vs closed AI models
  • Seedance 2.5: 30-second 4K AI video
  • Gemini Omni Flash video generation

Official sources: technical report on arXiv, Kandinsky 6 GitHub repository, Hugging Face paper page.

Details reflect the repository and report as of October 6, 2026. Check the repository for updated requirements, checkpoints, and license files.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Aug 11, 2026

antirez Ported MiniMax H3 to Apple Silicon — in C and Metal

Salvatore Sanfilippo (antirez) shipped h3.c on August 10, 2026 — MiniMax H3 video generation running natively in C and Metal on Apple Silicon, MIT licensed. MiniMax called it proof that "you can't hire this, you can only open-source and let it happen." The awkward part: H3's own license excludes the EU from local deployment.

Jul 31, 2026

MiniMax H3: Open Video Model — Locked Out of the US and EU

What started as a July 30 teaser became a confirmed open-weight release on August 3, 2026: a 33B omni-modal video model with native audio that tops Artificial Analysis's editing leaderboard. The catch is the license — it excludes the US, EU, UK, and South Korea from running the weights locally.

Jul 26, 2026

Why explainx.ai Supports Open-Source AI

Two letters landed on Washington's desk in July 2026 arguing opposite sides of the same question: should open-weight AI models stay legal to download and build on? Here's explainx.ai's own position, backed by the download, pricing, and adoption numbers — and what's actually at stake for the 350,000+ people we've taught to build with AI.