explainx / blog / topics
AI Image, Video and Voice
Generative media models make images, video, music, and voices. New releases arrive constantly, along with questions about rights, detection, and quality.
This page tracks the image, video, and voice models and tools we covered.
81 stories · latest Oct 9, 2026
Start here
The Storm Dance Trend: What It Is and How to Make Your Own AI Version
A rapid-fire arm-sweep-and-spin sequence timed to a bass drop is taking over short-form feeds under the name "the Storm." explainx.ai breaks down what actually defines the format, why it's spreading this fast, and gives a full workflow for filming your own or building an AI-enhanced version with Sora, Veo, Kling, and Runway.
How to make viral AI stunt videos like the Sylvester Stallone skate clips
Creator George Wu's hyper-realistic AI videos of Sylvester Stallone skating ramps at 80 blew up on X — while Jim Chuong's "Finally, something that isn't AI" quote-tweet hit 9.6M views as the punchline. explainx.ai walks through how these clips are made in PixVerse, how to label them honestly, and how viewers can tell real from generated without trusting vibes alone.
Build With MiniMax H3 Max: API, Local Setup, and 8 Project Ideas
H3 Max is fast enough to move AI video from a background render queue into an interactive product loop. This developer guide shows the hosted API path, explains what can and cannot run locally, and turns the speed gain into eight concrete applications.
Using Gemini to Spot Fake Cosmetics: A Multimodal LLM Case Study
Prof. William Grover photographed three packages of Rhode Peptide Lip Tint — two $5 eBay fakes and one genuine Sephora unit — and asked Google Gemini 3.6 Flash whether each was authentic. The model nailed both counterfeits by cross-referencing typos and mismatched compliance data across photos, then confidently declared the real one a fake. A clean look at where multimodal models help, where photo artifacts break them, and how to prompt around it.
Digital Camouflage: The Shirt That Makes AI Cameras Blind to You
Simon Weckert's "Digital Camouflage" shirt uses an adversarial pattern to make people-detection cameras fail to register a human figure — while remaining perfectly visible to anyone standing next to you. It's a working demo of a well-documented computer-vision weakness, staged in front of a real surveillance camera in Berlin.
Awesome GPT-Image-2: Prompt-as-Code Engine with 530+ Cases and Agent Skills
awesome-gpt-image-2 turns scattered community image prompts into Prompt-as-Code — atomic schemas, gallery cases, industrial templates, and an installable skill synced with gpt-image2.canghe.ai. Bilingual repo (English/Chinese) for production image APIs.
Timeline
October 2026
Oct 9
“Gods Don’t Give Gifts” Is the First AI Movie With an MPA RatingZack London, known on YouTube as Gossip Goblin, made an AI feature that the filmmakers say is the first AI film to get an MPA rating. It is rated R, moves to December 4, and will go for Best Animated Feature. Here is what trade press confirms and what rests on the producers’ word.
Oct 9
Iris-3B: An Open Pixel-Space Image Model With No VAE, and an Honest Negative ResultSperid Labs released Iris-3B on October 8, 2026, a 3B text-to-image diffusion transformer that outputs every pixel directly, with no VAE and no latent space. The paper reports the model matches Qwen-Image on OneIG. It also reports that pixel space did not beat latent models on depth or restoration.
Oct 9
Voyager: The "Codex for Creative Work" Harness for 100+ Tools, ExplainedVoyager launched on October 8, 2026 as an "open harness" for creative work. The founder calls it the Codex for creative work. We read the launch thread, the FAQ and the pricing page. One surprise: the app is not open source.
Oct 6
Kandinsky 6.0 Video: A 29B Open Model That Makes Sound and Picture TogetherKandinsky Lab open-sourced Kandinsky 6.0 Video on October 6, 2026: two diffusion models that generate five-second clips with synchronized 44 kHz audio, released with code and weights under the MIT license. This post covers the architecture, the claimed results, the GPU requirements, and where it fits next to other open video models.
September 2026
Sep 29
Kling 4.0 Preview Leak: Discord Screenshots, Not a LaunchOn September 24, 2026, Dev Mode Discord screenshots showed VIDEO 4.0 Preview and VIDEO 4.0 Flash Preview inside Kling's creator app, with a duration slider from 3 to 30 seconds. Kuaishou has not announced the model, published a model card, or listed 4.0 in the API changelog. This news post separates leaked UI from what you can ship on today.
Sep 26
The Storm Dance Trend: What It Is and How to Make Your Own AI VersionSep 26
How to make viral AI stunt videos like the Sylvester Stallone skate clipsSep 24
Fish Audio Drama 3 Preview: Directing AI Voice in Plain Language Instead of Audio TagsFish Audio introduced Drama 3, a preview text-to-speech model it calls the most controllable ever: describe tone and character in simple language, change voice mid-sentence, render multi-character scenes and regenerate a single word. Access is gated, pricing is unpublished. Here is what is confirmed, what to test, and how it compares.
Sep 20
AI Poster Slop Isn't a Model Problem — It's a Prompting ProblemA viral blog post and its 750-comment Hacker News argument prove that "AI slop" posters aren't a model limitation — they're what happens when nobody tells the model which design language to use. The same fix applies to AI-generated code and prose.
Sep 17
ElevenLabs Reception: The Questions Small Business Owners Actually AskedElevenLabs announced Reception on September 16, 2026, and the post drew 1.2 million views. The most useful thing in the thread was not the demo. It was small business owners asking the three questions the announcement did not answer: what happens off-script, what happens on interruption, and what happens to trust.
Sep 15
Brain Implant + AI Voice: What the Paralyzed-Speech Demo Actually ShowsPolymarket's X account shared a demo of a paralyzed woman with severely impaired speech holding real-time conversation through a brain implant and an AI-generated voice. The "mind reading" framing is wrong — this is residual speech-signal decoding plus personalized voice synthesis. Here's the corrected explanation and the applied-ML problem underneath it.
Sep 10
Suno Launches v6 AI Music Models With Warner and BMG After Licensing SettlementSuno launched v6, a new generation of its AI music generation models, built on licensing agreements with Warner Music Group and BMG rather than the unlicensed training data approach that triggered lawsuits from major labels. explainx.ai covers what changed in Suno's approach, why the labels settled rather than continued litigating, and what it means for the broader AI music generation landscape.
Sep 10
This Tennis AI Coach Was Built With Roboflow Agent and Claude CodeA single builder trained a computer vision system that tracks their own tennis strokes — ball speed, forehand vs. backhand, shot placement, and body position at contact — from plain iPhone video. The pipeline pairs Roboflow Agent for auto-labeling and fine-tuning with Claude Code driving the terminal, plus MediaPipe for pose estimation. explainx.ai breaks down the exact stack and why this kind of project is now a weekend build instead of a multi-week one.
Sep 9
HyperFrames: HeyGen’s Open-Source HTML-to-Video Framework for AI AgentsHyperFrames is HeyGen's open-source framework for turning plain HTML into frame-accurate MP4 video, shipped with 20 agent skills for Claude Code, Cursor, Gemini CLI, and Codex. This guide covers the composition model, the skills-router architecture, the explicit Remotion comparison, and what you can actually build with it today.
Sep 7
VoiceStudio: The Open-Source, Fully-Local ElevenLabs AlternativeVoiceStudio (formerly OmniVoice Studio) is an open-source, AGPL-3.0 desktop app that clones voices, dubs video into other languages, dictates system-wide, and produces audiobooks — all running locally, with 16 TTS engines and 11 ASR engines to choose from.
Sep 3
Ato Is a Screen-Free AI Device for Seniors. 2.9M Views and One Hard Question.Ato announced its second-generation device on September 2 after more than 2,500 older adults used it in beta over a year. It is voice-only, has no camera, and positions itself as a coordination layer between relatives, caregivers and providers. The post did 2.9 million views — and the top reply was somebody telling the audience to go visit their parents instead.
Sep 2
Developer Builds Endless AI TV With MiniMax H3 MaxDeveloper Rehan Sheikh turned faster-than-playback H3 Max generations into an always-on “interdimensional cable” stream influenced by viewer prompts. The demo proves continuous generative television is technically possible—and exposes brutal economics, weak memory, moderation, and copyright problems.
Sep 2
Google Pics: Workspace's New AI Image Editor ExplainedGoogle Workspace announced Google Pics on September 1, 2026 — an AI image creation and editing app built on Nano Banana 2 that lets teams edit individual objects, translate in-image text, and collaborate on images the way they already do on Docs and Slides. Here's what it actually does, why commenters immediately asked how it differs from Nano Banana and Imagen, and whether it's a real threat to Canva.
Sep 2
Build With MiniMax H3 Max: API, Local Setup, and 8 Project Ideas
August 2026
Aug 31
LeVJEPA: Video Pretraining at 20× Less Compute Than V-JEPA 2Lukas Kuhn et al. posted LeVJEPA (arXiv:2608.27395): video representation learning without EMA targets, stop-gradients, or pixel decoders. At ViT-S it uses up to 20.8× less compute than V-JEPA 2 at matched epochs — relevant for robotics and world-model builders.
Aug 31
Motion.so: URL-to-Launch-Video Agent for Motion DesignMotion (motion.so) from Mosaic (YC W25) is now generally available — an AI agent for motion design that accepts product links, X threads, YouTube videos, and DESIGN.md context to storyboard and render launch films you can edit in place.
Aug 31
OpenShot 4.0: Local AI Masking, Color Grading, and Screen RecordingOn August 30, 2026, OpenShot 4.0 shipped the biggest workflow upgrade in the project's history: a dedicated Color View with scopes and LUTs, a Recording View for screen/webcam/mic capture, 10 new effects, and local AI object masks via ONNX — no subscription. A 319-point Hacker News thread compared it to DaVinci Resolve, Kdenlive, and LosslessCut. explainx.ai breaks down what builders get from a free GPLv3 editor that runs YOLO on your machine.
Aug 29
Using Gemini to Spot Fake Cosmetics: A Multimodal LLM Case StudyAug 29
Diffusion Studio's open-source video editor turns every edit into codeOn August 28, 2026, Diffusion HQ (YC F24) open-sourced a video editor built on one idea: every edit is code, not an opaque render. The pitch is "code is the new database" — an agent can read, diff, and re-run a timeline the way it works a codebase. explainx.ai looks at the manual-edit-to-reusable-skill workflow, how it compares to ViMax and OpenCut, and whether editing-as-code actually fixes agent context loss.
Aug 29
MiniMax Fast H3 v1: real-time open video on BlackwellMiniMax announced Fast H3 v1 around August 29, 2026 — a faster inference variant of the H3 video model that the company says hits roughly a 14x speedup on NVIDIA Blackwell, aimed at real-time and faster-than-real-time open video generation. Details are thin. explainx.ai covers what real-time video unlocks for builders, how Fast H3 sits next to H3 Max and H3C, and the caveats that come with a provider-reported number.
Aug 28
Digital Camouflage: The Shirt That Makes AI Cameras Blind to YouAug 28
fal's H3 Max generates video faster than you can watch itfal Research shipped H3 Max on August 26 — a post-trained MiniMax H3 that renders a 5-second 768p clip with synced audio in under three seconds. Ethan Mollick called it a line being crossed: generation now takes less time than watching the result. explainx.ai covers the benchmarks, the $0.08/second pricing, and what breaks when video generation becomes interactive.
Aug 28
PhoneLLM: An Open Voice-Agent Model Claiming GPT-5.6 Terra Quality at 1/18th the CostDaily, the team behind the open-source Pipecat voice-AI framework, released PhoneLLM Alpha 1 on August 27, 2026 — an open-weights fine-tune of NVIDIA Nemotron 3 Nano claiming GPT-5.6 Terra-level quality on phone-support tasks at roughly 1/18th the cost and faster first-token latency. Here is what it actually claims, how the numbers stack up, and why this is an early alpha, not an independently verified benchmark.
Aug 27
World-First AI-Assisted Brain Surgery: What the AI Actually DidOn August 27, 2026, UCLH announced that a 48-year-old man became the first person in the world to have a brain tumour removed while an AI system analysed the live surgical video feed. The viral version of this story implies a robot did the operation. It did not. This is what the system actually does, why "live" is the hardest word in the sentence, and who really audits the code.
Aug 26
Awesome GPT-Image-2: Prompt-as-Code Engine with 530+ Cases and Agent SkillsAug 25
UK Gets Ukraine Avengers AI Labs: 5M Battlefield Frames for Model TrainingOn August 24, 2026, Prime Minister Andy Burnham and President Volodymyr Zelenskyy signed a UK-Ukraine AI partnership giving Britain first foreign access to Avengers AI Labs — a platform built on five million annotated battlefield frames from DELTA sensors and drone feeds. explainx.ai breaks down what the dataset actually contains, what UK teams can build with it, and what the declaration does not legally bind either side to.
Aug 22
Nari Labs Hits Sub-50ms TTS at $2 per Million CharactersNari Labs, the team behind the open TTS model Dia, published a technical breakdown of how they pushed Qwen3-TTS to 10 requests/second and sub-50ms time-to-first-audio on a single H100 — at roughly $2 per million characters. explainx.ai walks through the five serving techniques and the Hacker News practitioner Q&A that followed.
Aug 19
BGRemover.video: AI Video Background Removal, No Green Screen NeededBGRemover.video is an AI tool that strips or replaces a video's background without a green screen, exporting transparent WebM/MP4 in three steps. We cover how the AI works, pricing tiers, batch processing, and an honest note on what re-encoding does to embedded metadata like C2PA content credentials.
Aug 18
Cartesia Sonic-3.6: #1 on Both Artificial Analysis TTS BoardsThree months after Sonic-3.5, Cartesia released Sonic-3.6 in beta. It leads Artificial Analysis on both provider-voice and controlled-voice streaming leaderboards. Hear the English and Hinglish demos, then read what the Elo numbers actually mean for a production voice stack.
Aug 18
Roboflow Benchmark: GPT-5.6 Sol Is OpenAI's Best Vision Model — Gemini Still WinsRoboflow ML engineer Piotr Skalski published a VLM benchmark showing GPT-5.6 Sol is a massive leap for OpenAI on object detection and counting — up from 13.8 to 46.2 mAP@50 — but Gemini 3.5 Flash still beats it on most vision tasks at roughly a third of the cost. The post hit #1 on Hacker News twice, and Skalski himself now says Gemini 3.7 Flash is the better pick.
Aug 18
Stable Audio 3.0 Gets a DAW Plugin and a Rebuilt Web AppStability AI released two new ways to work with Stable Audio 3.0 on August 18, 2026 — a DAW plugin that puts generation directly inside Ableton and Logic, and a rebuilt StableAudio.com built for iterating on a track rather than generating once and leaving. Both run on Stability's commercially safe, fully licensed models. Here's what shipped, how licensing actually works, and how it compares to open alternatives.
Aug 9
Higgsfield Offers 33 Days of Unlimited Seedance 2.5 VideoHiggsfield launched a limited-time offer on August 7, 2026: unlimited generations on ByteDance's Seedance 2.5, no per-clip credit deduction, for up to 33 days depending on your plan. The catch is the fine print — resolution is capped at 720p, clip length varies by tier, and you have to manually flip an "Unlimited" toggle or it silently burns your credits anyway.
Aug 5
Bland Speech v3: Inside the "Human Speech Engine" LaunchOn August 4, 2026, phone-agent company Bland launched Speech v3, a standalone voice model it calls the "world's first Human Speech Engine." The centerpiece is a case study restoring a stroke survivor's voice — here's what the benchmark claim actually rests on and what the launch means for Bland's business.
Aug 1
RF-DETR: Roboflow's Real-Time Detection Transformer, ExplainedRF-DETR is a real-time detection transformer from Roboflow built on a DINOv2 backbone, spanning Nano to 2XLarge across detection, segmentation, and keypoint tasks. It hit ICLR 2026, and Roboflow now runs its architecture search directly on the platform. Here's what it is, how it benchmarks, and how to run it.
July 2026
Jul 31
Hugging Face Speech-to-Speech: Build Open-Source Voice AgentsJul 31
MiniMax H3: Open Video Model — Locked Out of the US and EUJul 30
Google Earth + Nano Banana 2: Reimagine Any PlaceJul 29
Fish Audio Raises $52M and Launches S2.1 Pro Voice AIJul 26
Inflect-Micro-v2: Full Local TTS Under 10M ParamsJul 23
FLUX 3: Black Forest Labs Unifies Video, Audio, and Robot ActionJul 19
LingBot-Map: Streaming 3D Reconstruction at 20 FPS — Robbyant GCT Guide (2026)Jul 19
Skyroot Vikram-1 Mission Aagaman: How AI Was Used — Onboard, in Engineering, and to Understand the LaunchJul 15
Overtone: Hinge Founder's $18M AI Matchmaker With No Profiles or SwipesJul 10
Reve 2.1: #2 Text-to-Image Arena, Top 4K Model, and Layout-First Visual IntelligenceJul 9
Google Photos Video Remix: Gemini Omni AI Video Editing in the Create TabJul 8
Silent Speech with Ultrasound: Aleph Neuro's 15.6% WER Demo ExplainedJul 8
Kokoro TTS: Local CPU-Friendly Speech at 82M Parameters (HN Guide, July 2026)Jul 7
X iOS Video Editor: Overlay Captions, Green Screen, and In-App Recording (July 2026)Jul 4
Seedance 2.0 Korean Neighborhood Prompt: The 12M-View Recipe ExplainedJul 3
Can Claude or LLMs Watch a Video? Here's How to Make It Work
June 2026
Jun 27
AI for Creative Hobbies: Music, Art, Writing, and the Question of What's Still YoursJun 27
How Diffusion Models Work: Complete Guide to AI Image Generation (2026)Jun 25
Krea 2 Technical Report: Open-Weights Image Foundation Model Built for Creative ExplorationJun 23
Moebius: 0.2B Parameters, 10B-Level Inpainting, 15× Faster Than FLUXJun 23
Seedance 2.5: ByteDance's 30-Second 4K AI Video ModelJun 22
HappyHorse 1.1: Alibaba Upgrades Its Top-Ranked Video Model With Native Audio and Multi-ReferenceJun 21
Palmier Pro: The Open Source Video Editor Where Claude Edits the Timeline With YouJun 21
Voicebox: The Free, Open Source AI Voice Studio That Replaces ElevenLabs and WisprFlow in One AppJun 20
Ideogram 4.0: Open-Weight Image Generation — How to Run, API & JSON Prompts (2026)Jun 19
"Bathed in Golden Light": What Experts and the Internet Actually Think About Midjourney MedicalJun 18
Midjourney Medical: Full-Body Ultrasonic CT Scanner, 60 Seconds, No Radiation — Everything from the Official AnnouncementJun 17
Midjourney Medical: The Full-Body Scanner That Was Actually Announced (Plus the Pre-Event Speculation)Jun 16
What Is Multimodal AI? Text, Image, Audio, and Video Models ExplainedJun 13
FIFA World Cup 2026: How AI Is Running the Tournament From Kickoff to Final WhistleJun 4
Miso One: 110ms Real-Time TTS Voice Model Guide 2026
May 2026
May 31
VoxCPM2: The 2B Parameter Tokenizer-Free TTS Model That Does Voice Design, Multilingual Speech, and True-to-Life Cloning (2026)May 27
OpenCut Rewrite: Open Source Video Editor Gets Plugins, Headless Mode, MCP Server, and Multi-Platform SupportMay 26
LongCat: MIT-Licensed Talking Avatar Model Revolutionizes AI Video GenerationMay 24
Frigate NVR: The Ultimate Open-Source AI-Powered Camera System for Home Assistant in 2026May 22
Runway Aleph 2.0: Professional Video Editing vs. Google Gemini OmniMay 21
How to Remove Objects from Videos with AI: Complete Guide to Video Object Removal 2026May 20
ViMax: Agentic Video Generation - Director, Screenwriter & Producer All-in-One (2026)May 4
Runway Characters: real-time conversational video agents from one image