Gradium's 155ms TTS, Z-Image Turbo, and the 2026 Race for Sub-Second AI Generation

TL;DR
- The shift: Independent benchmarks, not vendor press releases, now decide who leads real-time AI generation.
- Why it matters: Voice broke 200ms and image broke one second — teams can treat both as real-time infrastructure, not a demo trick.
- What’s next: Video is still stuck in seconds, not milliseconds, and closing that gap is the next fight.
A voice model just answered in 155 milliseconds — faster than most people notice a pause in conversation. An image generator now renders in eight steps, under a second, on hardware a hobbyist can buy. Video is the piece still missing.
The Benchmark That Beat the Marketing Page
Thesis: Real-Time AI Generation is no longer judged by vendor claims — independent benchmarks just became the deciding currency, and that shift is reshaping who gets believed.
For two years, “real-time” was a marketing word. That changed when Coval, an independent voice-AI evaluation platform, started publishing a live leaderboard instead of asking anyone to take a vendor’s word for it.
Gradium’s TTS model posted a 155ms median Time To First Audio on that benchmark, captured May 4, 2026 — first of nine models tested. Nine days later, Gradium’s own blog cited 158ms on the same ranking: a live dashboard refreshing roughly every 30 minutes, drifting because it’s real, not wrong (Gradium Blog).
A benchmark you can re-run beats a benchmark you have to trust. You’re either submitting to that test or explaining why you won’t.
Three Numbers That Define the New Speed Floor
Group the evidence by what it proves, not when it happened. Voice crossed one line, image crossed another, video hasn’t crossed either.
Voice: Gradium’s 155ms P50 came with a 2ms IQR and a 3.3% word error rate, fastest of nine models Coval tested (Gradium Blog) — possible because Streaming Inference emits audio as it’s generated instead of waiting for the full clip.
Image: few-step generation has been a multi-year target. Early Latent Consistency Model work proved it possible, SDXL Turbo shipped it using Adversarial Diffusion Distillation, and Z-Image Turbo is the newest entrant: a six-billion-parameter transformer from Alibaba’s Tongyi Lab, distilled to eight steps, running sub-second on H800 GPUs, trained for roughly $630K in compute (Z-Image model card; arXiv), under an Apache 2.0 license.
Video: still nowhere close. Fast-tier models — seedance-fast, pixverse-v5.6, Minimax-Hailuo-2.3-Fast — deliver clips in single-digit seconds. Research prototypes like StreamDiffusionV2 hit half-a-second time-to-first-frame at up to 58 frames per second, on a research rig, not a shipped product (arXiv).
Two modalities cleared the bar. The third is still warming up.
Who Gets to Sell Speed
Gradium is the most direct winner: a Paris-based startup spun out of Kyutai in late 2025, founded by Neil Zeghidour and Kyutai colleagues. It closed a $70 million seed round in December 2025, led by Firstmark and Eurazeo (TechCrunch) — proof a buyer can verify without asking, worth more than the round itself.
Coval wins too: every vendor measured on its leaderboard instead of its own blog post hands the referee authority the players don’t have.
Alibaba’s Tongyi Lab wins as well. Z-Image Turbo’s Apache 2.0 license means the fastest model and the cheapest model just became the same model, for image generation at least — changing the math for anyone building Generative Media APIs on it.
The fastest path to market just stopped requiring a closed deal.
It changes the plumbing too: real Generative Media Pipelines now run on a cheaper, faster foundation, provided delivery keeps up — output increasingly needs a persistent Websocket connection, not a buffered request-response call.
Security & compatibility notes:
- Gradium TTS SDK (Pipecat):
params=/voice=/model=are deprecated as of v0.0.105 — usesettings=instead.
Who’s Stuck Defending Marketing Numbers
ElevenLabs and Cartesia have a number problem. Promotional pages quoted Flash v2.5 near 75ms and Sonic under 100ms — numbers that don’t match Coval’s independent run: 288ms and 188ms. Marketing latency and measured latency are no longer the same number; only one is checkable.
Cartesia’s median lands second-fastest, at 188ms — but its results swing across a 100ms range, fifty times wider than Gradium’s by Coval’s own numbers. Sometimes fast, sometimes slow is harder to build around than reliably mid-pack.
Deepgram spent years marketing itself as the latency leader; on the same May 2026 snapshot, Aura-2 sits behind both Gradium and Cartesia, at 313ms. OpenAI’s TTS-1-HD is the clearest laggard: 2,295ms, slowest of the nine tested — batch generation wearing a real-time label.
You’re either built for streaming response, or built for someone to wait.
What Happens Next
Base case (most likely): Voice and image optimize the long tail — consistency, error rate, cost — instead of chasing lower medians, since both already sit near what a human can perceive. Independent benchmarks become the standard credibility test. Signal to watch: More vendors submitting to open, Coval-style benchmarks instead of self-reported numbers. Timeline: Through the rest of 2026.
Bull case: Streaming-diffusion video research like StreamDiffusionV2 moves from single-GPU papers into shipped products, repeating the path image generation took from research distillation to production speed. Signal: A production vendor, not a research lab, ships sub-second video at scale. Timeline: Tied to whichever lab commercializes it first.
Bear case: Generation speed outruns provenance tooling. Near-zero-latency synthetic voice and image, without a matching leap in AI Watermarking And Content Provenance, makes verifying what’s real harder right as faking it gets easier. Signal: Reports of real-time misuse outpacing watermarking adoption. Timeline: Already underway, compounding with every latency drop.
Frequently Asked Questions
Q: Which companies are using real-time AI generation in production in 2026? A: Gradium, Cartesia, ElevenLabs, and Deepgram ship production real-time text-to-speech. Alibaba’s Tongyi Lab released Z-Image Turbo as an open-source image model. Video relies on fast-tier options like seedance-fast and Minimax-Hailuo-2.3-Fast — none sub-second yet.
Q: How did Gradium achieve 155ms latency for AI speech generation? A: Gradium streams audio as it generates rather than waiting for a full clip — streaming inference — which is how its time-to-first-audio reached 155ms median on Coval’s benchmark, fastest of nine TTS models tested.
Q: What is the future of real-time AI generation beyond 2026? A: Expect the fight to shift from median latency, already near the limit of human perception, toward consistency, cost per generation, and provenance tooling — as independent benchmarks replace marketing as the trust mechanism.
Q: Will real-time AI video generation ever match image and audio generation speeds? A: Not yet, and not soon in production. Image reached 8-step, sub-second models only after several rounds of distillation. Video carries far more data per frame; today’s sub-second results exist only in research papers, not shipped products.
The Bottom Line
Voice and image generation crossed the sub-second line in 2026, and independent benchmarks — not vendor press releases — are what proved it. Video hasn’t crossed yet. Watch whether the next vendor to claim “real-time” submits to a benchmark like Coval’s, or just writes a blog post.
AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors