Three tools shipped on October 1, each addressing a different sensory modality: Black Forest Labs released FLUX 3 Image for structured image editing, Microsoft launched MAI-Transcribe-2-Streaming for real-time speech-to-text, and Tavus introduced Griffin for live video conversation.
Pixel-perfect control
Black Forest Labs, the Freiburg studio behind the original Stable Diffusion models, completed the image arm of the multimodal FLUX 3 family announced in July. FLUX 3 Image introduces bounding-box layout control: creators assign each element a position on a 0–1000 coordinate grid, feed up to ten reference images into a single generation request, and receive output at native 4K resolution without a separate upscaling pass. The multi-reference system maintains consistency across generations, which supports tasks such as preserving a character’s likeness across a product catalog or applying a single art direction to an entire campaign.
FLUX 3 Image preserves pixels outside an edited region with bit-identical fidelity. Prior inpainting approaches re-rendered masked regions by sampling from surrounding context, a process that introduced subtle drift in neighboring pixels even when only the masked area was supposed to change. Across multiple editing passes, each iteration corrupted pixels that the previous pass had left intact. FLUX 3 eliminates this degradation, and because its coordinate system accepts structured JSON, an LLM agent can issue successive edits to different bounding boxes across multiple API calls without accumulated visual drift. Open weights will follow within weeks.
Listening at machine speed
Artificial Analysis ranked MAI-Transcribe-2-Streaming first among 38 models for both final and partial transcript accuracy, at a 2.5% word error rate with final text arriving 0.13 seconds after the end of speech. The model begins producing partial transcripts just over 100 milliseconds after receiving audio, which allows an agent to start reasoning or calling tools before a speaker finishes a sentence.
MAI-Transcribe-2-Streaming supports 60 languages with continuous automatic detection, meaning that a session switching languages mid-stream requires no upfront declaration. Introductory pricing sits at $0.54 per audio hour. Microsoft claims that the model runs 55% faster and 60% cheaper than leading competitors. Microsoft has designated all three speech models as public previews, without production SLAs.
MAI-Transcribe-2-Streaming shipped alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash, a pair of text-to-speech models that support 23 languages and clone voices from five seconds of reference audio. The Flash variant targets voice agents specifically, with model inference latency of approximately 150 milliseconds at $15 per million characters. Together, the transcription and voice models form a complete listen-think-speak pipeline for conversational agents in which the listen and speak stages operate fast enough that the reasoning stage becomes the binding constraint.
Passing for human
Tavus, a San Francisco company with 150,000 developers on its existing avatar platform, calls Griffin a Human Interaction Model. That term describes a full-duplex video-to-video system that sees, hears, and speaks simultaneously while generating every pixel in every frame, including background, shadows, and body language, from a single reference image. In Tavus’s own blind study of 54 participants on one-minute video calls, 26 believed that they had spoken with a real person. That 48% rate represents a twentyfold increase over the 2.4% that the previous system, Phoenix-4.5, achieved across 41 participants.
Griffin renders 720p video at 25 frames per second in 320-millisecond chunks, with an average audio-to-video latency of 0.43 seconds on NVIDIA H100 hardware. On NVIDIA’s independently scored VideoFDB benchmark for full-duplex conversation, Griffin placed first among 15 models at 3.83 out of 5, against a human reference score of 3.92 and a next-best system at 2.80.
Tavus responded to these results by restricting Griffin to a research preview available only to trusted testers, citing disclosure and safety work that must precede wider release. The company has outlined plans for persistent on-screen labels, watermarking of generated video, and verified identity for meeting bots. The broader real-time video AI field has not built comparable disclosure infrastructure.
Quality is no longer a concern
FLUX 3 Image, MAI-Transcribe-2-Streaming, and Griffin have each solved their respective fidelity problems. Integration, economics, and trust now determine which of these capabilities reaches production. FLUX 3’s pixel-perfect editing reaches deployment only if agent pipelines adopt its JSON-based coordinate system. Microsoft’s speech-to-text model pushes transcription toward commodity pricing, which makes the reasoning layer between listening and speaking the remaining competitive surface. Griffin may be the likeliest of the three to remain restricted, because the disclosure infrastructure required for real-time video AI that passes for human does not yet exist.


