What happened this week
Full-duplex stopped being a speech-only idea, and the transport layer underneath it got rebuilt.
Two systems, not two papers
SeedRealtime from ByteDance Seed fuses audio, video, and text in one end-to-end model with no external VAD, adds proactive speaking driven by continuous environmental awareness, and is described as already deployed at scale. There is no technical report and no weights, so treat the pacing claims as vendor-reported. Alongside it, OpenAI published an engineering account of the GPT-Live stack. The headline for everyone else is not the model: it is WARP, a WebRTC profile that cuts media and data startup from six network round trips to one, already implemented in libwebrtc and Pion and proposed at the IETF. The media frontend also moved from Python asyncio to Go, which put the new system's p95 where the old system's p50 was.
The research side contracted
Full-duplex paper volume has now gone zero, five, two across three weeks, and this window produced no papers at all on backchannels, turn-taking prediction, endpointing, or VAD. What did land is unusually practical. PACE names generative context mis-anchoring, the failure where a barge-in leaves the model reasoning over its own words the user never heard, and takes accuracy on its benchmark from 25.0% to 96.3%. Aero Realtime puts every output slot on an 80ms grid where the model predicts a lexical token or a silence token, so when to speak and what to say are learned by one objective. One late catch: JoyAI-Talker, a full-duplex empathetic voice model from JD, was submitted on 2 August and so belongs to last week's window, where we missed it.
Platforms, and week two of Article 50
livekit-agents shipped turn-detection plumbing twice this week, and huggingface/speech-to-speech gave the open stack WebRTC for the Realtime API with semantic endpointing on by default. Qwen Audio Agent shipped six releases in seven days, adding a native speech-to-speech frontend, wake-word gating, and session memory.
On compliance, the second week after the EU AI Act's Article 50 took effect looks like the first: no voice-agent platform shipped a disclosure prompt, toggle, or audit log. The single piece of disclosure code we found is a contribution to LiveKit's Anam avatar plugin, written by Anam.
Japan, not Brussels
Brussels published nothing on Article 50 this window. Japan's Ministry of Justice did the more consequential thing, releasing on 7 August the final report of its study group on civil liability for unauthorised use of likeness and voice. It is an interpretive guideline under existing law and case law, not legislation, and it addresses whether generative AI use infringes publicity rights, the scope of damages and the availability of injunctions, and whether the Unfair Competition Prevention Act applies. Anyone licensing voices for the Japanese market should read the report itself.
Money
HappyRobot raised $150M at $1.2B for freight phone operations, Omilia raised $67M on stated ARR above $60M, and Yellow.ai announced a $550M SPAC merger naming its voice agent as its fastest-growing product. SoundHound posted record quarterly revenue of $61.9M, up 45%, and raised full-year guidance while explicitly excluding LivePerson, whose shareholders vote on the merger on 20 August.