What happened this week
Two arguments broke out. One is about whether end-to-end full-duplex is the right architecture at all. The other is about what Article 50 compliance actually means in practice.
The modular counter-argument got loud
After a month of full-duplex launches, the two most interesting turn-taking papers this window both argue for keeping the stack separable. VoiceChat-TTS from NVIDIA drives streaming synthesis directly off the LLM's token stream, emits silence when there is no text, and takes explicit interruption control tokens so a caller can cut in mid-utterance without a KV-cache reset. The stated pitch is duplex-like responsiveness without the speech-quality cost of end-to-end duplex. X2-Turn does the mirror image on the input side, predicting turn state frame-by-frame beside the ASR head on a Voxtral Realtime backbone and separating interruptions from ignorable backchannels from genuine completion.
Our running count of full-duplex and turn-taking papers now reads zero, five, two, three across four weeks, and for the second window running no newly named end-to-end full-duplex dialogue model appeared. ByteDance's SeedRealtime, announced two weeks ago, still has no technical report, no weights, no paper and no API.
Measuring it, and feeding it
DuplexWorld is the first voice-agent benchmark we have seen that treats turn-taking as a scored axis rather than a footnote, across six task worlds and 350+ hours. On the data side, a Korean full-duplex corpus surfaced as a 100-conversation preview of a claimed 2,000 hours, which would be the largest full-duplex set in any language outside English and Chinese. The SLT 2026 SmartGlasses Challenge released 106 hours of four-channel egocentric multi-talker audio the same week sensiBel and Aizip announced optical-MEMS microphones with edge beamforming: the same problem from the hardware end.
The output side turned competitive
Deepgram's Flux TTS went GA with an interruption contract that reports what the caller actually heard. Two days earlier Cartesia shipped keyterm prompting and configurable turn detection for Ink-2, publishing a 6.2% keyword miss rate against 13.5% for Flux. Open weights moved too: IndexTTS-2.5 from Bilibili added Japanese, Spanish and Arabic, FireRedTTS3 arrived Apache-2.0 with voice design and speech editing, dots.tts shipped one-step mean-flow checkpoints, and IIT Madras released SPRING_F5 across 24 Indian languages.
Detection shipped. Disclosure did not.
Four weeks after Article 50 took effect, no voice-agent platform has shipped a disclosure prompt, toggle, or audit log. What did ship was the other half: Resemble's DETECT-World and Google's Credentio, an open-source C2PA library that validates audio credentials but cannot yet create them. California's AB 853 requires covered providers to offer a free detection tool; the EU's marking and disclosure duties are the ones going unanswered. The market is building to the requirement that is easier to satisfy.
Meanwhile audio safety produced four independent papers, including an acoustic denial-of-service attack, automated red-teaming, and an inaudible low-frequency attack with a matching defense.
Rules and money
The UK's DRCF told Parliament that businesses remain liable for what their AI agents say, and Ofcom said it will publish research on consumer impacts of AI in telecoms later this year. Its own call for input on authentication and watermarking closed on 14 August with no accompanying statement. In Japan, no legislative follow-up to the Ministry of Justice voice-rights report. Both ISS and Glass Lewis now back the SoundHound acquisition of LivePerson, whose shareholders vote on 20 August. And two weeks after losing to GEMA in Munich, Suno announced a global alliance with BMG covering a model built with the music industry and opt-in participation for artists. Litigation and licensing landed in the same month.