Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework
Introduces LPS-TC (Lightweight Proactive Speech Turn Controller), a plug-and-play module that enables half-duplex models to perform full-duplex turn-taking and enhances existing full-duplex models with finer timing control Presents WildTurn, a large-scale real-world English dataset of ~2,981 hours of multi-turn stereo conversations with annotations for five turn-taking and five backchanneling styles Proposes a two-tier evaluation scheme measuring both chunk-level timing precision and turn-level
Analysis
TL;DR
- Introduces LPS-TC (Lightweight Proactive Speech Turn Controller), a plug-and-play module that enables half-duplex models to perform full-duplex turn-taking and enhances existing full-duplex models with finer timing control
- Presents WildTurn, a large-scale real-world English dataset of ~2,981 hours of multi-turn stereo conversations with annotations for five turn-taking and five backchanneling styles
- Proposes a two-tier evaluation scheme measuring both chunk-level timing precision and turn-level interaction quality under realistic streaming constraints
- Demonstrates superior performance in timing appropriateness and response quality when integrated with models like Qwen2.5-Omni (half-duplex) and Freeze-Omni (full-duplex)
- Achieves fine-grained style controllability and strong generalizability for more natural, human-like spoken interactions
Why It Matters
This work addresses a critical gap in conversational AI: the transition from rigid half-duplex turn-taking to natural full-duplex dialogue where agents can interrupt, backchannel, and respond proactively in real time. For AI practitioners building voice assistants or conversational agents, this framework provides a practical, modular solution to achieve human-like interaction dynamics without requiring complete model retraining.
Technical Details
- LPS-TC Architecture: A lightweight, plug-and-play proactive speech turn controller featuring a fine-grained action space that covers both reactive and proactive turn behaviors, bridging the capability gap between half-duplex and full-duplex systems
- WildTurn Dataset: Approximately 2,981 hours of filtered multi-turn stereo conversations sourced from face-to-face and telephone interactions, annotated with five distinct turn-taking styles and five backchanneling styles
- Two-Tier Evaluation Scheme: Assesses chunk-level timing precision (when to speak/interrupt) and turn-level interaction quality (how natural and appropriate the response is) under realistic streaming constraints
- Integration Experiments: Successfully integrated with Qwen2.5-Omni (half-duplex baseline) and Freeze-Omni (full-duplex baseline), demonstrating improved timing appropriateness and response quality across both model types
- Style Controllability: The framework enables fine-grained control over conversational styles, allowing agents to adapt their turn-taking behavior to different interaction contexts
Industry Insight
- The plug-and-play nature of LPS-TC means existing half-duplex conversational AI systems can be upgraded to full-duplex capabilities without complete architectural overhauls, offering a cost-effective path for companies looking to enhance their voice agents
- The introduction of WildTurn addresses a significant data scarcity in realistic full-duplex training, and its release could accelerate research in natural turn-taking dynamics across the community
- As real-time spoken AI becomes increasingly competitive, the two-tier evaluation framework provides a more rigorous standard for measuring conversational naturalness, pushing the industry beyond static benchmark metrics toward dynamic, streaming-aware assessment
Disclaimer: The above content is generated by AI and is for reference only.