Tencent introduced Gander, a research model for simultaneous conversation and agent tasks
Tencent introduced Gander, a research model that holds conversations in real time while also handling background tasks using a separate “cerebellum” and a replaceable “brain”. In tests, it had the best response timing but lower task accuracy.
The Tencent Hunyuan Speech team, together with researchers from several universities, introduced Gander, a model designed to hold fluid conversations in real time while simultaneously processing more complex tasks in the background. According to the technical report, the model processes speech, images and text simultaneously, and users can interrupt it at any time.
The architecture divides tasks between two modules named after human anatomy: the “cerebellum” controls the conversation itself second by second, while the replaceable “brain” handles reasoning and complex agent tasks in the background. According to the report, the brain can be replaced with other agent systems, such as Codex or Claude Code, without retraining the conversational component; in the tests, this role was performed by an unspecified model from the GPT-5.6 family from OpenAI.
In Full-Duplex-Bench v3, Gander had the best timing of all the systems compared, according to the published results – it started speaking at the right moment in all 100 scenarios and interrupted users in only 8 percent of cases, compared with 13.5 percent for GPT-Realtime and almost 48 percent for the weakest competitor (the commercial systems compared also included Gemini Live and Grok). However, the model lagged slightly in task completion accuracy and, according to the authors, performed worse than its base model in understanding video and audio, which they attribute to training focused on conversational fluency at the expense of accurate perception.
The model was trained on approximately 2.7 million examples. Tencent plans to release the weights and training data once the open release process is complete; the code is already available on GitHub, along with sample demos on the project page.
Why it matters
The model addresses a practical problem with voice assistants – the inability to speak while working in the background without unnecessarily interrupting users. According to the published tests, Gander does this better than competing commercial systems, but at the cost of lower accuracy in completing the tasks themselves, so for now it is a research trade-off, not a deployment-ready solution.
Two audiences, two different impacts
What this means
For individuals
Developers working on voice or conversational agents have a concrete architecture design available (a separate fast module for conversation and a replaceable module for reasoning), which they could try or use as inspiration once the weights and code are released.
For a business
Companies developing voice assistants and conversational agents have further evidence that separating a fast conversational layer from a slower reasoning model reduces inappropriate interruptions, but according to the published tests, this currently comes at the expense of task completion accuracy.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.