Nuance Labs is developing AI for face-to-face conversations, with intended applications in customer service, sales, coaching and education. The Seattle company announced a $50 million Series A on September 14, ahead of a public research preview planned for later this year.
The proposed interaction runs in both directions at once: a person speaks and moves while the model processes those signals and generates vocal and facial responses. Nuance calls this full-duplex operation. Its goal is an AI that can acknowledge someone while listening, including through nods and interjections, instead of waiting for a completed turn.
The technical work extends beyond producing a convincing face. In its earlier investment description, Accel said Nuance compresses human video into tokens—units a model can process—and extends language-model architectures to audio and visual information. The intended output is an avatar whose expression, body movement and tone change together as the conversation unfolds.
For prospective users, the value would be a system that responds to how something is said as well as the words themselves. Investor Lightspeed has envisaged live coaching that recognizes someone losing confidence during a presentation. That is an illustration of the intended application, rather than evidence of an operating customer deployment.
Lightspeed Venture Partners led the Series A, with Accel, South Park Commons, NVIDIA and Define Ventures participating. Nuance says the proceeds will support model development, research hiring and its first public research preview.
That preview will give people a chance to try the conversation model directly. The practical question is whether its responses help users express themselves and complete a task more effectively, especially when words, tone and facial expression convey different signals.