Conversational video AI — teaching machines the art of being human

AI Video & Visuals


This is perhaps the largest art project in human history, and aims to teach machines the art of being human.

“You can’t actually teach a machine to understand humans unless you also teach it to understand human emotions,” said Hassaan Raza, co-founder and CEO of Tavus, a San Francisco-based AI research lab and developer platform. Mr. Raza shared his insights through an interview.

Much of the human emotion that Raza talks about centers on the human face, where dozens of muscles interact in myriad ways to create complex facial expressions. Along with vocal intonation and gestures, meaning is conveyed in ways far beyond words.

“A lot of our research focuses on AI’s ability to see and recognize gestures and facial expressions, where the AI ​​agent determines, ‘Does this person look happy, sad, or tired?'” Raza said. “The system can measure emotion, tonality, and reciprocity and use them as signals to respond appropriately.”

However, this technology is still relatively broad. Small cues such as blink frequency, pupil dilation, and subtle microexpressions can be missed. And such cues can vary widely across cultures. While in Western cultures silence can signal discomfort, in many Asian cultures silence can convey acceptance and agreement.

What is conversational video AI?

Conversational video AI enables real-time, face-to-face video conversations with AI agents that can interact just like humans. Not all systems exist as humans. There are robots, animal-like avatars, and object avatars. But what is most in demand are human-centered avatars who can build genuine human relationships.

The global conversational AI market continues to grow. According to Fortune Business Insights, a global market research and consulting firm, the company is valued at $14.79 billion in 2025, and is expected to reach $17.97 billion in 2026 and $82.46 billion in 2034.

Why AI still struggles with natural human conversation

At present, there is still a huge gap between surface-level imitation and true understanding.

The most obvious indicators of this gap are what Quickblox founder and CEO Nate Mareach calls “hesitation indicators and buried pauses,” which are subtle audio cues that humans use to indicate that they’re still thinking. London-based Quickblox is a cloud-based communications platform that provides AI capabilities. These cues allow for a near-instantaneous handoff from speaker to speaker in 200 to 500 milliseconds, Malich said in an email response to questions. “If it’s not completely finished,[humans]tend to fill in the gaps with ‘thinking’ noise, but most AI systems freeze in the process.”

The challenge for conversational AI is to go beyond replicating the superficial appearance of emotional expressions and develop what researchers call a “reasoning theory of mind,” or the ability to model another human’s emotional expression. maybe what you think, what you intend, and even About To think.

“‘Humanness’ is not about perfectly synchronized micro-expressions, but the ability to have common intentionality, a common goal, context, and moral grounding in interaction,” Ravikumar Bhuvanagiri said in an emailed response to questions. Bhuvanagiri is an independent researcher at the McCombs School of Business at the University of Texas at Austin. “Current systems map correlations, not relationships,” he said.

Top leaders in interactive conversational video AI include Tavus and D-ID. HeyGen and Synthesia are focused on studio/scripted video for training and marketing, but both are exploring interactive features. Infrastructure and platform providers include Microsoft, Google, Meta, Nvidia, and OpenAI.

Tavus recently launched a PAL (Personal Agent Layer) that looks human-like and can see, hear, and retain conversations. Users can video chat with these “adaptive companions” and also interact with them via text and email.

Rather than create one universal assistant, Tavus designed five different personalities to suit different user needs and interaction styles.

Agent Dominic is billed as an “old-fashioned British butler” who is adept at sorting out the chaos in users’ lives. Chloe combines emotional support with increased productivity. Tavus said Noah speaks hard truths to those looking for honest advice and true friendship — “like the older brother that never existed.” Gossippy Ashley (The Terminal Online) is a media junkie who supports creative projects and tracks pop culture trends. Charlie is a technologist who works passionately with users on a variety of tasks.

The Tavus AI agents (all of whom appear to be in their early 30s, except Dominic, who is 50) evolve and change over time, learning your communication styles and preferences. The agent can observe the user and observe and comment on facial, body, and environmental cues. “You look a little sad today, what’s wrong with you?” or “I like the Art Deco lamp by your window. Where did you find it?”

Tavus has also developed an AI Santa Claus with the same functionality as PAL, which has proven popular with both adults and children during the holidays.

Together with Santa, Tavus recently announced Sparrow-1, a conversation flow control model that brings human-level timing to real-time voice and video AI.

Still, today’s AI agents may not be able to recognize something much more subtle than an Art Deco lamp: humor, sarcasm, and other tonal changes, “leading to robot responses that miss the point,” Marichi said. Even when an AI agent detects an emotion like anger, “it often lacks enough depth to understand the root cause or provide a nuanced response,” he says. “Distinguishing between these based on trigger words and non-verbal cues and developing thoughtful escalation paths will be an ongoing challenge into 2026.”

Being human involves admitting mistakes, and developers are increasingly building this capability into their AI systems.

“Epistemic humility models explicitly track uncertainty and force AI to admit what it doesn’t know rather than bluffing about consistency,” Jigyasa Grover said in an email response to questions. Grover is a machine learning engineer and author of the following books: Sculpting data for ML: The first act of machine learning.

Glover believes there are challenges and opportunities to move “from hyper-realistic mimicry to AI that can sustain joint attention, reason through ambiguity, and not just look like a human, but actively participate in a conversation like a human. That’s where the next leap in human-like interaction lies.”



Source link