Measure and close the realism gap in user simulators

Machine Learning


Modern conversational AI agents can typically handle complex tasks that span multiple turns, such as asking clarifying questions and actively assisting users. However, they often struggle with long interactions and often forget constraints or generate irrelevant responses. Improving these systems requires continuous training and feedback, but relying on the “gold standard” of live human testing is notoriously expensive, time-consuming, and difficult to scale.

The AI ​​research community is increasingly looking at it as a scalable alternative. user simulator — Agents powered by LLM are explicitly instructed to role-play as human users. However, modern LLM-based simulators can still suffer from significant problems.realism gapexhibit an unusual level of patience, or an unrealistic, sometimes encyclopedic knowledge of the field. Think of it like being a pilot using a flight simulator. The best simulators are as realistic as possible with unpredictable weather, sudden gusts of wind, and even birds flying into the engine. To close the realism gap of LLM-based user simulators, it must be quantified.

In our recent paper, we introduce: kelp apparela new dataset of human-AI conversations designed to do just that. ConvApparel exposes hidden flaws in today’s user simulations and provides a path to building trusted AI-based testers. To capture the full range of human behavior, from gratification to profound annoyance, we employed a unique dual-agent data collection protocol, randomly assigning participants to either a helpful “good” agent or an intentionally unhelpful “bad” agent. This setup, combined with a three-pronged validation strategy that includes population-level statistics, human-likeness scoring, and counterfactual verification, allows us to go beyond simple surface-level mimicry.



Source link