Towards Automated Evaluation of Multi-Turn Agentic AI Systems
Start Date
June 2025
End Date
January 2026
Status
completed
Description
Today's AI systems can plan, act, and adapt on their own to get a task done. Ask such an "agentic" AI to "book a restaurant that suits both my partner and me," and it works toward that goal through a real back-and-forth conversation, much like a personal assistant would. These exchanges are called multi-turn interactions: the AI and the user go back and forth over several steps, rather than settling everything in a single question and answer.
But the more we entrust to these systems, the more pressing one question becomes: are they working reliably and safely?
That is surprisingly hard to check. Unlike conventional software, AI doesn't always behave the same way. The same request can lead to different results, and no two conversations are quite alike. The space of possible situations is almost limitless. Today, this means people have to inspect each step by hand: did the AI understand the task? Did it complete it correctly? How does it handle unusual or tricky cases? This is costly, slow, and rarely thorough.
Our vision is to let AI test AI. We are building a framework in which a large language model (LLM), the same kind of AI that powers today's chatbots, steps into the role of the human user and holds realistic conversations with the system being tested. Starting from just a handful of real examples, it can generate a rich variety of lifelike dialogues that mirror how people genuinely behave, automatically explore countless scenarios, and measure not only whether a task succeeded, but how well the system progressed along the way.
The result is testing with far greater breadth and depth than manual work could ever reach, faster, more consistent, and more affordable, and with it, a crucial foundation for businesses to deploy intelligent assistants with confidence.
But the more we entrust to these systems, the more pressing one question becomes: are they working reliably and safely?
That is surprisingly hard to check. Unlike conventional software, AI doesn't always behave the same way. The same request can lead to different results, and no two conversations are quite alike. The space of possible situations is almost limitless. Today, this means people have to inspect each step by hand: did the AI understand the task? Did it complete it correctly? How does it handle unusual or tricky cases? This is costly, slow, and rarely thorough.
Our vision is to let AI test AI. We are building a framework in which a large language model (LLM), the same kind of AI that powers today's chatbots, steps into the role of the human user and holds realistic conversations with the system being tested. Starting from just a handful of real examples, it can generate a rich variety of lifelike dialogues that mirror how people genuinely behave, automatically explore countless scenarios, and measure not only whether a task succeeded, but how well the system progressed along the way.
The result is testing with far greater breadth and depth than manual work could ever reach, faster, more consistent, and more affordable, and with it, a crucial foundation for businesses to deploy intelligent assistants with confidence.
Leader contributor(s)
Member contributor(s)
Partner(s)
Calvin Risk
Funder
Division(s)