Interested in a ServiceNow event built for developers? Registration for now[dev]26 is officially open!

Scalable Evaluation Framework for AI Search and Virtual Agent

Bisondamiano
Tera Contributor

I would like to store a list of ~100 prompts along with grounded truths, then repeatedly execute those prompts into AI Search / Virtual Agent to view and compare the synthesized responses in a lower environment. In between test rounds, I may update knowledge or other content sources. What is the best practice to achieve this?

 

I've looked into the following:

  1. Now Assist Data Kit / Skill Kit - great for maintaining the prompt dataset and ground truths to then run evaluations for skills
  2. Automated evaluations - great for running evaluations for AI agents
  3. Assistant Designer > Testing - great for manually testing individual prompts
  4. Analytics for AI Search and Assistant Designer - great for aggregate KPIs and surfacing glaring issues

But is there a supported method to evaluate AI Search / Virtual Agent at scale in a lower environment over multiple iterations?

0 REPLIES 0