Scalable Evaluation Framework for AI Search and Virtual Agent
Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
10 hours ago
I would like to store a list of ~100 prompts along with grounded truths, then repeatedly execute those prompts into AI Search / Virtual Agent to view and compare the synthesized responses in a lower environment. In between test rounds, I may update knowledge or other content sources. What is the best practice to achieve this?
I've looked into the following:
- Now Assist Data Kit / Skill Kit - great for maintaining the prompt dataset and ground truths to then run evaluations for skills
- Automated evaluations - great for running evaluations for AI agents
- Assistant Designer > Testing - great for manually testing individual prompts
- Analytics for AI Search and Assistant Designer - great for aggregate KPIs and surfacing glaring issues
But is there a supported method to evaluate AI Search / Virtual Agent at scale in a lower environment over multiple iterations?
0 REPLIES 0
