ai contact center · stub

Testing and evaluating AI virtual agents

Verified 2026-09-30 · 60 sources · tier 1–2 · 2 disputed

Stub. This topic has 60 sources and no published article. The sources below are everything recorded so far.

Disputed: 2 sources below carry conflicting evidence.

So far, no summary has been generated for this topic. The sources below are everything recorded so far.

See also

Related to

  • Amazon Connect Contact Lens — Amazon Connect performance evaluations and generative-AI evaluation answers run on conversational analytics (Contact Lens) transcripts; this packet covers only their use on bot and AI agent interactions. The target topic id was seen only as a sibling directory name and its content was not read.

Sources

  1. 1
    Amazon Connect performance evaluations support automated evaluations of both human agent interactions and automated interactions handled by bots or AI agents.
    Evaluate agent and self-service interaction performance in Connect Customer · third paragraph after the Tip box · Checked 2026-09-30
  2. 2
    Amazon Connect defines faithfulness score as the proportion of sessions in which the Orchestration AI agent's responses remain faithful to the conversational context, including messages and tool call results.
    AI agent traces using Contact search and Contact details · AI agent performance metrics > Faithfulness score · Checked 2026-09-30
  3. 3
    In Amazon Connect, generative AI can answer up to 10 evaluation questions per contact, both through Ask AI and in automatically submitted evaluations; the limit does not apply to automation that uses contact categories or metrics.
    Evaluate agent performance in Connect Customer using generative AI · Use Ask AI... step 2.2; Set up automated evaluations using generative AI on the evaluation form · Checked 2026-09-30
  4. 4
    When the Amazon Connect generative AI evaluator cannot identify which answer option fits, it selects Not Applicable even if the question is not optional, and the question is excluded from scoring with its weight redistributed to other questions.
    Evaluate agent performance in Connect Customer using generative AI · Provide additional criteria for answering evaluation form questions using generative AI > AI answer mapping and scoring · Checked 2026-09-30
  5. 5
    AWS states that AI-generated evaluations in Amazon Connect are not 100% accurate and recommends reviewing a sample of evaluations before acting on AI outputs.
    Evaluate agent performance in Connect Customer using generative AI · Process to automate evaluations using generative AI > Note · Checked 2026-09-30
  6. 6
    AWS advises against using Amazon Connect generative AI evaluation for questions that need information outside the conversation transcript or that depend on tone of voice, stating that it cannot analyse screen recordings, access CRM or other external systems, evaluate across multiple contacts, or determine tone.
    Evaluate agent performance in Connect Customer using generative AI · Guidelines to improve generative AI accuracy > Selecting questions to be answered by generative AI > Don'ts · Checked 2026-09-30
  7. 7
    AWS states that transcription accuracy affects the accuracy of Amazon Connect generative AI evaluations, and that conversations in more than one language or with multiple parties speaking at once are not transcribed accurately.
    Evaluate agent performance in Connect Customer using generative AI · Limitations of generative AI-powered performance evaluations · Checked 2026-09-30
  8. 8
    Amazon Connect defines goal success rate as the proportion of sessions in which the Orchestration AI agent successfully resolved customer issues.
    AI agent traces using Contact search and Contact details · AI agent performance metrics > Goal success rate · Checked 2026-09-30
  9. 9
    AWS recommends keeping a manual evaluation process in place alongside AI-filled evaluations in Amazon Connect, to catch drift between AI-filled and manager-filled evaluations over time.
    Evaluate agent performance in Connect Customer using generative AI · Process to automate evaluations using generative AI > Note · Checked 2026-09-30
  10. 10
    Amazon Connect advises allowing up to 30 minutes after a contact ends for the Automated Interaction Log, and the toggle that shows flow and trace details, to become available.
    AI agent traces using Contact search and Contact details · Navigate to AI agent trace details > Note · Checked 2026-09-30
  11. 11
    Amazon Connect states that its AI agent performance metrics become available after 24 hours.
    AI agent traces using Contact search and Contact details · AI agent performance metrics > Note · Checked 2026-09-30
  12. 12
    For an AI agent prompt invocation, Amazon Connect trace details show model reasoning and tool call information, the input parameters passed to each tool call, and the latency of each span.
    AI agent traces using Contact search and Contact details · AI agent trace details > AI agents · Checked 2026-09-30
  13. 13
    Amazon Connect AI agent traces can carry an ESC label marking where the AI agent escalated to a human agent and an ERR label marking a span that received an error status, such as AI barge in or timeout.
    AI agent traces using Contact search and Contact details · AI agent trace details > AI agents, 'The following labels might appear in the trace' · Checked 2026-09-30
  14. 14
    Amazon Connect documents four AI agent performance metrics alongside AI agent traces: completeness score, faithfulness score, goal success rate and tool utilization accuracy, each a value between 0 and 1.
    AI agent traces using Contact search and Contact details · AI agent performance metrics · Checked 2026-09-30
  15. 15
    Amazon Connect instances that enabled 'Bot Analytics, Transcripts, and AI Agent Traces in Amazon Connect' before June 5, 2026 must disable and re-enable that setting to activate AI agent traces.
    AI agent traces using Contact search and Contact details · Enable AI agent trace details > step 4 Note · Checked 2026-09-30
  16. 16
    Amazon Connect AI agent trace details are documented as currently available for the voice channel.
    AI agent traces using Contact search and Contact details · AI agent trace details > Note · Checked 2026-09-30
  17. 17
    Copilot Studio's production Answer quality measure uses AI on a sample of answered questions, assessing completeness, relevance and groundedness, and labels each answer Good or Poor.
    Monitor conversational agents - Microsoft Copilot Studio · Use > Generated answer rate and quality · Checked 2026-09-30
  18. 18
    Copilot Studio evaluation runs can be triggered programmatically through REST APIs or Power Platform connectors, which Microsoft describes as enabling integration into CI/CD pipelines.
    About agent evaluation - Microsoft Copilot Studio · Integrate evaluations into automated flows · Checked 2026-09-30
  19. 19
    Copilot Studio's Compare meaning test method has a default passing score of 50, and a test case without an expected answer produces an Invalid result for that method.
    Choose evaluation methods - Microsoft Copilot Studio · Compare meaning · Checked 2026-09-30
  20. 20
    The Copilot Studio conversational test set article says conversation test sets can use the General quality, Keyword match, Tool use or Custom test methods.disputed
    Create a conversational test set - Microsoft Copilot Studio · Create a conversation test set > step 5, first sentence · Checked 2026-09-30
  21. 21
    The Copilot Studio test method table marks four methods as available for conversation test sets (General quality, Content safety, Keyword match and Custom) and marks Tool use as single response only.disputed
    Choose evaluation methods - Microsoft Copilot Studio · test method table, 'Test set type' column; section openers for Content safety and Tool use · Checked 2026-09-30
  22. 22
    A Copilot Studio conversation test set supports up to 20 test cases, each with up to 12 total messages (6 question and answer pairs).
    Create a conversational test set - Microsoft Copilot Studio · Create a conversation test set > step 3 Note · Checked 2026-09-30
  23. 23
    Copilot Studio's Custom test method is configured with evaluation instructions and two or more labels, each label assigned a Pass or Fail result that counts toward the test set pass rate.
  24. 24
    Copilot Studio splits escalated sessions into System intended, System unintended and User requested, and describes System intended escalation as an expected outcome that does not need investigation.
    Monitor conversational agents - Microsoft Copilot Studio · Effectiveness > Conversation outcomes, outcome table rows under 'Escalated' · Checked 2026-09-30
  25. 25
    Copilot Studio's General quality test method uses a large language model to score responses on relevance, groundedness, completeness and abstention, does not require expected answers, and is added to every test set by default.
    Choose evaluation methods - Microsoft Copilot Studio · General quality · Checked 2026-09-30
  26. 26
    Microsoft states that generating Copilot Studio test questions from the agent's existing knowledge sources or topics is good for testing how the agent uses what it already has but is not good for testing information gaps.
    Create a single response test set - Microsoft Copilot Studio · Generate a single response data set from knowledge or topics, first paragraph · Checked 2026-09-30
  27. 27
    Copilot Studio counts an engaged session as Abandoned when it times out after 30 minutes without reaching a resolved or escalated state.
    Monitor conversational agents - Microsoft Copilot Studio · Effectiveness > Conversation outcomes, outcome table row 'Abandoned' · Checked 2026-09-30
  28. 28
    Copilot Studio analytics counts a session as Escalated when the Escalate topic is triggered or a Transfer to agent node runs, whether or not the conversation actually transfers to a live agent.
    Monitor conversational agents - Microsoft Copilot Studio · Effectiveness > Conversation outcomes, outcome table row 'Escalated' · Checked 2026-09-30
  29. 29
    Copilot Studio's General quality method can flag a correct refusal for improvement because its abstention criterion checks whether the agent attempted to answer; Microsoft directs users to a Custom test method to treat an appropriate refusal as a pass.
    Choose evaluation methods - Microsoft Copilot Studio · General quality, paragraph beginning 'If a test case expects the agent to refuse to answer' · Checked 2026-09-30
  30. 30
    Copilot Studio counts a session as Resolved implied without any user confirmation; under generative AI orchestration this happens when the session times out with no remaining active plans.
    Monitor conversational agents - Microsoft Copilot Studio · Effectiveness > Conversation outcomes, outcome table row 'Resolved implied' · Checked 2026-09-30
  31. 31
    Copilot Studio evaluation test results are available in the product for 89 days; Microsoft directs users to export results to CSV to keep them longer.
    Create a single response test set - Microsoft Copilot Studio · introduction > Important · Checked 2026-09-30
  32. 32
    Microsoft states that Copilot Studio's content safety evaluators (hatefulness and unfairness, sexual content, violence, self-harm) do not guarantee that an agent is safe or appropriate for every scenario.
    About agent evaluation - Microsoft Copilot Studio · Why use automated testing?, final paragraph · Checked 2026-09-30
  33. 33
    Microsoft presents re-running the same Copilot Studio test set after agent changes as the way to get an objective standard for comparing performance before and after.
    About agent evaluation - Microsoft Copilot Studio · How agent evaluation works, test set bullet list · Checked 2026-09-30
  34. 34
    A Copilot Studio single response test set can contain up to 100 test cases, and an imported file up to 100 questions of at most 1,000 characters each.
    Create a single response test set - Microsoft Copilot Studio · introduction; Create a test set file to import > Important · Checked 2026-09-30
  35. 35
    In Microsoft Copilot Studio agent evaluation, a test case is a single simulated user interaction, either one question or an entire conversation, and a group of test cases is a test set.
    About agent evaluation - Microsoft Copilot Studio · How agent evaluation works · Checked 2026-09-30
  36. 36
    Copilot Studio agent evaluation offers eight test methods: General quality, Content safety, Compare meaning, Tool use, Keyword match, Text similarity, Exact match and Custom.
    Choose evaluation methods - Microsoft Copilot Studio · test method table at top of page · Checked 2026-09-30
  37. 37
    Dialogflow CX continuous deployment runs the same set of verification tests before a flow version is deployed to an environment, with the documented aim of preventing a bad version from going live.
    Continuous tests and deployment | Dialogflow CX | Google Cloud Documentation · page introduction, continuous deployment sentence · Checked 2026-09-30
  38. 38
    The Dialogflow CX continuous tests feature automatically runs a set of test cases configured for an environment to verify the behaviour of the flow versions in that environment.
    Continuous tests and deployment | Dialogflow CX | Google Cloud Documentation · page introduction, continuous tests definition · Checked 2026-09-30
  39. 39
    In Dialogflow CX legacy test cases, a difference in agent dialogue between the golden test case and the latest run produces a warning and does not prevent the test from passing.
    Test cases | Dialogflow CX | Google Cloud Documentation · Legacy test cases > Run test cases, 'Agent dialogue' item · Checked 2026-09-30
  40. 40
    In Dialogflow CX legacy test cases, the matched intent must be the same on every turn for the test to pass.
    Test cases | Dialogflow CX | Google Cloud Documentation · Legacy test cases > Run test cases, matched intent item · Checked 2026-09-30
  41. 41
    In Dialogflow CX legacy test cases, the active page must be the same on every turn for the test to pass.
    Test cases | Dialogflow CX | Google Cloud Documentation · Legacy test cases > Run test cases, current page item · Checked 2026-09-30
  42. 42
    Dialogflow CX playbook evaluations report three built-in metrics: semantic similarity, tool call accuracy and latency.
    Playbook evaluations | Dialogflow CX | Google Cloud Documentation · metrics description (heading text not captured) · Checked 2026-09-30
  43. 43
    The Dialogflow CX playbook semantic similarity metric compares the agent's conversation with a golden response and takes a value of 0, 0.5 or 1.
    Playbook evaluations | Dialogflow CX | Google Cloud Documentation · metrics description, semantic similarity · Checked 2026-09-30
  44. 44
    Google documents the Dialogflow CX built-in test case feature as a way to uncover bugs and prevent regressions.
    Test cases | Dialogflow CX | Google Cloud Documentation · page introduction, before the first section heading · Checked 2026-09-30
  45. 45
    Because Dialogflow CX legacy test cases pass when only the agent's dialogue text differs, they do not by themselves catch a regression that changes what the agent says without changing the intent, page or tracked parameters.inferred
    Test cases | Dialogflow CX | Google Cloud Documentation · Legacy test cases > Run test cases, comparison list · Checked 2026-09-30
  46. 46
    The Dialogflow CX playbook tool call accuracy metric reflects how faithfully the conversation includes the tools expected to be invoked, on a range of 0 to 1.
    Playbook evaluations | Dialogflow CX | Google Cloud Documentation · metrics description, tool call accuracy · Checked 2026-09-30
  47. 47
    In Dialogflow CX legacy test cases, if tracking parameters were added when the test case was created, the test fails on missing, unexpected or mismatched session parameter values.
    Test cases | Dialogflow CX | Google Cloud Documentation · Legacy test cases > Run test cases, session parameters item · Checked 2026-09-30
  48. 48
    The Genesys Cloud Bot Performance Summary view separates bot sessions that ended in a disconnect into Customer Disconnect, Bot Disconnect and System Error Disconnect.
    View available columns in performance views by category - Genesys Cloud Resource Center · Bot Performance Summary view > Disconnects · Checked 2026-09-30
  49. 49
    In the Genesys Cloud Bot Performance Summary view, the User Exit column counts bot sessions where the user requested to speak to an agent.
    View available columns in performance views by category - Genesys Cloud Resource Center · Bot Performance Summary view > Exits > User Exit · Checked 2026-09-30
  50. 50
    In a Lex V2 Test Workbench conversation test, an intent mismatch causes the remaining turns of that conversation to be skipped, whereas a slot value mismatch or a transcription mismatch lets the remaining turns continue.
    Test results details in Test Workbench · failure error message table, rows 'Intent Mismatch', 'Slot value mismatch', 'Transcription Mismatch' · Checked 2026-09-30
  51. 51
    In Lex V2 Test Workbench conversation intent failure metrics, a successful intent does not mean the whole conversation succeeded; the metric considers each intent's value regardless of the intents before or after it.
    Test results details in Test Workbench · Conversation results tab > Conversation intent failure metrics · Checked 2026-09-30
  52. 52
    A Lex V2 test set is generated either by uploading a CSV file or from conversation logs, and can contain audio or text input.
    Generate a test set for Test Workbench · first paragraph under the page title · Checked 2026-09-30
  53. 53
    The Amazon Lex V2 Test Workbench establishes baseline bot performance covering intent and slot performance, for utterances given as single inputs or as conversations.
    Evaluating Lex V2 bot performance with the Test Workbench · third paragraph under the page title · Checked 2026-09-30
  54. 54
    NIST AI RMF subcategory MEASURE 2.3 calls for AI system performance or assurance criteria to be measured, qualitatively or quantitatively, and demonstrated for conditions similar to the deployment setting.
    Measure - AIRC (NIST AI RMF Playbook) · MEASURE 2.3 subcategory statement · Checked 2026-09-30
  55. 55
    NIST AI RMF subcategory MEASURE 2.4 calls for the functionality and behaviour of an AI system and its components to be monitored when in production.
    Measure - AIRC (NIST AI RMF Playbook) · MEASURE 2.4 subcategory statement · Checked 2026-09-30
  56. 56
    NIST AI RMF subcategory MEASURE 2.5 calls for the AI system to be demonstrated valid and reliable before deployment and for limits on its generalisability beyond development conditions to be documented.
    Measure - AIRC (NIST AI RMF Playbook) · MEASURE 2.5 subcategory statement · Checked 2026-09-30
  57. 57
    The vendor outcome measures documented here are defined differently (Copilot Studio session outcomes with implied resolution by timeout, Amazon Connect goal success rate per session, Genesys exit and disconnect counts), so a containment or resolution figure from one platform is not directly comparable with another's.inferred
    Monitor conversational agents - Microsoft Copilot Studio · Effectiveness > Conversation outcomes, outcome table · Checked 2026-09-30
  58. 58
    Webex AI Agent Studio provides a chat preview and a voice preview in which an administrator interacts with the AI agent as an end user and observes its responses.
    Webex AI Agent Studio Administration guide · preview sections (chat preview; voice preview) · Checked 2026-09-30
  59. 59
    The Webex AI Agent Studio Sessions page offers filters for hiding test sessions, for sessions where agent handover happened, for sessions where an error occurred, and for downvoted sessions.
    Webex AI Agent Studio Administration guide · Sessions section, filter list · Checked 2026-09-30
  60. 60
    Webex AI Agent Studio stores session transcripts for 90 days, after which they are automatically purged.
    Webex AI Agent Studio Administration guide · Sessions section, transcript retention sentence · Checked 2026-09-30

Documents

tier 1 standards and regulators

Measure - AIRC (NIST AI RMF Playbook)

National Institute of Standards and Technology · accessed 2026-09-30

tier 2 current vendor documentation

About agent evaluation - Microsoft Copilot Studio

Microsoft · 2026-09-16 · accessed 2026-09-30

tier 2 current vendor documentation

AI agent traces using Contact search and Contact details

Amazon Web Services · accessed 2026-09-30

tier 2 current vendor documentation

Choose evaluation methods - Microsoft Copilot Studio

Microsoft · 2026-09-16 · accessed 2026-09-30

tier 2 current vendor documentation

Continuous tests and deployment | Dialogflow CX | Google Cloud Documentation

Google Cloud · 2026-09-24 · accessed 2026-09-30

tier 2 current vendor documentation

Create a conversational test set - Microsoft Copilot Studio

Microsoft · 2026-03-12 · accessed 2026-09-30

tier 2 current vendor documentation

Create a single response test set - Microsoft Copilot Studio

Microsoft · 2026-08-28 · accessed 2026-09-30

tier 2 current vendor documentation

Evaluate agent and self-service interaction performance in Connect Customer

Amazon Web Services · accessed 2026-09-30

tier 2 current vendor documentation

Evaluate agent performance in Connect Customer using generative AI

Amazon Web Services · accessed 2026-09-30

tier 2 current vendor documentation

Evaluating Lex V2 bot performance with the Test Workbench

Amazon Web Services · accessed 2026-09-30

tier 2 current vendor documentation

Generate a test set for Test Workbench

Amazon Web Services · accessed 2026-09-30

tier 2 current vendor documentation

Monitor conversational agents - Microsoft Copilot Studio

Microsoft · 2026-06-01 · accessed 2026-09-25

tier 2 current vendor documentation

Playbook evaluations | Dialogflow CX | Google Cloud Documentation

Google Cloud · 2026-09-24 · accessed 2026-09-30

tier 2 current vendor documentation

Test cases | Dialogflow CX | Google Cloud Documentation

Google Cloud · 2026-09-24 · accessed 2026-09-30

tier 2 current vendor documentation

Test results details in Test Workbench

Amazon Web Services · accessed 2026-09-30

tier 2 current vendor documentation

View available columns in performance views by category - Genesys Cloud Resource Center

Genesys · accessed 2026-09-30

tier 2 current vendor documentation

Webex AI Agent Studio Administration guide

Cisco Systems, Inc. (Webex Help Center) · 2026-09-03 · accessed 2026-09-04

Cite this page

APA

WarmTransfer. (2026, September 30). Testing and evaluating AI virtual agents. WarmTransfer. https://warmtransfer.net/knowledge/ai-agent-evaluation

BibTeX

@misc{warmtransfer-ai-agent-evaluation,
  title  = {Testing and evaluating AI virtual agents},
  author = {{WarmTransfer}},
  year   = {2026},
  url    = {https://warmtransfer.net/knowledge/ai-agent-evaluation},
  note   = {Verified 2026-09-30}
}