Part 1: Introduction
As AI becomes central to contact center operations, powering every customer engagement channel, evaluation is no longer a back-office technical exercise. Evaluation is a critical business capability directly impacting customer experience, operational effectiveness, and business outcome
But the evaluation is not one-dimensional. Organizations must think about when, how, by whom, and on what data AI systems are evaluated. This blog explores the key facets of AI evaluation and how they apply specifically to contact center environments.
1. Development-Stage vs. Production-Stage Evaluation
Definition: Evaluation at development time refers to tests and assessments performed during the AI software creation phase. Production time evaluation occurs after deployment, when the AI software is live and serving real users.
Implications: Development-time evaluation provides a controlled environment to identify and address issues before AI reaches customers. This helps reduce risk and prevent costly failures. However, it cannot fully capture real-world complexity. Production-time evaluation reflects actual customer behavior and operating conditions. While it offers critical insight into true performance and experience, manage it carefully to avoid customer impact when issues surface.
How contact centers should think about this:
- Use development-time evaluation to prevent issues. These include incorrect intent detection, poor prompt behavior, broken escalation flows, non-compliant responses, or unacceptable latency before they ever reach customers.
- Use production-time evaluation to detect and measure real customer impact, like drops in containment, rising transfers to human agents, customer frustration, regional or channel-specific issues, and performance degradation caused by real traffic patterns
2. Manual vs. Automated Evaluation
Definition: Manual execution involves running evaluation tasks at human command. Automated evaluation is run at predetermined time or based on triggers.
Implications: Manual evaluation brings human judgment, context, and nuance that automation alone cannot capture, making it especially valuable when evaluation needs are unpredictable, environmental changes are not captured by automated triggers, or when automated runs would be too costly. Automated evaluation complements this, providing consistent, scalable coverage as AI systems evolve, reliably reevaulating systems after known changes, such as releases or configuration updates through CI/CD pipelines.
How contact centers should think about this:
- Use automation for baseline quality, regression, and continuous monitoring
- Use manual evaluation for exceptions, deep dives, and human judgment
- The best strategy combines both. Automation ensures assessment of critical changes, while human evaluators interpret results, investigate anomalies, and adapt to unexpected conditions.
3. Evaluations Run by the Platform vs. by Customers
Definition: Evaluations can be conducted internally by the developer organization (such as Microsoft) or externally by customers using the software in their own environments.
Implications: Developer-run and customer-run evaluations each provide distinct and necessary value. Internal evaluations establish a consistent baseline for quality, safety, and compliance. Customer-led evaluations surface real-world behaviors, operational constraints, and usage patterns that cannot be fully anticipated during development. Relying on only one limits visibility and can leave gaps in reliability or usability.
How contact centers should think about this:
- Rely on platform evaluations to establish a trusted baseline. This ensures core capabilities—such as accuracy, safety and compliance, latency, escalation behavior, and failure handling—meet enterprise standards before features are rolled out broadly.
- Platform providers should partner closely with customers. This enables them to run their own evaluations and deeply understand AI performance within their specific domains, workflows, and operating environments. This collaboration helps surface both expected and edge-case behaviors across positive and negative scenarios.
➡️ Next week in Part 2, we will look at the different types of data being handled by a contact centers.
Source: Microsoft Dynamics 365
Article: By Raimondas Lencevicius (Principal Research Scientist)