Part 2: Introduction
Last week in Part 1 we discussed the fact that as AI becomes central to contact center operations, powering every customer engagement channel, evaluation is no longer a back-office technical exercise. Evaluation is a critical business capability directly impacting customer experience, operational effectiveness, and business outcome.
Today in Part 2, we will look at the difference between synthetic data and production traffic, as well as the evaluation after test runs or real-time executions in the contact centers.
4. Synthetic Data vs. Production Traffic
Definition: Synthetic data refers to artificially generated datasets designed to simulate specific scenarios. Production traffic comprises actual user interactions and data generated during live operation.
Implications: Data Fidelity and Risk: Synthetic data enables safe, repeatable evaluation without exposing sensitive information or impacting real users. However, it may lack the complexity and unpredictability of production data. Production traffic delivers high-fidelity insights but carries risks of data leakage, performance degradation, or user impact. Relevance: Synthetic data is valuable for early-stage, edge-case, or privacy-sensitive evaluations. Production traffic is essential for verifying AI system behavior under real-world conditions.
How contact centers should think about this:
- Begin with synthetic data to evaluate safely and iterate quickly, especially when testing new scenarios, edge cases, or changes.
- Leverage production data to validate performance at scale, ensuring AI behaves as expected under real customer traffic and operating conditions.
- Treat production evaluation as a continuous monitoring and learning loop, focused on measuring impact and improving quality, rather than experimenting on live customers.
5. Evaluation After vs. During Execution
Definition: Post-execution evaluation analyzes the results after a process or test run finishes, while in-execution (real-time) evaluation monitors and assesses behavior as it unfolds.
Implications: post-execution evaluation enables deep analysis and long-term improvement, while in-execution evaluation allows faster detection and mitigation of issues. Using both helps contact centers balance insight with real-time protection of customer experience.
How contact centers should think about this:
- Post-conversation evaluation can provide a large amount of information about correctness, groundedness, and resolution effectiveness across completed AI interactions.
- Real-time evaluation of empathy and sentiment enables timely intervention, such as escalating to a human agent or allowing supervisor guidance during the interaction
Together, these approaches form a core part of AI evaluation in the contact center, helping organizations balance deep analysis with real‑time protections.
Final Thoughts: A Modern Evaluation Mindset
There is no single “right” way to evaluate AI systems. Instead, evaluation should be viewed as a multi-dimensional strategy that evolves alongside your AI systems.
By thoughtfully strategizing across evaluation dimensions, organizations can build AI systems that are not only intelligent, but also trustworthy, resilient, and customer-first.
Evaluation is no longer optional; it is how modern organizations ensure AI delivers on its promise, every day.
Source: Microsoft Dynamics 365
Article: By Raimondas Lencevicius (Principal Research Scientist)