
As agents, having access to our call recordings and QA audit details is an invaluable resource. These tools not only allow us to assess where we currently stand in terms of call quality, but also provide insights on how we can continuously improve. Reviewing our calls helps us identify strengths and areas for growth, particularly in enhancing our communication skills and refining our approach during customer interactions Review collected by and hosted on G2.com.
AI evaluation is a helpful tool, and I understand that this feature in Level AI is still evolving. As with any new technology, there’s ongoing development to make it more accurate and effective.
At present, AI tends to evaluate calls in a very literal manner. For example, in the hold etiquette rubric, if the system does not detect the exact word "hold," it may mark the agent down even if the agent clearly asked permission using alternative phrases such as, "Can you stay on the line while I try to reach out to your caterer?" This shows how the current AI model may not yet fully capture context or variations in language.
Also, AI sometimes flags calls that go beyond 12 minutes, regardless of complexity. In reality, certain calls involve uncontrollable factors or intricate issues that naturally require more time. For this reason, it might be more effective if AHT (Average Handle Time) is not strictly used as a markdown criterion, as each interaction varies depending on the situation.
I believe with continuous refinement, AI evaluation will become more flexible and accurate in recognizing context, tone, and intent, helping both agents and quality teams achieve more balanced assessments. Review collected by and hosted on G2.com.