We used to grade call quality by a supervisor randomly picking a few calls a week and going through them by ear, which meant most calls never got looked at and grading was honestly a bit inconsistent depending on who was doing it that week. Automated scorecards against our own criteria go through way more calls now, consistent scoring every time instead of depending on which supervisor happened to be reviewing that day. Caught a pattern where one agent consistently skipped a required disclosure step, something that had been happening for weeks without anyone noticing since it never landed in the small random sample being manually checked. Review collected by and hosted on G2.com.
Setting the actual criteria weights took real back and forth, our first version was too strict on minor phrasing differences and flagged agents for things that honestly didn't matter much in the bigger picture. Review collected by and hosted on G2.com.