A VP asks: "Does your support agent work?" The team responds: "It feels like it's working really well." The VP wants metrics. The team needs to put together a dashboard.
What should the dashboard surface?
Why did you pick that answer? Two or three sentences. The act of articulating it is what builds the judgment — not the click that follows.
"Does it work?" is multiple questions. A VP asking that wants to know: is it correct, does it know its limits, do customers find it useful, is it affordable, is it fast. Each is a different metric measuring a different dimension. A dashboard that answers them separately lets stakeholders zoom into the dimension they care about. A single bundled metric flattens all the questions into one number that no one fully trusts.
A "success rate" metric bundles correctness, escalation behavior, and resolution speed into one number that responds to all three. When it changes, you can't tell why. Stakeholders end up not trusting it because it doesn't answer their specific question.
Customer satisfaction is important but lagging — it tells you about user perception weeks after the fact, not about correctness or operational health. By the time satisfaction drops, the underlying problems have already produced bad outcomes. Lead with correctness; satisfaction is one of several outcomes.
The agent's self-reported confidence measures how the agent feels about its answers — not how correct they are. Self-reported confidence is poorly calibrated, especially on cases the agent gets wrong: by definition, the model doesn't know it's wrong, or it wouldn't be reporting that confidence. Reporting it as a quality metric is misleading because the metric correlates poorly with actual correctness.