Your customer support agent has been running for six months at 87% first-contact resolution. The board has been impressed by the metrics. Your VP of Operations is under pressure to free up reviewer capacity for a new initiative and proposes reducing human review on all "high-confidence" cases — those where the agent self-reports confidence above 0.9.
Aggregate accuracy on high-confidence cases is 97%. The proposal would cut review costs significantly. Your VP wants to roll this out next month.
What's your most defensible position?
Why did you pick that answer? Two or three sentences. The act of articulating it is what builds the judgment — not the click that follows.
Aggregate accuracy is one of the most dangerous metrics to act on without stratification. A system that's 97% overall might be 99% on simple cases that dominate the volume and 73% on complex cases where errors are most costly. Before reducing review, you need to understand the distribution of errors — by case type, dollar amount, customer segment. Often you'll find that aggregate metrics are masking concentrated risk in exactly the segments where you most need human review. The defensible position is: "the 97% looks great, and before we act on it we need to verify the errors aren't concentrated where they hurt most."
This is the genuinely tempting answer under organizational pressure — the metrics look great, the VP needs the capacity, delay looks like timidity. The misconception is treating aggregate accuracy as if it tells you about the distribution of errors. It doesn't. A system can have excellent aggregate accuracy while concentrating its failures on the highest-stakes cases. "97% is good enough" is a reasoning style that produces predictable, expensive incidents.
Phasing the rollout is operationally responsible but doesn't address the actual question: do you know whether the 97% is uniform or concentrated? If errors are concentrated in high-stakes cases, phasing just spreads the problem across two months instead of one. Phasing is a deployment strategy; stratification is an analysis prerequisite. The phased rollout would be appropriate after stratification confirms the metrics are uniformly good.
Refusing on principle without doing the analysis is no more defensible than approving on principle. "Never automate this" forfeits real capacity gains where they're warranted, and it doesn't model the kind of analytical thinking the situation requires. The right answer is conditional on what stratification reveals.