Domain 08 · Eval & Testing
Scenario

An invoice extractor outputs a confidence score per field — {vendor: "Acme Corp" (0.89), date: "2024-05-04" (0.92), amount: 1287.42 (0.71), invoice_number: "INV-9938" (0.68)}. The team currently routes any extraction with any field below 0.85 to human review. About half the time, those routed extractions turn out to be entirely correct. Meanwhile, several auto-approved extractions (all fields above 0.85) have turned out to be wrong.

What's the right fix?

Choose your answer
Before you reveal — write your reasoning

Why did you pick that answer? Two or three sentences. The act of articulating it is what builds the judgment — not the click that follows.

Saved locally to your device.
Pick an option and write a sentence of reasoning to enable.