04
Domain 04
Prompt Engineering & Structured Output
Schema constraints over instructions, categorical criteria over adjectives, examples for format variation.
22 questions
- Q16What's the most reliable fix?Your invoice extractor returns malformed JSON 5% of the time, crashing the downstream parser.
- Q17What's the right schema-level fix?A discontinued_date field that doesn't exist in most source documents — and the model is fabricating values to fill it.
- Q18What's the most effective revision?A code review agent flagging false positives — the prompt politely asks it to "be conservative."
- Q19Which intervention will most improve it?Extracting research methodologies from papers with wildly different document structures.
- Q20What's the right call?50% cost savings via the Message Batches API — for which workloads is that the right call?
- Q41What's the highest-leverage improvement?Invoice extraction fails 8% of the time. Blind retries recover 60%. Can we do better?
- Q42Will retry-with-feedback help here?A receipt photo is genuinely cut off, hiding the merchant name. Will retry-with-feedback help?
- Q43How can the schema design itself surface these errors at extraction time?Line items don't sum to the stated total. Days later, accounting catches it as an anomaly.
- Q44What's the best schema improvement?20% of expenses end up in "Other" and the analyst can't tell what they actually were.
- Q45What's the right architectural fix?A 1,200-line proposal review missing cross-section inconsistencies and contradicting itself by page 30.
- Q66What's the right schema improvement?The code review agent's findings are dismissed 30% of the time — and the same patterns keep showing up.
- Q67What's the right submission cadence to keep your SLA safe?Anthropic's batch SLA is 24 hours. Your team's SLA is 30 hours. You submit one batch nightly.
- Q68What's the right approach?10,000 invoices submitted; 200 failed. The current plan: re-submit all 10,000 to ensure consistency.
- Q69What's the cleaner fix?An invoice extractor returns dates in whatever format appeared in the source — and downstream parses 5/4 sometimes as April, sometimes as May.
- Q70What's the most defensible approach?Independent verification absolutely catches errors. The question is whether to verify all 50,000 invoices when it doubles the cost.
- Q101What's the right framing?The same Claude that summarizes earnings reports brilliantly fails at arithmetic any calculator gets right. Why?
- Q102What's the most accurate explanation of what happened?A confident answer about top-three March 2024 movies — two of which didn't exist. What just happened?
- Q103What's the most likely cause?A consultant gets Q4 2024 numbers from Claude that turn out to be Q4 2023, presented as current.
- Q104Which revision applies CRAFT correctly?A junior analyst writes "Write a marketing email about our new product." The output is generic. CRAFT it.
- Q105When does few-shot beat zero-shot for this task?Eight feedback categories with overlap. "App crashes during upgrade" — bugs? pricing? features? Zero-shot or few-shot?
- Q106What's the right addition?Six style rules and the model still forgets some on long documents. What's the one move that integrates them?
- Q107For which task is chain-of-thought most beneficial?Invoice extraction vs refund eligibility. Both could use chain-of-thought — but only one actually benefits.