08
Domain 08
Evaluation, Testing & LLM-as-Judge
Building test suites, calibrating LLM-as-Judge, measuring consistency, and answering "does it work?" with data.
8 questions
- Q84What's the most likely cause?After a month, the LLM-as-Judge's scores have drifted upward — same agent, more 8s and 9s than before.
- Q85What should they do before deploying?A team is deploying LLM-as-Judge into production KPI reporting. The VP will see the dashboard. What should they do first?
- Q86What's the problem with this test suite?20 test cases all passing. The team declares the agent ready for production. What's missing?
- Q87What's the right framing?Same query × 5 runs, 5 slightly different responses. Same recommendation, different framings. Problem or not?
- Q97What's the right addition?$0.02 to $2.50 per query. Monthly bills swing wildly. The CFO wants to know what's driving the variance.
- Q98What's the right fix?Half the routed-to-review extractions turn out correct; several auto-approved ones turn out wrong. Single global threshold.
- Q99What should the dashboard surface?VP asks "does your agent work?" The team says "feels like it's working well." The VP wants metrics.
- Q100What's the most defensible pick?Two weeks to launch. Six things to do, time for two. Pick the two that prevent the worst outcomes.