Your team uses the Message Batches API for nightly invoice extraction across 50,000 documents. Anthropic's SLA on batch processing is up to 24 hours. Your team's SLA to internal stakeholders is 30 hours from data availability. You currently submit one batch per night.
What's the right submission cadence to keep your SLA safe?
Why did you pick that answer? Two or three sentences. The act of articulating it is what builds the judgment — not the click that follows.
Plan for the SLA, not the median. Anthropic's 24-hour batch SLA is a worst-case guarantee — most batches finish much faster, but you can't plan around "usually fast." If your downstream SLA is 30 hours, you need a submission cadence that keeps you safe under the worst case. Submitting every 4–6 hours gives each batch its full SLA window plus padding before your downstream commitment.
"Most batches finish quickly" isn't a plannable property of a workload. The day a batch takes 22 hours is the day your stakeholders find out you were depending on luck.
Real-time API calls cost twice as much per token and don't address the actual question (matching submission cadence to SLA). Switching APIs without exhausting the cheaper tuning is premature.
Setting your internal SLA equal to the upstream SLA leaves zero buffer for your own processing, retry, or hand-off time. You'd miss the SLA almost every night even when batches finish on time.