Anthropic’s cost-per-task evidence
Check the effort settings and evaluated tasks behind the savings claim. Include unsuccessful attempts and review when repeating the comparison.
Compute College
Measure the total workload cost required to produce an accepted AI result.
Cost per successful outcome is what a workload spends to produce one accepted result. It includes model calls, tools, retries, review, and relevant infrastructure, so it can tell a different story from the price of a single request.
Memory trick: The product pays for accepted work, not attempted calls.
A low token price can produce an expensive workflow if the model needs retries or human correction. A more capable model can be economical if it raises first-pass acceptance enough to reduce downstream work. The denominator should reflect the product result users actually value.
A workload makes 1,000 requests at $0.04 each, 180 retries at $0.02, and 80 minutes of review valued at $0.50 per minute. If 920 results are accepted, the $56 total is about $0.061 per successful outcome, not $0.04 per request.
Example figures are illustrative calculations, not current quoted market prices.
Current example
Sonnet 5.5 keeps Sonnet 5’s listed standard rates of $2 per million input tokens and $10 per million output tokens. Anthropic nevertheless reports up to 30% lower task costs in its testing, describing fewer tokens and changes in tool use. The teaching point is that the bill depends on how much work the model performs. Those reported savings do not predict your application’s result.
Check the effort settings and evaluated tasks behind the savings claim. Include unsuccessful attempts and review when repeating the comparison.
Use the billing categories that apply to your traffic, rather than treating all tokens as one rate.
Sources reviewed and recorded Oct 6, 2026. “Up to 30%” describes Anthropic’s reported tests; it is not a guaranteed saving or an independent result. Your denominator is accepted outcomes, and your numerator includes the whole workload.
Provider spend divided by request count leaves out review, retries, and failed outcomes. That shortcut makes an unreliable system look cheaper than it is.
Practical takeaway
Define the accepted outcome and build a cost ledger for one workload. Include model calls, tools, retries, review, fixed capacity, and the number of accepted results.
Decision check: does the metric include every material cost required to turn a request into the result the product promises?
Compute College learning path
Step 48 of 48: Cost per successful outcome