Learn AI compute, then follow the market
← Back to Compute College

Compute College

GPT-6.1 Sol: benchmarks, pricing, and workload fit

Understand GPT-6.1 Sol through task quality, standard API pricing, cached context, reasoning effort, and cost per accepted result.

Plain-English definition

GPT-6.1 Sol is OpenAI’s September 29, 2026 model release for coding, computer use, and professional work. Its benchmark results describe performance under particular test conditions. To choose it for an application, compare accepted results, token use, and response time on the work your users actually need.

Memory trick: A model’s sticker price buys tokens. Your product needs finished work.

Why it matters

OpenAI lists standard short-context input at $2 per million tokens and output at $10 per million. Astra’s corresponding rates are $10 and $50. That price gap makes Sol worth testing for work you currently send to Astra. It does not establish that Sol produces the same answers, takes the same number of steps, or meets the same deadline. A cheaper call can still need a costly retry.

  • A model that meets your quality threshold at a lower total cost can make a previously expensive feature affordable. Test the difficult cases as well as the routine ones before moving traffic.
  • OpenAI reports improvements across coding, document work, and computer use. Treat those results as evidence for a shortlist. The prompts, tools, effort settings, and grading rules affect what a score means.
  • Cached input has a different price from fresh input. Repeated instructions may reduce spend when they qualify for caching, while changing documents and newly generated answers continue to create paid work.

Simple example

Consider a request with 100,000 fresh input tokens and 20,000 output tokens. At the reviewed standard short-context rates, Sol costs $0.20 for input and $0.20 for output: $0.40. Astra costs $1 plus $1: $2. If both produce an accepted result on the first attempt, Sol has a lower token bill. If Sol fails and the system sends the full task to Astra, that path costs $2.40 before other charges.

  • This is illustrative arithmetic using published rates, not a measured Sol-versus-Astra workload. It excludes cache writes, tools, retries beyond the stated fallback, and reviewer time.
  • For an eligible cache read of all 100,000 input tokens, Sol’s listed $0.10-per-million cached-input rate gives a $0.01 input charge. Check how your actual service creates and bills its cache.
  • Keep short-context and long-context prices separate. Batch, Flex, Fast, and Ultrafast are also distinct pricing modes; do not apply a standard rate to traffic billed under another mode.

Example figures are illustrative calculations, not current quoted market prices.

Current example

What the September release establishes

OpenAI’s launch page describes GPT-6.1 Sol as approaching Astra on selected evaluations at lower cost. It discusses DeepSWE, GDP.pdf, AutomationBench, and OSWorld, among other tests. These assess different kinds of work, so their scores cannot be combined into a single success rate for your application. The release also announced Sol Ultrafast as forthcoming; that announcement alone does not establish current access.

API pricing by mode and context

Verify fresh input, cached input, cache writes, and output rates before building a budget.

Deployment safety evaluation

Read the system-card addendum for safety tests and their conditions. A test result is not a permission policy for your agent.

Release and pricing reviewed and recorded Oct 6, 2026. Benchmark comparisons are OpenAI’s published evidence, not independent Compute College measurements. Standard short-context prices are used here; other modes and context bands differ.

Common mistake

One-fifth of the token price does not mean one-fifth of your production bill. An agent can generate more reasoning, repeat tool calls, or escalate to another model. A benchmark can also use a different environment from your application. Keep the whole task in the cost record, including failed attempts and any premium fallback.

Practical takeaway

What you can do with this

Choose one workflow with a clear acceptance rule and replay a representative set of saved tasks on Sol and your current model. Keep the instructions, available tools, and review process consistent. Record the model version and effort setting beside each result. Compare accepted-task cost and completion time, then inspect failures before deciding which traffic to move.

  • Include ambiguous cases, long documents, and tool errors in the sample. A model that looks economical on easy prompts may behave differently when the task requires several steps.
  • Count the entire trajectory: fresh and cached tokens, generated tokens, tool charges, retries, and escalation. Keep human review visible even if it sits outside the provider invoice.
  • Start with a bounded task category and preserve an evaluated fallback. Expand only when the quality threshold and response-time requirement hold under realistic traffic.

Can Sol finish this workload within the agreed quality, time, and cost limits after failures and fallback calls are counted?

Compute College

Follow model releases as market signals

Follow model releases as AI compute market signals in the ComputeTape Market Brief.

Read the Market Brief

Compute College track

Model Benchmarks & AI Compute Economics

Step 15 of 27: GPT-6.1 Sol: benchmarks, pricing, and workload fit