GPT-6.1 Sol: Specs, Price, and Modelflare Access

A source-backed account of gpt-6.1-sol: published specs, short and long context prices, how to read the 29 September 2026 launch comparisons, and the Modelflare groups and user prices returned on 30 September 2026.

OpenAI released GPT-6.1 Sol on 29 September 2026. The API model ID is gpt-6.1-sol. The launch page describes it as an upgrade to gpt-6-sol: close to gpt-6-astra on complex coding, computer use, and professional work, at one-fifth of Astra's standard input and output prices. Cached input is $0.10 per million tokens, 5% of standard input and half of the gpt-6-sol cache read. For the hardest scientific research, the launch page still points to Astra. gpt-6-sol stays available.

This article keeps three layers apart. The specification and the prices come from OpenAI's launch page and model page. The benchmark comparisons copy the deltas, costs, and conditions the launch page actually states. Modelflare did not rerun them. Groups and user prices come from GET https://modelflare.dev/api/pricing read on 30 September 2026. A vendor phrase such as "about one-fifth of the cost" describes OpenAI's own harness.

What the sources support

The short-context list price is $2 input, $0.10 cache read, $2.50 cache write, and $10 output per million tokens. Above 272,000 input tokens, input and cache are doubled and output is multiplied by 1.5 for the whole request, which is $4 / $0.20 / $5 / $15. The Modelflare price row from the same day writes that same pair in two pricing_rules tiers. request_rules on that row is empty, so the catalog unit prices are not multiplied again by Fast, Batch, Flex, or a regional premium.

The public selectable groups are openai-award (0.03) and openai-stable (0.08). openai-premium is a public group, and this model is not attached to it. The model row also names openai-enterprise and non-public openai-first-topup. The public response does not include an enterprise ratio, so this article does not print an enterprise user price.

Published specification

Item GPT-6.1 Sol
Model ID gpt-6.1-sol
Release 29 September 2026, OpenAI API
Context window 1,050,000 tokens
Maximum input 922,000 tokens
Maximum output 128,000 tokens
Input to output Text and images to text
Knowledge cutoff 30 April 2026
reasoning.effort low, medium (default), high, xhigh, max
Efforts that are not accepted none, minimal
Tool calling Responses API
Chat Completions Supported, without tools
Endpoints listed Chat Completions, Responses, Batch
Stated as unsupported on the model page Realtime, Assistants, fine-tuning, embeddings, image generation, video, speech

The model page also says the model supports US and EU data residency, and that Fast is unavailable with EU data residency. The Modelflare price row read for this article has no residency switch. Read residency from the model page, and keep it separate from a gateway parameter.

Maximum input of 922,000 lines up with the 1,050,000 context window and the 128,000 maximum output: 1,050,000 − 128,000 = 922,000. Reasoning tokens occupy context and are billed as output tokens. max_output_tokens caps reasoning tokens, visible output, and formatting tokens together. When that cap is hit, the response status is incomplete and the reason is max_output_tokens.

Official prices

Short context and long context

Amounts are US dollars per million tokens. The short tier is the model page. The long tier is the rule stated there: above 272,000 input tokens, input and cache are doubled, output is multiplied by 1.5, and the whole request uses the long tier. The Sol and Astra columns come from those models' pricing_rules in the Modelflare pricing response of the same day. Their short tiers match the published list prices.

Per 1M tokens 6.1 short 6.1 long Sol short Sol long Astra short Astra long
Input $2 $4 $2 $4 $10 $20
Cache read $0.10 $0.20 $0.20 $0.40 $1 $2
Cache write $2.50 $5 $2.50 $5 $12.50 $25
Output $10 $15 $10 $15 $50 $75

The integer 272,000 stays on the short tier: the price-row condition is input_context_tokens <= 272000. From 272,001 upward the long tier applies, including the first 272,000 tokens. The 6.1 cache read is half of Sol. Standard input and output are one-fifth of Astra. The cache read is not one-fifth of Astra. Astra's short-context cache read is $1.

Other official tiers

The model page states three further official multipliers: Fast is 2× Standard, Batch and Flex are 50% below Standard, and regional processing adds 10% where it is available. The launch page also says that Ultrafast will be offered in Codex in the coming days, with up to 8× the token generation of standard speed. On the Modelflare price row of 30 September 2026, request_rules is empty, has_request_rules is false, and there is no Ultrafast model ID. Budget a Modelflare call with the two tiers above. Whether an official multiplier has reached a bill is answered by the usage record.

How to read the benchmarks

The launch page publishes deltas and costs, not a table of absolute scores that can be copied out. The table below keeps only comparisons the page states. The cost sentences describe OpenAI's harness. Modelflare did not rerun the tests.

Evaluation Comparison stated on the launch page Condition
DeepSWE v1.1 Matches Astra; 6.4 percentage points above Sol's best score Lower reasoning effort, at about one-fifth of the cost
GDP.pdf Above Opus 5.5 with fallbacks; approaches Astra Across the reasoning settings tested; under half the cost of Opus, about one-fifth the cost of Astra
AutomationBench 2.2 percentage points above Opus 5.5; 4.8 percentage points above Sol at the same setting medium; about one-third the cost of Opus
OSWorld 2.0 offline set 7 percentage points above Sol; within 2.1 percentage points of Astra max; partial reward, v2026.08.08; under half the cost of Sol, about one-seventh the cost of Astra
Terminal-Bench Science 0.1 More than doubles Sol's score; Astra remains highest, at 68.1% max; average cost per task $5.47 for 6.1, $23.21 for Opus 5.5, $23.80 for Astra
Factuality Share of answers with at least one factual error falls from 11.4% to 7.7% low; within 1.9 percentage points of Astra

The factuality sample is de-identified conversations where a user had flagged an earlier error. The launch page says these prompts are deliberately difficult and are not typical traffic. The Terminal-Bench Science dollars are OpenAI's published average cost per task, not a Modelflare group price. Astra's 68.1% is still the highest scientific score, and the launch page says to keep the hardest scientific research on Astra.

The alignment tests on the same page add a set of deliberately difficult failure rates. When the search tool is broken, the share of max-effort answers that do not tell the user is 2.1% for 6.1, 4.9% for Sol, 1.5% for Astra, and 28.7% for Luna. The launch page says no attempt to bypass an automated safety reviewer was observed, matching Astra and Sol. The detail is in the GPT-6.1 Sol system-card addendum. These rates are not failure rates on ordinary traffic.

Modelflare prices and groups

On 30 September 2026, gpt-6.1-sol has input_price 2, completion_ratio 5, cache_ratio 0.05, and create_cache_ratio 1.25. billing_mode is tiered_expr. Short-tier unit prices are $2 input, $0.10 cache read, $2.50 cache write, and $10 output. The long tier is $4 / $0.20 / $5 / $15. The endpoint types are openai, openai-response, and openai-response-compact.

The public selectable groups that include this model are the two below. The group is chosen on the key. Both groups call the same model ID.

Price table

Group Ratio Tier Input / 1M Cache read / 1M Cache write / 1M Output / 1M
openai-award 0.03 Input up to 272,000 $0.06 $0.003 $0.075 $0.30
openai-award 0.03 Input above 272,000, whole request $0.12 $0.006 $0.15 $0.45
openai-stable 0.08 Input up to 272,000 $0.16 $0.008 $0.20 $0.80
openai-stable 0.08 Input above 272,000, whole request $0.32 $0.016 $0.40 $1.20

The catalog description of openai-award is price-priority routing for budget-sensitive, retryable development, testing, and flexible workloads, with a stated 65% daily cache-hit protection for eligible requests and compensation for a shortfall. The openai-stable description is everyday development, longer projects, and production use where stability matters, with the stated protection at 75% daily. Both sentences are the group-definition text. This article did not measure the hit rate separately.

The model row also lists openai-enterprise. The same public response has no ratio for that group in group_definitions or group_ratio, so the table above leaves it out. Non-public openai-first-topup is ratio 0.015, described as unlocked when a single top-up reaches $20. It is not a public selectable group. When a key also has fallback groups, a request that leaves the campaign group is billed on the group that serves it. Public openai-premium is ratio 0.13, and this model's enable_groups does not include it.

Three examples

The first two rows fall in the short tier. The third row is 280,000 input tokens, so the whole request uses the long tier. The cache row assumes the usage record counts those input tokens as a cache hit.

Request List price openai-award openai-stable
100,000 uncached input, 5,000 output $0.2500 $0.0075 $0.0200
100,000 cache read, 5,000 output $0.0600 $0.0018 $0.0048
280,000 uncached input, 10,000 output $1.2700 $0.0381 $0.1016

The first row is 0.1 × $2 + 0.005 × $10 = $0.25. The second is 0.1 × $0.10 + 0.005 × $10 = $0.06. The third is 0.28 × $4 + 0.01 × $15 = $1.27, and the group cells multiply by 0.03 or 0.08. Applying the short-tier unit prices to those 280,000 tokens would produce $0.0198 on award, which is not the price of this row. The live page is the gpt-6.1-sol pricing page. The full catalog is Models and pricing.

Retryable development and testing use openai-award. For everyday production and longer projects, the catalog describes openai-stable as that kind of default. The two ratios are 3% and 8% of the short-tier and long-tier catalog prices. They are not 3% and 8% of a price that has already been multiplied for Fast, Batch, Flex, or a regional premium.

Calling it on Modelflare

Create a key on openai-award or openai-stable. Send the model ID gpt-6.1-sol exactly. The canonical base URL is https://modelflare.dev/v1. The same API is also published at https://cf.modelflare.dev/v1 and https://origin.modelflare.dev/v1. The root domain is the canonical site.

Tool calling uses Responses. The default effort is already medium. The request below sets it explicitly.

export MODELFLARE_API_KEY="YOUR_MODELFLARE_API_KEY"

curl -sS https://modelflare.dev/v1/responses \
  -H "Authorization: Bearer $MODELFLARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6.1-sol",
    "reasoning": {"effort": "medium"},
    "input": "Review this migration plan and list the three assumptions that most need a test."
  }'

Chat Completions can send text without tools. The field name is reasoning_effort.

curl -sS https://modelflare.dev/v1/chat/completions \
  -H "Authorization: Bearer $MODELFLARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6.1-sol",
    "reasoning_effort": "medium",
    "messages": [
      {
        "role": "user",
        "content": "Review this migration plan and list the three assumptions that most need a test."
      }
    ]
  }'

POST /v1/responses/compact is listed on this model. It is a compact operation on gpt-6.1-sol, not a second model ID. Image input uses the ordinary content array on a user message. The difference between the two entrances is covered in Responses API versus Chat Completions. How usage meets the bill is covered in AI API cost tracking.

Request differences from GPT-6 Sol

The short-tier input and output list prices match gpt-6-sol, and the cache read moves from $0.20 to $0.10. Long-tier output is $15 on both, and the long-tier cache read moves from $0.40 to $0.20. gpt-6-sol remains in the catalog. Its public groups on this reading include openai-premium, which 6.1 does not. When switching to 6.1, pin the model name on the completed result. A result that lands on another ID is not this switch.

Shapes that fail

reasoning.effort: none
reasoning.effort: minimal
tools carried on a Chat Completions request

none and minimal return HTTP 400. When moving from gpt-6-sol or gpt-6-luna, both of which accept none, compare again at low and remove none from the request. Tools, computer use, file search, and web search go through Responses. Chat Completions carries requests that have no tools.

Checks before shifting traffic

  1. Pin gpt-6.1-sol. The model name on the completed result has to be that ID as well.
  2. Estimate the public price on openai-award or openai-stable, and choose the key group by the cost of failure.
  3. Replay one non-sensitive set at low, medium, high, xhigh, and max. none and minimal return 400.
  4. Put tools on Responses. The Chat Completions example carries no tools.
  5. Price one prompt at 272,000 input tokens and one at 272,001. The second uses long-tier unit prices for the whole request.
  6. Budget with the two catalog tiers. Fast, Batch, Flex, and the regional premium on the model page follow the amount that actually appears on the usage record.
  7. Read the cached-token count on the usage record before using the cache example above.
  8. Keep the hardest scientific research on gpt-6-astra. gpt-6-sol stays available.

Sources

Checked on 30 September 2026. Modelflare did not independently reproduce the benchmarks.