The same repository request can take several attempts: an agent proposes a patch, runs tests, and then fixes errors. Low token prices alone don’t answer the main question: how much does it cost to get a change that’s ready to accept.
As of October 2, 2026, the answer is this: GPT‑6.1 Sol’s standard API rates for input and output tokens are five times lower than GPT‑6 Astra’s. OpenAI also says Sol matches Astra on one agentic coding benchmark. But that doesn’t prove overall parity in Codex or guarantee fivefold savings on a finished patch.
First, distinguish the API from a Codex subscription
The API is billed by tokens used. Codex through ChatGPT has an included usage limit on the applicable plan, so API prices don’t apply to it. OpenAI explains that limit usage depends on the model, task, and settings, while API usage is billed separately, in its Work and Codex usage guide.
The table shows standard API rates per million tokens for requests with input context of up to 272,000 tokens:
| Model | Standard input | Cached input | Cache write | Output |
|---|---|---|---|---|
| GPT‑6.1 Sol | $2 | $0.10 | $2.50 | $10 |
| GPT‑6 Sol | $2 | $0.20 | $2.50 | $10 |
| GPT‑6 Astra | $10 | $1 | $12.50 | $50 |
| GPT‑6 Luna | $0.10 | $0.01 | $0.125 | $0.50 |
At standard input and output rates, Sol costs five times less than Astra, while cached input costs ten times less. Compared with GPT‑6 Sol, however, the new version is no cheaper for standard input and output: its pricing advantage is cached input, whose price has been halved. Luna costs far less than all three, but its low price alone says nothing about how well it will handle a difficult task. These rates are listed in OpenAI’s API pricing.
What does one equivalent attempt cost?
For a concrete example, consider a hypothetical request with 20,000 new input tokens, 80,000 cached tokens, and 10,000 output tokens. Assume the cache is hit; cache writes, tool calls, and retries are excluded.
| Model | Calculation | Cost |
|---|---|---|
| GPT‑6.1 Sol | $0.040 + $0.008 + $0.100 | $0.148 |
| GPT‑6 Sol | $0.040 + $0.016 + $0.100 | $0.156 |
| GPT‑6 Astra | $0.200 + $0.080 + $0.500 | $0.780 |
| GPT‑6 Luna | $0.002 + $0.0008 + $0.005 | $0.0078 |
With this token mix, one attempt with Sol costs about 5.3 times less than one with Astra. Compared with GPT‑6 Sol, the difference is $0.008, or about 5.1% of the attempt’s cost. This is an arithmetic example, not a measurement of a typical Codex task: models may differ in token usage, number of tool calls, and number of retries.
There is also a pricing threshold: for requests with more than 272,000 input-context tokens, input and cache rates double, while output rates increase by 1.5 times—for the entire request. So the calculation above does not apply to long contexts without adjustment; the terms are listed on the GPT‑6.1 Sol model page.
What does the comparison with Astra actually establish?
In its GPT‑6.1 Sol announcement, OpenAI said that Sol matched Astra on DeepSWE v1.1, a benchmark of complex engineering tasks in real-world codebases, at about one-fifth the cost. The precise wording is “on DeepSWE v1.1, according to OpenAI.”
On OSWorld 2.0, which evaluates computer-based task completion, Sol trailed Astra by 2.1 percentage points at the highest reasoning level. OpenAI estimated that the cost per task was about seven times lower. The company cautions that its evaluations were conducted in a research environment or through the API and may differ from production interfaces because of system instructions, tools, and settings.
These are results published by the provider itself, not an independent comparison of the models in the same Codex workflow. The materials reviewed for this article do not include a comparable independent test of Sol and Astra on identical Codex tasks.
What developers are saying
Early discussions after launch are a reason to test the model in your own environment, not a representative evaluation. In one Reddit thread about speed and usage limits, users complain about slow generation. One author says a simple action or context compaction took them at least two minutes. These are isolated observations without identical tasks or controlled conditions.
In another discussion about Sol’s speed, participants disagree: some consider slow generation a problem, while others note that output-token speed alone does not show how long the entire task will take. These are different metrics, and the thread provides a controlled measurement of neither.
There is also an isolated negative review of an attempt to build a game demo: the author says the model missed requirements and did not fix major shortcomings after clarification. This does not measure the model’s overall quality or directly compare it with Astra.
Finally, in a GitHub issue about model selection in the VS Code extension, a user reported on September 30 that Sol appeared in the Codex desktop app but was absent from the extension’s model list on Windows. This describes a specific configuration; it does not establish either the cause or the scale of the issue.
The authors of these posts use pseudonyms, so no direct quotations attributed to verified real names are included here. The more modest practical takeaway from these early reviews is that quality, time to task completion, and usage-limit consumption should be measured separately.
A brief timeline
- September 3, 2026. OpenAI introduced GPT‑6 Astra.
- September 22, 2026. GPT‑6 Sol and GPT‑6 Luna launched in the API.
- September 29, 2026. OpenAI announced the release of GPT‑6.1 Sol for the API, ChatGPT Work, and Codex.
- September 30, 2026. A GitHub report said Sol was missing from the VS Code extension’s model list on Windows.
Release dates are listed in the OpenAI API changelog, while the VS Code report appears in a specific GitHub discussion. This is an early timeline, not an assessment of availability to all users.
When to choose Sol and when to choose Astra
For Plus and Standard Business, OpenAI gives an estimated range of 15–160 local messages per five-hour window for Sol and 5–45 for Astra. These are not guaranteed task counts: usage depends on the model, task, and settings; cloud tasks may use more of the limit, and weekly limits may also apply. These estimates should not be treated as a savings multiplier or directly compared with API prices; the terms are explained in the Work and Codex limits guide.
The practical choice depends on the cost of an error and the time you’re willing to wait:
- Sol is worth testing on repeatable tasks with clear ways to check the result—for example, when tests and acceptance criteria are well defined. Its standard API rates are significantly lower than Astra’s.
- Astra may be a better fit for complex or risky tasks where result quality matters more than the cost of one attempt. How much better it handles your team’s tasks needs to be tested in practice.
- Luna may work for simple, repeatable operations if its quality is sufficient. The lowest token price does not prove that complex work will cost less with this model.
To compare them, run the same small set of tasks on each model and decide in advance what counts as an acceptable result—for example, passing tests, meeting requirements, and receiving approval in review. Record the number of attempts, token usage, tool calls, and time to a finished patch. That lets you estimate the cost of the result, not just the price of an answer.
Verdict: Sol really does cost five times less than Astra for standard API input and output tokens; on one agentic coding benchmark, OpenAI reports an equal result. But there isn’t enough evidence to establish overall parity in Codex. API savings don’t apply directly to subscriptions, and early developer reviews are still scattered and contradictory. Sol is a credible candidate for a more economical option, but whether it should replace Astra depends on how both models perform on your team’s tasks.