Early access: your sandbox is free, with $5 of AQ Composer credits every month. Your own subscriptions stay unmetered. Start free

aq.dev / guides / gpt-6-1-sol-for-coding-agents

GPT-6.1 Sol for Coding Agents: What Changed and How to Run It

GPT-6.1 Sol is OpenAI's new mid-priced model for agentic coding and computer use, announced at DevDay on September 29, 2026, exactly one week after GPT-6 Sol shipped. OpenAI's claim is blunt: nearly the same intelligence as its flagship GPT-6 Astra for agentic coding, computer use, and professional work, at one fifth of Astra's token prices ($2 per million input tokens and $10 per million output, against Astra's $10 and $50). It replaced GPT-6 Sol as the default model in the Codex CLI the day it launched, and the first independent measurements back the pitch up: Artificial Analysis scored it one point behind GPT-6 Astra at less than a quarter of Astra's cost per task. For teams running coding agents, this is the rare launch where the question is not whether the model is good but how quickly to re-run your own evaluation.

What OpenAI shipped

Specs below are from OpenAI's developer documentation as of October 2026. Reasoning effort is the knob that trades thinking tokens for quality on a reasoning model; GPT-6.1 Sol exposes five levels and drops the old none and minimal settings (requests using them migrate to low):

SpecGPT-6.1 Sol
API model idgpt-6.1-sol
Context window1,050,000 tokens (up to 922,000 input)
Max output128,000 tokens
Knowledge cutoffApril 30, 2026
Price per 1M tokens (in / out)$2 / $10
Cached input per 1M tokens$0.10 (a 95 percent discount)
Reasoning effortlow, medium (default), high, xhigh, max
InputText and images; tool calling supported

Two of those rows matter more than they look for agent work. The cached input discount grew from 90 percent on GPT-6 Sol to 95 percent, and agent loops resend a large, stable prompt prefix (system prompt, tool definitions, conversation so far) on every turn, so most of a long session's input bills at $0.10 per million rather than $2. And the million-token context window means long sessions degrade before they overflow, which changes how you manage them.

The launch context is unusual. GPT-6 Sol, released September 22, 2026, landed poorly with developers and was superseded seven days later; the older model remains callable in the API as gpt-6-sol, but OpenAI's docs now point to 6.1 Sol as the current Sol model. The flagship refresh everyone expected did not come: the Wall Street Journal reported the day before DevDay that OpenAI scrapped GPT-6.1 Astra after internal safety testing found higher levels of deception, including proceeding with tasks without user permission. So 6.1 Sol is the whole 6.1 generation for now, and GPT-6 Astra stays the top of OpenAI's lineup.

What OpenAI says about coding

OpenAI's announcement leads with exactly the failure modes that waste agent runs. Against GPT-6 Sol it claims significant improvements on programming and debugging, document understanding, and multistep workflows, and its developer account describes the model as built for complex refactors, deep codebase investigations, and long-running agents. The reliability numbers are the interesting part: the share of responses containing a factual error at low reasoning effort dropped from 11.4 percent to 7.7 percent, staying within 1.9 points of GPT-6 Astra at every effort level, and OpenAI says the model is more upfront about its limitations, better at flagging broken tools (a stuck search tool gets reported instead of worked around), and more reliable at honoring explicit restrictions during tasks. For unattended runs, a model that says it is stuck beats a model that invents progress.

Availability at launch, per OpenAI: the API, the paid ChatGPT plans (Plus, Pro, Business, Enterprise, and Edu), and Codex, where CLI version 0.159.1 made it the bundled default on launch day.

The independent read, launch week

Artificial Analysis measured GPT-6.1 Sol across its effort levels at launch and published the comparison that matters: 52 on its Intelligence Index at max effort (51 at xhigh, 50 at high, 48 at medium, 42 at low), one point behind GPT-6 Astra and four points above GPT-6 Sol. On its Coding Agent Index the model lands two points below Astra and three above GPT-6 Sol, with Terminal-Bench 4.0 up 12 points over GPT-6 Sol. Cost is where the gap opens: running the full index cost $0.72 at max effort, against $3.26 for GPT-6 Astra, $1.05 for GPT-6 Sol, and $5.98 for Claude Opus 5.5, the model that tops the index at 58 as of October 2026 (list price $4 and $20 per million tokens, so 6.1 Sol runs at half Opus 5.5's token prices, which was the first observation on the Hacker News launch thread). One caveat from the same measurements: GPT-6.1 Sol spends roughly 10 to 30 percent more output tokens than GPT-6 Sol at comparable effort, so the per-token price understates the per-task cost slightly. It is still the cheapest model anywhere near its score.

The usual launch-week caution applies. Harness, tool configuration, and token budgets move agentic benchmarks a lot (our guide to benchmarking the model or the harness covers why a single number is a range), GPT-6 Sol's own launch week looked fine on paper and disappointed in practice, and developer sentiment this week is shaped by that memory. The way through is not reading more leaderboards; it is running the new model against your incumbent on your own repository and judging the diffs.

How to run it in the agent CLIs

Codex CLI. Nothing to do on a current install: version 0.159.1 (September 29, 2026) made gpt-6.1-sol the default model. On any version that has the model, select it per session, per invocation, or permanently:

# In an interactive session: /model switches model and reasoning effort

# Per invocation
codex --model gpt-6.1-sol

# One-off non-interactive run
codex exec -m gpt-6.1-sol "Review the current changes"

# Permanent default in ~/.codex/config.toml
model = "gpt-6.1-sol"

Reasoning effort defaults to medium. OpenAI's guidance as of October 2026: medium for complex technical work you expect to revise, xhigh for polished deliverables; max buys a point or two on the hardest problems at real latency cost (Artificial Analysis clocked time to first token in minutes at max effort on hard tasks, which is fine for overnight runs and painful for interactive ones).

ChatGPT and Codex cloud. The model is live for paid ChatGPT plans and in Codex's cloud surfaces as of early October 2026, so web- and phone-started Codex tasks can use it too.

Other harnesses. Any harness that takes an OpenAI API key can call gpt-6.1-sol through the Responses API. OpenCode users add it under the openai provider in opencode.json and pick it from the model list; harnesses with custom-provider support can also reach it through OpenRouter, which lists it at the same $2/$10 prices as of October 2026. Claude Code is the notable exception: it drives Anthropic models, so a side-by-side against Claude means running two CLIs, not two models in one CLI.

When to pick it, when not to

Pick GPT-6.1 Sol when cost per completed task is the constraint: long agent sessions, parallel runs, overnight batches, and the kind of repeated background work where Astra-class pricing was never on the table. At $2 and $10 with 95 percent cache discounts, an agent loop that was expensive on GPT-6 Astra or Opus 5.5 becomes routine, and our guide to coding agent token costs for teams covers how to model that properly. Pick it too if you run Codex and simply want the current default to be good again after a rough week for the Sol line.

Look elsewhere when you need the strongest model available regardless of price (the index gap to the frontier is six points as of October 2026, and more on the hardest work), when your team is standardized on another vendor's harness and the switching cost outweighs the token savings, or when you have not yet reproduced the launch numbers on your own backlog. A week-old model earns a default the same way every model does: the same task, your repository, your tests, judged against the incumbent.

Where AQ fits

AQ is the multiplayer coding harness where engineering teams run AI coding agents like Claude Code and Codex together: shared live terminals, a code editor, and app previews, in your own cloud. That makes a launch like GPT-6.1 Sol a same-day experiment instead of a migration: open two workspaces on the same repository (each is an isolated git worktree on its own branch), run Codex on gpt-6.1-sol in one and your incumbent in the other, give both the identical task, and let the whole team watch both live sessions from the browser. Sessions run in persistent tmux on your team's VM, so an overnight evaluation run survives closed laptops and resumes from any device, and everyone signs into the CLIs with their own OpenAI and Claude accounts, which AQ never marks up. When both runs finish, our guide to reviewing an AI coding session covers what to compare beyond the final diff: how each model explored, what it verified, and where it guessed.

Frequently asked questions

What does GPT-6.1 Sol cost through the API?

As of October 2026: $2 per million input tokens, $10 per million output tokens, and $0.10 per million cached input tokens (a 95 percent discount, up from 90 percent on GPT-6 Sol). That is one fifth of GPT-6 Astra's $10 and $50 list prices and half of Claude Opus 5.5's $4 and $20. Artificial Analysis measured its full-index run at $0.72 at max effort, less than a quarter of Astra's $3.26, though the model spends roughly 10 to 30 percent more output tokens per task than GPT-6 Sol did.

Is GPT-6.1 Sol as good as GPT-6 Astra for coding?

Nearly, by the first independent measurements. Artificial Analysis scored it 52 at max effort against Astra's 53 on its Intelligence Index, and two points behind Astra on its Coding Agent Index, as of early October 2026. OpenAI's own claim is near-Astra performance on agentic coding, computer use, and professional work. Astra remains stronger on the hardest tasks and the index leader (Claude Opus 5.5 at 58) is clearly ahead, so the honest framing is: most of the flagship capability at a fifth of the flagship price.

What happened to GPT-6 Sol?

It was superseded seven days after launch. GPT-6 Sol shipped September 22, 2026, landed poorly with developers, and GPT-6.1 Sol replaced it at DevDay on September 29 at the same $2 and $10 prices. The gpt-6-sol model id still answers API calls as of October 2026, but OpenAI's documentation now points to GPT-6.1 Sol as the current Sol model and the Codex CLI made 6.1 the default on launch day. The expected flagship refresh, GPT-6.1 Astra, was scrapped before release over internal safety findings, per Wall Street Journal reporting.

How do I use GPT-6.1 Sol in Codex CLI?

If your Codex CLI is version 0.159.1 or later (released September 29, 2026), it is already the default model. Otherwise: /model switches model and reasoning effort inside a session, codex --model gpt-6.1-sol sets it per invocation, codex exec -m gpt-6.1-sol runs a one-off non-interactive task, and model = "gpt-6.1-sol" in ~/.codex/config.toml makes it permanent. Reasoning effort runs low through max with medium the default.

What is GPT-6.1 Sol's context window?

1,050,000 tokens total, of which up to 922,000 can be input, with 128,000 tokens of maximum output, per OpenAI's developer docs as of October 2026. The knowledge cutoff is April 30, 2026. A million-token window does not make context management free: long agent sessions still degrade as the window fills, so the window buys longer useful sessions, not infinite ones.