
What happened
GPT-6.1 Sol is now generally available on Amazon Bedrock. According to OpenAI, it matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth the cost per task and beats GPT-6 Sol's best score by 6.4 percentage points.
Why it matters
Cheaper reasoning per task could lower the total cost of running AI agents, since fewer wrong turns mean fewer tool calls and less human intervention. The claim is OpenAI's own and hinges on whether real workloads see the same results.
What to watch
The cost comparison rests on a single benchmark, DeepSWE v1.1, not production workloads, so savings may vary by task. Note that classifier-flagged traffic is retained by AWS for up to 30 days unless zero data retention is requested.
WHO IT HITSEnterprise development teams building AI agents on Amazon Bedrock, and the platform and security engineers who govern model access through IAM policies, CloudTrail, and VPC endpoints.
Summaries like this, in your inbox every morning.
OpenAI's GPT-6.1 Sol is now generally available on Amazon Bedrock, running on an inference engine built for performance, security, and reliability at scale. AWS describes it as a major upgrade to GPT-6 Sol, aimed at agentic coding, computer use, and professional work that runs frequently.
The pitch rests on how reasoning quality shapes the economics of an agent task. AWS argues that a wrong turn can add model interactions, tool calls, latency, and human intervention, so token prices and reasoning quality both feed into the total cost of completing a task. On that logic, GPT-6.1 Sol's claimed efficiency on software engineering workflows, where an agent must understand an unfamiliar repository, trace dependencies, and validate a change, is presented as the lever that cuts cost per task.
Beyond coding, AWS positions the model for interpreting complex documents, improving multistep workflows across business tools, and recognizing when a tool has failed or an action is restricted. On Bedrock, governance and security controls such as IAM policies, CloudTrail auditing, VPC endpoints, and hardware-isolated inference are framed as complementary safeguards so people stay involved at consequential decisions. The open question is whether the benchmark-level cost advantage translates to production workloads, which is likely to depend on the specific tasks customers run.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
HBM is approaching half the cost of a GPU-HBM CoWoS package, prompting the question of whether memory remains…

DeepSeek is bringing more of the software it uses to develop its AI models to Huawei Technologies' Ascend 950…

Among respondents at companies with 1,001+ employees, 50.0% said AI is used company-wide, and 46.0% flagged AI…

Nvidia released the Open Agent Safety Platform on September 28, days after CEO Jensen Huang called warnings fr…

Oracle invoked "force majeure" to delay payment on its Project Jupiter data center, and its 2056 bonds then tr…

McDonald’s is increasingly using AI to guide menu prices in the U.S
