AIToday
Large Language ModelsOpenAI BlogPublished: Aug 11, 2026, 01:01 JST

GPT-5.6 Sol completes finance decks in 100% of tests

GPT-5.6 Sol completes finance decks in 100% of tests

3 Key Points

  1. What happened

    Model ML, a finance AI startup, evaluated OpenAI's GPT-5.6 Sol on its Composite benchmark for financial services. In PowerPoint workflows, GPT-5.6 Sol completed the task in 100% of test cases, compared with 76% for Opus 5, and cleared Model ML's professional-readiness gate in 43.3% of cases, versus 26.7% for Opus 5.

  2. Why it matters

    Model ML's agents automate finance work from research to finished, editable PowerPoint decks and Excel workbooks. The company found that GPT-5.6 Sol gets agents 'far closer to the final output' than earlier models, reducing manual rebuilding of analysis for finance professionals.

  3. What to watch

    Model ML says the results gave it evidence to expand GPT-5.6 Sol in production, including some workflows previously handled by Opus 4.8. At one global asset manager, a bespoke tearsheet that took an analyst about an hour now takes about five minutes.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Model ML's evaluation is notable for focusing on the 'last mile' of finance work—the point where analysis must be turned into editable, traceable, and review-ready files. The startup's Composite benchmark tests not just whether a model can produce a plausible document, but whether that document holds up to scrutiny, with correct numbers, traceable sources, and editable charts. GPT-5.6 Sol's high completion rate on PowerPoint generation, paired with a higher pass rate on the 'professional-readiness' gate than Opus 5, positions it as a strong tool for this specific, high-stakes use case.

The context of this comparison is a shift in how finance professionals work. Model ML's customers are moving toward browser-based outputs that stay connected to the underlying models and source material. The company's co-founder argues that PowerPoint, Excel, and Word were 'designed for a world where creating knowledge work was manual,' and that 'AI has changed that assumption.' This suggests document creation is becoming less about manual formatting and more about the agent's ability to plan, reason, and maintain context over a long workflow.

Whether GPT-5.6 Sol becomes a default choice in finance likely hinges on how its results generalize beyond Model ML's specific harness. The benchmark shows efficiency gains—for example, using about 21% fewer tokens than Fable 5 on PowerPoint—but also some trade-offs, such as slightly lower visual quality scores than Opus 5 on specific sub-metrics. For Model ML customers, the immediate impact is tangible: deliverables that took an hour can now take minutes, shifting the professional's role from rebuilding analysis to refining judgment.

FAQ
How does GPT-5.6 Sol compare to Opus 5 on PowerPoint creation?
GPT-5.6 Sol completed the PowerPoint workflow in 100% of test cases, compared with 76% for Opus 5. It also passed the professional-readiness gate 43.3% of the time, versus 26.7% for Opus 5.
What is Model ML's 'Composite' benchmark?
Composite is Model ML's evaluation benchmark for AI in financial services. It follows an assignment from an initial brief through research and calculations to an editable deck or spreadsheet, then checks numbers, sources, formulas, structure, and visual quality.
How much faster is finance work with GPT-5.6 Sol?
At one global asset manager, a bespoke tearsheet that took an analyst about an hour to assemble now takes about five minutes. In another workflow, agents processed virtual data rooms with more than 100,000 rows and hundreds of files in one pass.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Tesla's Ara bot forecast: FSD V15, Unsupervised driving soonTop Companies AI · 3h ago
  • Disney taps Karandeep Anand as first-ever CTOTop Companies AI · 3h ago
  • Zenity Labs finds zero-click bugs in Salesforce AgentforceTop Companies AI · 3h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle adds AI agents to Ads and Analytics for faster marketing insights