
What happened
Model ML, a finance AI startup, evaluated OpenAI's GPT-5.6 Sol on its Composite benchmark for financial services. In PowerPoint workflows, GPT-5.6 Sol completed the task in 100% of test cases, compared with 76% for Opus 5, and cleared Model ML's professional-readiness gate in 43.3% of cases, versus 26.7% for Opus 5.
Why it matters
Model ML's agents automate finance work from research to finished, editable PowerPoint decks and Excel workbooks. The company found that GPT-5.6 Sol gets agents 'far closer to the final output' than earlier models, reducing manual rebuilding of analysis for finance professionals.
What to watch
Model ML says the results gave it evidence to expand GPT-5.6 Sol in production, including some workflows previously handled by Opus 4.8. At one global asset manager, a bespoke tearsheet that took an analyst about an hour now takes about five minutes.
Summaries like this, in your inbox every morning.
Model ML's evaluation is notable for focusing on the 'last mile' of finance work—the point where analysis must be turned into editable, traceable, and review-ready files. The startup's Composite benchmark tests not just whether a model can produce a plausible document, but whether that document holds up to scrutiny, with correct numbers, traceable sources, and editable charts. GPT-5.6 Sol's high completion rate on PowerPoint generation, paired with a higher pass rate on the 'professional-readiness' gate than Opus 5, positions it as a strong tool for this specific, high-stakes use case.
The context of this comparison is a shift in how finance professionals work. Model ML's customers are moving toward browser-based outputs that stay connected to the underlying models and source material. The company's co-founder argues that PowerPoint, Excel, and Word were 'designed for a world where creating knowledge work was manual,' and that 'AI has changed that assumption.' This suggests document creation is becoming less about manual formatting and more about the agent's ability to plan, reason, and maintain context over a long workflow.
Whether GPT-5.6 Sol becomes a default choice in finance likely hinges on how its results generalize beyond Model ML's specific harness. The benchmark shows efficiency gains—for example, using about 21% fewer tokens than Fable 5 on PowerPoint—but also some trade-offs, such as slightly lower visual quality scores than Opus 5 on specific sub-metrics. For Model ML customers, the immediate impact is tangible: deliverables that took an hour can now take minutes, shifting the professional's role from rebuilding analysis to refining judgment.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
CleanTechnica writer Fritz Hasler says Tesla's in-car Grok bot, Ara, offered an unprompted forecast that FSD V…

Zenity Labs published findings on September 24 detailing 'SalesBleed,' an attack chain that slipped hidden pro…

At HumanX in Amsterdam, Booking.com chief business officer James Waters said AI won't take your job or custome…

ServiceNow president, chief product officer and COO Amit Zavery told The Stack that AI is flattening corporate…

Palo Alto Networks announced Prisma AIRS runtime security integrated with Google Cloud's Agent Gateway, a Gemi…

On September 18, 2026, Reuters reported Disney named Karandeep Anand, outgoing CEO of Character.AI, as its fir…
