AIToday
Large Language ModelsTHE DECODERPublished: Apr 26, 2026, 19:00 JST1 min read

Investment bankers reject all AI outputs in first major real-world benchmark — even GPT-5.4 fails to meet client standards

Investment bankers reject all AI outputs in first major real-world benchmark — even GPT-5.4 fails to meet client standards

3 Key Points

  1. Handshake AI and McGill University tested nine top AI models (GPT-5.4, Claude Opus 4.6, Gemini 2.5 Pro, and others) against 100 realistic investment banking tasks using BankerToolBench, an open-source benchmark created with input from 500 current and former bankers at Goldman Sachs, JPMorgan, Morgan Stanley, and other major firms. Not a single model output was ready to send to a client as-is; 41% needed major rework and 27% were completely unusable.

  2. The benchmark grades actual deliverables junior bankers produce — Excel financial models with working formulas, PowerPoint presentations, PDFs, and Word memos — against 150 individual criteria covering technical correctness, compliance, and client readiness. GPT-5.4 scored highest at 58.1 out of 100 but still failed nearly half the criteria. Claude Opus 4.6 produces polished-looking outputs that hide a critical flaw: key numbers are hardcoded as fixed values rather than linked to formulas, making it impossible to run scenario analyses (e.g., changing a purchase price updates nothing).

  3. For investment banking teams, this means AI cannot yet replace junior analyst work or streamline client deliverables, despite vendor claims. The benchmark also revealed why: models make four recurring mistakes — buggy code generation (41% of failures), broken business logic like adding costs to the revenue line (27%), aborted data queries (18%), and fabricating missing numbers and presenting them as sourced (13%).

  4. The full benchmark, including test data, rubrics, and an AI verifier called Gandalf, is publicly available and can also be used to train models via reinforcement learning. Early experiments with smaller models like Qwen showed five- to thirteen-fold performance gains, though from very low baselines.

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • askpolly raises $3M to turn social media chatter into market researchSiliconANGLE AI · 33m ago
  • ELYZA secures patent for AI app generationITmedia AI+ · 33m ago
  • German AI startup Atira raises $17.5M to automate industrial quotingFortune AI · 33m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleFactory managers should treat decades-old machinery as strategic assets, not liabilities — hybrid approaches combining legacy hardware with modern data tools cut costs while avoiding e-waste