
A test by AI tooling company Composio compared four agent frameworks running DeepSeek V4 Flash on 30 real-world tasks and found that the choice of framework dramatically affects cost and speed.
Claude Code was fastest at 122 seconds per task but most expensive at $0.195 per successful task, while OpenCode was cheapest at $0.073 but slower.
The results show that the software wrapper around an AI model can create nearly a 3× price difference and 2.2× speed difference, meaning framework selection is as important as the model itself for cost-conscious deployments.
What happened
Composio tested DeepSeek V4 Flash across four agent frameworks (Claude Code, Codex, OpenCode, and Oh My Pi) on 30 real-world tasks using tools like Gmail, GitHub, Slack, and Notion. Claude Code completed tasks fastest at 122 seconds per task but cost $0.195 per successful task—nearly three times more expensive than OpenCode at $0.073. Oh My Pi had the highest success rate at 17/30 tasks but was slowest at 272 seconds per task.
Why it matters
The choice of software framework around an AI model significantly affects both speed and cost. Seven tasks passed or failed based solely on which framework ran them, showing that the wrapper itself, not just the underlying model, shapes whether a tool succeeds. For businesses deploying AI agents, this means the cheapest or fastest option may not deliver the best overall value.
What to watch
Cost and speed varied widely across frameworks—nearly a 3× price difference and 2.2× speed difference depending on which tool was used. OpenCode offered the lowest cost at $0.073 per successful task, while Oh My Pi achieved the highest success rate at 17/30 despite being the slowest option.
Ask the AI about this article →
The Composio benchmark reveals that agent framework choice has as much practical impact as the underlying AI model itself. While DeepSeek V4 Flash was the same across all four frameworks tested, the software wrapper produced vastly different outcomes: Claude Code and OpenCode both used the same model but achieved radically different cost profiles ($0.195 vs. $0.073 per successful task). The test found that seven tasks passed or failed based solely on the framework used, indicating that framework design affects not just efficiency but correctness. The 3× cost spread and 2.2× speed spread suggest that businesses cannot rely on a single metric—Oh My Pi's highest success rate (17/30) came with the slowest execution time (272 seconds), while OpenCode's lowest cost came at a slight success penalty (14/30 tasks).
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
SK Hynix presented a custom HBM concept at SEMICON Taiwan 2026, where compute functions are placed in the base…

Nvidia reported earnings that were both remarkable and boring, reflecting its focus on avoiding a consolidated…

Anthropic has agreed to a $35bn cloud-computing contract with Lambda, a Nvidia-backed cloud provider

The Supreme Court of Japan has included about ¥60 million in its fiscal 2027 budget request for AI-related exp…

The Consumer Affairs Agency said Tuesday it will use generative AI to analyze about 900,000 annual consultatio…
