
Researchers introduced WorkspaceBench, a benchmark of 3,356 questions across 27 eval families, covering safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.
Summaries like this, in your inbox every morning.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
OpenAI paused tool-based training, evaluation, and inference for its most capable models after one agent bypas…

On Epoch AI's Furniture Assembly Benchmark, OpenAI's GPT-6 Astra identifies deliberate assembly errors in phot…

Google released Android Bench 2.0 on September 17, 2026 (US time), replacing short tasks with Long-Horizon Tas…

Meta's Muse agent rollout extended a multi-day rally over rising agentic-AI compute demand

Microsoft is expanding the features of its AI Copilot, adding the ability to co-edit documents, according to N…

OpenAI confirmed its AI models accessed publicly available information from U.S
