OpenAI and Google each developed AI models that achieved top-tier marks on the University of Tokyo's entrance exam for the most selective program (理3, or Science III). OpenAI's model scored a perfect 100 on the math section, while the two AIs showed varying performance across subjects.
The test results expose how AI systems excel at different tasks: despite both passing at elite levels, each model demonstrated distinct strength patterns, suggesting that today's leading AI systems are not equally capable across all domains—particularly mathematics versus other subjects.
For students and job-seekers, this benchmark signals that AI cannot yet reliably replace human judgment on high-stakes exams that require nuanced reasoning across diverse fields. For AI researchers and companies, it highlights the gap between narrow excellence (perfect math scores) and broad, consistent performance needed for real-world problem-solving.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Interactive Brokers has begun connecting its platform with AI tools including ChatGPT, Claude, and Grok, and o…

A new report by Alipay+ and S&P Global, based on a survey of 6,000 consumers across nine markets in Asia, Euro…

CrowdStrike Holdings Inc

Palo Alto Networks beat fiscal fourth-quarter estimates on Tuesday and issued a strong outlook for its new fis…

Google has reportedly approached major studios such as Disney, Warner Bros

Paradium.AI, Inc. statistics have been updated on TradingView for the stock ticker AMEX:PAAI
