AIToday
AI Business & IndustryITmedia AI+Published: Jul 6, 2026, 19:00 JST2 min read

AIモデルはもう十分賢い、ボトルネックは「評価」に Databricks研究者

AIモデルはもう十分賢い、ボトルネックは「評価」に Databricks研究者

Key takeaway

  • Databricks Chief AI Scientist Jonathan Frankle argues that today's AI models are intelligent enough; the real bottleneck for adoption is evaluating and governing AI outputs — not building smarter models.

  • Unlike humans, who need one licensing exam, AI systems require orders of magnitude more rigorous checks because a single software flaw can affect thousands of users simultaneously.

  • Frankle says defining what "good work" means and translating that into evaluation checklists is unexpectedly hard and may take over a decade to solve.

3 Key Points

  1. What happened

    Jonathan Frankle, Chief AI Scientist at Databricks (and co-founder of acquired AI firm MosaicML), argues that today's AI models are already intelligent enough. The real challenge for AI adoption is no longer model capability, but rather evaluation, governance, and cost-efficiency — determining whether AI is doing good work, building cost-effective agents, and controlling AI systems.

  2. Why it matters

    Frankle contends that even if AI model performance improvements stopped today, decades of work remain in figuring out how to use existing models well. Critically, AI outputs require evaluation far more rigorous than human performance reviews: a single flaw in self-driving software could cascade across many vehicles, whereas a human driver's license test suffices for one person. This difference means AI demands "orders of magnitude more rigorous evaluation" than humans do. For businesses struggling with lower-than-expected AI output quality, addressing how to define and measure "good work" may be more pressing than waiting for smarter models.

  3. What to watch

    Frankle emphasizes that translating human standards for quality into detailed checklists is unexpectedly difficult — "without a mind-reading machine, humans accurately describing their own expectations remains the near-term bottleneck." He believes AI evaluation is "far harder and more important than building the next massive model" and may take over 10 years to solve. Model performance depends on providers like OpenAI and Anthropic, but users must define what "good work" means for their own use cases.

Ask the AI about this article →

FAQ

Why is AI evaluation harder than building better AI models?
A single flaw in AI software can cascade across many systems simultaneously (for example, in self-driving cars), whereas humans need only pass one test. This means AI requires orders of magnitude more rigorous evaluation than humans do. Additionally, turning implicit human standards for quality into detailed checklists is surprisingly difficult, and without a mind-reading machine, humans accurately describing their own expectations remains a near-term bottleneck.
What should companies do if their AI outputs are lower quality than expected?
Frankle suggests investing first in evaluation and governance rather than waiting for smarter models. By rigorously assessing whether AI outputs match instructions and avoiding harmful outputs, then adjusting prompts, harnesses, and context management based on those assessments, even lower-performing models can produce high-quality outputs. Users should define what "good work" means for their own use cases, since this cannot be dictated by model providers.
How long might it take to solve AI evaluation challenges?
Frankle believes AI evaluation is "far harder and more important than building the next massive model" and may take over 10 years to solve.

Get the latest AI Business & Industry news every morning

For example, today's edition would include:

  • AI agents outpace human security teams, forcing endpoint defensesSiliconANGLE AI · 2h ago
  • PlayNitride eyes Micro LED optical comms in 2 yearsDIGITIMES Asia · 2h ago
  • Anthropic launches Enterprise Frontier Safeguards for secure AI monitoringITmedia AI+ · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSaints & Masters acquires Xencia to boost Microsoft cloud business