
Databricks Chief AI Scientist Jonathan Frankle argues that today's AI models are intelligent enough; the real bottleneck for adoption is evaluating and governing AI outputs — not building smarter models.
Unlike humans, who need one licensing exam, AI systems require orders of magnitude more rigorous checks because a single software flaw can affect thousands of users simultaneously.
Frankle says defining what "good work" means and translating that into evaluation checklists is unexpectedly hard and may take over a decade to solve.
What happened
Jonathan Frankle, Chief AI Scientist at Databricks (and co-founder of acquired AI firm MosaicML), argues that today's AI models are already intelligent enough. The real challenge for AI adoption is no longer model capability, but rather evaluation, governance, and cost-efficiency — determining whether AI is doing good work, building cost-effective agents, and controlling AI systems.
Why it matters
Frankle contends that even if AI model performance improvements stopped today, decades of work remain in figuring out how to use existing models well. Critically, AI outputs require evaluation far more rigorous than human performance reviews: a single flaw in self-driving software could cascade across many vehicles, whereas a human driver's license test suffices for one person. This difference means AI demands "orders of magnitude more rigorous evaluation" than humans do. For businesses struggling with lower-than-expected AI output quality, addressing how to define and measure "good work" may be more pressing than waiting for smarter models.
What to watch
Frankle emphasizes that translating human standards for quality into detailed checklists is unexpectedly difficult — "without a mind-reading machine, humans accurately describing their own expectations remains the near-term bottleneck." He believes AI evaluation is "far harder and more important than building the next massive model" and may take over 10 years to solve. Model performance depends on providers like OpenAI and Anthropic, but users must define what "good work" means for their own use cases.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CrowdStrike extends its Falcon platform to police AI agents at the endpoint, treating each agent as an asset w…
PlayNitride Inc., a Micro LED maker, expects its technology to enter commercial optical communications applica…

McKinsey's 2025 survey found that while 65% of companies continuously use generative AI, fewer than 5% have ac…

Anthropic announced Enterprise Frontier Safeguards (EFS) on September 1, offering enterprise customers privacy…

Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1 on September 1

New Goldman Sachs analysis finds that currencies of South Korea, Taiwan, and Malaysia are outperforming those…
