
Apple researchers have published work on measuring how confident large language models are when they make function calls—a key capability for AI systems that autonomously use tools to solve tasks.
The research is important because incorrect function calls with irreversible consequences, such as moving money or deleting files, could cause serious harm.
By quantifying the model's uncertainty before execution, organizations can decide whether an AI is confident enough to proceed with a task, making autonomous AI systems safer for real-world use.
What happened
Apple's machine learning team published research on Uncertainty Quantification (UQ) methods for LLMs that make function calls—a standard approach for giving AI models tool-use capabilities. The work focuses on measuring how confident an LLM is that a function call will solve a task correctly before executing it.
Why it matters
When LLMs autonomously execute tasks with irreversible effects—such as transferring money or deleting data—incorrect function calls can cause severe harm. Confidence measurement allows organizations to evaluate whether an LLM is reliable enough to proceed with a task before it runs, reducing the risk of costly or destructive errors in real-world deployments.
What to watch
This research addresses a critical gap in AI safety for autonomous systems. As LLMs become more widely deployed to handle sensitive business operations, methods that quantify AI confidence may become a key safeguard before handing over control of high-stakes tasks.
Ask the AI about this article →
Apple's research addresses a fundamental challenge in deploying large language models for autonomous task execution: the need to know whether the model is actually confident in its proposed action before that action is taken. The function-calling paradigm has become a standard way to expand LLM capabilities beyond text generation, enabling models to interact with external systems and tools. However, this same capability creates new risks. An LLM may propose a function call with high surface-level plausibility while being fundamentally uncertain about whether it solves the user's task correctly. In domains where task execution is irreversible—financial transfers, data deletion, system configuration changes—such errors can be costly or destructive. Uncertainty Quantification methods provide a way to attach a confidence score to the model's proposed action, giving downstream systems and human operators a signal about whether execution should proceed. This work is part of a broader industry movement toward making autonomous AI systems more trustworthy and transparent.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Google DeepMind chief Koray Kavukcuoglu said being at the frontier of AI is the only thing that matters to the…

John Deere is testing an AI assistant called “JD” that answers farmers' questions on topics like equipment set…

Google has launched Google Pics, a new suite of creative design tools for Workspace users, built around Gemini…

OpenAI said today that it is integrating ChatGPT Health with Epic's electronic health record (EHR) system, whi…

Google is launching Google Pics, an AI-powered image creation and editing tool that will be part of Google Wor…

Google DeepMind launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite
