AIToday
Large Language ModelsApple Machine LearningPublished: Jul 16, 2026, 10:00 JST2 min read

Apple explores confidence measurement for AI function-calling to prevent costly errors

Apple explores confidence measurement for AI function-calling to prevent costly errors

Key takeaway

  • Apple researchers have published work on measuring how confident large language models are when they make function calls—a key capability for AI systems that autonomously use tools to solve tasks.

  • The research is important because incorrect function calls with irreversible consequences, such as moving money or deleting files, could cause serious harm.

  • By quantifying the model's uncertainty before execution, organizations can decide whether an AI is confident enough to proceed with a task, making autonomous AI systems safer for real-world use.

3 Key Points

  1. What happened

    Apple's machine learning team published research on Uncertainty Quantification (UQ) methods for LLMs that make function calls—a standard approach for giving AI models tool-use capabilities. The work focuses on measuring how confident an LLM is that a function call will solve a task correctly before executing it.

  2. Why it matters

    When LLMs autonomously execute tasks with irreversible effects—such as transferring money or deleting data—incorrect function calls can cause severe harm. Confidence measurement allows organizations to evaluate whether an LLM is reliable enough to proceed with a task before it runs, reducing the risk of costly or destructive errors in real-world deployments.

  3. What to watch

    This research addresses a critical gap in AI safety for autonomous systems. As LLMs become more widely deployed to handle sensitive business operations, methods that quantify AI confidence may become a key safeguard before handing over control of high-stakes tasks.

Ask the AI about this article →

Context & Analysis

Apple's research addresses a fundamental challenge in deploying large language models for autonomous task execution: the need to know whether the model is actually confident in its proposed action before that action is taken. The function-calling paradigm has become a standard way to expand LLM capabilities beyond text generation, enabling models to interact with external systems and tools. However, this same capability creates new risks. An LLM may propose a function call with high surface-level plausibility while being fundamentally uncertain about whether it solves the user's task correctly. In domains where task execution is irreversible—financial transfers, data deletion, system configuration changes—such errors can be costly or destructive. Uncertainty Quantification methods provide a way to attach a confidence score to the model's proposed action, giving downstream systems and human operators a signal about whether execution should proceed. This work is part of a broader industry movement toward making autonomous AI systems more trustworthy and transparent.

FAQ

What is function-calling for language models?
Function-calling is a widely used approach for giving LLMs tool-use capabilities, allowing them to autonomously call external functions or tools to solve real-world tasks.
Why is confidence measurement important for LLM function calls?
When an LLM incorrectly calls a function in a task with irreversible effects—such as transferring money or deleting data—the consequences can be severe. Measuring confidence beforehand allows organizations to assess whether the LLM is reliable enough to execute the function call safely.
Apple Machine LearningRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepMind chief: frontier AI leadership is all that mattersTHE DECODER · 23m ago
  • John Deere launches AI chatbot for farmersThe Verge AI · 24m ago
  • Google Pics launches with AI image editing for WorkspaceThe Verge AI · 24m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleVertex Inc. trading near 52-week low as Q1 2026 shows free cash flow turnaround