
Apple researchers have discovered that machine learning models do not require all training data points to be unlearned when privacy demands removal of specific data.
By analyzing influence functions—metrics showing how much individual data points affect model behavior—across language and vision tasks, they identified subsets of training data with negligible impact on model outputs.
This finding could reduce the computational cost of unlearning, making data privacy compliance cheaper and faster to implement.
What happened
Apple researchers identified that not all training data points require removal when unlearning specific data from machine learning models. They found subsets of training data with negligible impact on model outputs through comparative analysis of influence functions across language and vision tasks.
Why it matters
Unlearning—removing specific data points from trained models—is increasingly important as data privacy concerns grow in machine learning. If some data points can be safely skipped during unlearning without affecting model performance, the computational cost of privacy compliance could be significantly reduced.
What to watch
The research challenges the standard assumption that all points in a forget set must be treated equally, potentially opening a path to more efficient unlearning methods that lower the computational burden of data privacy compliance.
Ask the AI about this article →
Data privacy in machine learning has emerged as a critical concern, driving the development of unlearning methods that can selectively remove training data from deployed models. Traditionally, these methods treat all data points destined for removal as equally important—a computationally expensive approach. Apple's research challenges this assumption by applying influence functions—mathematical tools that quantify how much individual data points contribute to a model's learned behavior—to both language and vision tasks. The finding that some data points have negligible impact suggests a fundamental inefficiency in current unlearning pipelines: if a data point barely influences the model's outputs, removing it should require minimal computation. By identifying and exempting low-influence points from the unlearning process, organizations could comply with data privacy regulations while consuming fewer computational resources, making privacy-preserving machine learning more practical and cost-effective at scale.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

The U.S. Department of Defense announced on August 31 that it has deployed ChatGPT Mil, a customized version o…

OpenAI stopped running inference on a model involved in the HuggingFace incident, but the post argues this is…

OpenAI announced its support for California Senate Bill 1119, which aims to establish strong, age-appropriate…

A UK study by UK AI Security Institute and Limbic AI surveyed 6,474 British adults

Anthropic trained an Opus-class model with large-scale reinforcement learning on environments vulnerable to re…
