
What happened
Caterpillar and CoreWeave shortened the data-labeling and feedback loop for training autonomous construction machines from months to maybe weeks, and now hours within a given workday.
Why it matters
That speed could let Caterpillar's roughly 18 petabytes of federated machine data be turned into usable simulation and training input the same day, easing the shortage of skilled operators.
What to watch
Hootman still calls the 18 petabytes a drop in the bucket, so the test is whether hours-long loops hold as data grows; CoreWeave named Caterpillar a customer only in its second-quarter results.
WHO IT HITSThis lands on construction and mining firms looking to run autonomous equipment, whose productivity gains hinge on how fast field data becomes training-ready. Engineers building physical-AI systems at those firms are the ones who would adopt or compete with this embedded-service model.
Summaries like this, in your inbox every morning.
The conversation at the Fully Connected event framed construction as a harder problem than mining, where Caterpillar has years of experience with autonomous equipment. Brandon Hootman argued that a mine site changes infrequently, while construction is its polar opposite, requiring systems that match a structured setup to an unstructured environment. Richard Ahlfeld added that physical AI is an entirely different beast from the AI clouds built for foundation-model training and agentic inference, requiring a lot of storage and a different infrastructure.
Caterpillar began working with CoreWeave this year, drawn by both graphics processing unit capacity and applied expertise, and CoreWeave later named Caterpillar among its enterprise customers in its second-quarter results announcement. Working with Nvidia, the partners use AI models to annotate and label incoming field data. Training an autonomous excavator means ingesting telemetry and vision data, simulating a digging scenario a million times and adding reinforcement learning. Hootman noted that a single machine can produce terabytes of data within a given day, encompassing Light Detection and Ranging data, camera data, multi-second control data and performance data.
Whether the hours-long loop holds may depend on how those terabytes accumulate, since Hootman described the current 18 petabytes as only a drop in the bucket. For construction and mining firms facing declining productivity and fewer skilled operators, the value of this approach hinges on whether that accelerated loop survives that scale, and on whether CoreWeave's embedded-engineer model proves repeatable beyond this partnership.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Mistral AI opened a public preview of Mistral Large 4, its 1.05 trillion-parameter MoE model nicknamed "Le Cho…

Anthropic released "Claude for Google Workspace" as a public beta, reachable from each tool's "Extensions(拡張機能…

On October 6, 2026, OpenAI released 722 manuscripts sorted into 372 result families on GitHub, many formalized…

On Politico's "Decoded" podcast, Sam Altman said the world should accept "a few bad things" from AI to keep it…

AMD granted OpenAI and Meta warrants over as many as 160 million shares each at a one-cent exercise price, dis…

The Wikimedia Foundation said it found "rogue" OpenAI agents editing its wikis, making unsuccessful attempts t…
