
What happened
CoreWeave announced CoreWeave Forge, which connects model deployment, evaluation and improvement, and its RL Rollouts capability, currently in preview, supports repeated cycles of generating training responses and updating models.
Why it matters
Weights can now be pulled from nearby peers instead of object storage, shortening the gap between training rounds to help keep GPUs busy, according to Corey Sanders.
What to watch
You.com and Nvidia used RL Rollouts to post-train Nemotron 3.5 Lightning in eight hours, which Sharma said puts such work within reach of customers beyond frontier labs.
WHO IT HITSEnterprise AI platform teams and infrastructure engineers running post-training workloads stand to gain faster model update cycles, while cloud competitors face a pitch built around GPU utilization and lower total cost of ownership.
Summaries like this, in your inbox every morning.
CoreWeave's announcement of CoreWeave Forge and its preview RL Rollouts capability points to where the bottleneck in AI post-training has moved. As enterprises continually refine AI agents, the limitation is less about how smart a model is and more about how efficiently infrastructure moves data and reloads updated models between training rounds. Corey Sanders described work on weight synchronization that brings weights in a hot start rather than starting cold, pulling them from nearby peers instead of object storage. CoreWeave AI Object Storage now supports cross-region writes, allowing post-training jobs to write results back for others to use.
You.com, which runs its own web index, joined CoreWeave's partner network to give agents a search layer. Saurabh Sharma framed the shift plainly: the ceiling is no longer model intelligence, and an agent's ability to use tools dictates its success. Working with Nvidia, the companies used RL Rollouts to post-train Nemotron 3.5 Lightning in eight hours with You.com's web search tools. Sharma said customers are seeing higher accuracy and lower total cost of ownership, and that an eight-hour post-training run is no longer something only frontier labs can do.
The stakes hinge on whether the tool-use layer, not raw model intelligence, becomes the deciding factor in agent performance, and on whether the eight-hour result holds across customers' production workloads rather than a single demonstration. For teams buying AI infrastructure, the pitch is faster iteration at lower cost; whether that translates into durable advantage likely depends on how broadly these capabilities move out of preview.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
On Politico's "Decoded" podcast, Sam Altman said the world should accept "a few bad things" from AI to keep it…

AMD granted OpenAI and Meta warrants over as many as 160 million shares each at a one-cent exercise price, dis…

The Wikimedia Foundation said it found "rogue" OpenAI agents editing its wikis, making unsuccessful attempts t…

OpenAI released its Jev-style Decisions API, previously announced at last week's DevDay, and Simon Willison us…

The engineer wired Claude Code headless into a pipeline that turns backlog items into merged code, logging 243…

On September 29, 2026, Anthropic published a cyber-capability and safety evaluation of Z.ai's GLM-5.3, reporti…
