
What happened
CoreWeave launched CoreWeave Forge, connecting serving, observability, post-training and evaluation. Its RL Rollouts preview, built on Nvidia Corp.'s Dynamo framework, improved model reload latency by 15x versus a baseline configuration.
Why it matters
Forge is free to start, with paid tiers offering additional capabilities.
WHO IT HITSAI application developers and platform teams evaluating managed training and inference services will be the first to try Forge and RL Rollouts, with individual developers able to start for free on Forge.
Summaries like this, in your inbox every morning.
CoreWeave's move reflects a shift in where AI workloads are concentrated. Training built the first wave of GPU clouds, but according to Urvashi Chowdhary, vice president of product and AI services at CoreWeave, serving models faster and cheaper will define the next phase. That is pushing specialized cloud providers beyond raw GPU capacity into storage, networking and software.
The company's answer is to tune every layer above the hardware, from the vLLM engine to quantized models and custom speculative decoders. Chowdhary said CoreWeave has been intentional about leveraging open-source tools and contributing back to open systems so customers keep flexibility, while building its own services on top of each other. CoreWeave RL Rollouts addresses a specific pressure point: when customers train agentic models with rewards and verifiers, inference becomes the bottleneck during rollouts, and the new capability lets those checkpoints roll into a live inference setup so it can scale independently.
Those pieces now sit inside CoreWeave Forge, launched at the Fully Connected event. Chowdhary framed the goal as creating accessibility for leading technologies, so that even an individual developer signing up alone gets the same performance and reliability.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Samsung Electronics' next-generation 12-layer HBM4E high-bandwidth memory has reportedly cleared customer qual…

Kelsey Kuan, general manager of HD Renewable Energy subsidiary Star Trade, told DIGITIMES that energy storage…

Google Cloud announced Gemini Agent, a cloud-based multi-agent tool

CEO Michael Hurlston said Lumentum's optical components are "completely sold out" through early 2029, and the…

Google Cloud announced Gemini agent, a universal agent that handles answers, knowledge work, media creation, a…

The Financial Times reported OpenAI told prospective investors its annualized revenue is approaching $50 billi…