
What happened
Apple researchers and NUS co-authors introduced RISED, which uses rubrics for online data selection and policy supervision, achieving the highest mean pass rate across environments and ranking first or second in every individual environment across model backbones.
Why it matters
The approach offers richer textual feedback than scalar rewards, which lack cross-environment relationship information and within-group contrast when rewards are identical.
What to watch
The gains hinge on the LLM judge's tagging using a predefined shared rubric vocabulary; watch whether this vocabulary scales to new environments without manual redefinition.
WHO IT HITSThis research could matter to AI teams training agents across multiple simulated environments, offering a method to select data and supervise policies when scalar rewards fail.
Summaries like this, in your inbox every morning.
Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without considering relationships between current rollouts across environments. As environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals.
To address these limitations, the researchers repurposed rubrics beyond their use as reward. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide data selection to align with the overall behavioral composition of the mixed-environment batch while limiting overlap with already-selected data. Positive rubrics provide privileged context for an on-policy self-distillation teacher, and negative rubrics guide subsequent rollout generation away from recurring failure modes.
The outcome hinges on whether the rubric vocabulary can generalize across new environments without manual redefinition, and on the LLM judge's tagging accuracy. For AI teams training agents across multiple simulated environments, this approach could offer a way to improve performance when scalar rewards provide insufficient contrast.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Meta, Walmart, Stripe and Sierra Technologies are creating the Personal Agent Protocol, introduced by Sierra c…
Mistral AI opened a public preview of Mistral Large 4, its 1.05 trillion-parameter MoE model nicknamed "Le Cho…

Reflection AI started early access to Beam, its first open-weights model, with 501 billion total and 23 billio…

Anthropic is releasing an expanded Cyber Verification Program with three access levels — Defense, Red Team, an…

Google began rolling out "Simple Guide" in Gemini Live on Android, letting users share camera or screen views…

Anthropic launched Claude for Google Workspace as a public beta for paid Claude plans, adding Claude to Google…
