Skip to content
worth noting AI agents

New ARISE-RL framework for agent self-evolution using reinforcement learning

only one source so far

A research team introduced the ARISE-RL framework for training agents using RL with rubric-mediated co-evolution of Generator and Solver, complemented by the RG-SED method and ECR-Bench benchmark; according to the authors, it achieves state-of-the-art results.

A research team published the ARISE-RL (Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning) framework on arXiv, designed to train open agents using reinforcement learning. According to the authors, it addresses the lack of verifiable correct answers and scalable evaluation criteria (rubrics) for long, open-ended agentic tasks, which leads to unstable and weak rewards and degrades fine-grained optimization signals during group-based policy learning.

The framework combines two components in a process the authors call rubric-mediated co-evolution: Generator, which creates tasks and rubrics grounded in real tool observations and is rewarded for creating valid tasks of moderate difficulty matched to the current capability of Solver, and Solver, which learns from fine-grained signals of rubric fulfillment through multi-step reasoning and tool use.

The team also introduced Reward-Gated Self-Evolution Distillation (RG-SED) — a procedure that selectively distills a memory-augmented variant of the same policy back into the policy itself, but only when memory actually improves the reward. This is intended to reduce distribution mismatch and prevent blind imitation of noisy guidance.

For evaluation, the authors introduced ECR-Bench, a benchmark suite with expert-calibrated rubrics covering single-tool deep research and multi-tool travel planning tasks. According to the authors, ARISE-RL consistently achieves the best (state-of-the-art) results across all tested benchmarks; this is a claim by the study authors, and the source does not report independent verification. The source contains no information about the institution of the authors, code, a license, or planned availability.

What changed

Why it matters

This is research addressing a specific technical problem — reward instability when training agents on long, open-ended tasks involving tool use. For companies and developers building or training their own agentic systems using RL, it offers a potential method and benchmark to evaluate, but it is neither a finished product nor a tool with stated availability or a license.

Two audiences, two different impacts

What this means

01

For individuals

Researchers and developers working on reinforcement learning for agents have a new method and benchmark (ECR-Bench) available to use when designing their own experiments.

More practical updates →
02

For a business

Teams developing agentic systems with RL training gain another published approach to compare with their own agent training pipeline, especially for tasks requiring tools (tool use).

Development More business impacts →
agents ARISE-RL reinforcement learning Reward design rubrics self-evolution

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
arXiv cs.AI (Artificial Intelligence) research source · first detected ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning