Overview
SkillForge introduces an agentic reinforcement learning (RL) methodology designed to enhance the ability of large language model (LLM) agents to address complex, long-horizon tasks. This method focuses on the dynamic management of skill libraries, which function as a form of memory for these agents. The core concept involves the co-evolution of skills and the agent's policy through a fitness-driven skill lifecycle, where skills transition through trial, active, stable, and retired states.
Research Context
Memory-augmented reinforcement learning is a technique employed to bolster LLM agents' capacity for solving intricate, extended tasks. Within this context, skills are defined as instruction pairings accompanied by an applicability condition, which is associated with specific task types. A challenge in this area is the indiscriminate retention of every skill as the agent's policy progresses. This practice can lead to the accumulation of obsolete or detrimental skill entries, potentially misleading the agent during task execution. The research addresses this issue by proposing a system for selective skill management.
Approach
SkillForge implements a dynamic skill lifecycle to compile and evolve the skill library. This lifecycle is fitness-driven, guiding skills through distinct states: trial, active, stable, and retired. This process facilitates the co-evolution of the skills and the LLM agent's policy throughout the training period.
The methodology incorporates several sequential and iterative phases:
- Pre-RL Evaluation: Initially, a pre-reinforcement learning evaluation phase utilizes the base model's own rollouts. During this phase, skills with low fitness are pre-retired. This action results in a filtered skill library, which then serves as the foundation for supervised fine-tuning (SFT).
- Reinforcement Learning Integration: Following the SFT phase, reinforcement learning commences from this checkpoint.
- Iterative Skill Management: At each iteration of the RL process, the system continues to engage in selective retirement, stabilization, and LLM-guided mutation of skills. These actions collectively contribute to the ongoing forging of the skill library, operating in conjunction with policy optimization.
Findings
Across multiple interactive agent benchmarks, SkillForge demonstrated efficacy in improving task performance. Key findings include:
- SkillForge achieved the highest aggregate success rate among tested methods.
- It delivered a relative improvement of up to 7.8% over the strongest baseline.
- The system maintained a compact skill library throughout the training process.
Why This Matters
The development of SkillForge addresses the challenge of managing skill libraries in memory-augmented reinforcement learning for LLM agents. By dynamically evolving and pruning skills, the method aims to prevent the accumulation of obsolete or harmful entries that could negatively impact agent performance. This systematic approach to skill management contributes to the development of more effective and reliable LLM agents for complex, long-horizon tasks.
Potential Applications
The research introduces SkillFurnace, a dataset comprising over 5,000 annotated records. This dataset bundles several components:
- Retirement-filtered SFT trajectories.
- Evolved skill libraries, including fitness annotations.
- Retirement events, with human-annotated failure categories.
SkillFurnace is designed to support further research into skill quality and lifecycle management within the domain of LLM agents.