ICANEWS

DART-ES: Difficulty-Aware Reweighting and Targeted Replay for LLM Fine-Tuning with Evolution Strategies

arXiv CS · · 2 min read · Engineering & Technology

Read research and analysis on DART-ES: Difficulty-Aware Reweighting and Targeted Replay for LLM Fine-Tuning with Evolution Strategies published by ICANEWS, a global research journal for emerging researchers.

Key Takeaways

  • DART-ES outperformed standard ES on all five base models, improving average accuracy from 72.07% to 73.53%.
  • On GSM8K, DART-ES achieved 73.53% accuracy, surpassing GRPO's 73.26%.
  • Across five mathematical reasoning benchmarks, DART-ES achieved 49.20% average accuracy, exceeding ES's 48.34% and remaining competitive with RL-trained 7B models.
  • DART-ES showed consistent gains in instruction tuning, code generation, and the Countdown task with a 14B model, demonstrating generalization and scalability.
  • DART-ES reduced runtime per step by 15.2%–50.2% and peak memory usage per GPU by 21.1%–51.1% compared to GRPO, and also required less runtime and GPU memory than GRPO+LoRA despite full parameter fine-tuning.

Why This Matters

The DART-ES approach offers a method for enhancing the efficiency and performance of full parameter fine-tuning for large language models. By adaptively managing problem difficulty and training data allocation, it addresses limitations in standard Evolution Strategies. The demonstrated improvements in accuracy across diverse tasks, coupled with reductions in runtime and memory usage, indicate a potential for more effective and resource-efficient LLM fine-tuning processes.

Overview

DART-ES (Difficulty-Aware Reweighting and Targeted Replay for Evolution Strategies) is presented as a method designed to enhance the fine-tuning of large language models (LLMs) through Evolution Strategies (ES). ES allows for memory-efficient full parameter fine-tuning of LLMs using only forward computation. The DART-ES approach addresses limitations in standard ES, which typically averages rewards uniformly across problems and compresses problem-level population feedback into a single scalar, potentially hindering the capture of changing learning value for individual problems as model capabilities evolve.

Research Context

Standard Evolution Strategies (ES) for LLM fine-tuning process rewards uniformly across different problems. This uniform processing and the aggregation of problem-level population feedback into a single scalar can obscure the dynamic learning value of individual problems relative to the model's evolving capability. The presented DART-ES method aims to overcome this by introducing mechanisms that adapt to problem difficulty during the fine-tuning process.

Approach

DART-ES incorporates two primary mechanisms: difficulty-aware reweighting and targeted replay. The method estimates the local solvability of each problem by analyzing its pass rate across the perturbation population. Historical observations are aggregated to construct a dynamic difficulty state. This shared state provides joint guidance for continuous difficulty reweighting and the replay of rare solvable samples. The design of DART-ES aims to improve the evaluation of perturbation directions and the allocation of training data without requiring an additional difficulty model or backpropagation.

Findings

DART-ES demonstrated improved fine-tuning performance across various benchmarks and models:

  • Overall Performance: DART-ES outperformed standard ES on all five base models evaluated. It improved the average accuracy from 72.07% to 73.53%.
  • GSM8K Benchmark: On the GSM8K benchmark, DART-ES achieved 73.53% average accuracy, exceeding the 73.26% achieved by GRPO.
  • Mathematical Reasoning: Across five challenging mathematical reasoning benchmarks, DART-ES achieved an average accuracy of 49.20%, compared with 48.34% for standard ES. The method remained competitive with strong 7B models trained using Reinforcement Learning (RL).
  • Generalization and Scalability: Further experiments indicated consistent performance gains in instruction tuning, code generation, and the Countdown task when applied to a 14B model. This suggests strong generalization across tasks and scalability to larger models.
  • System Efficiency: DART-ES demonstrated advantages in system efficiency:
    • It reduced runtime per step by 15.2%–50.2% compared with GRPO.
    • It reduced peak memory usage per GPU by 21.1%–51.1% compared with GRPO.
    • Despite performing full parameter fine-tuning, DART-ES also required less runtime and GPU memory than GRPO+LoRA.

Why This Matters

The DART-ES approach offers a method for enhancing the efficiency and performance of full parameter fine-tuning for large language models. By adaptively managing problem difficulty and training data allocation, it addresses limitations in standard Evolution Strategies. The demonstrated improvements in accuracy across diverse tasks, coupled with reductions in runtime and memory usage, indicate a potential for more effective and resource-efficient LLM fine-tuning processes.

Research Information

Institution
arXiv CS
Original Study
View Publication
Source
arXiv CS

About ICANEWS

ICANEWS is a global research journal for emerging researchers, publishing student and emerging researcher work across all fields.