Overview
This study systematically investigates capability scaling-down laws for Large Language Model (LLM) compression. The research aims to understand how different compression methods—pruning, quantization, and distillation—affect model capabilities. It specifically measures capability loss across tasks such as mathematics, code generation, and question answering. The work develops predictive relations that link these capability measurements to various factors, including model size, training stage, compression settings, data availability, and training exposure. The accuracy, measurement efficiency, and generalization of these predictive relations to unseen configurations and model states are evaluated.
Research Context
LLM compression is employed to reduce inference costs and memory requirements. However, selecting an appropriate compression method and its configuration often relies on empirical trials. This empirical approach is necessary because comparable resource reductions can lead to different levels of capability loss. The research addresses this challenge by seeking to establish systematic relationships between compression parameters and the resulting capability changes.
Approach
The research framework measures capability loss across three distinct task domains: mathematics, code generation, and question answering. It correlates these measured losses with several model and compression-related variables:
- Model size
- Training stage
- Specific compression settings
- Data availability
- Training exposure
Simple predictive relations are developed to model these relationships. The evaluation of these predictors focuses on their accuracy, the efficiency with which they reduce the required measurements, and their ability to generalize. The study employs independent evaluations across two model families.
Findings
The investigation yielded several key findings regarding capability scaling-down laws for LLM compression:
-
Density Response and Pruning Prediction: Sharing the density response across different pruning levels can halve the configuration measurements required to fit a pruning predictor. When applied to new Pythia states, pre-registered OLMo-2 test states, and under Wanda pruning, the compact relation achieved a match within 0.020 nats per token compared to a regression fitted with all measurements on math and code tasks. This required refitting coefficients for each specific setting.
-
Distillation Experiment Insights: Controlled distillation experiments indicated that the cost associated with heavy data reuse recurs across various question-answering distributions. The net benefit of distillation, in these cases, was observed to depend on the specific evaluation distribution used.
-
Decision Value of Predictions: The research evaluated the decision value of these predictions by comparing numerical selection methods with configuration medians and fixed method priorities. For question answering tasks within the tested candidate sets, selection using these predictions captured most of the available cross-method benefit. In these instances, a fixed method priority attained the same regret. However, for mathematics and code tasks, the opportunities for such benefits were smaller.
-
Predictive Scope Clarification: The results clarify the predictive scope of capability scaling-down laws and their utility in the selection of compression methods.
Why This Matters
This research provides systematic insights into the relationship between LLM compression techniques and their impact on model capabilities. By developing predictive relations, the study offers a method to reduce the empirical overhead associated with selecting and configuring compression approaches. The findings suggest that a more principled, predictive approach can inform decisions, potentially leading to more efficient selection of compression methods and configurations for specific tasks.