Overview
As neural networks continue to increase in scale, the development of model compression techniques becomes increasingly pertinent for achieving efficient inference, particularly within environments characterized by limited computational resources. Existing structured pruning methods reduce the network by removing neurons or channels deemed less significant; however, this approach risks discarding potentially useful information contained within these units. The present research explores the consolidation of trained networks through a coarse-graining perspective, specifically addressing the question of which information should be retained when multiple neuronal degrees of freedom are consolidated.
Research Context
The imperative for model compression arises from the growing scale of neural networks, which necessitates more efficient inference. Traditional structured pruning, while effective in reducing network size, operates by eliminating units. This study investigates an alternative approach: neuron merging, framed as a coarse-graining process, to consolidate neuronal degrees of freedom while preserving essential information.
Approach
The study discusses cluster-based merging methods designed for the compression of trained neural networks. One contribution involves a data-free method that utilizes contribution-weighted averaging. Additionally, the research proposes neuron-merging methods where neuron responses are transformed back into the pre-activation space through the application of the inverse activation function. Subsequently, the weights and biases corresponding to each representative neuron are estimated using the least-squares method.
The investigation also incorporated two distinct strategies for evaluating these methods:
- Data-assisted strategy: This approach utilizes actual training inputs during the merging process.
- Data-free strategy: This alternative employs randomly generated inputs for the merging process.
These comparisons were conducted to provide empirical insights into the efficacy and characteristics of the proposed neuron-merging techniques.
Findings
Empirical evidence derived from the tested sigmoid networks indicates specific utilities for different types of information within the neuron-merging process. The comparisons suggest that:
- Weight information is particularly useful for the clustering phase of the neuron-merging procedure.
- Activation information proves useful for the reconstruction of representative neurons following the merging process.
Why This Matters
The ongoing growth in the scale of neural networks underscores the increasing importance of model compression for enabling efficient inference. Addressing resource limitations, this research contributes methods for consolidating trained networks, aiming to retain valuable information that might be lost in other compression techniques. These insights contribute to the field of efficient neural network deployment.