Overview
Research on capability transfer in vision-language models (VLMs) has explored model merging as a training-free approach to adapt reasoning capabilities initially developed in language models. This method, however, often involves endpoint-based transfer, which can complicate the distinction between pre-existing model differences and specific changes acquired during post-training reasoning. A new formulation addresses this by focusing on the training-stage update, specifically isolating parameter changes induced by reinforcement learning (RL).
While isolating these RL updates provides a clearer understanding of acquired capabilities, transferring the complete update between models can be suboptimal. An observation from this work indicates that components within the RL update exhibit varying degrees of cross-model transferability. Specifically, dominant directions within the update transfer more effectively than the entire update.
Based on this finding, a method termed Selective-RL was developed. This approach isolates the RL-stage update and then retains its dominant matrix-wise directions while preserving their original magnitudes. These selected components are subsequently transferred to the language modules of a recipient VLM.
Research Context
The broader context for this research involves transferring reasoning capabilities from established language models to VLMs. Model merging represents one strategy for this transfer, aiming to integrate capabilities without extensive retraining. However, traditional endpoint-based transfer methods, which transfer final model states, can obscure the specific contributions of post-training interventions by conflating them with inherent architectural or foundational differences between the source and target models.
This work proposes an alternative perspective by focusing on the 'training-stage update.' This involves examining the incremental parameter modifications that occur specifically during reinforcement learning, rather than the complete, merged model state. The aim is to isolate the specific changes that confer enhanced reasoning capabilities, thereby providing a more granular understanding of capability acquisition and transfer.
Approach
The core approach of this research involves isolating the parameter changes generated during the reinforcement learning (RL) phase. This contrasts with endpoint-based transfer, which merges final model states. Following the isolation of the RL-stage update, the methodology identifies and retains its dominant matrix-wise directions. A key aspect of this retention is the preservation of the original magnitudes associated with these dominant directions. These specific, magnitude-preserved, dominant components of the RL update are then transferred exclusively to the language modules of a target VLM.
The effectiveness of Selective-RL was evaluated against full-update interpolation. This comparison aimed to determine if selectively transferring components of the update, rather than the entire update, yielded superior performance. Matched controls were employed to further validate the method, ensuring that any observed gains were attributable to the selective transfer mechanism and not simply due to update magnitude or arbitrary low-rank approximations alone.
Findings
- The components of a reinforcement learning (RL) update differ substantially in their cross-model transferability.
- Dominant directions within an RL update transfer more effectively across models compared to transferring the complete update.
- Selective-RL, which isolates RL-stage updates, retains dominant matrix-wise directions with magnitude preservation, and transfers them to VLM language modules, improved performance.
- Across three model families and five visual-reasoning benchmarks, Selective-RL outperformed full-update interpolation in 12 out of 15 comparisons.
- A specific gain of 8.55 percentage points was observed on the MathVision benchmark for the Qwen recipient when using Selective-RL.
- Matched controls demonstrated that the observed gains were not reproducible by merely applying arbitrary low-rank approximations or simply preserving update magnitude without selectivity.
- The results indicate a distinction between capabilities acquired during post-training and those that are transferable across different models.
Why This Matters
This research offers a training-stage perspective on cross-model capability transfer, distinguishing what is acquired during post-training from what remains effectively transferable between models. By demonstrating that dominant components of RL updates transfer more effectively, the findings refine the understanding of how reasoning capabilities can be efficiently imparted to new models. This provides a more targeted approach for improving VLM performance in visual reasoning tasks by focusing on specific, high-impact parameter changes rather than relying on full-model transfers.
Potential Applications
The method described, Selective-RL, offers a strategy for enhancing the visual reasoning capabilities of vision-language models (VLMs) by specifically targeting the transfer of effective reinforcement learning (RL) updates. This could lead to more efficient capability transfer without requiring extensive retraining of the recipient VLM. The approach could be applied to improve performance across various visual reasoning benchmarks by selectively integrating refined reasoning abilities developed in other language models.
Key Limitations Mentioned by Researchers
The abstract does not explicitly mention limitations of the study. However, it does highlight that transferring the full RL update is suboptimal and that component transferability varies, implying a complexity in full-update transfer that Selective-RL aims to mitigate.