Overview
This research investigates weak-to-strong (W2S) generalization within preference learning, particularly under conditions of zero-shot distribution shift. The study highlights that strong models fine-tuned with weak preference labels can exhibit successful performance within their training distribution but fail to generalize when presented with different preference datasets. A specific representational failure mode is identified, wherein weak-supervised fine-tuning processes can cause the strong model to converge on source-domain features, thereby hindering the development of broadly transferable preference representations.
Research Context
Weak-to-strong generalization is posited as a framework for scalable oversight in artificial intelligence. However, existing evaluations of W2S systems frequently test 'students' (strong models) using train-test distributions that are matched. This study addresses a gap by examining W2S preference learning when zero-shot distribution shift is introduced, where the test distribution differs from the training distribution without any direct overlap or prior exposure during training.
Findings
- Strong models trained with weak preference labels can demonstrate success within their original training distribution (in-distribution performance).
- These same models can fail to transfer effectively when applied across different preference datasets, particularly under zero-shot distribution shift.
- The identified failure mode is representational: weak-supervised fine-tuning can lead the strong model to prioritize source-domain features, thereby preventing the retention of broadly transferable preference representations.
- To counter this, a regularizer named Representation Anchoring (Anchor) is proposed. Anchor functions by constraining excessive deviation from the pretrained strong model's original representation space during the fine-tuning process.
- Anchor facilitates task-relevant adaptation while mitigating representational drift.
- Evaluation across various preference domains, datasets, and model families indicates that Anchor consistently enhances out-of-distribution transfer.
- Concurrently, Anchor maintains competitive performance in-distribution.
Why This Matters
The described evaluation protocol, along with the introduced transfer-aware metrics and the proposed Representation Anchoring method, reveals inherent brittleness in current W2S reward modeling approaches when confronted with distribution shifts. These elements collectively offer a practical direction toward achieving more robust preference transfer capabilities in weak-to-strong generalization systems.