Overview
The research introduces CapField-OPD, an on-policy distillation (OPD) framework designed to consolidate the capabilities of multiple specialized teachers into a single student model. This framework addresses identified limitations in current multi-teacher OPD approaches for flow-based generative models, particularly concerning the implicit binding of desired capabilities to prompt content and the vulnerability to prompt perturbations.
CapField-OPD enables continuous adjustment of desired capability strength during inference by integrating teachers into a continuous capability field using explicit capability coordinates. This method is explored across compositional generation, text rendering, and visual aesthetics, aiming to maintain or improve performance while offering flexible capability control.
Research Context
Prior work in reward-specialized post-training has yielded powerful expert models for flow-based generative processes. Multi-teacher on-policy distillation (OPD) has been employed to merge the distinct capabilities of these experts into a unified student model. However, existing OPD methodologies typically route each prompt to a single teacher based on its semantic category. This approach establishes an implicit link between the desired capability and the prompt's content, which can lead to several challenges. Specifically, this coupling makes the invocation of a particular capability susceptible to prompt perturbations and restricts users from explicitly manipulating the intensity of that capability during inference. The development of CapField-OPD seeks to overcome these specific limitations by decoupling capability invocation from semantic content and enabling continuous control.
Approach
CapField-OPD operates by integrating multiple teacher models into a continuous capability field. This integration is achieved through the use of explicit capability coordinates. The teacher models themselves serve as anchors within this constructed field. The specific coordinates dictate how the outputs from these anchor teachers are combined to generate a consolidated output. This mechanism ensures that each distinct capability configuration receives a unique supervision target during the training process.
A key aspect of this approach is that capability control becomes independent of prompt semantics. Furthermore, the framework incorporates a profiling step for the learned field. This profiling is conducted on a small calibration set. The coordinate that demonstrates the highest mean reward during this calibration is designated as the recommended default. Additionally, coordinates that are frequently identified as optimal are presented as a promising candidate set, intended for test-time scaling.
Findings
- CapField-OPD consolidates multiple specialized teachers into a single student model.
- The student model either preserves or surpasses the performance of the individual specialized teachers.
- The framework reliably invokes desired capabilities even under prompt variations that preserve semantics.
- CapField-OPD supports continuous capability control.
- The method facilitates coordinate-based test-time scaling.
- Extensive experiments were conducted across three distinct domains: compositional generation, text rendering, and visual aesthetics, consistently demonstrating these capabilities.