Overview
Research explored the internal representations within text-to-song generation models, specifically investigating how artist identity is encoded. Previous studies have documented behavioral instances of these models imitating specific artists or reproducing training data. This work focused on understanding the underlying internal mechanisms.
Research Context
Prior interpretability efforts in generative audio have largely concentrated on identifying semantic concepts, such as genre or time signature, within model activations. The current study differentiates itself by probing for artist identity representations. The motivation stems from documented behavioral phenomena where text-to-song generation models can be prompted to imitate particular artists or to regurgitate entire songs from their training data. Little is understood about the internal representations that might underpin these behaviors.
Approach
The researchers utilized a controlled case study methodology. Their investigation focused on a specific trained model, ACE-Step 1.5. The dataset for this study comprised 2,000 songs, attributed to 100 distinct artists. The core of the approach involved probing the model to ascertain if linearly decodable representations of artist identity could be extracted solely from song lyrics. This process excluded any additional identifiers beyond the lyrics themselves.
Findings
- A trained model can be probed to reveal linearly decodable representations of artist identity within its internal activations.
- These artist identity representations are extractable from song lyrics alone, without requiring additional identifiers.
- The artist associated with a given set of lyrics can be identified within the model's internal activations.
- This artist-level conditioning signal propagates from the lyric encoder component of the model to its diffusion backbone during the inference process.
- Lyrics constitute an artist-level conditioning channel.
Why This Matters
The findings indicate that lyrics function as an artist-level conditioning channel, a factor not addressed by prompt-side replication safeguards. More broadly, the work illustrates that latent-space analysis offers a method to audit what generative music models have implicitly learned from their training data. This audit capability extends to understanding how specific attributes, like artist identity, are encoded and processed internally.