Overview
Research indicates that fine-tuning with benign data can compromise safety alignment in Audio Large Language Models (LLMs), paralleling observations in text and vision modalities. This study presents the first systematic investigation into the safety implications of benign fine-tuning specifically within Audio LLMs. It identifies a vulnerability where benign samples, adjacent to harmful content via spoken words or sound characteristics, can lead to elevated Jailbreak Success Rates (JSRs).
Research Context
Prior work established that benign data fine-tuning degrades safety alignment in text and vision LLMs. However, it remained unclear whether the distinct properties of audio input differentially influence this vulnerability. Audio presents a unique challenge where benign samples can exhibit proximity to harmful content through two dimensions: what is explicitly said (semantic content) and how it sounds (acoustic properties).
Approach
The study evaluated three state-of-the-art Audio LLMs. It employed a proximity-based framework designed to decompose embedding-space distance. This framework segregated the distance into three axes: semantic, acoustic, and mixed. This methodology aimed to discern which input properties drive the observed degradation in safety alignment.
Findings
- Benign fine-tuning significantly elevated the Jailbreak Success Rate (JSR) across the evaluated Audio LLMs. JSR increased from single-digit percentages to as high as 87%.
- The dominant axis of vulnerability was determined to be architecture-conditioned. This axis is specifically influenced by how each model's encoder and projector components transform audio input into the backbone LLM's input space.
- Depending on the encoder design, the most damaging proximity axis shifted from semantic to acoustic.
- Mechanistically, benign fine-tuning was observed to selectively suppress late-layer refusal circuits within the models. Simultaneously, frozen encoders preserved upstream representations of harmful content. This suggests a recognition-refusal dissociation, where the model still detects harmful content but ceases to refuse it.
- Two practical defense strategies were identified:
- Filtering training data to maximize distance from harmful embeddings.
- Implementing a textual system prompt during inference.
- These defense mechanisms were effective in reducing JSR to near-zero without requiring architectural modifications to the models.
Why This Matters
The findings indicate that safety evaluation protocols for LLMs should incorporate considerations for modality-specific characteristics and architectural designs. Audio LLMs serve as a valuable testbed for understanding the broader fragility of alignment mechanisms in LLMs, providing insights into how distinct input properties and model architectures contribute to safety vulnerabilities.