Overview
This research investigates the challenge of data contamination in large language models (LLMs), particularly when the contaminating content is presented in a different language than the evaluation benchmark. Data contamination can lead to inflated benchmark scores by allowing models to memorize evaluation content rather than exhibiting genuine generalization capabilities. The study highlights that auditing for such contamination becomes difficult when exposed content differs in language from the evaluation benchmark.
Research Context
The core problem addressed is the invalidation of benchmark evaluations due to data contamination. While contamination is known to enable models to benefit from memorized evaluation content, its detection is complicated when the exposed content's language differs from the evaluation benchmark. Traditional post-hoc contamination probes, typically designed for English content, may fail in multilingual scenarios.
Approach
The researchers employed a controlled experimental setup to study this failure mode. They deliberately exposed four open-weight instruction-tuned LLMs to Arabic translations of evaluation items from the MMLU and XQuAD benchmarks. The exposure levels were varied systematically. Following exposure, the models were evaluated on the original English tasks. This controlled setup was framed as a proxy for contamination, not a reconstruction of real-world pretraining leakage.
Two English-centric post-hoc probes, TS-Guessing and Min-K%++, were initially tested. The study then introduced Translation-Aware Contamination Detection (TACD), a training-data-free diagnostic. TACD is based on two principles: cross-lingual prediction consistency and choice reordering.
Findings
- English-centric Probe Performance: The signals from the two English-centric post-hoc probes, TS-Guessing and Min-K%++, largely disappeared under translated exposure. TS-Guessing remained weak, with an exception for model-specific positional recall observed on MMLU. Min-K%++ consistently performed at or below chance levels.
- Impact on English MMLU Performance: Despite the absence of a strong English contamination signal from the probes, English MMLU performance increased with Arabic exposure. This finding indicates that the lack of an English contamination signal does not equate to the absence of an exposure effect.
- Translation-Aware Contamination Detection (TACD) Results: Cross-lingual consistency, as measured by TACD, was substantially higher than an independence baseline. This consistency generally increased relative to the clean condition. However, the magnitude of this increase was model-dependent and not strictly monotonic.
- Implication for Contamination Detection: These results suggest that translation can obscure contamination-related effects from probes designed solely for English content.
Why This Matters
The study highlights a critical vulnerability in current LLM evaluation practices: multilingual data contamination can evade detection by English-only probes. The observed increase in English MMLU performance following exposure to translated content, despite the probes' failure, indicates that contamination effects can still be present and impact model behavior. The introduction of TACD provides a diagnostic approach that offers evidence of contamination-consistent behavior, framing it as such rather than a definitive membership test. This work motivates the development of multilingual diagnostics to detect contamination in LLMs trained on diverse datasets.