Overview
Research introduces GAPA (Gender Associations of Physical Attributes), a dataset designed to investigate whether physical descriptions, often recommended as alternatives to explicit identity labels in AI systems, achieve gender-neutral communication. The study empirically examines human interpretation of such descriptions and evaluates the alignment of Large Language Models (LLMs) with these human associations. Findings indicate that physical descriptions retain systematic gender associations for human readers and that LLMs exhibit specific biases in recovering these associations.
Research Context
Recent work in AI fairness, accessibility, and ethics has suggested that foundation models should avoid inferred identity labels (e.g., "she," "his") when describing people. Instead, these guidelines recommend employing seemingly "objective" physical descriptions (e.g., "short hair," "a defined jawline"). The empirical question of whether such descriptive language genuinely achieves gender-neutral communication remained unaddressed prior to this study.
Approach
The study developed GAPA, a dataset comprising 316 common physical attributes. These attributes were gathered from diverse sources. Human gender-association ratings were collected for these attributes from 304 US-based annotators, resulting in 14,706 individual ratings.
Subsequently, 16 LLMs were evaluated against these human ratings. The selected LLMs represented various model families, sizes, and post-training variants. The evaluation aimed to assess the models' ability to recover human gender associations from descriptive language.
Finally, the researchers developed and released a proxy model. This model was trained to predict human gender associations of descriptive language. Its utility was demonstrated through a sociolinguistic analysis applied to character descriptions within the LitBank corpus.
Findings
- Human readers assign structured and graded gender associations to physical descriptions.
- These gender associations are more consistent and distinctive for descriptions associated with women and men compared to those associated with non-binary identities.
- Large Language Models partially recover the gender associations observed in human ratings.
- LLMs exhibit systematic alignment biases when processing these associations.
- Specific model biases include compressed rating distributions, indicating a narrower range of association strengths than humans.
- Alignment between LLM predictions and human associations was weaker for attributes associated with men.
- Models demonstrated asymmetric abstention, disproportionately targeting the non-binary category by declining to provide associations.
- The study provides the first empirical evidence that ostensibly "objective" physical descriptions can maintain systematic gender associations in human interpretation.
- Systematic patterns of model-human misalignment were uncovered.
Why This Matters
The findings challenge the assumption that replacing explicit gender labels with physical descriptions necessarily results in gender-neutral communication. This has implications for human-AI interaction, highlighting downstream challenges in using physical descriptions to communicate subjective identity categories within these systems.
Potential Applications
The released proxy model, trained to predict human gender associations of descriptive language, was demonstrated to be useful through a sociolinguistic analysis of character descriptions in LitBank. This suggests its potential application in analyzing and understanding gender representations in textual data.