Overview
A reproducible classification framework has been developed for programming language (PL) resources, categorizing 646 languages into four distinct tiers. This initiative addresses a gap in systematic resource-tier taxonomies for code, contrasting with existing classifications for natural languages (Joshi et al., 2020). The proposed taxonomy is designed to support the development and evaluation of large language models (LLMs) that generate code.
Research Context
The field of natural language processing (NLP) benefits from systematic categorization of languages based on resource availability, acknowledging the wide disparity among the world's 7,000+ languages. A similar disparity exists among programming languages. The increasing capabilities of LLMs in code generation underscore the necessity for such a resource-tier taxonomy within the programming language domain.
Approach
The research established the first reproducible resource classification for programming languages. This involved grouping 646 languages into four tiers based on their resourcefulness. The study then analyzed the distribution of these categorized languages across seven major code corpora.
Findings
- The taxonomy classified 646 programming languages into four tiers.
- Analysis revealed that 1.9% of the classified languages, specifically those in Tier 3 (High), accounted for 74.6% of all tokens present in seven major corpora.
- Conversely, 71.7% of the languages, designated as Tier 0 (Scarce), contributed only 1.0% of the total tokens across the same corpora.
- Statistical analyses, including assessments of within-tier inequality, dispersion, and distributional skew, confirmed that the observed imbalance in resource distribution among programming languages is both extreme and systematic.
Why This Matters
The presented taxonomy offers a principled framework for two key areas: dataset curation and the tier-aware evaluation of multilingual LLMs. This structured classification provides a basis for understanding and addressing the resource disparities among programming languages, particularly as LLMs become more central to code generation tasks.