Overview
Research introduces Mawqif-XT, an Arabic benchmark dataset for cross-target stance detection. This dataset comprises 996 manually annotated Arabic tweets, collected from three distinct public targets: Women Driving, E-Cars, and Trimester System. The primary function of Mawqif-XT is to serve as a held-out evaluation set. It is specifically designed to assess the generalization capabilities of models when confronted with both semantically related and previously unseen targets within the domain of Arabic stance detection. The annotation process for Mawqif-XT followed the established Mawqif annotation scheme, applying labels for stance, sentiment, and sarcasm to each tweet. This new extension, when combined with the original Mawqif dataset (intended for training and development), forms a comprehensive benchmark, termed Mawqif-v2 Extension, for evaluating cross-target generalization in Arabic.
Research Context
The landscape of publicly available Arabic datasets specifically tailored for target-specific stance detection remains limited. A particular gap identified is the scarcity of resources suitable for evaluating cross-target generalization. This limitation directly impedes the development and assessment of models capable of transferring learned stance detection abilities across different topics. The introduction of Mawqif-XT directly addresses this identified need by providing a dedicated resource to advance research in this specific area of natural language processing for Arabic.
Approach
The construction of Mawqif-XT involved several key steps:
- Data Collection: A total of 996 Arabic tweets were gathered.
- Target Selection: The tweets were sourced from discussions pertaining to three public targets: Women Driving, E-Cars, and Trimester System.
- Annotation Scheme: Annotation was conducted following the established Mawqif annotation scheme.
- Labeling: Each collected tweet received manual annotations for three distinct attributes: stance, sentiment, and sarcasm.
- Dataset Role: Mawqif-XT is designated as a held-out evaluation set. This design choice is intended to facilitate the assessment of model generalization.
- Complementary Dataset: The original Mawqif dataset is specified for use in the training and development phases of model evaluation.
To establish a foundational understanding and provide a basis for future comparisons, the research also included the establishment of baseline results. This involved the use of various model types:
- Transformer Models: Both Arabic and multilingual transformer models were employed.
- Zero-shot LLMs: Large Language Models (LLMs) operating in a zero-shot configuration were also utilized.
This approach aims to provide a reproducible evaluation framework for the community.
Findings
The primary outcome of this research is the development and presentation of the Mawqif-XT dataset. This dataset consists of 996 manually annotated Arabic tweets. Each tweet within Mawqif-XT carries annotations for stance, sentiment, and sarcasm, adhering to the original Mawqif annotation scheme. The dataset's specific utility lies in its role as a held-out evaluation set, designed for assessing model generalization across targets that are both semantically related and previously unseen. This facilitates the combined use of Mawqif-XT with the original Mawqif dataset, which is designated for training and development, to form the Mawqif-v2 Extension. Furthermore, the research established baseline results employing several Arabic and multilingual transformer models, alongside zero-shot large language models (LLMs).
Why This Matters
The provision of Mawqif-XT addresses a noted limitation in the availability of Arabic datasets for target-specific stance detection, particularly concerning the evaluation of cross-target generalization. By offering a designated held-out evaluation set, it enables researchers to rigorously test how well models adapt to new or related topics without prior training on those specific targets. This contributes to establishing a standardized benchmark for reproducible evaluation within Arabic stance detection research.
Potential Applications
The Mawqif-XT dataset, in conjunction with the original Mawqif dataset, provides a benchmark for evaluating cross-target generalization in Arabic stance detection. This enables the assessment of models designed to generalize across different topics. The established baseline results, derived from Arabic and multilingual transformer models as well as zero-shot LLMs, offer points of comparison for future research efforts.