Overview
Research into role-specialized Question Answering (QA) pipelines has focused on the transfer of rationales from a reasoner component to a verifier component. The utility of this rationale message, specifically whether it leads to improved answers, more robust support assessment, or introduces new failure modes, has been unclear. This study introduced a message-intervention diagnostic to explore these aspects.
Research Context
Role-specialized QA pipelines are characterized by the transfer of rationales between distinct components. The precise impact of this inter-component communication on downstream tasks, particularly on the verifier's function, constituted the primary research question. The study aimed to clarify the effect of varying rationales on verifier judgments when evidence and candidate answers were held constant.
Approach
A message-intervention diagnostic was developed to isolate the effect of the rationale. This diagnostic involved fixing the evidence and candidate answer components while systematically varying only the rationale transmitted across the reasoner-to-verifier boundary. The study utilized DeepSeek as both the generator and verifier. Evaluations were conducted on a dataset comprising 400 examples drawn from MuSiQue, HotpotQA, and 2WikiMultiHopQA.
Intervention Types
- Faithful rationales: These represented valid or correct rationales.
- Corrupted rationales: These were altered or incorrect rationales.
- Harmless paraphrases: These involved slight linguistic variations of rationales that were expected to have minimal impact.
Prompting Strategies
- Blind verifier prompt: The verifier processed information without explicit instructions to scrutinize the rationale.
- Explicit rationale-checking prompt: The verifier was specifically instructed to check the provided rationale.
Human Audits
Human evaluations were performed to assess the model's performance in specific scenarios. This involved auditing instances of corrupted rationales to understand human perception and model behavior.
Findings
The study yielded several key observations regarding rationale communication:
Impact on Answer Accuracy
- Faithful rationales added almost no answer accuracy when compared to scenarios where no rationale was provided.
Impact on Support Judgments
- Corrupted rationales strongly altered support judgments made by the verifier.
- Under a blind verifier prompt, harmless paraphrases shifted support judgments by 0–2.5%.
- Corrupted rationales, under a blind verifier prompt, shifted support judgments by 10–22%.
- An explicit rationale-checking prompt amplified this pattern, with corrupted rationales shifting support judgments by 34–55%.
Impact on Final Answers
- Final answers moved less significantly than support judgments, showing shifts of 2–30%.
- Only 2.9–35.3% of corrupted support flips co-occurred with changes in the final answer.
Human Audit Results
- Human audits revealed that 16 out of 42 valid corruptions constituted cases of corruption-overtrust by the model.
- Blind humans rejected or marked as unclear 9 out of 10 audited corrupted rationales that the model accepted.
Channel Activity Across Models and Tasks
Cross-model and task-boundary checks indicated conditions under which the rationale channel was active, amplified, inert, or folded into the task label.
Why This Matters
The study's findings indicate that rationale sharing within QA pipelines should primarily be evaluated as a mechanism for verification messages, rather than solely as a means to achieve higher answer accuracy. The differential impact of rationales on support judgments versus final answers, coupled with the observed corruption-overtrust by models compared to human discernment, highlights the complexities of reliable rationale integration.