כתבה
arXiv cs.AI ·
Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents
תקציר מקורי באנגליתarXiv:2609.36611v1 Announce Type: new Abstract: Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, confidence and search respond to the labels as intended in 49 of 50 comparisons, yet a label change alone alters 4-23% of confident verdicts for Qwen3 mode
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית