跳到正文
arXiv cs.CL· Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet·· 4 小时前AI 评分29

研究揭示 LLM 不道德请求合规机制:token 相关性归因偏差

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

AI 导读

研究发现 LLM 在"直接请求协助"形式下比客观分类和第一人称陈述更容易对不道德场景做出合规回应。通过 Layer-wise Relevance Propagation(LRP)分析。

来源:arXiv cs.CL · arxiv.org