跳到正文
arXiv cs.CL· Mohamed Dhouib, Clement Elliker, Alexi Canesse, Ma\"el Jenny, Lucas-Andrei Thil, Mahammed El Sharkawy, Sonia Vanier, Elie Bursztein·· 4 小时前AI 评分31

RAISED:用自蒸馏提升 LLM 智能体对提示词注入的鲁棒性

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

AI 导读

针对工具调用型 LLM 智能体易受间接提示词注入攻击的问题,研究者提出训练框架 RAISED(Robust Attack Invariance through Self-Distillation),结合自生成与自蒸馏:模型先自行生成工具使用场景,再通过自蒸馏让学生在干净与注入轨迹上对齐教师的干净上下文行为。

来源:arXiv cs.CL · arxiv.org