arXiv cs.AI· Yuhan Chen, Siyuan Zhang, Nan Wang, Feiyang Kang, Ruoxi Jia·· 12 小时前AI 评分29
Hybrid Latent Attention:为循环语言模型压缩 KV cache
Hybrid Latent Attention for Looped Language Models
AI 导读
针对循环语言模型重复堆叠层导致 KV cache 扩大 T 倍的问题,研究者提出 Hybrid Latent Attention(HLA),在滑动窗口内保留精确 key/value,窗口外旧 token 压缩为 latent 供各轮循环的 query 直接读取。
来源:arXiv cs.AI · arxiv.org