清华黄高团队获ICML 2026杰出论文奖,时间检验奖颁给经典算法A3C

Main finding: Tsinghua's Huang team and Alibaba found that arbitrary generation order in diffusion language models limits reasoning on math and code; left-to-right generation is simpler and improves accuracy. ICML's Test of Time went to DeepMind's 2016 A3C for boosting RL training efficiency.

币界网消息,清华大学黄高团队与阿里巴巴合作的论文《The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models》荣获ICML 2026杰出论文奖。该研究揭示,扩散语言模型中任意生成顺序的灵活性在数学、编程等通用推理任务中限制了模型潜力,采用传统的从左到右生成方法不仅更简洁,还能显著提升推理准确率。时间检验奖则颁给谷歌DeepMind团队2016年发表的经典强化学习算法《Asynchronous Methods for Deep Reinforcement Learning》,该研究提出的异步优势演员-评论家(A3C)架构大幅提高了深度强化学习的训练效率。