publications

2026

  1. ACCORD: Action-Conditioned Contextual Grounding for Language Agents
    Lai Jiang, Cheng Qian, Zhenhailong Wang, Pan Lu, Heng Ji, and Hao Peng
    2026
  2. MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning
    Yuhao Shen, Lang Cao, Simo Du, Yuqing Wang, Juexiao Zhou, Hao Peng, and Yue Guo
    2026
  3. CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists
    Junlin Yang, Dylan Zhang, Xiangchen Song, Qirun Dai, Xiao Liu, Yuen Chen, Aniket Vashishtha, Jing Shi, Chenhao Tan, and Hao Peng
    2026
  4. Towards a Universal Causal Reasoner
    Qirun Dai, Xiao Liu, Jiawei Zhang, Dylan Zhang, Hao Peng, and Chenhao Tan
    2026
  5. RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
    Yufeng Du, Phillip Harris, Minyang Tian, Eliu A. Huerta, Srikanth Ronanki, Subendhu Rongali, Aram Galstyan, and Hao Peng
    2026
  6. Useful Memories Become Faulty When Continuously Updated by LLMs
    Dylan Zhang, Yanshan Lin, Zhengkun Wu, Yihang Sun, Bingxuan Li, Dianqi Li, and Hao Peng
    2026
  7. Improving Clinical Diagnosis with Counterfactual Multi-Agent Reasoning
    Zhiwen You, Xi Chen, Aniket Vashishtha, Simo Du, Gabriel Erion-Barner, Hongyuan Mei, Hao Peng, and Yue Guo
    2026
  8. Tackling Distractor Documents in Multi-Hop QA with Reinforcement and Curriculum Learning
    Jerry Huang, Siddarth Madala, Risham Sidhu, Cheng Niu, Hao Peng, Julia Hockenmaier, and Tong Zhang
    In Findings of the Association for Computational Linguistics: EACL 2026, 2026
  9. Process Reward Models That Think
    Muhammad Khalifa, Rishabh Agarwal, Lajanugen Logeswaran, Jaekyeom Kim, Hao Peng, Moontae Lee, Honglak Lee, and Lu Wang
    Transactions on Machine Learning Research (TMLR), 2026
  10. Countdown-Code: A Testbed for Studying the Emergence and Generalization of Reward Hacking in RLVR
    Muhammad Khalifa, Zohaib Khan, Omer Tafveez, Hao Peng, and Lu Wang
    2026
  11. Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion
    Pengcheng Jiang, Judith Yue Li, Moonkyung Ryu, R. Lily Hu, Kun Su, Zhong Yi Wan, Liam Hebert, Hao Peng, Jiawei Han, Dima Kuzmin, and 1 more author
    In Proceedings of the International Conference on Machine Learning (ICML), 2026
  12. Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
    Yangyi Chen, Hao Peng, Tong Zhang, and Heng Ji
    Transactions on Machine Learning Research (TMLR), 2026
  13. Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
    Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Sungryull Sohn, Yunxiang Zhang, Moontae Lee, Hao Peng, Lu Wang, and Honglak Lee
    2026
  14. Generalization of RLVR Using Causal Reasoning as a Testbed
    Brian Lu, Hongyu Zhao, Shuo Sun, Hao Peng, Rui Ding, and Hongyuan Mei
    In Proceedings of the International Conference on Learning Representations (ICLR), 2026
  15. oral
    Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
    Sagnik Mukherjee, Lifan Yuan, Pavan Jayasinha, Dilek Hakkani-Tür, and Hao Peng
    In Proceedings of the International Conference on Machine Learning (ICML), 2026
  16. Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning
    Dylan Zhang, Yufeng Xu, Haojin Wang, Qingzhi Chen, and Hao Peng
    In Proceedings of the International Conference on Machine Learning (ICML), 2026
  17. RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
    Zhiyuan Zeng, Hamish Ivison, Yiping Wang, Lifan Yuan, Shuyue Stella Li, Zhuorui Ye, Siting Li, Jacqueline He, Runlong Zhou, Tong Chen, and 7 more authors
    In Proceedings of the International Conference on Machine Learning (ICML), 2026
  18. From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones
    Lifan Yuan, Weize Chen, Yuchen Zhang, Ganqu Cui, Hanbin Wang, Ziming You, Ning Ding, Zhiyuan Liu, Maosong Sun, and Hao Peng
    In Proceedings of the International Conference on Learning Representations (ICLR), 2026
  19. Executable Counterfactuals: Improving LLMs’ Causal Reasoning Through Code
    Aniket Vashishtha, Qirun Dai, Hongyuan Mei, Amit Sharma, Chenhao Tan, and Hao Peng
    In Proceedings of the International Conference on Learning Representations (ICLR), 2026
  20. oral
    mCLM: A Modular Chemical Language Model that Generates Functional and Makeable Molecules
    Carl Edwards, Chi Han, Gawon Lee, Thao Nguyen, Sara Szymkuć, Chetan Kumar Prasad, Bowen Jin, Jiawei Han, Ying Diao, Ge Liu, and 4 more authors
    In Proceedings of the International Conference on Learning Representations (ICLR), 2026
  21. Process Reinforcement through Implicit Rewards
    Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Yuchen Zhang, Jiacheng Chen, Wendi Li, Bingxiang He, Yuchen Fan, Tianyu Yu, and 15 more authors
    Transactions on Machine Learning Research (TMLR), 2026

2025

  1. Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
    Pengcheng Jiang, Jiacheng Lin, Zhiyi Shi, Zifeng Wang, Luxi He, Yichen Wu, Ming Zhong, Peiyang Song, Qizheng Zhang, Heng Wang, and 24 more authors
    2025
  2. A Survey of Data Attribution: Methods, Applications, and Evaluation in the Era of Generative AI
    Junwei Deng, Yuzheng Hu, Pingbang Hu, Ting-Wei Li, Shixuan Liu, Jiachen T. Wang, Dan Ley, Qirun Dai, Benhao Huang, Jin Huang, and 18 more authors
    2025
    Available at SSRN
  3. Scaling Diffusion Language Models via Adaptation from Autoregressive Models
    Shansan Gong, Shivam Agarwal, Yizhe Zhang, Jiacheng Ye, Lin Zheng, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han, and 2 more authors
    In Proceedings of the International Conference on Learning Representations (ICLR), 2025
  4. Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
    Minhui Zhu, Minyang Tian, Xiaocheng Yang, Tianci Zhou, Penghao Zhu, Eli Chertkov, Shengyan Liu, Yufeng Du, Lifan Yuan, Ziming Ji, and 55 more authors
    2025
  5. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
    Yufeng Du, Minyang Tian, Srikanth Ronanki, Subendhu Rongali, Sravan Babu Bodapati, Aram Galstyan, Azton Wells, Roy Schwartz, Eliu A Huerta, and Hao Peng
    In Findings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025
  6. spotlight
    The Best Instruction-Tuning Data are Those That Fit
    Dylan Zhang, Qirun Dai, and Hao Peng
    In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2025
  7. The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
    Shivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han, and Hao Peng
    In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2025
  8. Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
    Sagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur, and Hao Peng
    In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2025
  9. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
    Ganqu Cui, Yuchen Zhang, Jiacheng Chen, Lifan Yuan, Zhi Wang, Yuxin Zuo, Haozhan Li, Yuchen Fan, Huayu Chen, Weize Chen, and 7 more authors
    2025
  10. Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities
    Qirun Dai, Dylan Zhang, Jiaqi W. Ma, and Hao Peng
    In Findings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025
  11. LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language
    Yubin Ge, Neeraja Kirtane, Hao Peng, and Dilek Hakkani-Tür
    2025
  12. Free Process Rewards without Process Labels
    Lifan Yuan, Wendi Li, Huayu Chen, Ganqu Cui, Ning Ding, Kaiyan Zhang, Bowen Zhou, Zhiyuan Liu, and Hao Peng
    In Proceedings of the International Conference on Machine Learning (ICML), 2025
  13. FactCheckmate: Preemptively Detecting and Mitigating Hallucinations in LMs
    Deema Alnuhait, Neeraja Kirtane, Muhammad Khalifa, and Hao Peng
    In Findings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025
  14. A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
    Suyu Ge, Xihui Lin, Yunan Zhang, Jiawei Han, and Hao Peng
    In Proceedings of the International Conference on Learning Representations (ICLR), 2025
  15. oral
    Retrieval Head Mechanistically Explains Long-Context Factuality
    Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, and Yao Fu
    In Proceedings of the International Conference on Learning Representations (ICLR), 2025
  16. Advancing LLM Reasoning Generalists with Preference Trees
    Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, and 5 more authors
    In Proceedings of the International Conference on Learning Representations (ICLR), 2025
  17. OpenHands: An Open Platform for AI Software Developers as Generalist Agents
    Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, and 14 more authors
    In Proceedings of the International Conference on Learning Representations (ICLR), 2025
  18. Eliminating Position Bias of Language Models: A Mechanistic Approach
    Ziqi Wang, Hanlin Zhang, Xiner Li, Kuan-Hao Huang, Chi Han, Shuiwang Ji, Sham M. Kakade, Hao Peng, and Heng Ji
    In Proceedings of the International Conference on Learning Representations (ICLR), 2025

2024

  1. Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals
    Yanai Elazar, Bhargavi Paranjape, Hao Peng, Sarah Wiegreffe, Khyathi Chandu, Vivek Srikumar, Sameer Singh, and Noah A. Smith
    In Findings of the Association for Computational Linguistics: EMNLP 2024, 2024
  2. ActionIE: Action Extraction from Scientific Literature with Programming Languages
    Xianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong, Tingfeng Luo, Qirong Ho, Hao Peng, Heng Ji, and Jiawei Han
    In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024
  3. S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
    Xihui Lin, Yunan Zhang, Suyu Ge, Liliang Ren, Barun Patra, Vishrav Chaudhary, Hao Peng, and Xia Song
    2024
  4. SciCode: A Research Coding Benchmark Curated by Scientists
    Minyang Tian, Luyu Gao, Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, and 19 more authors
    In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024
  5. A Single Transformer for Scalable Vision-Language Modeling
    Yangyi Chen, Xingyao Wang, Hao Peng, and Heng Ji
    Transactions on Machine Learning Research (TMLR), 2024
  6. PLUM: Preference Learning Plus Test Cases Yields Better Code Language Models
    Dylan Zhang, Shizhe Diao, Xueyan Zou, and Hao Peng
    arXiv preprint, 2024
  7. Source-Aware Training Enables Knowledge Attribution in Language Models
    Muhammad Khalifa, David Wadden, Emma Strubell, Honglak Lee, Lu Wang, Iz Beltagy, and Hao Peng
    In Proceedings of the Conference on Language Modeling (COLM), 2024
  8. Language Models Hallucinate, but May Excel at Fact Verification
    Jian Guan, Jesse Dodge, David Wadden, Minlie Huang, and Hao Peng
    In Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL), 2024
  9. best paper
    honorable mention
    LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
    Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang
    In Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL), 2024
  10. Executable Code Actions Elicit Better LLM Agents
    Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji
    In Proceedings of the International Conference on Machine Learning (ICML), 2024
  11. Data Engineering for Scaling Language Models to 128K Context
    Yao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Yoon Kim, and Hao Peng
    2024
  12. Examining LLMs’ Uncertainty Expression Towards Questions Outside Parametric Knowledge
    Genglin Liu, Xingyao Wang, Lifan Yuan, Yangyi Chen, and Hao Peng
    arXiv preprint, 2024
  13. spotlight
    TRAM: Bridging Trust Regions and Sharpness Aware Minimization
    Tom Sherborne, Naomi Saphra, Pradeep Dasigi, and Hao Peng
    In Proceedings of the International Conference on Learning Representations (ICLR), 2024
  14. MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
    Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji
    In Proceedings of the International Conference on Learning Representations (ICLR), 2024
  15. CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets
    Lifan Yuan, Yangyi Chen, Xingyao Wang, Yi R. Fung, Hao Peng, and Heng Ji
    In Proceedings of the International Conference on Learning Representations (ICLR), 2024
  16. LeTI: Learning to Generate from Textual Interactions
    Xingyao Wang, Hao Peng, Reyhaneh Jabbarvand, and Heng Ji
    In Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL), 2024

2023

  1. FiLM: Fill-in Language Models for Any-Order Generation
    Tianxiao Shen, Hao Peng, Ruoqi Shen, Yao Fu, Zaid Harchaoui, and Yejin Choi
    arXiv preprint, 2023
  2. Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback
    Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata
    arXiv preprint, 2023
  3. Efficiency Pentathlon: A Standardized Arena for Efficiency Evaluation
    Hao Peng, Qingqing Cao, Jesse Dodge, Matthew E. Peters, Jared Fernandez, Tom Sherborne, Kyle Lo, Sam Skjonsberg, Emma Strubell, Darrell Plessas, and 4 more authors
    arXiv preprint, 2023
  4. Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models’ Reasoning Performance
    Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot
    arXiv preprint, 2023
  5. oral
    Specializing Smaller Language Models towards Multi-Step Reasoning
    Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, and Tushar Khot
    In Proceedings of the International Conference on Machine Learning (ICML), 2023
  6. Complexity-Based Prompting for Multi-step Reasoning
    Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot
    In Proceedings of the International Conference on Learning Representations (ICLR), 2023
  7. Transparency Helps Reveal When Language Models Learn Meaning
    Zhaofeng Wu, William Merrill, Hao Peng, Iz Beltagy, and Noah A. Smith
    Transactions of the Association for Computational Linguistics (TACL), 2023

2022

  1. How Much Does Attention Actually Attend? Questioning the Importance of Attention in Pretrained Transformers
    Michael Hassid, Hao Peng, Daniel Rotem, Jungo Kasai, Ivan Montero, Noah A. Smith, and Roy Schwartz
    In Findings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022
  2. Modeling Context With Linear Attention for Scalable Document-Level Translation
    Zhaofeng Wu, Hao Peng, Nikolaos Pappas, and Noah A. Smith
    In Findings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022
  3. Twist Decoding: Diverse Generators Guide Each Other
    Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Hao Peng, Ximing Lu, Dragomir Radev, Yejin Choi, and Noah A. Smith
    In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022
  4. ABC: Attention with Bounded-memory Control
    Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, and Noah A. Smith
    In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2022
  5. Tailor: Generating and Perturbing Text with Semantic Controls
    Alexis Ross, Tongshuang Wu, Hao Peng, Matthew Peters, and Matt Gardner
    In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2022

2021

  1. Finetuning Pretrained Transformers into RNNs
    Jungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen, and Noah A. Smith
    In In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021
  2. spotlight
    Random Feature Attention
    Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong
    In Proceedings of the International Conference on Learning Representations (ICLR), 2021
  3. Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
    Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah Smith
    In Proceedings of the International Conference on Learning Representations (ICLR), 2021
  4. Contextualized Perturbation for Textual Adversarial Attack
    Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan
    In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2021
  5. Infusing Finetuning with Semantic Dependencies
    Zhaofeng Wu, Hao Peng, and Noah A. Smith
    Transactions of the Association for Computational Linguistics (TACL), 2021

2020

  1. A Mixture of h - 1 Heads is Better than h Heads
    Hao Peng, Roy Schwartz, Dianqi Li, and Noah A. Smith
    In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2020

2019

  1. PaLM: A Hybrid Parser and Language Model
    Hao Peng, Roy Schwartz, and Noah A. Smith
    In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2019
  2. RNN Architecture Learning with Sparse Regularization
    Jesse Dodge, Roy Schwartz, Hao Peng, and Noah A. Smith
    In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2019
  3. Text Generation with Exemplar-based Adaptive Decoding
    Hao Peng, Ankur Parikh, Manaal Faruqui, Bhuwan Dhingra, and Dipanjan Das
    In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2019

2018

  1. Rational Recurrences
    Hao Peng, Roy Schwartz, Sam Thomson, and Noah A. Smith
    In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2018
  2. best paper
    honorable mention
    Backpropagating through Structured Argmax using a SPIGOT
    Hao Peng, Sam Thomson, and Noah A. Smith
    In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2018
  3. Learning Joint Semantic Parsers from Disjoint Data
    Hao Peng, Sam Thomson, Swabha Swayamdipta, and Noah A. Smith
    In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2018
  4. "You Are No Jack Kennedy": On Media Selection of Highlights from Presidential Debates
    Chenhao Tan, Hao Peng, and Noah A. Smith
    In Proceedings of The Web Conference (WWW), 2018

2017

  1. Deep Multitask Learning for Semantic Dependency Parsing
    Hao Peng, Sam Thomson, and Noah A. Smith
    In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2017

2016

  1. A Convolutional Attention Network for Extreme Summarization of Source Code
    Miltiadis Allamanis, Hao Peng, and Charles Sutton
    In Proceedings of the International Conference on Machine Learning (ICML), 2016

2015

  1. Discriminative Neural Sentence Modeling by Tree-Based Convolution
    Lili Mou, Hao Peng, Ge Li, Yan Xu, Lu Zhang, and Zhi Jin
    In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2015
  2. Classifying Relations via Long Short Term Memory Networks along Shortest Dependency Paths
    Yan Xu, Lili Mou, Ge Li, Yunchuan Chen, Hao Peng, and Zhi Jin
    In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2015