publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

2026

  1. chasing-public-score.png
    Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
    Hardy Chen, Nancy Lau, Haoqin Tu, and 8 more authors
    arXiv preprint arXiv:2604.20200, 2026
  2. reasoning-while-asking.png
    Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers
    Xin Chen, Feng Jiang, Yiqian Zhang, and 5 more authors
    arXiv preprint arXiv:2601.22139, 2026

2025

  1. lmr-bench.png
    LMR-BENCH: Evaluating LLM Agent’s Ability on Reproducing Language Modeling Research
    Shuo Yan, Ruochen Li, Ziming Luo, and 12 more authors
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Nov 2025
  2. idea.png
    Idea: Enhancing the rule learning ability of large language model agent through induction, deduction, and abduction
    Kaiyu He, Mian Zhang, Shuo Yan, and 2 more authors
    In Findings of the Association for Computational Linguistics: ACL 2025, 2025
  3. mllm-bench.png
    Mllm-bench: evaluating multimodal llms with per-sample criteria
    Wentao Ge, Shunian Chen, Hardy Chen, and 8 more authors
    In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025