About Me

Hi! I’m Ni Mu (η‰Ÿε€ͺ in Chinese), a fourth-year Ph.D. student at Department of Automation, Tsinghua University, where I am fortunate to be advised by Prof. Qing-Shan Jia and working closely with Dr. Yiqin Yang.

Before that, I received my Bachelor’s degree from Gifted-Young Class and School of Computer Science and Engineering at Southeast University in 2023.


Research Interests

My research is primarily focused on reinforcement learning (RL) and large language models (LLMs), driven by a single vision: building agents that unleash their full capabilities in the real world, while remaining faithfully aligned with human values.

I am particularly interested in:

  • Preference-based reinforcement learning, making human guidance genuinely useful: identifying human-evaluable trajectory segments (CLARIFY), ensuring comparisons within the same task semantic (STAIR), evaluating along skill axes (S-EPOA, COLLIE), and extending preferences to multiple objectives (PbMORL) and multiple agents (MAGPIE) with theoretical guarantees.
  • LLMs for decision-making: benchmarking long-horizon reasoning (SunTzu) and orchestrating interpretable policy libraries (IFAC 2026, CASE 2026).
  • Real-world RL: Deploying robust and safe RL in complex environments, including data-center cooling control (CASE 2024) and autonomous driving (Safe-GPI, IEEE CASE 2025 WiRA Paper).

News

  • πŸŽ‰ June, 2026. Two papers on LLMs for industrial control were accepted to IFAC 2026 and IEEE CASE 2026.
  • 🐢 May, 2026. COLLIE was accepted to ICML 2026.
  • πŸ“Ά September, 2025. STAIR was accepted to NeurIPS 2025.
  • πŸ’ August, 2025. Honored to receive the Outstanding WiRA Student Paper Award at IEEE CASE 2025.
  • πŸ“„ July, 2025. PbMORL was accepted to IEEE T-ASE, which is the first PbRL framework for multi-objective settings with theoretical guarantees.
  • πŸ”Ž May, 2025. CLARIFY was accepted to ICML 2025.
  • πŸš€ April, 2025. S-EPOA was accepted to IJCAI 2025, which is the first work to tackle indistinguishable behaviors in PbRL.