CV

Ni Mu (็‰Ÿๅ€ช)

mn23@mails.tsinghua.edu.cn
Beijing, , CN

Summary

Ph.D. student in the Department of Automation at Tsinghua University, working on reinforcement learning and large language models.

Education

  • Control Science and Engineering (Ph.D.)
    present
    Tsinghua University, Department of Automation
    Courses: Reinforcement Learning, Large Language Models, Preference-based RL, Skill Discovery, Multi-objective RL
  • Computer Science (B.Eng.)
    2023-06
    Southeast University, Youth Class (ๅฐ‘ๅนด็ญ)
    GPA: 3 / 275 (top 1.5%)

Publications

  • CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
    2025
    ICML 2025
    Selects human-comparable segment pairs via contrastive learning in the trajectory embedding space to improve label efficiency in preference-based RL.
  • STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
    2025
    NeurIPS 2025
    Minimizes temporal distance between segment pairs so that preference comparisons stay within the same stage, improving learning efficiency in multi-stage tasks.
  • COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space
    2026
    ICML 2026
    Builds a semantically coherent latent space and constructs training-free guidance signals so that safe, diverse skills can be learned from very few human labels.
  • S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning
    2025
    IJCAI 2025
    Uses unsupervised skill discovery to select segments that differ widely in skill, improving human label efficiency in preference-based RL.
  • Preference-based Multi-Objective Reinforcement Learning
    2025
    IEEE Transactions on Automation Science and Engineering (T-ASE)
    A theoretically grounded multi-objective preference learning paradigm where optimizing the multi-objective preference reward model is equivalent to solving the convex Pareto front in policy space.
  • MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
    2026
    AAAI 2026
    Extracts scenario information via meta-state regularization and aligns model optimization with multi-scenario policy learning via meta-value regularization to build a cross-scenario world model.
  • Safety-Guaranteed Policy Composition Via Generalized Policy Improvement for Autonomous Vehicles
    2025
    IEEE CASE 2025 (Outstanding WiRA Student Paper Award)
    A Generalized-Policy-Improvement-based policy composition with safety and performance guarantees that resolves the performance-safety-sample-efficiency tradeoff for autonomous driving.
  • SunTzu: A Long-Horizon Decision Making Benchmark for Large Language Models
    2025
    Preprint (submitted to NeurIPS 2026)
    A benchmark for evaluating LLMs on long-horizon decision making in StarCraft II, plus a hierarchical agent baseline that self-corrects with expert knowledge and self-improves via SFT.
  • Covering the Pareto Frontier with LLM-Coordinated Interpretable Policy Library
    2026
    IFAC 2026 (23rd IFAC World Congress)
    Covers the Pareto frontier by coordinating an interpretable policy library with an LLM.
  • Learning Interpretable Index Policies via LLM-Driven Zero-Order Optimization
    2026
    IEEE CASE 2026
    Learns interpretable index policies by driving zero-order optimization with an LLM.
  • LLM-Empowered Knowledge Graph Construction for Optimization Formalization in Smart Manufacturing
    2025
    CAC 2025 (China Automation Congress)
    Uses an LLM to build knowledge graphs that formalize optimization problems in smart manufacturing.
  • GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent System
    2026
    ICLR 2026
    A state diffusion process that handles partial observability in multi-agent systems.
  • DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning
    2025
    NeurIPS 2025
    Moves beyond task ambiguity in language-conditioned reinforcement learning.
  • A Transformer-based Thermal Surrogate Model for Cooling Control in Data Centers
    2025
    IEEE Robotics and Automation Letters (RA-L)
    A transformer-based thermal surrogate model enabling efficient cooling control in data centers.
  • E-MAPP: Efficient Multi-Agent Reinforcement Learning with Parallel Program Guidance
    2022
    NeurIPS 2022
    Improves sample efficiency in multi-agent RL by guiding learning with parallel programs.
  • MAGPIE: Utilizing Agent-Specific Preferences for Multi-Agent Reinforcement Learning
    2025
    CAC 2025 (China Automation Congress)
    Leverages agent-specific preferences for multi-agent reinforcement learning.
  • Preference-Based Multi-Objective Reinforcement Learning with Explicit Reward Modeling
    2024
    CAC 2024 (China Automation Congress)
    Extends preference-based multi-objective RL with explicit reward modeling.
  • Large-scale data center cooling control via sample-efficient reinforcement learning
    2024
    IEEE CASE 2024
    Applies sample-efficient reinforcement learning to large-scale data-center cooling control.
  • Integrating mechanism and data: Reinforcement learning based on multi-fidelity model for data center cooling control
    2023
    CAC 2023 (China Automation Congress)
    Combines mechanism models with data via a multi-fidelity model for RL-based data-center cooling control.
  • Addressing Coupling in Restless Multi-Armed Bandits by Finetuning Whittle Index
    2025
    IEEE CASE 2025
    Addresses arm coupling in restless multi-armed bandits by finetuning the Whittle index.

Languages

Interests

  • Research
    Reinforcement learning, Preference-based RL, Skill discovery, Multi-objective RL, Model-based RL, Multi-agent RL, Large language models for decision-making