CV
Ni Mu (็ๅช)
Summary
Ph.D. student in the Department of Automation at Tsinghua University, working on reinforcement learning and large language models.
Education
- Control Science and Engineering (Ph.D.)presentTsinghua University, Department of AutomationCourses: Reinforcement Learning, Large Language Models, Preference-based RL, Skill Discovery, Multi-objective RL
- Computer Science (B.Eng.)2023-06Southeast University, Youth Class (ๅฐๅนด็ญ)GPA: 3 / 275 (top 1.5%)
Publications
- CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries2025ICML 2025Selects human-comparable segment pairs via contrastive learning in the trajectory embedding space to improve label efficiency in preference-based RL.
- STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning2025NeurIPS 2025Minimizes temporal distance between segment pairs so that preference comparisons stay within the same stage, improving learning efficiency in multi-stage tasks.
- COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space2026ICML 2026Builds a semantically coherent latent space and constructs training-free guidance signals so that safe, diverse skills can be learned from very few human labels.
- S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning2025IJCAI 2025Uses unsupervised skill discovery to select segments that differ widely in skill, improving human label efficiency in preference-based RL.
- Preference-based Multi-Objective Reinforcement Learning2025IEEE Transactions on Automation Science and Engineering (T-ASE)A theoretically grounded multi-objective preference learning paradigm where optimizing the multi-objective preference reward model is equivalent to solving the convex Pareto front in policy space.
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios2026AAAI 2026Extracts scenario information via meta-state regularization and aligns model optimization with multi-scenario policy learning via meta-value regularization to build a cross-scenario world model.
- Safety-Guaranteed Policy Composition Via Generalized Policy Improvement for Autonomous Vehicles2025IEEE CASE 2025 (Outstanding WiRA Student Paper Award)A Generalized-Policy-Improvement-based policy composition with safety and performance guarantees that resolves the performance-safety-sample-efficiency tradeoff for autonomous driving.
- SunTzu: A Long-Horizon Decision Making Benchmark for Large Language Models2025Preprint (submitted to NeurIPS 2026)A benchmark for evaluating LLMs on long-horizon decision making in StarCraft II, plus a hierarchical agent baseline that self-corrects with expert knowledge and self-improves via SFT.
- Covering the Pareto Frontier with LLM-Coordinated Interpretable Policy Library2026IFAC 2026 (23rd IFAC World Congress)Covers the Pareto frontier by coordinating an interpretable policy library with an LLM.
- Learning Interpretable Index Policies via LLM-Driven Zero-Order Optimization2026IEEE CASE 2026Learns interpretable index policies by driving zero-order optimization with an LLM.
- LLM-Empowered Knowledge Graph Construction for Optimization Formalization in Smart Manufacturing2025CAC 2025 (China Automation Congress)Uses an LLM to build knowledge graphs that formalize optimization problems in smart manufacturing.
- GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent System2026ICLR 2026A state diffusion process that handles partial observability in multi-agent systems.
- DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning2025NeurIPS 2025Moves beyond task ambiguity in language-conditioned reinforcement learning.
- A Transformer-based Thermal Surrogate Model for Cooling Control in Data Centers2025IEEE Robotics and Automation Letters (RA-L)A transformer-based thermal surrogate model enabling efficient cooling control in data centers.
- E-MAPP: Efficient Multi-Agent Reinforcement Learning with Parallel Program Guidance2022NeurIPS 2022Improves sample efficiency in multi-agent RL by guiding learning with parallel programs.
- MAGPIE: Utilizing Agent-Specific Preferences for Multi-Agent Reinforcement Learning2025CAC 2025 (China Automation Congress)Leverages agent-specific preferences for multi-agent reinforcement learning.
- Preference-Based Multi-Objective Reinforcement Learning with Explicit Reward Modeling2024CAC 2024 (China Automation Congress)Extends preference-based multi-objective RL with explicit reward modeling.
- Large-scale data center cooling control via sample-efficient reinforcement learning2024IEEE CASE 2024Applies sample-efficient reinforcement learning to large-scale data-center cooling control.
- Integrating mechanism and data: Reinforcement learning based on multi-fidelity model for data center cooling control2023CAC 2023 (China Automation Congress)Combines mechanism models with data via a multi-fidelity model for RL-based data-center cooling control.
- Addressing Coupling in Restless Multi-Armed Bandits by Finetuning Whittle Index2025IEEE CASE 2025Addresses arm coupling in restless multi-armed bandits by finetuning the Whittle index.
Languages
Interests
- ResearchReinforcement learning, Preference-based RL, Skill discovery, Multi-objective RL, Model-based RL, Multi-agent RL, Large language models for decision-making