Publications
* Equal contribution.
Journal Articles
- IEEE Transactions on Automation Science and Engineering (T-ASE), 2025
Overview
A theoretically grounded multi-objective preference learning paradigm: optimizing the multi-objective preference reward model is equivalent to solving the convex Pareto front in policy space.
- IEEE Robotics and Automation Letters (RA-L), 2025
Overview
A transformer-based thermal surrogate model that enables efficient cooling control in data centers.
Conference Papers
- IEEE CASE 2026, 2026
Overview
Learns interpretable index policies by driving zero-order optimization with a large language model. Under review (IEEE CASE 2026).
- ICML 2026, 2026
Overview
Unsupervised skill discovery often yields task-irrelevant or hazardous behaviors. COLLIE leverages dense unsupervised data to construct a semantically coherent skill latent space, enabling reliable, training-free guidance from sparse online human feedback.
- ICLR 2026, 2026
Overview
A state diffusion process that handles partial observability in multi-agent systems.
- AAAI 2026, 2026
Overview
Model-based RL world models generalize poorly across scenarios. MrCoM extracts scenario information via meta-state regularization and aligns model optimization with multi-scenario policy learning via meta-value regularization.
- IFAC 2026 (23rd IFAC World Congress), 2026
Overview
Covers the Pareto frontier by coordinating an interpretable policy library with a large language model.
- CAC 2025 (2025 China Automation Congress), 2025
Overview
Leverages agent-specific preferences for multi-agent reinforcement learning.
- [7] LLM-Empowered Knowledge Graph Construction for Optimization Formalization in Smart ManufacturingCAC 2025 (2025 China Automation Congress), 2025
Overview
Uses a large language model to construct knowledge graphs that formalize optimization problems in smart manufacturing.
- NeurIPS 2025, 2025
Overview
Moves beyond task ambiguity in language-conditioned reinforcement learning.
- NeurIPS 2025, 2025
Overview
In multi-stage tasks, preference-based RL suffers from stage misalignment: comparing behaviors from different stages is hard for humans and uninformative for policy optimization. STAIR minimizes the temporal distance between segment pairs so comparisons stay within a stage.
- IEEE CASE 2025, 2025
Overview
Addresses arm coupling in restless multi-armed bandits by finetuning the Whittle index.
- [11] Safety-Guaranteed Policy Composition Via Generalized Policy Improvement for Autonomous VehiclesIEEE CASE 2025 (Outstanding WiRA Student Paper Award), 2025
Overview
A Generalized-Policy-Improvement-based policy composition with safety and performance guarantees that resolves the performance–safety–sample-efficiency "impossible triangle" for autonomous vehicles. Outstanding WiRA Student Paper Award.
- ICML 2025, 2025
Overview
In preference-based RL, humans struggle to compare behaviorally similar trajectory segments, which hurts label efficiency. CLARIFY learns a contrastive objective in the trajectory embedding space to select segment pairs that humans can actually compare.
- IJCAI 2025, 2025
Overview
To address the indistinguishability of behaviorally similar segments in preference-based RL, S-EPOA uses unsupervised skill discovery to pick segments that differ widely in skill, improving human label efficiency.
- CAC 2024 (2024 China Automation Congress), 2024
Overview
Extends preference-based multi-objective reinforcement learning with explicit reward modeling.
- IEEE CASE 2024, 2024
Overview
Applies sample-efficient reinforcement learning to large-scale data-center cooling control.
- CAC 2023 (2023 China Automation Congress), 2023
Overview
Combines mechanism models with data via a multi-fidelity model for reinforcement-learning-based data-center cooling control.
- NeurIPS 2022, 2022
Overview
Improves sample efficiency in multi-agent reinforcement learning by guiding learning with parallel programs.
Preprints
- Preprint, 2025
Overview
A benchmark for evaluating LLMs long-horizon decision making in StarCraft II, together with a hierarchical agent baseline that self-corrects via expert knowledge and self-improves via supervised fine-tuning.
