Preference-based Multi-Objective Reinforcement Learning

Published in IEEE Transactions on Automation Science and Engineering (T-ASE), 2025

Reward-function design is notoriously hard in multi-objective reinforcement learning. We propose a multi-objective preference learning paradigm with theoretical guarantees: optimizing a multi-objective preference reward model is equivalent to solving the convex Pareto front in policy space.

Links: arXiv

Paper