S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning
Published in IJCAI 2025, 2025
In preference-based RL, humans struggle to compare behaviorally similar trajectory segments. S-EPOA designs an unsupervised skill-discovery method that selects segments differing widely in โskillโ for comparison, improving human label efficiency.
