S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

Published in IJCAI 2025, 2025

In preference-based RL, humans struggle to compare behaviorally similar trajectory segments. S-EPOA designs an unsupervised skill-discovery method that selects segments differing widely in โ€œskillโ€ for comparison, improving human label efficiency.

Links: arXiv ยท Code

Paper | Code