S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based RL
Date:
In preference-based RL, humans struggle to compare behaviorally similar trajectory segments, which hurts label efficiency. S-EPOA introduces an unsupervised skill-discovery method that selects segments differing widely along skill dimensions, making each preference query easy for humans to judge and highly informative.
Related paper: S-EPOA (arXiv 2408.12130) ยท ๐ Slides (PDF)
