S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based RL

Date:

In preference-based RL, humans struggle to compare behaviorally similar trajectory segments, which hurts label efficiency. S-EPOA introduces an unsupervised skill-discovery method that selects segments differing widely along skill dimensions, making each preference query easy for humans to judge and highly informative.

Related paper: S-EPOA (arXiv 2408.12130) ยท ๐Ÿ“„ Slides (PDF)

Direct Link