STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
Published in NeurIPS 2025, 2025
In multi-stage tasks, preference-based RL suffers from a stage-misalignment problem: comparing behaviors drawn from different stages is hard for humans to judge accurately and provides little useful signal for policy optimization. Motivated by a theoretical analysis, STAIR minimizes the temporal distance between segment pairs to ensure preference comparisons happen within the same stage, improving learning efficiency.
