Procedural Fairness Failures in RLHF from Preference Averaging
M P V S Gopinadh, Karthik Kamuju, Kummari Avinash, Muppana John Joshua, Srinivasa Raju Rudraraju
Problem
Standard RLHF trains a single reward model on aggregated pairwise preferences. That design assumes one shared preference distribution, even though raters routinely disagree about what a good response looks like. Some want concise answers, some want detail, and some want a technical register, and those preferences can conflict on the same prompt. When those groups are averaged into one reward, influence scales with how common a preference is in the data. Majority preferences can dominate the learned objective while minority preferences are systematically under-represented. The problem is procedural: this failure can come from the structure of reward learning rather than from a later policy-optimization bug. The paper tests whether separating preference modes at the reward-learning stage reduces it.
Method
We evaluate this in a controlled setting built to isolate preference aggregation from annotation noise and demographic confounds. The dataset contains 971 pairwise comparisons from 60 simulated raters across a pool of 20 prompts. Raters are programmatically assigned to three groups of 20, each locked to a distinct preference profile: concise responses of 20 to 35 words at 85% within-group consistency, detailed responses of 60 to 90 words at 85% consistency, and technical or formal responses of 40 to 60 words at 80% consistency. Each rater evaluates 15 to 18 randomly sampled prompts rather than the full pool. Each feedback instance is a prompt, two candidate responses, and a binary preference label. Data are split by rater with a 75/25 train-test partition, and ground-truth group assignments are used only at evaluation time. Preference-Aware RLHF (PA-RLHF) does not train on those ground-truth labels. Prompt-response pairs are embedded with frozen all-MiniLM-L6-v2. Clustering is run on user feature vectors, the mean preferred-response embedding concatenated with a response-length scalar, into k = 3 inferred modes by K-Means (silhouette 0.199, ARI 0.443 on held-out data). A separate logistic reward model is then trained inside each cluster on the prompt embedding concatenated with the difference of the two response embeddings, approximating Bradley-Terry pairwise preference. The base language model is held fixed, and there is no full policy optimization, so any change in alignment can be attributed to how preferences are aggregated rather than to later RL updates or extra model capacity.
Results
Alignment accuracy is agreement with group-specific preference judgments. Under standard preference averaging, overall accuracy is 46.9%, the majority mode reaches 56.2%, and the two minority modes reach only 41.2% and 40.3%, for a fairness gap of 15.9 percentage points between the best-aligned and worst-aligned groups. Under PA-RLHF, overall accuracy rises to 67.9%. The minority modes move to 68.8% and 73.1% (gains of 27.6 and 32.8 percentage points). The majority mode also rises, to 63.5%, but only by 7.3 percentage points, so the gain is not a transfer of accuracy from the majority to the minorities. The fairness gap falls to 9.6 percentage points, a 40% reduction. The residual gap shows that clustering-based separation does not remove all group-level misalignment, but it is enough to show that the original imbalance came from averaging, not from the minorities being inherently harder to fit.