A Contextual Quality Reward Model For Reliable and Efficient Best of N Sampling

Hyung Gyu Rho

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Modern preference alignment techniques, such as Best-of-N (BoN) sampling, rely on reward models trained with pairwise comparison data. While effective at learning relative preferences, this paradigm fails to capture a signal of response acceptability, leaving systems vulnerable to selecting the least bad of many unacceptable options. This is particularly problematic for hard prompts, where the risk of such false acceptances increases with the number of samples. In this paper, we address this critical reliability gap by introducing a new data collection and modeling framework. By augmenting preference data with an outside option, inspired by discrete choice models, we train a reward model that can distinguish not just what is better, but what is good enough. We leverage this capability to create an adaptive inference strategy, best of mini-N in-loop, which partitions the generation budget into sequential loops with a calibrated, early-exit condition. Our experiments show that when tuned as an alignment guardrail, it reduces reliability failures by 70%, and when tuned as an inference accelerator, it improves average inference speed by over 22% in IMBD-sentiment setting. We thus provide a principled and flexible framework for practitioners to explicitly manage the trade-off between reliability and computational efficiency.

Version published to 10.21203/rs.3.rs-7594024/v1 on Research Square
Sep 12, 2025

EL-MIATTs: Evaluation and Learning with Multiple Inaccurate True Targets

This article has 1 author:
1. Yongquan Yang
This article has no evaluationsLatest version Jan 12, 2026
A Survey on Robust Sequential Recommendation: Fundamentals, Challenges, Taxonomy, and Future Directions

This article has 5 authors:
1. Yatong Sun
2. Xiaochun Yang
3. Bin Wang
4. Yan Wang
5. Zhu Sun
This article has no evaluationsLatest version Jan 28, 2026
DDPO: Diversity-Driven Preference Optimization for Machine Translation Enhancing Robustness and Generalization

This article has 2 authors:
1. Donald Martin
2. Blake Bowman
This article has no evaluationsLatest version Dec 30, 2025

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

EL-MIATTs: Evaluation and Learning with Multiple Inaccurate True Targets

A Survey on Robust Sequential Recommendation: Fundamentals, Challenges, Taxonomy, and Future Directions

DDPO: Diversity-Driven Preference Optimization for Machine Translation Enhancing Robustness and Generalization