SIGIR 2026 · Full Paper

When More
Reformulations
Hurt

Avoiding Drift using
Ranker Feedback

Venktesh V  ·  Mandeep Rathee  ·  Avishek Anand
Stockholm University  |  L3S Research Center  |  TU Delft
0.60 0.57 0.54 0.51 0.48 0.45 3 5 10 15 25 # of reformulations Recall@100 ReformIR GenQR BM25 Recall@100 vs. # Reformulations (TREC DL19)

Abstract

The core challenge is,
adaptive selection of reformulations and navigating the combinatorial reformulation-document search space.

Modern retrieval pipelines increasingly rely on query reformulation and neural reranking to improve effectiveness, but this introduces a fundamental tradeoff between recall and query drift. Generating many reformulated queries can substantially increase recall, yet naively merging or exhaustively reranking their results is prohibitively expensive.

We propose ReformIR, a budget-aware retrieval framework that treats query reformulations as first-class features and performs online relevance estimation using a strong reranker as a teacher. Under a fixed reranking budget, a lightweight surrogate model adaptively prioritizes both reformulations and documents, suppressing drift through online feature selection.

🎯

Drift-Suppression via Feedback

A bandit-style loop uses a teacher reranker anchored to the original query, actively downweighting reformulations that drift from intent.

⚡

3.3–4.5× More Efficient

Far cheaper than LLM-based reranking. ReformIR uses LLMs where they're cheapest — rephrasing queries, not scoring documents.

🔌

Training-free Adapter

Drop ReformIR on top of any existing reformulation method (GenQR, QA-Expand, HyDE…) with no retraining required.

🔍

Interpretable Weights

Learns explicit reformulation weights, revealing which query variants drive relevance — a window into retrieval behavior.


The ReformIR Pipeline

An animated walk-through of the algorithm. Each step corresponds directly to Algorithm 1 from the paper. Click a step pill to jump to it.

REFORMIR PIPELINE ORIGINAL QUERY "sacraments of service…" LLM REFORMULATOR GenQR / QA-Expand / HyDE Q₁: Two sacraments of… w₁ = 0.48 ↑ Q₂: Purpose of sacraments… w₂ = 0.41 ↑ Q₃: Difference from other… w₃ = 0.11 ↓ drift! CANDIDATE POOL d₄₃ · d₁₂ · d₃ d₆₇ · d₉₀ · d₅₆ d₂₂ · d₃₃ · d₁₁ d₇₈ · d₅₅ · d₈₈ d₁₄ · d₉₅ · d₀₁ d₂₉ · d₄₄ · d₆₂ BM25 retrieval (each Qᵢ) SURROGATE MODEL s(d;w) = wᵀxd reformulations as features online feature selection linUCB / bandit-style sampling TEACHER RERANKER ψ(Q, d) — budget τ calls anchored to original query ↺ update w FINAL RANKING d₄₃ Q₁Q₂ ████ 0.91 d₁₂ Q₁Q₂ ███ 0.87 d₉₀ Q₂ ██ 0.74 d₃ Q₃ ▌ 0.21 d₅₆ Q₃ ▏ 0.14 Q₃ downweighted (drift) Q₁,Q₂ upweighted ✓ ACTIVE STEP Hover or click a step to explore Step 0 Step 1 Step 2 Steps 3–4 Output
0 · Original Query
1 · Generate Reformulations
2 · BM25 Candidate Pool
3 · Surrogate + Bandit
4 · Teacher Reranker Feedback
5 · Final Ranking

Algorithm

Budget-aware Online Optimization

1Input: Query Q, reformulations 𝒬={Q₁,…,Qₘ}, budget τ, surrogate w
2Input: teacher reranker ψ(Q,·), retrieval function ret(·)
3 
4// Step 1 — Candidate pool construction
5𝒟 ← ⋃i ret(Qᵢ)  // BM25 for each reformulation
6𝐱d ← features(d, 𝒬)  // reformulation retrieval signals per doc
7 
8// Steps 2–4 — Online bandit loop (budget τ)
9for t = 1 to τ do
10   d* ← argmaxd∈𝒟 s(d; w) + β·ucb(d)  // UCB selection
11   y* ← ψ(Q, d*)  // teacher reranker, anchored to Q
12   w ← update(w, 𝐱d*, y*)  // online feature selection
13   𝒟 ← 𝒟 \ {d*}  // remove scored doc
14end for
15 
16// Output — final ranking by surrogate score
17return rank(𝒟 ∪ scored, s(·; w))
Key insight: The teacher reranker on line 11 always evaluates relevance with respect to the original query Q — not the reformulations. This anchoring is the mechanism by which drift is suppressed: reformulations that retrieve documents irrelevant to Q receive low scores, causing their weights w to decay through the online update.

Results

Experiments on TREC DL19–DL22

Evaluated on MSMARCO passage corpora and TREC Deep Learning benchmarks. ReformIR serves as a training-free adapter applied on top of existing reformulation baselines.

NDCG@10 — TREC DL19

ReformIR+GenQR
0.742
GenQR (baseline)
0.556
ReformIR+QA-Exp
0.701
BM25
0.506
+33.5% gain over GenQR baseline

Recall@100 — Scaling Reformulations

ReformIR (stable ↑) GenQR (drift ↓) 3 5 10 15 25
ReformIR remains stable; GenQR degrades with more reformulations

Efficiency vs LLM Reranking

ReformIR latency
1×
LLM-reranker latency
3.3–4.5×
Same or better effectiveness at a fraction of LLM inference cost

Generator-agnostic gains

+ GenQR
✓
+ QA-Expand
✓
+ HyDE
✓
+ Query2Doc
✓
Consistent improvements across all reformulation generators

Citation

BibTeX

@inproceedings{venktesh2026reformir,
  title     = {When More Reformulations Hurt:
               Avoiding Drift using Ranker Feedback},
  author    = {Venktesh, V and Rathee, Mandeep
               and Anand, Avishek},
  booktitle = {Proceedings of the 49th International
               ACM SIGIR Conference on Research and
               Development in Information Retrieval},
  year      = {2026},
  doi       = {10.48550/arXiv.2605.00560}
}