ReFeri: Training-free LLM Verification via Recycling Few-shot Examples

Dongseok Lee1, Jimyung Hong1, Dongyoung Kim2, Jaehyung Kim1
1Yonsei University     2KAIST
Yonsei University Machine and Language Learning Lab (ML3)
EMNLP 2026 (Main)
ICML 2025 Workshop ES-FoMo-III · Spotlight Presentation (14/146)
Paper Code Citation
Overview of ReFeri: candidates are scored with a forward confidence score minus a backward reconstruction penalty

News


[2026] Accepted to EMNLP 2026 Main.
[2025] Spotlight at ICML 2025 Workshop ES-FoMo-III.

Overview


Response selection is a central bottleneck in test-time scaling. LLMs can generate multiple plausible reasoning paths for the same query, but majority voting is limited to outputs that can be normalized and aggregated, while learned verifiers require additional supervision and training.

ReFeri recycles the few-shot examples already available in the prompt as a training-free verifier. The same demonstrations used to guide generation are reused to evaluate which candidate follows the intended task while avoiding demonstration-specific shortcuts.

ReFeri combines two complementary likelihood signals. The forward confidence score measures how well a candidate follows the few-shot guidance, while the backward reconstruction penalty identifies candidates that are overly specific to the demonstrations. Their difference provides a lightweight verification signal without task-specific verifier training.

Results


ReFeri consistently improves training-free response selection. Across three generation LLMs and seven reasoning benchmarks, ReFeri achieves an 8.2% average relative gain over random selection and a 2.6% average relative improvement over the strongest prior training-free baseline.

Methods MuSR-taAcc. MuSR-opAcc. GPQAAcc. MATH500Acc. DROPEM / F1 HotpotQAEM / F1 MMLU-PROAcc. Avg.
GPT-4o
Zero-shot CoT67.5 ±0.862.1 ±0.449.5 ±0.877.1 ±0.674.2 ±0.8 / 84.9 ±0.437.8 ±0.3 / 50.3 ±0.474.0 ±0.263.2 ±0.2
Few-shot CoT87.2 ±0.669.6 ±0.647.3 ±1.775.5 ±0.180.4 ±0.2 / 89.0 ±0.244.9 ±0.4 / 58.6 ±0.373.7 ±0.268.4 ±0.3
LEAP88.1 ±1.668.0 ±1.247.8 ±2.875.2 ±0.481.0 ±0.5 / 89.4 ±0.444.5 ±0.6 / 57.8 ±0.573.9 ±0.268.4 ±0.4
USC85.9 ±1.971.5 ±0.748.2 ±1.677.3 ±0.381.8 ±0.3 / 90.2 ±0.145.9 ±0.3 / 60.2 ±0.574.9 ±0.569.4 ±0.5
CoT-WP88.3 ±0.267.5 ±1.449.5 ±1.878.1 ±0.683.3 ±0.1 / 90.5 ±0.846.3 ±0.9 / 59.6 ±0.974.4 ±0.369.6 ±0.0
Self-Certainty88.9 ±0.571.1 ±2.150.2 ±0.877.7 ±0.481.1 ±0.8 / 89.4 ±0.343.9 ±0.4 / 57.7 ±0.174.6 ±0.469.6 ±0.4
ReFeri90.8 ±0.473.7 ±2.251.8 ±0.378.5 ±0.683.6 ±0.2 / 90.9 ±0.646.9 ±0.4 / 61.0 ±0.775.4 ±0.371.5 ±0.4
ReFeri1B90.7 ±0.274.1 ±3.851.0 ±1.378.3 ±0.683.3 ±0.5 / 90.7 ±0.146.5 ±0.8 / 60.5 ±0.275.1 ±0.171.3 ±0.4
GPT-4o-mini
Zero-shot CoT57.6 ±1.358.9 ±2.741.8 ±1.075.8 ±0.677.2 ±0.7 / 85.1 ±0.831.5 ±0.4 / 41.6 ±0.463.0 ±0.158.0 ±0.1
Few-shot CoT77.5 ±0.560.3 ±0.842.4 ±1.174.7 ±0.776.5 ±0.2 / 82.9 ±0.233.6 ±0.3 / 44.9 ±0.263.0 ±0.261.1 ±0.2
LEAP74.9 ±3.260.3 ±3.243.6 ±0.374.5 ±0.175.4 ±0.4 / 82.5 ±0.533.3 ±0.6 / 44.5 ±0.563.1 ±0.260.7 ±0.1
USC76.5 ±2.560.5 ±1.044.6 ±1.276.4 ±1.278.8 ±1.7 / 85.0 ±1.035.1 ±0.2 / 47.0 ±0.364.2 ±0.562.3 ±0.5
CoT-WP79.3 ±1.358.5 ±2.641.9 ±0.977.0 ±0.776.9 ±0.6 / 82.7 ±0.634.7 ±0.9 / 45.9 ±0.964.6 ±0.461.8 ±0.5
Self-Certainty81.7 ±1.660.1 ±0.641.6 ±2.876.8 ±1.076.7 ±0.4 / 83.1 ±0.634.7 ±0.1 / 45.9 ±0.463.9 ±0.862.2 ±0.1
ReFeri83.1 ±0.262.0 ±0.644.6 ±2.678.2 ±0.479.1 ±0.5 / 84.7 ±0.835.7 ±0.4 / 47.4 ±0.664.9 ±0.463.9 ±0.5
ReFeri1B82.8 ±0.462.9 ±1.445.3 ±1.277.7 ±0.278.3 ±0.8 / 83.9 ±1.035.1 ±0.5 / 46.6 ±0.564.6 ±0.563.8 ±0.2
LLaMA-3.1-8B
Zero-shot CoT43.3 ±1.551.9 ±1.220.9 ±0.742.9 ±1.359.7 ±0.7 / 65.9 ±0.515.7 ±0.6 / 21.7 ±0.540.1 ±0.439.2 ±0.2
Few-shot CoT64.5 ±2.253.4 ±0.823.7 ±0.342.4 ±0.561.1 ±0.7 / 66.5 ±0.919.3 ±0.3 / 25.4 ±0.439.2 ±0.443.4 ±0.5
LEAP64.7 ±3.953.3 ±5.126.8 ±1.841.9 ±0.557.0 ±1.1 / 62.6 ±1.319.9 ±0.2 / 26.7 ±0.136.5 ±0.742.9 ±1.2
USC67.7 ±4.554.8 ±2.228.1 ±3.148.5 ±1.268.6 ±0.9 / 74.3 ±1.325.5 ±1.1 / 33.4 ±1.045.8 ±0.448.4 ±1.1
CoT-WP71.9 ±0.653.1 ±1.430.5 ±2.548.1 ±1.771.0 ±0.5 / 75.1 ±0.625.7 ±0.1 / 33.2 ±0.445.4 ±0.749.4 ±0.4
Self-Certainty72.4 ±3.655.9 ±0.630.8 ±2.650.7 ±2.270.8 ±1.1 / 76.2 ±1.025.2 ±0.5 / 32.9 ±1.044.5 ±0.850.0 ±1.1
ReFeri76.5 ±3.457.5 ±0.533.5 ±1.750.2 ±1.670.9 ±1.6 / 76.3 ±0.525.7 ±0.7 / 33.6 ±0.644.7 ±0.551.3 ±0.6
ReFeri1B76.9 ±3.255.5 ±0.831.8 ±2.050.1 ±1.770.9 ±1.4 / 75.8 ±1.025.3 ±0.5 / 32.9 ±0.644.2 ±0.550.7 ±0.5

ReFeri uses a LLaMA-3.1-8B estimator; ReFeri1B uses LLaMA-3.2-1B. Bold: best, underlined: second-best within each generation model.

ReFeri scales with the candidate pool. As the number of responses increases from 1 to 20, accuracy improves from 75.8 to 79.4 on MATH500, 41.4 to 45.5 on GPQA, and 75.6 to 86.0 on MuSR-ta, while several competing training-free selectors saturate or degrade.

Test-time scaling on MATH500
(a) MATH500
Test-time scaling on GPQA
(b) GPQA
Test-time scaling on MuSR-ta
(c) MuSR-ta

GPT-4o-mini generates N = 1–20 candidates with Few-shot CoT; each method selects one.

Analysis


The backward term captures demonstration mimicry. Ranking candidates by the raw backward term and reusing each as a one-shot demonstration, accuracy decreases with rank: candidates with larger backward terms better reconstruct the original few-shot answers, which is exactly the behavior ReFeri penalizes.

Accuracy by backward-term rank on MATH500
(a) MATH500
Accuracy by backward-term rank on MuSR-ta
(b) MuSR-ta

Demonstration fragments inflate the backward term. Prepending the first ρ% of the few-shot examples to each candidate raises the raw backward term (solid) much more sharply than the forward score (dashed).

Raw forward and backward scores under perturbation on MATH500
(a) MATH500
Raw forward and backward scores under perturbation on MuSR-ta
(b) MuSR-ta

This penalty keeps selection robust to injected artifacts. When demonstration fragments are prepended to a low-forward incorrect candidate, forward-only selection drops from 45.3% to 33.3%, while ReFeri reaches 47.5%.

Method ρ = 0 ρ = 50 ρ = 100
LLaMA-3.1-8B estimator
Forward45.2942.0333.31
ReFeri45.2947.4847.48
LLaMA-3.2-1B estimator
Forward45.4141.3830.71
ReFeri45.4747.6047.66

Selection accuracy (%) when the first ρ% of the retrieved demonstration answer is prepended to a low-forward incorrect candidate (GPT-4o, GPT-4o-mini, and LLaMA-3.1-8B generators; MATH500, GPQA, MuSR-ta).

Robustness and Efficiency


ReFeri remains effective across estimation models. Holding candidate pools fixed, ReFeri improves over Few-shot CoT with every tested estimator across model families and scales: LLaMA-3.2-1B, Qwen2.5-7B, LLaMA-3.1-8B, and LLaMA-3.1-70B.

Average accuracy of ReFeri with different estimation models versus Few-shot CoT

A lightweight estimator provides a favorable cost-accuracy trade-off. ReFeri with a 1B estimator remains on a strong trade-off frontier and runs approximately 2.6× faster than CoT-WP with an 8B estimator.

Cost-accuracy trade-off of response selection methods

Citation


@article{lee2025training,
  title={Training-free LLM Verification via Recycling Few-shot Examples},
  author={Lee, Dongseok and Hong, Jimyung and Kim, Dongyoung and Kim, Jaehyung},
  journal={arXiv preprint arXiv:2506.17251},
  year={2025}
}