FALSE DISCOVERY CONTROLLED HALLUCINATION MITIGATION

CORAL

Mitigating Object Hallucination in Large Vision-Language Models via False Discovery Controlled Visual Data Splitting

A principled, training-free framework for image-level hallucination control using uncertainty-aware visual data splitting, mirror statistics, and false discovery rate control.

✦ Accepted to NeurIPS 2026 ✦
Chang Liu Yu Tian Rui Xie*

University of Central Florida   ·   *Corresponding author

OVERVIEW

Controlling hallucinations at the image level

CORAL treats object hallucination as a false discovery control problem rather than evaluating each object query independently.

Overview of the CORAL framework
CORAL builds paired mirror views from the same visual input, computes mirror statistics from model responses, and applies a data-driven threshold to control false discoveries.

METHOD

Three components, one controlled decision rule

The framework combines uncertainty-aware visual data splitting, mirror statistics, and false discovery rate control.

01

Uncertainty-Aware Visual Data Splitting

Two mirror views are generated from a shared Gaussian noise source using opposite perturbation signs.

f+v = v + τvZv
f−v = v − τvZv
02

Mirror Statistic

The paired logit contrasts capture whether a generated token responds consistently to the two symmetric perturbations.

Δt = |Δ+t + Δ−t| − |Δ+t − Δ−t|
03

False Discovery Rate Control

The negative tail of the mirror statistic is used to estimate spurious visual discoveries and select a data-driven threshold.

Target FDR level in experiments: q = 0.10
Training-Free No retraining or additional supervision.
Image-Level Control Controls the fraction of false discoveries within an image.
High Power Preserves visually grounded objects while suppressing hallucinations.
Model-Flexible Evaluated across multiple large vision-language models.

RESULTS

Controlled FDR with strong power

Main MSCOCO results for LLaVA-OneVision-7B, Qwen2.5-VL-7B, and InternVL3-8B.

Random · Avg. FDR 0.0736

Below the target FDR level of 0.10.

Random · Avg. Power 90.97%

Retains most visually grounded objects.

POPE · Random Avg. F1 92.96

Average across the three main LVLMs.

MMBench · InternVL3-8B 94.50

Compared with 83.40 under regular decoding.

CORAL false discovery rate results
Overall false discovery rate across sampling settings.
CORAL power results
Overall power across sampling settings.
Evaluation scope. The paper evaluates CORAL with FDR and power together with POPE, MME, CHAIR, and MMBench. Experiments cover MSCOCO, A-OKVQA, and GQA, with aggregate evaluations reported over 3000 randomized trials.
CORAL POPE evaluation
POPE evaluation on MSCOCO.
CORAL FDR target ablation
Effect of the target FDR level q on POPE performance.

QUALITATIVE EXAMPLES

Suppressing unsupported visual content

Examples from the paper show how CORAL rejects hallucinated predictions while preserving visually grounded information.

Qualitative examples of CORAL hallucination mitigation
Hallucination mitigation examples from CORAL.

SCOPE

What CORAL controls — and where it is limited

Designed for grounded prediction control

CORAL operates on token-level model logits and uses mirror statistics to distinguish stable visual evidence from perturbation-driven noise before FDR-controlled selection.

Current limitation

CORAL cannot be directly applied to black-box commercial APIs that do not expose token-level logits. The paper also identifies relational, attribute, and open-ended generation settings as directions for future work.

CITATION

Cite CORAL

@inproceedings{liu2026coral,
  title     = {Mitigating Object Hallucination in Large Vision-Language Models via False Discovery Controlled Visual Data Splitting},
  author    = {Liu, Chang and Tian, Yu and Xie, Rui},
  booktitle = {40th Conference on Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}