deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Tri-PvP benchmark shows visual bias in omni-modal models starts in early layers

Tri-PvP, a new benchmark for omni-modal models, finds that images dominate answer selection, that the bias is detectable in early transformer layers, and that contrastive decoding only partially suppresses it.

Tri-PvP benchmark shows visual bias in omni-modal models starts in early layers

A benchmark built on perceptual–propositional conflicts

A write-up published on dev.to on October 10, 2026 describes Tri-PvP, a new benchmark designed to expose modality bias in omni-modal large language models — systems that take text, images and audio together. Its headline findings: visual evidence dominates how these models pick answers, the preference takes shape in the earliest layers of the network, and a leading inference-time fix only partially suppresses it.

The benchmark's name refers to its method: perceptual–propositional evidence conflicts. According to the dev.to post, earlier evaluations mixed perceptual cues such as raw images or audio with propositional statements like "this is a dog," so a wrong answer could never be traced with confidence to the visual stream. By setting perception and proposition against each other, Tri-PvP can attribute an error to a specific modality.

The scale of visual dominance

The write-up reports that visual bias now outweighs every other modality signal in the majority of omni-modal models. Citing the underlying paper's Figure 3, the author notes that across 20 model and evidence-type comparisons, the BIAS_IMAGE label dominates in 18, and its magnitude frequently passes 60% of the total bias signal. When these models answer incorrectly, the image input is usually the reason.

Bias shows up before the decision head

The most consequential result concerns where the bias lives. The write-up reports that modality-bias information becomes readable to a simple linear probe from the first few transformer layers, not only at the output stage. That matters because previous mitigation strategies — post-hoc output filtering and fine-tuning on balanced data — assumed that correcting behaviour at or near the output would be enough. If the preference is encoded this early, those approaches treat the symptom rather than the cause.

A mitigation that softens but does not erase

The write-up also describes a remedy: a low-disturbance contrastive decoding adjustment applied at inference time, with no parameter updates. It reduces bias while leaving task performance essentially intact — OmniBench accuracy slips a single point, from 38.4% to 37.4%, according to the reported figures.

The fix is not complete, though. The same linear probes that detect the original bias can still detect residual visual preference after the adjustment. The intervention weakens the bias rather than removing it, which, as the write-up concludes, points future work toward combining representation-level regularization with architecture-aware training objectives if full neutrality is the goal.

Why it matters

The result reframes multimodal fairness as a representational problem rather than an output problem. If bias is recoverable from early layers, output-level checks will systematically miss it, and models certified as neutral on end-to-end benchmarks may still carry a strong internal visual preference that surfaces in edge cases.

For practitioners, the dev.to post makes two concrete suggestions: run Tri-PvP as a standard pre-deployment check for omni-modal systems, and fold early-layer probing plus contrastive decoding into pipelines, with the aim of keeping visual bias below the roughly 60% dominance level while downstream scores remain essentially unchanged. More broadly, the benchmark gives the field a measurable definition of modality neutrality — one that today's models and mitigations do not yet fully meet.

  • #multimodal
  • #llms
  • #ai-bias
  • #benchmarks
  • #interpretability

Related posts