Papers
arxiv:2607.20092

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

Published on Jul 22
· Submitted by
Karan Goyal
on Jul 23
Authors:
,
,

Abstract

Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.

Community

Paper author Paper submitter

entrap_overview

Paper author Paper submitter

Contextual entrainment — a model's tendency to be pulled by auxiliary context regardless of whether that context is relevant, true, or meaningful — has been studied in unimodal language models but remains largely unexamined in vision-language models. We argue this multimodal setting is substantive rather than incremental: entrainment becomes a dual phenomenon drivable by both textual and visual context, and it opens a scene-relative veracity distinction with no counterpart in text-only work. We introduce ENTRAP-VL, a manually curated dataset of 1,500 items across eight categories, organized on two axes (association and veracity) and split into a textual-entrainment stream (eight conditions) and a visual-entrainment stream (three). We release the instrument, not measurements, so the community can investigate the phenomenon rigorously.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.20092
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.20092 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.20092 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.