Skip to main navigation Skip to search Skip to main content

Measuring the effects of visual salience in human and AI descriptions with image editing

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

How does our perception of the world influence the way we talk about it? Psycholinguistic studies have investigated whether visual salience correlates with entity mention and ordering, but often disregarded its effect on grammar or relied
on simplistic images or artificial cues. In this study, we explore the use of generative AI to better control for salience in visual stimuli while keeping them realistic, and to serve as a proxy for human participants in studying how different types of salience impact image descriptions. We consider three salience types: perceptual (e.g. relative size in the image), inherent (e.g. animacy), and relational (e.g. human–object interaction). We first analyze human- and AI-generated captions for natural images to examine how salience correlates with how early, and in what grammatical role, an entity is mentioned. We find strong correlations between models and humans in this observational study, justifying the use of AI models alone in a further causal study. For this second study, we created datasets composed of pairs of images, where we used an image-editing model to intervene on the salience of a target entity. We show that relational and perceptual salience lead to the entity being mentioned earlier in captions and being mapped to more prominent grammatical roles. The magnitude of this effect varies across entity types, with animate entities (high inherent salience) showing a particularly distinct pattern.
Original languageEnglish
Title of host publicationProceedings of the 30th Conference on Computational Natural Language Learning
PublisherAssociation for Computational Linguistics (ACL)
Pages1-17
Number of pages17
Publication statusAccepted/In press - 20 Apr 2026
EventThe 30th Conference on Computational Natural Language Learning - San Diego, United States
Duration: 3 Jul 20264 Jul 2026
Conference number: 30
https://conll.org/

Conference

ConferenceThe 30th Conference on Computational Natural Language Learning
Abbreviated titleCoNLL 2026
Country/TerritoryUnited States
CitySan Diego
Period3/07/264/07/26
Internet address

Fingerprint

Dive into the research topics of 'Measuring the effects of visual salience in human and AI descriptions with image editing'. Together they form a unique fingerprint.

Cite this