Abstract
How does our perception of the world influence the way we talk about it? Psycholinguistic studies have investigated whether visual salience correlates with entity mention and ordering, but often disregarded its effect on grammar or relied
on simplistic images or artificial cues. In this study, we explore the use of generative AI to better control for salience in visual stimuli while keeping them realistic, and to serve as a proxy for human participants in studying how different types of salience impact image descriptions. We consider three salience types: perceptual (e.g. relative size in the image), inherent (e.g. animacy), and relational (e.g. human–object interaction). We first analyze human- and AI-generated captions for natural images to examine how salience correlates with how early, and in what grammatical role, an entity is mentioned. We find strong correlations between models and humans in this observational study, justifying the use of AI models alone in a further causal study. For this second study, we created datasets composed of pairs of images, where we used an image-editing model to intervene on the salience of a target entity. We show that relational and perceptual salience lead to the entity being mentioned earlier in captions and being mapped to more prominent grammatical roles. The magnitude of this effect varies across entity types, with animate entities (high inherent salience) showing a particularly distinct pattern.
on simplistic images or artificial cues. In this study, we explore the use of generative AI to better control for salience in visual stimuli while keeping them realistic, and to serve as a proxy for human participants in studying how different types of salience impact image descriptions. We consider three salience types: perceptual (e.g. relative size in the image), inherent (e.g. animacy), and relational (e.g. human–object interaction). We first analyze human- and AI-generated captions for natural images to examine how salience correlates with how early, and in what grammatical role, an entity is mentioned. We find strong correlations between models and humans in this observational study, justifying the use of AI models alone in a further causal study. For this second study, we created datasets composed of pairs of images, where we used an image-editing model to intervene on the salience of a target entity. We show that relational and perceptual salience lead to the entity being mentioned earlier in captions and being mapped to more prominent grammatical roles. The magnitude of this effect varies across entity types, with animate entities (high inherent salience) showing a particularly distinct pattern.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 30th Conference on Computational Natural Language Learning |
| Publisher | Association for Computational Linguistics (ACL) |
| Pages | 1-17 |
| Number of pages | 17 |
| Publication status | Accepted/In press - 20 Apr 2026 |
| Event | The 30th Conference on Computational Natural Language Learning - San Diego, United States Duration: 3 Jul 2026 → 4 Jul 2026 Conference number: 30 https://conll.org/ |
Conference
| Conference | The 30th Conference on Computational Natural Language Learning |
|---|---|
| Abbreviated title | CoNLL 2026 |
| Country/Territory | United States |
| City | San Diego |
| Period | 3/07/26 → 4/07/26 |
| Internet address |
Fingerprint
Dive into the research topics of 'Measuring the effects of visual salience in human and AI descriptions with image editing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver