Umeå University's logo

umu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Measuring image-relation alignment: reference-free evaluation of VLMs and synthetic pre-training for open-vocabulary scene graph generation
Umeå University, Faculty of Science and Technology, Department of Computing Science.
Umeå University, Faculty of Science and Technology, Department of Computing Science.
Cnrs Irl 2010 Crossing, Australia; Imt Atlantique, France.
National Institute of Informatics, Japan.
2025 (English)In: 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025: proceedings, Institute of Electrical and Electronics Engineers (IEEE), 2025, p. 7664-7673Conference paper, Published paper (Refereed)
Abstract [en]

Scene Graph Generation (SGG) encodes visual relationships between objects in images as graph structures. Thanks to the advances of Vision-Language Models (VLMs), the task of Open-Vocabulary SGG has been recently proposed where models are evaluated on their functionality to learn a wide and diverse range of relations. Current benchmarks in SGG, however, possess a very limited vocabulary, making the evaluation of open-source models inefficient. In this paper, we propose a new reference-free metric to fairly evaluate the open-vocabulary capabilities of VLMs for relation prediction. Another limitation of Open-Vocabulary SGG is the reliance on weakly supervised data of poor quality for pre-training. We also propose a new solution for quickly generating high-quality synthetic data through region-specific prompt tuning of VLMs. Experimental results show that pre-training with this new data split can benefit the generalization capabilities of Open-Voc SGG models11Code and data available at https://github.com/Maelic/OpenVocSGG.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025. p. 7664-7673
Series
IEEE International Conference on Computer Vision Workshops, ISSN 2473-9936, E-ISSN 2473-9944
Keywords [en]
open-vocabulary, relationship prediction, scene graph generation, vision-language models
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:umu:diva-252865DOI: 10.1109/ICCVW69036.2025.00789Scopus ID: 2-s2.0-105035155213ISBN: 9798331589882 (electronic)ISBN: 9798331589899 (print)OAI: oai:DiVA.org:umu-252865DiVA, id: diva2:2058548
Conference
2025 IEEE/CVF International Conference on Computer Vision Workshops ICCV-W 2025, Honolulu, United States, 19-20 October 2025
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)Available from: 2026-05-07 Created: 2026-05-07 Last updated: 2026-05-07Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Neau, MaëlicFalomir, Zoe

Search in DiVA

By author/editor
Neau, MaëlicFalomir, Zoe
By organisation
Department of Computing Science
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 8 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf