Umeå University's logo

umu.sePublications
Change search
Link to record
Permanent link

Direct link
Neau, Maëlic
Publications (1 of 1) Show all publications
Neau, M., Falomir, Z., Buche, C. & Sugimoto, A. (2025). Measuring image-relation alignment: reference-free evaluation of VLMs and synthetic pre-training for open-vocabulary scene graph generation. In: 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025: proceedings. Paper presented at 2025 IEEE/CVF International Conference on Computer Vision Workshops ICCV-W 2025, Honolulu, United States, 19-20 October 2025 (pp. 7664-7673). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Measuring image-relation alignment: reference-free evaluation of VLMs and synthetic pre-training for open-vocabulary scene graph generation
2025 (English)In: 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025: proceedings, Institute of Electrical and Electronics Engineers (IEEE), 2025, p. 7664-7673Conference paper, Published paper (Refereed)
Abstract [en]

Scene Graph Generation (SGG) encodes visual relationships between objects in images as graph structures. Thanks to the advances of Vision-Language Models (VLMs), the task of Open-Vocabulary SGG has been recently proposed where models are evaluated on their functionality to learn a wide and diverse range of relations. Current benchmarks in SGG, however, possess a very limited vocabulary, making the evaluation of open-source models inefficient. In this paper, we propose a new reference-free metric to fairly evaluate the open-vocabulary capabilities of VLMs for relation prediction. Another limitation of Open-Vocabulary SGG is the reliance on weakly supervised data of poor quality for pre-training. We also propose a new solution for quickly generating high-quality synthetic data through region-specific prompt tuning of VLMs. Experimental results show that pre-training with this new data split can benefit the generalization capabilities of Open-Voc SGG models11Code and data available at https://github.com/Maelic/OpenVocSGG.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Series
IEEE International Conference on Computer Vision Workshops, ISSN 2473-9936, E-ISSN 2473-9944
Keywords
open-vocabulary, relationship prediction, scene graph generation, vision-language models
National Category
Computer Sciences
Identifiers
urn:nbn:se:umu:diva-252865 (URN)10.1109/ICCVW69036.2025.00789 (DOI)2-s2.0-105035155213 (Scopus ID)9798331589882 (ISBN)9798331589899 (ISBN)
Conference
2025 IEEE/CVF International Conference on Computer Vision Workshops ICCV-W 2025, Honolulu, United States, 19-20 October 2025
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Available from: 2026-05-07 Created: 2026-05-07 Last updated: 2026-05-07Bibliographically approved
Organisations

Search in DiVA

Show all publications