Measuring image-relation alignment: reference-free evaluation of VLMs and synthetic pre-training for open-vocabulary scene graph generation
2025 (Engelska)Ingår i: 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025: proceedings, Institute of Electrical and Electronics Engineers (IEEE), 2025, s. 7664-7673Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]
Scene Graph Generation (SGG) encodes visual relationships between objects in images as graph structures. Thanks to the advances of Vision-Language Models (VLMs), the task of Open-Vocabulary SGG has been recently proposed where models are evaluated on their functionality to learn a wide and diverse range of relations. Current benchmarks in SGG, however, possess a very limited vocabulary, making the evaluation of open-source models inefficient. In this paper, we propose a new reference-free metric to fairly evaluate the open-vocabulary capabilities of VLMs for relation prediction. Another limitation of Open-Vocabulary SGG is the reliance on weakly supervised data of poor quality for pre-training. We also propose a new solution for quickly generating high-quality synthetic data through region-specific prompt tuning of VLMs. Experimental results show that pre-training with this new data split can benefit the generalization capabilities of Open-Voc SGG models11Code and data available at https://github.com/Maelic/OpenVocSGG.
Ort, förlag, år, upplaga, sidor
Institute of Electrical and Electronics Engineers (IEEE), 2025. s. 7664-7673
Serie
IEEE International Conference on Computer Vision Workshops, ISSN 2473-9936, E-ISSN 2473-9944
Nyckelord [en]
open-vocabulary, relationship prediction, scene graph generation, vision-language models
Nationell ämneskategori
Datavetenskap (datalogi)
Identifikatorer
URN: urn:nbn:se:umu:diva-252865DOI: 10.1109/ICCVW69036.2025.00789Scopus ID: 2-s2.0-105035155213ISBN: 9798331589882 (digital)ISBN: 9798331589899 (tryckt)OAI: oai:DiVA.org:umu-252865DiVA, id: diva2:2058548
Konferens
2025 IEEE/CVF International Conference on Computer Vision Workshops ICCV-W 2025, Honolulu, United States, 19-20 October 2025
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)2026-05-072026-05-072026-05-07Bibliografiskt granskad