Umeå University's logo

umu.sePublikasjoner
Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Measuring image-relation alignment: reference-free evaluation of VLMs and synthetic pre-training for open-vocabulary scene graph generation
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap.
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap.
Cnrs Irl 2010 Crossing, Australia; Imt Atlantique, France.
National Institute of Informatics, Japan.
2025 (engelsk)Inngår i: 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025: proceedings, Institute of Electrical and Electronics Engineers (IEEE), 2025, s. 7664-7673Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Scene Graph Generation (SGG) encodes visual relationships between objects in images as graph structures. Thanks to the advances of Vision-Language Models (VLMs), the task of Open-Vocabulary SGG has been recently proposed where models are evaluated on their functionality to learn a wide and diverse range of relations. Current benchmarks in SGG, however, possess a very limited vocabulary, making the evaluation of open-source models inefficient. In this paper, we propose a new reference-free metric to fairly evaluate the open-vocabulary capabilities of VLMs for relation prediction. Another limitation of Open-Vocabulary SGG is the reliance on weakly supervised data of poor quality for pre-training. We also propose a new solution for quickly generating high-quality synthetic data through region-specific prompt tuning of VLMs. Experimental results show that pre-training with this new data split can benefit the generalization capabilities of Open-Voc SGG models11Code and data available at https://github.com/Maelic/OpenVocSGG.

sted, utgiver, år, opplag, sider
Institute of Electrical and Electronics Engineers (IEEE), 2025. s. 7664-7673
Serie
IEEE International Conference on Computer Vision Workshops, ISSN 2473-9936, E-ISSN 2473-9944
Emneord [en]
open-vocabulary, relationship prediction, scene graph generation, vision-language models
HSV kategori
Identifikatorer
URN: urn:nbn:se:umu:diva-252865DOI: 10.1109/ICCVW69036.2025.00789Scopus ID: 2-s2.0-105035155213ISBN: 9798331589882 (digital)ISBN: 9798331589899 (tryckt)OAI: oai:DiVA.org:umu-252865DiVA, id: diva2:2058548
Konferanse
2025 IEEE/CVF International Conference on Computer Vision Workshops ICCV-W 2025, Honolulu, United States, 19-20 October 2025
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)Tilgjengelig fra: 2026-05-07 Laget: 2026-05-07 Sist oppdatert: 2026-05-07bibliografisk kontrollert

Open Access i DiVA

Fulltekst mangler i DiVA

Andre lenker

Forlagets fulltekstScopus

Person

Neau, MaëlicFalomir, Zoe

Søk i DiVA

Av forfatter/redaktør
Neau, MaëlicFalomir, Zoe
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric

doi
isbn
urn-nbn
Totalt: 11 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf