Umeå University's logo

umu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Fusion in context: a multimodal approach to affective state recognition
KTH: The Royal Institute of Technology, Stockholm, Sweden.
Pal Robotics, Barcelona, Spain.
Umeå University, Faculty of Science and Technology, Department of Computing Science.ORCID iD: 0000-0003-2282-9939
KTH: The Royal Institute of Technology, Stockholm, Sweden.
Show others and affiliations
2025 (English)In: 2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), IEEE, 2025, p. 1049-1055Conference paper, Published paper (Refereed)
Abstract [en]

Accurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional expressions can be influenced by contextual factors, leading to misinterpretations if context is not considered. Multimodal fusion, combining modalities like facial expressions, speech, and physiological signals, has shown promise in improving affect recognition. This paper proposes a transformer-based multimodal fusion approach that leverages facial thermal data, facial action units, and textual context information for context-aware emotion recognition. We explore modality-specific encoders to learn tailored representations, which are then fused and processed by a shared transformer encoder to capture temporal dependencies and interactions. The proposed method is evaluated on a dataset collected from participants engaged in a tangible tabletop Pacman game designed to induce various affective states. Our results demonstrate improvements from incorporating contextual information and multimodal fusion, achieving 89% F1 score with our full model compared to 65% for action units alone and 30% for thermal data alone.

Place, publisher, year, edition, pages
IEEE, 2025. p. 1049-1055
Series
IEEE RO-MAN, ISSN 1944-9445, E-ISSN 1944-9437
Keywords [en]
computer vision, Human detection, social human-robot interaction
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:umu:diva-247946DOI: 10.1109/RO-MAN63969.2025.11217904Scopus ID: 2-s2.0-105024539281ISBN: 9798331587710 (electronic)OAI: oai:DiVA.org:umu-247946DiVA, id: diva2:2025516
Conference
34th IEEE International Conference on Robot and Human Interactive Communication, RO-MAN 2025, Eindhoven, Netherlands, August 25-29, 2025
Available from: 2026-01-07 Created: 2026-01-07 Last updated: 2026-01-07Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Güneysu Özgür, Arzu

Search in DiVA

By author/editor
Güneysu Özgür, Arzu
By organisation
Department of Computing Science
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 31 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf