Umeå universitets logga

umu.sePublikationer
Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
CIPHE: A Framework for Document Cluster Interpretation and Precision from Human Exploration
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap. (Foundations of Language Processing)ORCID-id: 0000-0002-4366-7863
Aeterna Labs, Sweden.ORCID-id: 0000-0001-6601-5190
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap. (Foundations of Language Processing)ORCID-id: 0000-0001-7349-7693
2024 (Engelska)Ingår i: Proceedings of the 4th international conference on natural language processing for digital humanities / [ed] Mika Hämäläinen; Emily Öhman; So Miyagawa; Khalid Alnajjar; Yuri Bizzoni, Association for Computational Linguistics, 2024, s. 536-548Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Document clustering models serve unique application purposes, which turns model quality into a property that depends on the needs of the individual investigator. We propose a framework, Cluster Interpretation and Precision from Human Exploration (CIPHE), for collecting and quantifying human interpretations of cluster samples. CIPHE tasks survey participants to explore actual document texts from cluster samples and records their perceptions. It also includes a novel inclusion task that is used to calculate the cluster precision in an indirect manner. A case study on news clusters shows that CIPHE reveals which clusters have multiple interpretation angles, aiding the investigator in their exploration.

Ort, förlag, år, upplaga, sidor
Association for Computational Linguistics, 2024. s. 536-548
Nyckelord [en]
document clustering, topic modeling, clustering, human evaluation, CIPHE, news articles
Nationell ämneskategori
Språkbehandling och datorlingvistik
Forskningsämne
datorlingvistik
Identifikatorer
URN: urn:nbn:se:umu:diva-231697Scopus ID: 2-s2.0-85216576924ISBN: 979-8-89176-181-0 (digital)OAI: oai:DiVA.org:umu-231697DiVA, id: diva2:1912204
Konferens
4th International Conference on Natural Language Processing for Digital Humanities, Miami, USA, November 15-16, 2024
Tillgänglig från: 2024-11-11 Skapad: 2024-11-11 Senast uppdaterad: 2025-04-02Bibliografiskt granskad
Ingår i avhandling
1. Evaluating document clusters through human interpretation
Öppna denna publikation i ny flik eller fönster >>Evaluating document clusters through human interpretation
2025 (Engelska)Doktorsavhandling, sammanläggning (Övrigt vetenskapligt)
Alternativ titel[sv]
Utvärdering av dokumentkluster genom mänsklig tolkning
Abstract [en]

Document clustering is a technique for organizing and discovering patterns in large collections of text, often used in applications such as news aggregation and contextual advertising. An example is the automatic grouping of news articles by theme, which is the focus of this thesis. For a clustering to be successful, typically the resulting clusters need to appear interpretable and coherent to a human. However, there is a lack of efficient methods to reliably assess the quality of a clustering in terms of human-perceived coherence, which is essential for ensuring its usefulness in real-world applications.

To address the lack of evaluation methods for document clustering focusing on human interpretation, we introduced Cluster Interpretation and Precision from Human Exploration (CIPHE). CIPHE tasks human evaluators to explore document samples from a cluster and collects their interpretation. The interpretation is collected through a standardized survey and then processed with the framework metrics to yield the cluster precision and characteristics. This thesis presents and discusses the development process of CIPHE. The feasibility of performing the exploratory tasks of CIPHE in a crowdsourcing environment was investigated, which resulted in insights on how to formulate instructions. Additionally, CIPHE was confirmed to identify characteristics other than the main theme such as the negative emotional response.

CIPHE was paired with a standard clustering pipeline to evaluate its capabilities and limitations. The pipeline is widely applied for its adaptability and conceptual simplicity, and also being part of the popular topic model BERTopic. The empirical results of applying CIPHE suggest that the pipeline, when integrated with a Transformer-based language model, generally yields coherent clusters.

Additionally, topic models have a similar aim as document clustering which is to automate the corpus processing and present the underlying themes to a human. Topic modeling has rich research on the human interpretation of topic coherence. In the thesis, the human interpretation collected with CIPHE was related to established research in topic coherence. Specifically, the human interpretation collected with CIPHE was used to highlight limitations with the keyword representations that topic coherence evaluation relies on.

Ort, förlag, år, upplaga, sidor
Umeå: Umeå University, 2025. s. 44
Serie
Report / UMINF, ISSN 0348-0542 ; 25.03
Nyckelord
document clustering, topic modeling, information retrieval, human evaluation, human-in-the-loop, news clustering, natural language processing, topic coherence, human interpretation
Nationell ämneskategori
Språkbehandling och datorlingvistik
Forskningsämne
datalogi
Identifikatorer
urn:nbn:se:umu:diva-236295 (URN)9789180706469 (ISBN)9789180706476 (ISBN)
Disputation
2025-04-03, Lindellhallen 3 (UB.A.230), Samhällsvetarhuset, Umeå, 13:15 (Engelska)
Opponent
Handledare
Forskningsfinansiär
Stiftelsen för strategisk forskning (SSF), ID19-0055
Tillgänglig från: 2025-03-13 Skapad: 2025-03-10 Senast uppdaterad: 2025-03-12Bibliografiskt granskad

Open Access i DiVA

fulltext(902 kB)215 nedladdningar
Filinformation
Filnamn FULLTEXT01.pdfFilstorlek 902 kBChecksumma SHA-512
bee83c49f7dc5d7d2f07010f818151fed16924d14df06965fb6918e8258d4bf3d7ac8bc71811cdda06f63e465c4cb4e7bcb8e9867e2fc5ca307011cb4392d568
Typ fulltextMimetyp application/pdf

Övriga länkar

ScopusPublisher's full text

Person

Eklund, AntonDrewes, Frank

Sök vidare i DiVA

Av författaren/redaktören
Eklund, AntonForsman, MonaDrewes, Frank
Av organisationen
Institutionen för datavetenskap
Språkbehandling och datorlingvistik

Sök vidare utanför DiVA

GoogleGoogle Scholar
Totalt: 218 nedladdningar
Antalet nedladdningar är summan av nedladdningar för alla fulltexter. Det kan inkludera t.ex tidigare versioner som nu inte längre är tillgängliga.

isbn
urn-nbn

Altmetricpoäng

isbn
urn-nbn
Totalt: 494 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf