Umeå University's logo

umu.sePublikasjoner
Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Comparing human-perceived cluster characteristics through the lens of CIPHE: measuring coherence beyond keywords
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap. Aeterna Labs, Sweden.ORCID-id: 0000-0002-4366-7863
Swedish University of Agricultural Sciences, Sweden.ORCID-id: 0000-0001-6601-5190
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap.ORCID-id: 0000-0001-7349-7693
2025 (engelsk)Inngår i: Journal of Data Mining and Digital Humanities, E-ISSN 2416-5999, Vol. NLP4DH, artikkel-id 32Artikkel i tidsskrift (Fagfellevurdert) Published
Abstract [en]

A frequent problem in document clustering and topic modeling is the lack of ground truth. Models are typically intended to reflect some aspect of how human readers view texts (the general theme, sentiment, emotional response, etc), but it can be difficult to assess whether they actually do. The only real ground truth is human judgement. To enable researchers and practitioners to collect such judgement in a cost-efficient standardized way, we have developed the crowdsourcing solution CIPHE -- Cluster Interpretation and Precision from Human Exploration. CIPHE is an adaptable framework which systematically gathers and evaluates data on the human perception of a set of document clusters where participants read sample texts from the cluster. In this article, we use CIPHE to study the limitations that keyword-based methods pose in topic modeling coherence evaluation. Keyword methods, including word intrusion, are compared with the outcome of the thorougher CIPHE on scoring and characterizing clusters. The results show how the abstraction of keywords skews the cluster interpretation for almost half of the compared instances, meaning that many important cluster characteristics are missed. Further, we present a case study where CIPHE is used to (a) provide insights into the UK news domain and (b) find out how the evaluated clustering model should be tuned to better suit the intended application. The experiments provide evidence that CIPHE characterizes clusters in a predictable manner and has the potential to be a valuable framework for using human evaluation in the pursuit of nuanced research aims.

sted, utgiver, år, opplag, sider
Centre pour la Communication Scientifique Directe (CCSD) , 2025. Vol. NLP4DH, artikkel-id 32
Emneord [en]
document clustering, topic modeling, topic modeling evaluation, news clustering, topic coherence, human evaluation methods, crowdsourced cluster validation, BERTopic, CIPHE
HSV kategori
Forskningsprogram
datalogi
Identifikatorer
URN: urn:nbn:se:umu:diva-236229DOI: 10.46298/jdmdh.15044OAI: oai:DiVA.org:umu-236229DiVA, id: diva2:1943168
Merknad

The code for the CIPHE platform is uploaded at https://github.com/antoneklund/CIPHE/

Tilgjengelig fra: 2025-03-09 Laget: 2025-03-09 Sist oppdatert: 2025-03-11bibliografisk kontrollert
Inngår i avhandling
1. Evaluating document clusters through human interpretation
Åpne denne publikasjonen i ny fane eller vindu >>Evaluating document clusters through human interpretation
2025 (engelsk)Doktoravhandling, med artikler (Annet vitenskapelig)
Alternativ tittel[sv]
Utvärdering av dokumentkluster genom mänsklig tolkning
Abstract [en]

Document clustering is a technique for organizing and discovering patterns in large collections of text, often used in applications such as news aggregation and contextual advertising. An example is the automatic grouping of news articles by theme, which is the focus of this thesis. For a clustering to be successful, typically the resulting clusters need to appear interpretable and coherent to a human. However, there is a lack of efficient methods to reliably assess the quality of a clustering in terms of human-perceived coherence, which is essential for ensuring its usefulness in real-world applications.

To address the lack of evaluation methods for document clustering focusing on human interpretation, we introduced Cluster Interpretation and Precision from Human Exploration (CIPHE). CIPHE tasks human evaluators to explore document samples from a cluster and collects their interpretation. The interpretation is collected through a standardized survey and then processed with the framework metrics to yield the cluster precision and characteristics. This thesis presents and discusses the development process of CIPHE. The feasibility of performing the exploratory tasks of CIPHE in a crowdsourcing environment was investigated, which resulted in insights on how to formulate instructions. Additionally, CIPHE was confirmed to identify characteristics other than the main theme such as the negative emotional response.

CIPHE was paired with a standard clustering pipeline to evaluate its capabilities and limitations. The pipeline is widely applied for its adaptability and conceptual simplicity, and also being part of the popular topic model BERTopic. The empirical results of applying CIPHE suggest that the pipeline, when integrated with a Transformer-based language model, generally yields coherent clusters.

Additionally, topic models have a similar aim as document clustering which is to automate the corpus processing and present the underlying themes to a human. Topic modeling has rich research on the human interpretation of topic coherence. In the thesis, the human interpretation collected with CIPHE was related to established research in topic coherence. Specifically, the human interpretation collected with CIPHE was used to highlight limitations with the keyword representations that topic coherence evaluation relies on.

sted, utgiver, år, opplag, sider
Umeå: Umeå University, 2025. s. 44
Serie
Report / UMINF, ISSN 0348-0542 ; 25.03
Emneord
document clustering, topic modeling, information retrieval, human evaluation, human-in-the-loop, news clustering, natural language processing, topic coherence, human interpretation
HSV kategori
Forskningsprogram
datalogi
Identifikatorer
urn:nbn:se:umu:diva-236295 (URN)9789180706469 (ISBN)9789180706476 (ISBN)
Disputas
2025-04-03, Lindellhallen 3 (UB.A.230), Samhällsvetarhuset, Umeå, 13:15 (engelsk)
Opponent
Veileder
Forskningsfinansiär
Swedish Foundation for Strategic Research, ID19-0055
Tilgjengelig fra: 2025-03-13 Laget: 2025-03-10 Sist oppdatert: 2025-03-12bibliografisk kontrollert

Open Access i DiVA

fulltext(2210 kB)92 nedlastinger
Filinformasjon
Fil FULLTEXT01.pdfFilstørrelse 2210 kBChecksum SHA-512
80e0e189738c3fde927f835e6d535a1e715f0f0d02968ae3a1d11952b7446ec0e4f85e217c17a72edd5ca6b6b3cf15c4dcda4a1d9f6abe11be56f6781fd2512e
Type fulltextMimetype application/pdf

Andre lenker

Forlagets fulltekst

Person

Eklund, AntonDrewes, Frank

Søk i DiVA

Av forfatter/redaktør
Eklund, AntonForsman, MonaDrewes, Frank
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar
Totalt: 93 nedlastinger
Antall nedlastinger er summen av alle nedlastinger av alle fulltekster. Det kan for eksempel være tidligere versjoner som er ikke lenger tilgjengelige

doi
urn-nbn

Altmetric

doi
urn-nbn
Totalt: 510 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf