Umeå University's logo

umu.sePublikasjoner
Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
ADCluster: Adaptive Deep Clustering for unsupervised learning from unlabeled documents
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap.ORCID-id: 0000-0002-6791-8284
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap.ORCID-id: 0000-0001-8820-2405
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap.ORCID-id: 0000-0002-9842-7840
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap.ORCID-id: 0000-0001-7349-7693
2023 (engelsk)Inngår i: Proceedings of the 6th International Conference on Natural Language and Speech Processing (ICNLSP 2023) / [ed] Mourad Abbas; Abed Alhakim Freihat, Association for Computational Linguistics, 2023, s. 68-77Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

We introduce ADCluster, a deep document clustering approach based on language models that is trained to adapt to the clustering task. This adaptability is achieved through an iterative process where K-Means clustering is applied to the dataset, followed by iteratively training a deep classifier with generated pseudo-labels – an approach referred to as inner adaptation. The model is also able to adapt to changes in the data as new documents are added to the document collection. The latter type of adaptation, outer adaptation, is obtained by resuming the inner adaptation when a new chunk of documents has arrived. We explore two outer adaptation strategies, namely accumulative adaptation (training is resumed on the accumulated set of all documents) and non-accumulative adaptation (training is resumed using only the new chunk of data). We show that ADCluster outperforms established document clustering techniques on medium and long-text documents by a large margin. Additionally, our approach outperforms well-established baseline methods under both the accumulative and non-accumulative outer adaptation scenarios.

sted, utgiver, år, opplag, sider
Association for Computational Linguistics, 2023. s. 68-77
Emneord [en]
deep clustering, adaptive, deep learning, unsupervised, data stream
HSV kategori
Forskningsprogram
datalogi; datorlingvistik
Identifikatorer
URN: urn:nbn:se:umu:diva-220260OAI: oai:DiVA.org:umu-220260DiVA, id: diva2:1833048
Konferanse
6th International Conference on Natural Language and Speech Processing (ICNLSP 2023), Online, December 16-17, 2023.
Tilgjengelig fra: 2024-01-31 Laget: 2024-01-31 Sist oppdatert: 2024-07-02bibliografisk kontrollert
Inngår i avhandling
1. Deep learning for news topic identification in limited supervision and unsupervised settings
Åpne denne publikasjonen i ny fane eller vindu >>Deep learning for news topic identification in limited supervision and unsupervised settings
2024 (engelsk)Doktoravhandling, med artikler (Annet vitenskapelig)
Alternativ tittel[sv]
Djup maskininlärning för identifiering av nyhetsämnen i inlärningssituationer med begränsad eller ingen övervakning
Abstract [en]

In today's world, following news is crucial for decision-making and staying informed. With the growing volume of daily news, automated processing is essential for timely insights and in aiding individuals and corporations in navigating the complexities of the information society. Another use of automated processing is contextual advertising, which addresses privacy concerns associated with cookie-based advertising by placing ads solely based on web page content, without tracking users or their online behavior. Therefore, accurately determining and categorizing page content is crucial for effective ad placements. The news media, heavily reliant on advertising to sustain operations, represent a substantial market for contextual advertising strategies.

Inspired by these practical applications and the advancements in deep learning over the past decade, this thesis mainly focuses on using deep learning for categorizing news articles into topics of varying granularity. Considering the dynamic nature of these applications and the limited availability of relevant labeled datasets for training models, the thesis emphasizes developing methods that can be trained effectively using unlabeled or partially labeled data. It proposes semi-supervised text classification models for categorizing datasets into predefined coarse-grained topics, where only a few labeled examples exist for each topic, while the majority of the dataset remains unlabeled. Furthermore, to better explore coarse-grained topics within news archives and streams and overcome the limitations of predefined topics in text classification the thesis suggests deep clustering approaches that can be trained in unsupervised settings. 

Moreover, to address the identification of fine-grained topics, the thesis introduces a novel story discovery model for monitoring event-based topics in multi-source news streams. Given that online news reporting often incorporates diverse modalities like text, images, video, and audio to convey information, the thesis finally initiates an investigation into the synergy between textual and visual elements in news article analysis. To achieve this objective, a text-image dataset was annotated, and a baseline was established for event-topic discovery in multimodal news streams. While primarily intended for news monitoring and contextual advertising, the proposed models can, more generally, be regarded as novel approaches in semi-supervised text classification, deep clustering, and news story discovery. Comparison with state-of-the-art baseline models demonstrates their effectiveness in addressing the respective objectives.

sted, utgiver, år, opplag, sider
Umeå: Umeå University, 2024. s. 62
Serie
Report / UMINF, ISSN 0348-0542 ; 24.04
Emneord
Topic Identification, Data Clustering, News Stream Clustering, Semi-Supervised Learning, Unsupervised Learning, Event Topics, News Stories, Multimodal News, Document Classification, Document Clustering, Deep Learning, Deep Clustering, Pre-trained Language Models
HSV kategori
Identifikatorer
urn:nbn:se:umu:diva-222534 (URN)9789180703420 (ISBN)9789180703437 (ISBN)
Disputas
2024-04-16, MIT.A.121, MIT-huset, Umeå, 13:15 (engelsk)
Opponent
Veileder
Tilgjengelig fra: 2024-03-26 Laget: 2024-03-20 Sist oppdatert: 2024-03-22bibliografisk kontrollert

Open Access i DiVA

fulltext(810 kB)185 nedlastinger
Filinformasjon
Fil FULLTEXT01.pdfFilstørrelse 810 kBChecksum SHA-512
7b99c4b3ac476439877d2cd654481c9c61321232bb0be6a40526fafbf07891fdef76446fdd35c35137eb29943b0a61a6150bdb830c1aa1570cad53e9bc49e0a9
Type fulltextMimetype application/pdf

Andre lenker

Publisher's full text

Person

Hatefi, ArezooVu, Xuan-SonBhuyan, Monowar H.Drewes, Frank

Søk i DiVA

Av forfatter/redaktør
Hatefi, ArezooVu, Xuan-SonBhuyan, Monowar H.Drewes, Frank
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar
Totalt: 186 nedlastinger
Antall nedlastinger er summen av alle nedlastinger av alle fulltekster. Det kan for eksempel være tidligere versjoner som er ikke lenger tilgjengelige

urn-nbn

Altmetric

urn-nbn
Totalt: 716 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf