Topic Modeling for Short Texts with Co-occurrence Frequency-based Expansion

Gabriel Pedrosa; Marcelo Pita; Paulo Bicalho; Anisio Lacerda; Gisele L. Pappa

doi:10.1235/bracis.vi.104

PDF

Published Dec 14, 2016

DOI: https://doi.org/10.1235/bracis.vi.104

Gabriel Pedrosa Marcelo Pita Paulo Bicalho Anisio Lacerda Gisele L. Pappa

Abstract

Short texts are everywhere on the Web, including messages in social media, status messages, etc, and extracting semantically meaningful topics from these collections is an important and difficult task. Topic modeling methods, such as Latent Dirichlet Allocation, were designed for this purpose. However, discovering high quality topics in short text collections is a challenging task. This is because most topic modeling methods rely on information coming from the word co-occurrence distribution in the collection to extract topics. As in short text this information is scarce, topic modeling methods have difficulties in this scenario, and different strategies to tackle this problem have been proposed in the literature. In this direction, this paper introduces a method for topic modeling of short texts that creates pseudo-documents representations from the original documents. The method is simple, effective, and considers word co-occurrence to expand documents, which can be given as input to any topic modeling algorithm. Experiments were run in four datasets and compared against state-of-the-art methods for extracting topics from short text. Results of coherence, NPMI and clustering metrics showed to be statistically significantly better than the baselines in the majority of cases.

How to Cite

PEDROSA, Gabriel et al. Topic Modeling for Short Texts with Co-occurrence Frequency-based Expansion. BRACIS, [S.l.], dec. 2016. Available at: <http://143.54.25.88/index.php/bracis/article/view/104>. Date accessed: 10 july 2026. doi: https://doi.org/10.1235/bracis.vi.104.

ABNT APA BibTeX CBE EndNote - EndNote format (Macintosh & Windows) MLA ProCite - RIS format (Macintosh & Windows) RefWorks Reference Manager - RIS format (Windows only) Turabian

Issue

2016: BRACIS

Section

Artigos

Article Sidebar

Main Article Content

Abstract

Article Details