Skip to main navigation Skip to search Skip to main content

Improving Topic Model Clustering of Newspaper Comments for Summarisation

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Online newspaper articles can accumulate comments at volumes that prevent close reading. Summarisation of the comments allows interaction at a higher level and can lead to an understanding of the overall discussion. Comment summarisation requires topic clustering, comment ranking and extraction. Clustering must be robust as the subsequent extraction relies on a good set of clusters. Comment data, as with many social media datasets, contains very short documents and the number of words in the documents is a limiting factors on the performance of LDA clustering. We evaluate whether we can combine comments to form larger documents to improve the quality of clusters. We find that combining comments with comments that reply to them produce the highest quality clusters.
Original languageEnglish
Title of host publicationProceedings of the 54th Annual Meeting of the Association for Computational Linguistics – Student Research Workshop
PublisherAssociation for Computational Linguistics
Pages43-50
Number of pages8
ISBN (Print)978-1-945626-02-9
DOIs
Publication statusE-pub ahead of print - 12 Aug 2016
Event54th Annual Meeting of the Association for Computational Linguistics – Student Research Workshop - Berlin, Germany
Duration: 7 Aug 201612 Aug 2016

Conference

Conference54th Annual Meeting of the Association for Computational Linguistics – Student Research Workshop
Abbreviated titleACL 2016
Country/TerritoryGermany
CityBerlin
Period7/08/1612/08/16

Fingerprint

Dive into the research topics of 'Improving Topic Model Clustering of Newspaper Comments for Summarisation'. Together they form a unique fingerprint.

Cite this