Abstract
Online newspaper articles can accumulate comments at volumes that prevent close reading. Summarisation of the comments allows interaction at a higher level and can lead to an understanding of the overall discussion. Comment summarisation requires topic clustering, comment ranking and extraction. Clustering must be robust as the subsequent extraction relies on a good set of clusters. Comment data, as with many social media datasets, contains very short documents and the number of words in the documents is a limiting factors on the performance of LDA clustering. We evaluate whether we can combine comments to form larger documents to improve the quality of clusters. We find that combining comments with comments that reply to them produce the highest quality clusters.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics – Student Research Workshop |
| Publisher | Association for Computational Linguistics |
| Pages | 43-50 |
| Number of pages | 8 |
| ISBN (Print) | 978-1-945626-02-9 |
| DOIs | |
| Publication status | E-pub ahead of print - 12 Aug 2016 |
| Event | 54th Annual Meeting of the Association for Computational Linguistics – Student Research Workshop - Berlin, Germany Duration: 7 Aug 2016 → 12 Aug 2016 |
Conference
| Conference | 54th Annual Meeting of the Association for Computational Linguistics – Student Research Workshop |
|---|---|
| Abbreviated title | ACL 2016 |
| Country/Territory | Germany |
| City | Berlin |
| Period | 7/08/16 → 12/08/16 |
Fingerprint
Dive into the research topics of 'Improving Topic Model Clustering of Newspaper Comments for Summarisation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver