Abstract / Description of output
SemEval 2021 Task 7, HaHackathon, was the first shared task to combine the previously separate domains of humor detection and offense detection. We collected 10,000 texts from Twitter and the Kaggle Short Jokes dataset, and had each annotated for humor and offense by 20 annotators aged 18-70. Our subtasks were binary humor detection, prediction of humor and offense ratings, and a novel controversy task: to predict if the variance in the humor ratings was higher than a specific threshold. The subtasks attracted 36-58 submissions, with most of the participants choosing to use pre-trained language models. Many of the highest performing teams also implemented additional optimization techniques, including task-adaptive training and adversarial training. The results suggest that the participating systems are well suited to humor detection, but that humor controversy is a more challenging task. We discuss which models excel in this task, which auxiliary techniques boost their performance, and analyze the errors which were not captured by the best systems.
Original language | English |
---|---|
Title of host publication | Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval 2021) |
Publisher | Association for Computational Linguistics (ACL) |
Pages | 105-119 |
Number of pages | 15 |
ISBN (Electronic) | 978-1-954085-70-1 |
DOIs | |
Publication status | Published - 1 Aug 2021 |
Event | 15th International Workshop on Semantic Evaluation - Online Duration: 5 Aug 2021 → 6 Aug 2021 https://semeval.github.io/SemEval2021/ |
Workshop
Workshop | 15th International Workshop on Semantic Evaluation |
---|---|
Abbreviated title | SemEval 2021 |
Period | 5/08/21 → 6/08/21 |
Internet address |