Projects per year
Abstract / Description of output
Like most other minority languages, Scottish Gaelic has limited tools and resources available for Natural Language Processing research and applications. These limitations restrict the potential of the language to participate in modern speech technology, while also restricting research in fields such as corpus linguistics and the Digital Humanities. At the same time, Gaelic has a long written history, is well-described linguistically, and is unusually well-supported in terms of potential NLP training data. For instance, archives such as the School of Scottish Studies hold thousands of digitised recordings of vernacular speech, many of which have been transcribed as paper-based, handwritten manuscripts. In this paper, we describe a project to digitise and recognise a corpus of handwritten narrative transcriptions, with the intention of re-purposing it to develop a Gaelic speech recognition system.
Original language | English |
---|---|
Title of host publication | Proceedings of the 4th Celtic Language Technology Workshop at LREC 2022 (CLTW 4) |
Editors | Theodorus Fransen, William Lamb, Delyth Prys |
Publisher | European Language Resources Association (ELRA) |
Pages | 60-70 |
Number of pages | 11 |
ISBN (Electronic) | 9791095546733 |
Publication status | Published - 15 Jun 2022 |
Event | The 4th Celtic Language Technology Workshop at LREC 2022 - Marseille, France Duration: 20 Jun 2022 → 20 Jun 2022 http://techiaith.bangor.ac.uk/celticlt/cltw/?lang=en |
Workshop
Workshop | The 4th Celtic Language Technology Workshop at LREC 2022 |
---|---|
Abbreviated title | CLTW 2022 |
Country/Territory | France |
City | Marseille |
Period | 20/06/22 → 20/06/22 |
Internet address |
Keywords / Materials (for Non-textual outputs)
- Scottish Gaelic
- Handwriting Recognition
- minority languages
- Low-Resource NLP
- Digital Humanities
Fingerprint
Dive into the research topics of 'Handwriting recognition for Scottish Gaelic'. Together they form a unique fingerprint.Projects
- 1 Finished
-
Machine-assisted Transcription for Handwritten Gaelic Narratives
Lamb, W., Loxley, J., Alex, B., Sinclair, M. & Muehlberger, G.
1/08/19 → 31/07/20
Project: University Awarded Project Funding
-
Developing automatic speech recognition for Scottish Gaelic
Evans, L., Lamb, W., Sinclair, M. & Alex, B., 15 Jun 2022, Proceedings of the 4th Celtic Language Technology Workshop at LREC 2022 (CLTW 4). Fransen, T., Lamb, W. & Prys, D. (eds.). European Language Resources Association (ELRA), p. 110-120 11 p.Research output: Chapter in Book/Report/Conference proceeding › Conference contribution
Open AccessFile -
Proceedings of the 4th Celtic Language Technology Workshop at LREC 2022 (CLTW 4)
Fransen, T. (ed.), Lamb, W. (ed.) & Prys, D. (ed.), 15 Jun 2022, Marseille: European Language Resources Association (ELRA). 133 p.Research output: Book/Report › Book
Open AccessFile