Handwriting recognition for Scottish Gaelic

Mark Sinclair, Will Lamb, Beatrice Alex

Research output: Contribution to conferencePaperpeer-review


Like most other minority languages, Scottish Gaelic has limited tools and resources available for Natural Language Processing research and applications. These limitations restrict the potential of the language to participate in modern speech technology, while also restricting research in fields such as corpus linguistics and the Digital Humanities. At the same time, Gaelic has a long written history, is well-described linguistically, and is unusually well-supported in terms of potential NLP training data. For instance, archives such as the School of Scottish Studies hold thousands of digitised recordings of vernacular speech, many of which have been transcribed as paper-based, handwritten manuscripts. In this paper, we describe a project to digitise and recognise a corpus of handwritten narrative transcriptions, with the intention of re-purposing it to develop a Gaelic speech recognition system.
Original languageEnglish
Publication statusAccepted/In press - 6 May 2022
EventThe 4th Celtic Language Technology Workshop at LREC 2022 - Marseille, France
Duration: 20 Jun 202220 Jun 2022


WorkshopThe 4th Celtic Language Technology Workshop at LREC 2022
Abbreviated titleCLTW 2022
Internet address


Dive into the research topics of 'Handwriting recognition for Scottish Gaelic'. Together they form a unique fingerprint.

Cite this