Efficient Language-Independent Retrieval of Printed Documents without OCR

Walid Magdy, Kareem Darwish, Motaz El-Saban

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Recent book digitization initiatives have facilitated the access and search of millions of books. Although OCR remains essential for retrieving printed documents, OCR engines remain limited in the languages they handle and are generally expensive to build. This paper proposes a language independent approach that enables search through printed documents in a way that combines image-based matching with conventional IR techniques without using OCR. While image-based matching can be effective in finding similar words, complementing it with efficient retrieval techniques allows for sub-word matching, term weighting, and document ranking. The basic idea is that similar connected elements in printed documents are clustered and represented with ID’s, which are then used to generate equivalent textual representations. The resultant representations are indexed using an IR engine and searched using the equivalent ID’s of the connected elements in queries. Though, the main benefit of the proposed approach lies in languages for which no OCR exists, the technique was tested on English and Arabic to ascertain the relative effectiveness of the approach. The approach achieves more than 61% relative effectiveness compared to using OCR for both languages. While the reported numbers are lower than that of OCR-based approaches, the proposed method is fully automated, does not require any supervised training, and allows documents to be searchable within a few hours.
Original languageEnglish
Title of host publicationString Processing and Information Retrieval
Subtitle of host publication16th International Symposium, SPIRE 2009, Saariselkä, Finland, August 25-27, 2009, Proceedings
PublisherSpringer
Pages334-343
Number of pages10
ISBN (Electronic)978-3-642-03784-9
ISBN (Print)978-3-642-03783-2
DOIs
Publication statusPublished - 2009

Publication series

NameLecture Notes in Computer Science (LNCS)
PublisherSpringer Berlin Heidelberg
Volume5721
ISSN (Print)0302-9743

Fingerprint

Dive into the research topics of 'Efficient Language-Independent Retrieval of Printed Documents without OCR'. Together they form a unique fingerprint.

Cite this