Share Email Print
cover

Proceedings Paper

Font group identification using reconstructed fonts
Format Member Price Non-Member Price
PDF $14.40 $18.00
cover GOOD NEWS! Your organization subscribes to the SPIE Digital Library. You may be able to download this paper for free. Check Access

Paper Abstract

Ideally, digital versions of scanned documents should be represented in a format that is searchable, compressed, highly readable, and faithful to the original. These goals can theoretically be achieved through OCR and font recognition, re-typesetting the document text with original fonts. However, OCR and font recognition remain hard problems, and many historical documents use fonts that are not available in digital forms. It is desirable to be able to reconstruct fonts with vector glyphs that approximate the shapes of the letters that form a font. In this work, we address the grouping of tokens in a token-compressed document into candidate fonts. This permits us to incorporate font information into token-compressed images even when the original fonts are unknown or unavailable in digital format. This paper extends previous work in font reconstruction by proposing and evaluating an algorithm to assign a font to every character within a document. This is a necessary step to represent a scanned document image with a reconstructed font. Through our evaluation method, we have measured a 98.4% accuracy for the assignment of letters to candidate fonts in multi-font documents.

Paper Details

Date Published: 24 January 2011
PDF: 8 pages
Proc. SPIE 7874, Document Recognition and Retrieval XVIII, 78740N (24 January 2011); doi: 10.1117/12.873398
Show Author Affiliations
Michael P. Cutter, Univ. of Kaiserslautern (Germany)
Joost van Beusekom, Univ. of Kaiserslautern (Germany)
German Research Ctr. for Artificial Intelligence (Germany)
Faisal Shafait, Univ. of Kaiserslautern (Germany)
German Research Ctr. for Artificial Intelligence (Germany)
Thomas M. Breuel, Univ. of Kaiserslautern (Germany)


Published in SPIE Proceedings Vol. 7874:
Document Recognition and Retrieval XVIII
Gady Agam; Christian Viard-Gaudin, Editor(s)

© SPIE. Terms of Use
Back to Top