Computer Science Faculty Publications and Presentations

Extending Page Segmentation Algorithms for Mixed-Layout Document Processing

Amy Winder, Boise State University
Tim Andersen, Boise State UniversityFollow
Elisa Barney Smith, Boise State UniversityFollow

Document Type

Conference Proceeding

Publication Date

9-18-2011

DOI

https://doi.org/10.1109/ICDAR.2011.251

Abstract

The goal of this work is to add the capability to segment documents containing text, graphics, and pictures in the open source OCR engine OCRopus. To achieve this goal, OCRopus' RAST algorithm was improved to recognize non-text regions so that mixed content documents could be analyzed in addition to text-only documents. Also, a method for classifying text and non-text regions was developed and implemented for the Voronoi algorithm enabling users to perform OCR on documents processed by this method. Finally, both algorithms were modified to perform at a range of resolutions. Our testing showed an improvement of 15-40% for the RAST algorithm, giving it an average segmentation accuracy of about 80%. The Voronoi algorithm averaged around 70% accuracy on our test data. Depending on the particular layout and idiosyncracies of the documents to be digitized, however, either algorithm could be sufficiently accurate to be utilized.

Copyright Statement

© 2011 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. DOI: 10.1109/ICDAR.2011.251

Publication Information

Winder, Amy; Andersen, Tim; and Barney Smith, Elisa. (2011). "Extending Page Segmentation Algorithms for Mixed-Layout Document Processing". 11th International Conference on Document Analysis and Recognition (ICDAR 2011), 1245 - 1249.

Link to Full Text

COinS

ScholarWorks

Computer Science Faculty Publications and Presentations

Extending Page Segmentation Algorithms for Mixed-Layout Document Processing

Document Type

Publication Date

DOI

Abstract

Copyright Statement

Publication Information

Browse

Links

Search

Author Corner

ScholarWorks

Computer Science Faculty Publications and Presentations

Extending Page Segmentation Algorithms for Mixed-Layout Document Processing

Authors

Document Type

Publication Date

DOI

Abstract

Copyright Statement

Publication Information

Share

Browse

Links

Search

Author Corner