Research/OCR

From Publication Station
Revision as of 13:56, 4 December 2015 by Andre (talk | contribs)

OCR (optical character recognition)

tesseract

https://code.google.com/p/tesseract-ocr/

Teassearct is OCR software. It was initially developed by HP Labs between 1985 and 1995 currently its development is sponsored by Google.

It is free software, released under the Apache License.

install

Debian:

aptitude install tesseract-ocr

Mac:

using homebrew need to run the commands:

brew install leptonica --with-libtiff
brew install tesseract --all-languages

https://gist.github.com/henrik/1967035


Run

prerequisites

source files should:

  • be in .tiff format
  • have at least 300dpi - otherwise the text recognition will be very sloppy
  • contain only one column text

command

tesseract input.tiff output

will result in OCRed file output.txt

Languages

By default tesseract is optimized to work with English language. This behavior can be change by installing extra packages required for other languages and by giving it the correct setting.