Dokumentationerreichbar
Tesseract, die klassische Erkennungsmaschine
Die älteste noch benutzte Erkennungsmaschine und der Beleg dafür, wie lange dieses Fach schon läuft. Sie entstand zwischen 1985 und 1994 in den Hewlett-Packard-Laboren in Bristol und in Greeley, Colorado; 2005 gab HP sie als offenen Quelltext frei, von 2006 bis August 2017 lag die Entwicklung bei Google. Das Paket enthält eine Bibliothek und ein Kommandozeilenprogramm und schreibt neben reinem Text auch hOCR, PDF, TSV, ALTO und PAGE.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 07.09.2026:
Tesseract was originally developed at Hewlett-Packard Laboratories Bristol UK and at Hewlett-Packard Co, Greeley Colorado USA between 1985 and 1994
bestätigt 24.09.2026In 2005 Tesseract was open sourced by HP. From 2006 until August 2017 it was developed by Google.
bestätigt 24.09.2026This package contains an OCR engine - libtesseract and a command line program - tesseract.
bestätigt 24.09.2026Tesseract supports various output formats: plain text, hOCR (HTML), PDF, invisible-text-only PDF, TSV, ALTO and PAGE.
bestätigt 24.09.2026