Home / OCR and text

OCR Redaction and Text Recognition

A scanned page has no text layer. A photograph of a form is an image. Text on a screen inside a video is neither. Redactor recognizes all of it, routes each region to the engine that reads its script, and makes the result both redactable and searchable.

TEXT LIFTED OFF THE PIXELSScanned and photographed pagesrotation correctedOn-screen text in videoframe by frameObjects embedded in PDFsimages inside documentsHandwritingICRNon-Latin scriptsPerso-Arabic and othersOCR thresholds set separately for video, image and document.

How OCR Redaction Works

Text appearing in video frames, images and scanned documents is recognized and indexed. Each text region is classified by script and routed to the recognition engine that handles it, so a document mixing scripts is read rather than partly missed.

Four OCR engines ship, and the engine is selected per content type, because the engine that suits a video frame is not the one that suits a document page.

What OCR Redaction Covers

Objects inside documents

Document pages are rasterised so the same detection that runs on images runs on the page content. A face, license plate or vehicle embedded as an image inside a PDF is detected and obscured the same way it would be in a standalone image. This is the gap most document redaction tools have. A text-only tool redacts character runs; it cannot see a photograph on page four.

Tables and spreadsheets

Spreadsheets and tables in documents are redacted by column and by row, so a whole field of sensitive values is removed in one action rather than cell by cell.

Searchable afterwards

Text that appears on screen or on a scanned page is searchable across the library. A result opens at the exact timestamp in a video or the exact page in a document.

Thresholds

OCR confidence thresholds are set separately for video, image and document, because the three have different noise floors.

Scripts routed

Latin

Including Welsh, which matters for UK public bodies under Welsh language duties

Perso-Arabic

Arabic, Urdu, Sindhi, Dari, Pashto

Devanagari

Cyrillic

Chinese and Japanese

Korean

Routing is per region, not per document, so a page mixing scripts is handled region by region.

FAQ

OCR Redaction questions, answered

Can text be redacted from scanned documents?

Yes. Scanned and photographed pages are read through OCR, including pages that arrive rotated or skewed, so a document that was never digitally native is redactable rather than set aside.

Is on-screen text in video detected?

Yes. Text visible in a video frame, such as a dashboard, a chat window or an application interface, is read and can be redacted, which matters for screen recordings and for body-worn footage that captures a terminal.

What about text inside images embedded in a PDF?

Objects inside a document are redacted, not only its text. A face, licence plate or vehicle appearing in an image embedded in a PDF is detected and obscured the same way it would be in a standalone image. A tool that reads only the text layer leaves that visible.

Are non-Latin scripts supported?

Each text region is classified by script and routed to the engine that handles it, covering Latin, Perso-Arabic, Devanagari, Cyrillic, Chinese and Japanese, and Korean. A document mixing scripts is read rather than partly missed.

Is handwriting recognised?

Handwritten content is handled through intelligent character recognition, which is what makes handwritten case notes and annotated forms tractable rather than manual.

Send Us Your Worst Scan

A skewed photocopy, a handwritten form, a page in mixed scripts. That is the useful test, not a clean PDF.