Home / OCR and text
OCR Redaction and Text Recognition
A scanned page has no text layer. A photograph of a form is an image. Text on a screen inside a video is neither. Redactor recognizes all of it, routes each region to the engine that reads its script, and makes the result both redactable and searchable.
How OCR Redaction Works
Text appearing in video frames, images and scanned documents is recognized and indexed. Each text region is classified by script and routed to the recognition engine that handles it, so a document mixing scripts is read rather than partly missed.
Four OCR engines ship, and the engine is selected per content type, because the engine that suits a video frame is not the one that suits a document page.
What OCR Redaction Covers
Scripts routed
| Script class | Covers |
|---|---|
| Latin | Including Welsh, which matters for UK public bodies under Welsh language duties |
| Perso-Arabic | Arabic, Urdu, Sindhi, Dari, Pashto |
| Devanagari | |
| Cyrillic | |
| Chinese and Japanese | |
| Korean |
Routing is per region, not per document, so a page mixing scripts is handled region by region.
Objects inside documents, not only text
Document pages are rasterised so the same detection that runs on images runs on the page content. A face, license plate or vehicle embedded as an image inside a PDF is detected and obscured the same way it would be in a standalone image.
This is the gap most document redaction tools have. A text-only tool redacts character runs; it cannot see a photograph on page four.
Tables and spreadsheets
Spreadsheets and tables in documents are redacted by column and by row, so a whole field of sensitive values is removed in one action rather than cell by cell.
Searchable afterwards
Text that appears on screen or on a scanned page is searchable across the library. A result opens at the exact timestamp in a video or the exact page in a document.
Thresholds
OCR confidence thresholds are set separately for video, image and document, because the three have different noise floors.
Where It Matters
- Screen recording redaction — on-screen text in dashboards and application windows
- Medical and clinical documents — scanned forms and handwritten notes
- eDiscovery — production sets of mixed scanned and native files
- Legal
What OCR Redaction Does Not Do
- Script routing governs text recognition. PII language coverage is a separate matter and a narrower set. Never read a script list as a PII-detection language list.
- Detection on embedded objects operates on the rendered page, so an object must be visible in the page as rendered.
- OCR quality follows scan quality.
How OCR Redaction Is Evaluated
- Text regions are classified by script and routed to the engine for that script, so non-English documents redact.
- Routing is per region, so a document mixing scripts is handled region by region.
- Perso-Arabic coverage spans Arabic, Urdu, Sindhi, Dari and Pashto.
- Four OCR engines ship, selected per content type by deployment configuration.
- Detection runs on the rendered page, so faces, plates and vehicles inside embedded images are covered, which extends beyond the character runs a text-only tool redacts.
- Spreadsheets redact by column or by row.
- OCR text search finds a word that appears on screen rather than spoken, landing on the timestamp or page.
Send Us Your Worst Scan
A skewed photocopy, a handwritten form, a page in mixed scripts. That is the useful test, not a clean PDF.