Redaction, Document Redaction, Redactor

Multilingual Document Redaction: Arabic, Right-to-Left and Mixed-Script PDFs

What does a redaction tool do with a witness statement written in Arabic, or a benefits letter in Urdu, when the models it runs were built and tested on English? Usually it does nothing visible, because personal information written in a script the tool can't read is personal information it can't find. The names pass through untouched, and no error reports that anything was missed.

Multilingual document redaction depends on three separate steps, and each can fail on its own. The page has to be read by a recognizer built for its script, right-to-left text has to be put into the order it was written, and detection has to run in the right language. Our guide to how to properly redact a PDF covers what removal must achieve once the words are found, and the sections below cover finding them.

Records in other languages reach many US offices as a matter of routine. The Census Bureau reports that the number of people who speak a language other than English at home "nearly tripled from 23.1 million (about 1 in 10) in 1980 to 67.8 million (almost 1 in 5) in 2019." Those residents write letters to agencies, fill in their forms and appear in their case files.

Scripts and languages are two different lists

Unicode's annex on the script property observes that scripts such as Latin and Arabic "are used for the representation of the writing systems of hundreds or even thousands of different languages." The Unicode Standard adds that the Arabic script "has been extended to represent a number of other languages, such as Persian, Urdu, Pashto, Sindhi, and Uyghur, as well as many African languages." Going the other way, the same annex notes that "The best known example is the Japanese language, whose writing system uses four scripts."

Reading and detection each depend on a different one of those two lists. Reading is a question of script, since the recognizer has to know the shapes of the letters. Detection is a question of language, since names, address formats and identity numbers follow the conventions of the language and country they come from. A tool can therefore read a script well and still detect poorly in one of the languages written in it, and a vendor's headline language count may describe either step. Asking which list applies to reading, which to detection and which to transcription settles the question quickly.

Reading each script with a reader built for it

Whatever an OCR engine fails to read never reaches detection, so a reader trained on Latin letters, handed an Arabic, Devanagari or Chinese page, produces no error and no redactions on those lines. Arabic is especially hard for a reader built around separate letters, and the W3C's Arabic layout requirements note that "Arabic script is a cursive writing system; i.e, letters can join to their neighboring letters." Multi-script readers tend to do better on English as well. A 2022 OCR benchmark in the Journal of Computational Social Science found that "Accuracy for English was considerably higher than for Arabic."

We ran into that gap ourselves when choosing how VIDIZMO Redactor should read pages written in Arabic script. A general multi-script recognizer, the simplest option to run, underperformed on real Arabic-script content, so we chose a recognizer built for connected, right-to-left writing. A reader for each script should be chosen by its measured quality on real material, and that is also the standard to hold any tool to.

Routing pages to the right reader can happen at two levels, and each suits a different kind of collection. Sending a whole document to the reader for its language works well for a batch of Arabic letters and badly for a page that mixes Arabic headers with English tables. Classifying each block of text on the page by script, and sending each block to a matching reader, lets a bilingual form be read region by region. That costs more work on every page and needs a model that recognizes scripts as well as letters. Mostly monolingual collections may not need it, while a mixed collection of immigration files or cross-border contracts usually does. Ask any vendor what happens to a script outside its routed list, such as Hebrew, Thai or Amharic, since a reader that doesn't know a script rarely returns anything useful from it.

Right-to-left text and the order it is stored in

A recognizer that decodes a line of pixels from left to right, the way it scans the image, produces an Arabic string in reverse. Stored that way, a name or an identity number never matches what a detector or a keyword search is looking for. The W3C's note on visual and logical ordering describes the same failure in browsers, where search boxes capture text in logical order, "which causes the search key not to match the text stored in visual order." It is also a common reason Arabic copied out of a PDF comes out scrambled.

Logical order is the one software expects, and Unicode's bidirectional algorithm states that "The Unicode Standard prescribes a memory representation order known as logical order." It also notes that "There are several scripts (such as Arabic or Hebrew) where the natural ordering of horizontal text in display is from right to left." Numbers complicate the picture, because the W3C layout requirements point out that "Numbers, even Arabic numbers, are written from left to right, as is text in a script that is normally left-to-right." Redactor stores every line of right-to-left text in logical reading order before anything searches it, whichever reader produced the line.

Reordering a line also moves its words away from the pixels they came from, and that makes right-to-left scans the place to check each redaction against the image itself. A box that sits one word off leaks the name it was meant to cover, while a box stretched across the whole line withholds more than the release requires.

Names and numbers that change form

Personal information in multilingual records rarely appears in only one form, and a search built for one form misses the others. A name written in Arabic in a letter may appear in a Latin transliteration on the envelope, the case sheet and the index. Transliteration follows no single standard, so the same person can be Mohammed, Muhammad and Mohamed within one file. A keyword list for a release needs every spelling the file actually uses, gathered from the documents themselves rather than guessed.

Numbers change form as well, in a way that pattern matching written for European digits never sees. The Unicode Standard's chapter on these scripts explains that the decimal system "was subsequently adopted in the Arabic world with a different appearance," and it encodes Arabic-Indic digits as a separate sequence. It adds another sequence "for Persian, Sindhi, and Urdu to account for the differences in appearance and directional treatment when rendering them." An identity number or a date written in those digits is a different string of characters from the same number in European digits, so a pattern written for 0 to 9 passes straight over it. Search each redacted file for every identity number and date in each digit form its documents use, and check how a tool's patterns treat those digits before relying on them.

Detection runs in one language at a time

On a mostly English document with Spanish names in it, or on a bilingual intake form, one language's detection model ends up reading the other language's names, and some of them slip through. The models that recognize a name, an address or an identity number learn from text in a particular language and follow its conventions for how those things are written. A detection run set to the wrong language is reading with the wrong expectations.

Most of the practical answer lies in how the work is organized before detection runs. Grouping batches by language lets each document be analyzed in the language it is written in, and flagging mixed-language pages for a reviewer covers what detection is least likely to catch. Access requests from speakers of other languages raise the same issue, as our article on GDPR redaction for data subject access requests discusses.

Testing on your own multilingual documents

The useful test of any tool is a small set of your own documents chosen to break things, and a good set includes pages like the four below, drawn from real requests where possible:

  1. A page in each language your office actually receives, in both a born-digital and a scanned version.
  2. A page that mixes two scripts, such as an Arabic letter with an English reference block.
  3. A right-to-left page with numbers, dates and an identity number embedded in the text.
  4. A form with names in a language other than the one the rest of the form uses.

On every page in the set, confirm that each redaction covers the intended words and nothing beside them, and search the redacted file for each removed name in every script it appeared in. Our articles on redacting scanned PDFs and on handwritten, rotated and skewed scans cover the scan problems that compound these, since a crooked Arabic fax is harder than either condition alone.

Which scripts Redactor reads, and how

Which reader handles a document in Redactor depends on its script and on two settings, the job's language and the deployment's reading mode.

Script How Redactor reads documents in it
Latin By default
Arabic, Urdu and Sindhi With the Arabic-script reader, when the job's language is set to one of them
Devanagari, Cyrillic, Chinese and Japanese, Korean When the deployment selects a script-aware mode, which routes each text region to a matching reader

PII detection then runs on the recognized text in one language per job. Its classes cover names, contact details and financial identifiers, along with country-specific identifiers for the United States, the United Kingdom, Spain, Italy, Poland, India, Australia and Singapore. Welsh reads on the default path, which matters to UK public bodies with Welsh language duties.

Supported formats, detection methods and non-Latin script handling are summarized in the document redaction overview.

TopicsRedactionDocument RedactionRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.