Home / Claims adjusting firm
Claim Files, Redacted Before Anyone Else Sees Them
A claims adjusting firm works on behalf of insurance companies. Every file it touches belongs to somebody else's policyholder, and almost every file has to be shared with somebody.
At a Glance
100 CALs
A central redaction team, one licence per user
3 regions
United States, United Kingdom and European Union
Hundreds a day
On an unpredictable daily load
AI plus rules
Inference and regex in one pass
The Organisation
This customer is a claims adjusting firm. Its licensed adjusters investigate and settle property and casualty claims on behalf of insurance companies, doing the work an insurer's own claims team would otherwise do.
It operates across three regulatory territories: the United States, the United Kingdom and the European Union, handling claims wherever the insurers it works for are writing policies. The book is general property and casualty, spanning commercial auto, business owners' policies, general liability, professional lines and medical malpractice.
That three-region footprint matters more than it might appear. It means one firm is answerable under US state privacy law, under UK GDPR and under EU GDPR at the same time, on files that look identical to an adjuster and are governed differently depending on whose policyholder the claimant is.
The relationship is the whole reason redaction matters here. The firm is not handling its own customers' data. It is handling data belonging to the policyholders of the insurers it works for, and it answers to each of those insurers for how that information is treated before it moves anywhere.
And claims work moves constantly. A file goes to the carrier, to a broker, to an expert, to counsel, sometimes to the claimant. Each of those audiences is entitled to a different amount of what the file contains.
Who Actually Does the Redaction
Redaction here is not something each adjuster does on their own files between other tasks. The firm runs it as a central team, holding a hundred Client Access Licenses, one per user.
That is a deliberate operating choice and it shapes what the software has to be good at. A central team builds judgement that an occasional user never does. They see every carrier's requirements, learn which documents recur, and develop a feel for what a given insurer will object to. They also become the single point where a mistake would happen, which is a much easier thing to govern than redaction happening in a hundred places at once.
It also means the people using the product are in it all day, every day. Software that is tolerable for someone who touches it twice a month is not good enough for a team whose entire job it is.
Why They Work in the Interface
Plenty of high-volume redaction runs headless. A pipeline takes files from one place, applies a policy and writes the output somewhere else, and nobody opens anything.
That is not this deployment. The firm's team works in the interface, because manual redaction is a standing part of the job rather than an exception.
A claim file regularly contains something no rule and no model should be asked to decide on its own. A handwritten note in a margin. A carrier that wants one category of information kept that another wants removed. A photograph where what identifies someone is not a face or a plate but the building behind them. Those are judgement calls, and the team makes them by looking at the document and drawing on it.
So the studio is the workplace, not the fallback. Automated detection does the volume and gets the file most of the way, and a person finishes it. The measure that matters to this firm is not how much the automation caught without help, but how quickly a trained operator can take a file from arrival to defensible.
What Arrives in a Claim File
Claims work is not tidy. A single file can carry a hospital bill, a police report, a bank statement, a set of damage photographs, dashcam footage and a recorded call with a claimant, all relating to the same incident.
Documents are the bulk of it, and the firm's volume is a few hundred documents a day. But a document-only tool would have covered part of the problem and left the rest, because the photographs and the recordings carry as much identifying information as the paperwork does.
Why the Documents Are the Hard Part
A claims document is rarely a clean digital PDF.
They arrive in several languages
With claims running across the United States, the United Kingdom and the European Union, a file can contain paperwork in more than one language, and sometimes more than one within the same document. A European claim can carry a local-language police report, an English-language policy schedule and a medical note in a third language, all in one submission.
They arrive at whatever angle the camera was held at
Claimants and adjusters photograph paperwork rather than scan it. Pages come in rotated, skewed, shot on a desk under bad light, or photographed at an angle that makes the text trapezoidal rather than rectangular.
They arrive in whatever format the sender had
PDFs, office documents, images of documents, scans, and email attachments carrying any of those.
They contain images as well as text
This is the part that catches most tools out. A claim document is frequently a wrapper around photographs: a damage report with pictures embedded in it, a proof-of-loss form with a photograph of an identity document pasted in, an estimate with site images. A tool that reads the text of a PDF and stops there will produce a file that looks redacted and still shows a face, a licence plate or a house number in an image halfway down page four.
What Has to Come Out
The identifiers in a claim file span two quite different kinds.
In the text there are names, addresses, dates of birth, account numbers, credit card numbers and policy identifiers. These behave like PII anywhere else, and a detector finds them by reading.
In the images there are street signs, house numbers, licence plates, badges, identity documents and faces. None of these can be found by reading, because none of them are text in the document. They are objects inside pictures inside pages.
Both sets had to go, in one pass, from the same file.
Detection Alone Was Not Enough
The firm had its own list of things that always had to be removed. Claim reference formats, policy number patterns, internal identifiers, the specific fields a given carrier insists are never shared.
Those are not a matter of judgement, and they are not something a general model should be asked to infer. They are rules.
So the deployment runs both. AI detection handles the open-ended work: finding the names nobody listed in advance, the faces in a photograph, the address on a sign behind a damaged vehicle. Rule-based detection handles the known quantities: regular expressions, context words and vocabulary lists that catch an identifier because it matches a defined shape rather than because a model recognised it.
Neither approach covers the other's ground. A regular expression will never find a face. A model should not be relied on to remember that this particular carrier's claim references always start with two letters and a hyphen.
Letting Outsiders Send Files Safely
Claims files do not all originate inside the firm. Claimants, repairers, medical providers and experts all send material in.
Rather than take that by email, the firm issues a secure upload link. The external party submits through it without needing an account, the submission arrives for triage rather than landing in an inbox, and what they send goes through malware scanning before anything else happens to it. Only then does it enter the detection and redaction pipeline.
That ordering matters more than it sounds. An upload route open to the public is an obvious way into an organisation, and a file that has not been scanned should never reach a processing pipeline, let alone a reviewer's screen.
A Load That Will Not Sit Still
The firm's daily volume is unpredictable by nature. Claims do not arrive evenly. A quiet day might bring a handful of documents. A day after a storm brings several hundred.
This is a large part of why the firm wanted this as a service rather than something it ran. It did not want to size infrastructure for a peak it could not forecast, or to operate capacity that sits idle between events. It wanted to send work and get work back.
So the deployment is shared SaaS. No servers were stood up, nothing was sized, and the load being ten files or four hundred is not the firm's problem to absorb.
Where a Person Still Decides
Automation handles the volume. It does not get the last word.
The team opens a redacted file, sees what was found, and corrects it: adding a region the detector missed, removing one it should not have obscured, adjusting what was drawn, or applying a decision that is specific to the carrier the file belongs to.
For a firm answerable to several different insurers about the same kind of file, the ability to make a specific correction on a specific document matters more than a high average. An average is a statement about a population. A claim file is a single document going to a named recipient, and it is either right or it is not.
Sharing, Once It Is Safe
Once a file is redacted it goes where it needs to go, to people inside the organisation and to people outside it, with access controlled rather than assumed. That is the point of the whole exercise. The firm is in the business of moving claim files between parties, and redaction is what makes each of those movements defensible.
Why It Worked
The team was equipped rather than replaced
The firm did not buy automation to remove people from the process. It put a hundred trained operators in a studio built for the work and gave them detection that does the mechanical part. The judgement stayed where judgement belongs.
One system covered all four formats
Documents are the bulk of the volume, but photographs, footage and recorded calls carry the same identifiers. Splitting that across tools would have meant separate workflows and separate audit trails for parts of one claim.
Objects inside documents were treated as objects
Reading a PDF's text and stopping there is the single most common way a redacted claim file still leaks, and it is the failure the embedded images were always going to cause.
Inference and rules ran together
The firm's known identifiers are rules and were treated as rules. Everything unpredictable was left to detection.
Nothing had to be built
For a workload with no stable shape, a service that absorbs the peak was worth more than infrastructure the firm would have had to size, run and justify.
Why This Customer Is Not Named
We do not have a marketing agreement covering the use of their name, so the deployment is described and the customer is not identified. Where a customer is not named here it is usually because they asked not to be, or because we have not asked, rather than because the deployment is small.
Ask for a Reference
For a serious evaluation we will arrange a call with an operation doing something close to what you are doing.