Orus Studio
Insurance

Using an LLM where it helps and rules where it must

An insurance intermediary processing about 3,000 claims a month

A claims operation drowning in unstructured documents, where a fully automated extraction would have been both unreliable and unacceptable to auditors.

Client
An insurance intermediary processing about 3,000 claims a month
Industry
Insurance
Duration
13 weeks
Team
3 engineers
Year
2025

The situation

Claims arrived as scans, phone photographs and PDFs of wildly varying quality, and were keyed in manually by a team of eleven.

Manual keying produced errors that surfaced weeks later during settlement, when they were expensive to correct.

Regulatory obligations meant an automated decision had to be explainable and attributable, ruling out an opaque end-to-end model.

An earlier off-the-shelf OCR attempt had failed on handwritten and low-quality documents, which were a large share of the volume.

What we did

01

Extraction assisted, decisions ruled

The language model extracts fields and cites the region of the document each value came from. It does not decide anything. Eligibility and settlement run on explicit, versioned rules. This split is what made the system acceptable to the compliance function, and it is the right split regardless — a model that decides is a model you cannot explain.

02

Confidence routing rather than full automation

Each extracted field carries a confidence score. Above a threshold it passes through; below it, the document is routed to a human with the uncertain field highlighted and the source region shown. Roughly 70% of documents pass without intervention, and the team's work shifted from keying everything to checking exceptions.

03

Citations as the review interface

Because every extracted value points at the pixels it came from, a reviewer verifies by glancing rather than re-reading the document. This detail did more for throughput than any accuracy improvement in the model.

04

A held-out evaluation set from real documents

We built an evaluation set of 400 manually labelled real documents, weighted toward the poor-quality cases that had defeated the previous attempt. Every prompt and model change was measured against it, so improvements were demonstrated rather than assumed.

The parts that were actually hard

Problem

Handwritten claim forms in regional languages defeated both the model and the earlier OCR.

How we handled it

We did not solve this. Those documents route directly to human entry, and the system tracks what proportion they represent. Claiming to handle them would have produced silent errors in exactly the cases least likely to be checked.

Problem

Model output occasionally drifted in format, breaking downstream parsing.

How we handled it

We constrained output to a strict schema with validation and a bounded retry, and treated a schema violation as low confidence rather than attempting to repair it. Repairing malformed output silently is how wrong values reach production.

What shipped

  • Document ingestion supporting scans, photographs and PDFs
  • LLM extraction with per-field confidence and source citations
  • Confidence-based routing between automatic processing and human review
  • Versioned, explainable rules engine for eligibility decisions
  • Held-out evaluation set with per-release accuracy reporting
  • Full audit trail linking every decision to its inputs and rule version

Outcomes

  • Around 70% of documents process without human intervention, with the remainder routed for targeted review rather than full keying
  • The claims team moved from data entry to exception handling without reduction in headcount
  • Compliance accepted the system on the basis that no decision is made by the model
  • Errors found at settlement fell substantially, since extraction is verified at intake rather than weeks later

Outcomes are described qualitatively where no clean measured baseline existed before the work started.

Built with

Next.jsPythonPostgreSQLClaude APIAWS

Have a problem shaped like this one?

Tell us what you are dealing with. We will come back with a scoped estimate and an honest view on whether we are the right fit.

Explore