BFSI · Insurance · AI automation
Human in the Loop
Designing human oversight into AI automation for insurance workflows
- 60%Drop in error rate

The problem
Insurance companies process huge volumes of claims documents every day. AI extraction takes most of the manual effort out of that work, but it can’t be 100% accurate, especially on unstructured documents. In insurance, a single wrong value can mean a financial loss or a compliance issue.
Three failure modes kept showing up:
- Inaccurate predictions that needed a person to verify them before anyone could act.
- Missing mandatory fields that the AI couldn’t find, which blocked the case.
- Slow, fragmented handling of these exceptions, which delayed client requests.
The real question wasn’t “how do we make the AI better.” It was how to design the handoff between AI and human so automation stays fast and every approved case can be trusted.
Who it’s for
Insurance professionals reviewing AI-processed cases:
- Claims analysts, who verify claim details and approve payouts
- Underwriters, who check policy data before a decision
- Risk assessors, who look for gaps and inconsistencies
All three share the same job: they aren’t doing data entry anymore. They’re checking the AI’s work, and they need to do it quickly without missing anything.
My role
I was the product designer for the HITL system, working with product analysts, delivery managers, and the backend and NLU teams to assess feasibility and define the use cases. I worked with the frontend team to keep the UI consistent with the rest of the platform, and kept stakeholders updated on progress and blockers.
The approach
Rather than replacing the AI or bolting on a generic review screen, we added a Human in the Loop layer directly into the automation pipeline. Every case the AI processes flows into a review queue where a person can:
- Confirm the AI’s predictions
- Check for missing mandatory fields
- Add or correct data against the source documents
- Approve or reject the case

Key decisions
Two paths: bulk approval and individual review
Not every case deserves the same scrutiny. High-confidence cases with complete data can be selected and approved in bulk from the list view. Anything flagged opens into a detailed review. This lets teams move fast on the easy majority and spend their attention where the AI is least sure.
Flag missing fields before the case is opened
The list view shows a signifier on cases with missing fields, along with filters for case type, process status, and approval status. Reviewers can triage the queue at a glance instead of opening every case to find out what’s wrong.

Fix the root cause: editable case type
When the AI misclassifies a case, for example tagging a pet insurance claim as travel insurance, every extracted field is wrong because the schema is wrong. So reviewers can change the case type itself, and the attribute list updates to match. One correction at the root replaces a dozen field-by-field fixes.
“Restart Process” instead of manual rework
If the predictions are too far off to fix by hand, the reviewer can send the case back to the AI to re-run extraction, for example after correcting the case type. This keeps the human in a reviewer role rather than turning them back into a data-entry operator.
Verify against the source, side by side
The detail view puts the extracted fields next to a document preview with page navigation, zoom, and a browser view option. Required fields are clearly marked, and empty ones show inline errors. Reviewers check each value against the original document without switching windows.

Required-field validation at approval
Approval checks whether all mandatory attributes are filled before a case can be completed.
Impact
- 60% drop in error rate, improving data accuracy and compliance.
- Greater trust in the AI, because human oversight stayed built into every decision.
- Smooth adoption by insurance teams, thanks to an interface that matched how they already worked.
What I learned
- Transparency builds trust. People were more comfortable with the AI when they could see exactly what it had predicted and where it was unsure.
- Design the collaboration, not just the screen. The value came from deciding what the AI does, what the human does, and how cleanly work passes between them.
- Iterate with real scenarios. Frequent feedback loops with the team exposed edge cases, like misclassified case types, that shaped the biggest decisions.