InsureBench
InsureBench is a benchmark that measures how language models perform on insurance work. It spans two core workflows — underwriting and claims & coverage — built from document-grounded cases that each resolve to a single verifiable answer. Models are evaluated pass@1 and scored against the recorded outcome, not the wording of the response.
An AI benchmark for insurance
InsureBench is an insurance AI benchmark. It measures how language models handle the document-grounded work the industry runs on: reading policies and supporting files, applying the terms, and producing a decision or a number that can be checked against a recorded outcome — drawn from real insurance work, not synthetic exam questions.
General-purpose benchmarks reward fluent prose. Insurance work rewards something else: tracing a clause through a long policy, reconciling conflicting documents, and arriving at a defensible decision an auditor could check. InsureBench is built around that gap.
- Scored pass@1 — one attempt, no retries
- Document-grounded — real policies, applications, claim files
- Verifiable outcomes — a decision, determination, or number
- Two core workflows — underwriting and claims
- A GDPval-style benchmark by Huzzle Labs
The two core workflows
Underwriting →
Models assess risk from application materials, decide whether to offer cover, and set terms such as limits, exclusions, and pricing inputs.
Claims & coverage →
Models read the policy and claim file, determine whether a loss is covered, find the controlling clauses, and calculate the amount payable.
How models are scored
Every case resolves to a single verifiable answer: a decision, a covered-or-not determination, or a number. Models run pass@1, one attempt per case with no retries. Scores reflect the recorded outcome, not the style or fluency of the response — the full rules are in the methodology.
A GDPval-style benchmark
InsureBench follows GDPval: evaluate models on real, economically valuable work instead of abstract puzzles. Where GDPval spans many occupations, InsureBench is a GDPval for insurance — built around the specific tasks underwriters and claims handlers carry out.
Explore
Frequently asked questions
What is InsureBench?
InsureBench is an AI benchmark for insurance. It measures how language models handle document-grounded insurance work: reading policies and supporting files, applying the terms, and producing a decision or a number that can be checked against a recorded outcome.
What does InsureBench measure?
It spans two core workflows — underwriting and claims & coverage. Each case is built from real insurance work rather than synthetic exam questions.
How are models scored on InsureBench?
Every case resolves to a single verifiable answer: a decision, a covered or not-covered determination, or a number. Models run pass@1, one attempt per case, and are scored against the recorded outcome rather than the style of the response. See the methodology for the full grading rules.
Is InsureBench like GDPval?
Yes. InsureBench follows the GDPval approach of evaluating models on real, economically valuable work instead of abstract puzzles. Where GDPval spans many occupations, InsureBench is a GDPval for insurance, focused on the tasks underwriters and claims handlers carry out.
Why does insurance need its own AI benchmark?
General reasoning and exam-style benchmarks don't capture insurance work, which turns on reading long, messy policy documents, applying specific clauses, and producing auditable decisions and numbers. InsureBench measures models on that real work instead of abstract questions.
Who builds InsureBench, and when does the leaderboard open?
InsureBench is built by Huzzle Labs, working with practising underwriters and claims handlers. The benchmark launches soon, when the public leaderboard of pass@1 scores for frontier models opens.
Backed by 10X Founders·Angel Invest·Emerge·a16z Scout Fund·Thomas Wolf Hugging Face·Bernd Heinemann Allianz·Yaser Khalighi Stanford