Parityhealth-plan operations benchmark

The work a health plan does, scored the same way for every model.

210 tasks of payer operations run against 28 frontier and open-weight models. Same prompts, same output contract, same graders.

210
tasks of payer work
10
families of work
28
models scored
19.9k
graded responses
2026-09-02
prices and models frozen

What capability costs

The dashed line is the Pareto frontier. Nothing below and to the right of it is worth buying for this work. Frontier models and the top four are labelled. Hover any other point for its name, price and interval.

Against
Score
AnthropicGoogleOpenAIxAIMoonshotZ.aiAlibabaTencentDeepSeekone generation back

Off the scale: DeepSeek V3.2 (57.2). More than twenty-five points below the field, so drawn on the rail beneath the axis instead of inside it. Including it in the scale would cost every other comparison on this chart the resolution that makes it readable. The score is printed beside the marker, and the row is in the table below, unaltered.

Leader
99.4
Claude Fable 5.1
95% CI 99.099.7
Cheapest inside five points
$6.44 /1k
Gemini 3.7 Flash at 98.8
8× under the leader
Open weights behind closed
4.6 pts
Claude Fable 5.1 over Kimi K3
the best of each

Equal weight to each of the 10 families

The headline is the unweighted mean of the family scores, not of the 210 items. Families differ in size, so weighting by task would let the largest set the ranking.

score, the interval around it, and how it holds up on the hard half
/
Vendor
Tier
Weights
Generation
Served on
28 of 28 models
ModelClassScore95% interval
7487100
HardBeats
1Claude Fable 5.1frontier 99.499.324
2Gemini 3.7 Flashfast 98.898.421
3GPT-5.5frontier n−198.698.820
4GPT-5.6 Solfrontier 98.398.120
5Grok 4.6frontier 98.297.620
6Grok 4.5frontier n−197.397.118
7Gemini 3.1 Profrontier 97.196.419
8Claude Opus 5frontier 96.696.415
9Gemini 3 Flashfast 96.395.214
10Kimi K3frontier open 94.894.513
11Gemini 2.5 Profrontier n−194.895.013
12GLM-5.2frontier open n−194.394.113
13Claude Opus 4.8frontier n−194.293.910
14Claude Sonnet 5balanced 94.293.511
15Claude Sonnet 4.6balanced n−192.591.07
16Qwen3.7 Maxfrontier open n−190.990.37
17GLM-5.3frontier open 90.692.07
18Qwen3.8 Maxfrontier open 90.489.67
19GPT-5.6 Terrabalanced 89.888.76
20GLM-5.3 Flashfast open 89.689.66
21Hunyuan 4frontier open 89.687.66
22Qwen3.8 Flashfast open 87.386.75
23Kimi K2.6balanced open n−184.282.23
24DeepSeek V4 Profrontier open 83.783.43
25GPT-5.6 Lunafast 81.580.41
26DeepSeek V4 Flashfast open 79.380.11
27Claude Haiku 4.5fast 78.178.31
28DeepSeek V3.2balanced open n−157.254.40
The score column is the unweighted mean of the family scores, one per family. The interval track is zoomed to 74100, because on a 0–100 axis every row here would be the same length. The dot is the point estimate and the line is a 95% cluster bootstrap over tasks. A ◀ means the interval runs off the left of the track. The number in the score column is always the unclipped one. Rows marked ×1 were sampled once, so their point estimate is a single draw. Read those against the interval rather than the number. Click any column to sort, or a model to open its page.

Where the differences live

One number per model hides two things. Which tasks still separate the field, and which model is best at the specific task you are automating. Those are different questions with different answers.

Not every family still tells models apart

Every model's score on every family, one row per family, widest spread at the top. A row bunched against the right edge is a task the field has finished. A row that spans the plot still separates models. The grey bar is the range from worst to best. Points that would cover each other are nudged up or down, never sideways.

AnthropicGoogleOpenAIxAIMoonshotZ.aiAlibabaTencentDeepSeek

The axis runs from 0 to 100. The eight original families sit in a knot at the right edge; the two expert-tier families, the plan-year ledger above all, are where the field still spreads. Hover a point for the model and its exact score.

The ranking is not one ranking

Score by model and family, models ordered worst to best overall. Read across a row for one model's shape, and down a column to pick a model for that queue.

≤80100
The colour scale is clipped at 80 so the band most of the grid sits in stays legible. Cells below 80, most of them in the ledger column, are all the darkest shade. The number in the cell is always the real score.

What survives contact with production

Two questions decide whether a model can run unattended. Does the answer hold when you ask again, and does it know the difference between work it must refuse and work it must do.

What survives being run again

Pale bar: mean score. Solid bar: the share of tasks right on every attempt. The gap between them is the part of a headline number that does not survive a rerun.

Only the 28 models sampled more than once appear here.

Over-refusal is the axis that actually varies

The bar is how often a model refused work a plan must carry out. The diamond is the other failure, doing something a plan is not permitted to do, which all but one model did on exactly zero tasks. Safety reporting that measures only that second direction gives every model here full marks, including the one that refuses 44% of legitimate work.

What more reasoning effort buys

Every other number on this site was produced at the vendor's default reasoning setting, which is what you get when you wire the model up the obvious way. These runs are the same models with that setting changed: 5 models across the settings below, over the suite as each row ran it: the whole suite for OpenAI and Google, the eight original families for Anthropic, whose variants were not rerun on the expert tier. OpenAI and Anthropic take a native effort setting and the harness passes it through. Google takes a thinking-token budget instead, mapped at low = 1,024, medium = 6,000, high = 20,000 tokens. Compare a model against itself down the table, and read the intervals: most of these gaps are not differences. Comparing across vendors compares their definitions of effort, which are not the same thing.

Effort against cost

One line per model, points ordered low effort, vendor default, high effort, plotted against what each setting cost per item. A line that runs flat and to the right is a model whose extra thinking bought nothing but a larger bill.

ModelSettingScore95% CIHard$/1k tasksMedianOutput tokensOf which thinking
GPT-5.6 Sollow97.695.899.197.2$12.232.9s40966%
medium98.396.999.498.1$9.752.9s29863%
high98.897.999.698.6$17.083.1s65579%
GPT-5.6 Terralow91.188.893.490.3$8.252.1s52375%
medium89.887.292.488.7$5.391.9s28964%
high94.692.296.793.5$10.462.4s70581%
Claude Opus 5low98.897.599.798.4$14.122.3s25617%
default98.997.999.698.7$18.233.3s45439%
high98.797.499.698.7$18.983.2s46340%
Claude Sonnet 5low95.492.598.093.4$5.392.7s23836%
default96.994.898.696.1$7.403.7s46263%
high97.094.898.895.2$7.213.9s44962%
Gemini 3.1 Prolow97.195.298.996.7$29.5710.5s2,01891%
default97.195.098.896.4$36.2111.4s2,65395%
high97.595.299.396.9$59.2413.4s4,49096%
Reasoning tiers, one sample per item, each group scored over the families every one of its rows ran. Rows within a model are directly comparable; rows across models are not.

One generation of movement

Each vendor's current model against the one it replaced, matched by price tier. Whether an upgrade is worth the migration is a different question from who leads. On this suite the answer is not uniformly yes.

Current generation against the one before it

One row per pair. The hollow dot is the model that was replaced, the filled one is its successor, and the number beside it is the move in points. Sorted by how much the upgrade bought.

Five findings worth acting on, and one caveat.

01

Single-step rule application is solved.

The top 9 models sit above 95 on the full suite, and on the easier families the leaders are separated by less than their own confidence intervals. That is a capability statement, not a ranking. For a claim adjudicated against one accumulator, a policy applied to one clean record, or a measure applied to one member, the model is no longer the constraint. The retrieval, the plumbing and the review path around it are.

02

Depth still separates the field.

The 144 tasks marked hard at authoring time are long claim chains across a household, coordination of benefits, contradictory plan documents, and near-miss pairs. On those, the spread widens from 42.2 points to 44.9. Read the hard column if the queue you are automating is not the easy half of the work.

03

Price and capability have come apart.

Gemini 3.7 Flash scores within five points of the leader at $6.44 per thousand tasks against the leader's $51.54, a factor of 8. At payer volumes that ratio decides whether a queue gets automated at all. That is why every ranking on this site carries a price.

04

Over-refusal is the larger compliance problem.

Across the field, models did something a plan is not permitted to do on 0.8% of the tasks they should have declined, and refused work a plan must carry out on 8.7% of the tasks they should have completed. The second number is the one that kills a deployment quietly. A refusal reads as caution in a pilot and as a useless tool in production. Reporting only the first number rewards a model that refuses everything.

05

Code-set recall is the weakest link, and retrieval fixes it.

On code tasks with the governing rule supplied in the prompt the field averages 97.0. On tasks where the model must answer from memory it averages 94.5. That gap is the argument for putting a code-set lookup in front of a model before pointing it at coding work, and here it is measured rather than assumed.

06

This is one run of a focused suite.

210 tasks, 3 samples each where the budget allowed and one where it did not. Confidence intervals are reported, and models whose intervals overlap should be read as close rather than ranked. Tasks were written and reviewed internally; external clinical and compliance review is planned. Synthetic material is cleaner than a real intake queue, so read these scores as an optimistic estimate. The methodology pagelists the known limitations.

10 families, 210 tasks

Ordered by how much the family separates the field. The ones at the top are where the choice of model still matters.

Read this before quoting a number

Tasks were run up to 3 times per model. The leaderboard averages a model's attempts. “Right every try” under Operations is the stricter number, and it carries the attempt count beside it because models did not all finish the same number of passes. Rows marked ×1 were run once and sit slightly higher than they would after resampling. Read those against the interval, not the point. 880 transport failures are excluded from scoring rather than counted as zeros.

Full methodology, including what this benchmark does not measure, is on the methodology page. Every response from the 28 ranked models, with its own reasoning where the vendor exposes it, is in transcripts. The 19,923 figure above covers the whole run, including the reasoning-tier variants, which are scored but not ranked.