The work a health plan does, scored the same way for every model.
210 tasks of payer operations run against 28 frontier and open-weight models. Same prompts, same output contract, same graders.
- 1Claude Fable 5.199.4
- 2Gemini 3.7 Flash98.8
- 3GPT-5.598.6
- 4GPT-5.6 Sol98.3
- 5Grok 4.698.2
- 6Grok 4.597.3
What capability costs
The dashed line is the Pareto frontier. Nothing below and to the right of it is worth buying for this work. Frontier models and the top four are labelled. Hover any other point for its name, price and interval.
Off the scale: DeepSeek V3.2 (57.2). More than twenty-five points below the field, so drawn on the rail beneath the axis instead of inside it. Including it in the scale would cost every other comparison on this chart the resolution that makes it readable. The score is printed beside the marker, and the row is in the table below, unaltered.
95% CI 99.0–99.7
8× under the leader
the best of each
Equal weight to each of the 10 families
The headline is the unweighted mean of the family scores, not of the 210 items. Families differ in size, so weighting by task would let the largest set the ranking.
| Model | Class | Score↓ | 95% interval 7487100 | Hard | Beats | |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | frontier | 99.4 | 99.3 | 24 | |
| 2 | Gemini 3.7 Flash | fast | 98.8 | 98.4 | 21 | |
| 3 | GPT-5.5 | frontier n−1 | 98.6 | 98.8 | 20 | |
| 4 | GPT-5.6 Sol | frontier | 98.3 | 98.1 | 20 | |
| 5 | Grok 4.6 | frontier | 98.2 | 97.6 | 20 | |
| 6 | Grok 4.5 | frontier n−1 | 97.3 | 97.1 | 18 | |
| 7 | Gemini 3.1 Pro | frontier | 97.1 | 96.4 | 19 | |
| 8 | Claude Opus 5 | frontier | 96.6 | 96.4 | 15 | |
| 9 | Gemini 3 Flash | fast | 96.3 | 95.2 | 14 | |
| 10 | Kimi K3 | frontier open | 94.8 | 94.5 | 13 | |
| 11 | Gemini 2.5 Pro | frontier n−1 | 94.8 | 95.0 | 13 | |
| 12 | GLM-5.2 | frontier open n−1 | 94.3 | 94.1 | 13 | |
| 13 | Claude Opus 4.8 | frontier n−1 | 94.2 | 93.9 | 10 | |
| 14 | Claude Sonnet 5 | balanced | 94.2 | 93.5 | 11 | |
| 15 | Claude Sonnet 4.6 | balanced n−1 | 92.5 | 91.0 | 7 | |
| 16 | Qwen3.7 Max | frontier open n−1 | 90.9 | 90.3 | 7 | |
| 17 | GLM-5.3 | frontier open | 90.6 | 92.0 | 7 | |
| 18 | Qwen3.8 Max | frontier open | 90.4 | 89.6 | 7 | |
| 19 | GPT-5.6 Terra | balanced | 89.8 | 88.7 | 6 | |
| 20 | GLM-5.3 Flash | fast open | 89.6 | 89.6 | 6 | |
| 21 | Hunyuan 4 | frontier open | 89.6 | 87.6 | 6 | |
| 22 | Qwen3.8 Flash | fast open | 87.3 | 86.7 | 5 | |
| 23 | Kimi K2.6 | balanced open n−1 | 84.2 | 82.2 | 3 | |
| 24 | DeepSeek V4 Pro | frontier open | 83.7 | 83.4 | 3 | |
| 25 | GPT-5.6 Luna | fast | 81.5 | 80.4 | 1 | |
| 26 | DeepSeek V4 Flash | fast open | 79.3 | 80.1 | 1 | |
| 27 | Claude Haiku 4.5 | fast | 78.1 | 78.3 | 1 | |
| 28 | DeepSeek V3.2 | balanced open n−1 | 57.2 | ◀ | 54.4 | 0 |
Where the differences live
One number per model hides two things. Which tasks still separate the field, and which model is best at the specific task you are automating. Those are different questions with different answers.
Not every family still tells models apart
Every model's score on every family, one row per family, widest spread at the top. A row bunched against the right edge is a task the field has finished. A row that spans the plot still separates models. The grey bar is the range from worst to best. Points that would cover each other are nudged up or down, never sideways.
The axis runs from 0 to 100. The eight original families sit in a knot at the right edge; the two expert-tier families, the plan-year ledger above all, are where the field still spreads. Hover a point for the model and its exact score.
The ranking is not one ranking
Score by model and family, models ordered worst to best overall. Read across a row for one model's shape, and down a column to pick a model for that queue.
What survives contact with production
Two questions decide whether a model can run unattended. Does the answer hold when you ask again, and does it know the difference between work it must refuse and work it must do.
What survives being run again
Pale bar: mean score. Solid bar: the share of tasks right on every attempt. The gap between them is the part of a headline number that does not survive a rerun.
Over-refusal is the axis that actually varies
The bar is how often a model refused work a plan must carry out. The diamond is the other failure, doing something a plan is not permitted to do, which all but one model did on exactly zero tasks. Safety reporting that measures only that second direction gives every model here full marks, including the one that refuses 44% of legitimate work.
What more reasoning effort buys
Every other number on this site was produced at the vendor's default reasoning setting, which is what you get when you wire the model up the obvious way. These runs are the same models with that setting changed: 5 models across the settings below, over the suite as each row ran it: the whole suite for OpenAI and Google, the eight original families for Anthropic, whose variants were not rerun on the expert tier. OpenAI and Anthropic take a native effort setting and the harness passes it through. Google takes a thinking-token budget instead, mapped at low = 1,024, medium = 6,000, high = 20,000 tokens. Compare a model against itself down the table, and read the intervals: most of these gaps are not differences. Comparing across vendors compares their definitions of effort, which are not the same thing.
Effort against cost
One line per model, points ordered low effort, vendor default, high effort, plotted against what each setting cost per item. A line that runs flat and to the right is a model whose extra thinking bought nothing but a larger bill.
| Model | Setting | Score | 95% CI | Hard | $/1k tasks | Median | Output tokens | Of which thinking |
|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | low | 97.6 | 95.8–99.1 | 97.2 | $12.23 | 2.9s | 409 | 66% |
| medium | 98.3 | 96.9–99.4 | 98.1 | $9.75 | 2.9s | 298 | 63% | |
| high | 98.8 | 97.9–99.6 | 98.6 | $17.08 | 3.1s | 655 | 79% | |
| GPT-5.6 Terra | low | 91.1 | 88.8–93.4 | 90.3 | $8.25 | 2.1s | 523 | 75% |
| medium | 89.8 | 87.2–92.4 | 88.7 | $5.39 | 1.9s | 289 | 64% | |
| high | 94.6 | 92.2–96.7 | 93.5 | $10.46 | 2.4s | 705 | 81% | |
| Claude Opus 5 | low | 98.8 | 97.5–99.7 | 98.4 | $14.12 | 2.3s | 256 | 17% |
| default | 98.9 | 97.9–99.6 | 98.7 | $18.23 | 3.3s | 454 | 39% | |
| high | 98.7 | 97.4–99.6 | 98.7 | $18.98 | 3.2s | 463 | 40% | |
| Claude Sonnet 5 | low | 95.4 | 92.5–98.0 | 93.4 | $5.39 | 2.7s | 238 | 36% |
| default | 96.9 | 94.8–98.6 | 96.1 | $7.40 | 3.7s | 462 | 63% | |
| high | 97.0 | 94.8–98.8 | 95.2 | $7.21 | 3.9s | 449 | 62% | |
| Gemini 3.1 Pro | low | 97.1 | 95.2–98.9 | 96.7 | $29.57 | 10.5s | 2,018 | 91% |
| default | 97.1 | 95.0–98.8 | 96.4 | $36.21 | 11.4s | 2,653 | 95% | |
| high | 97.5 | 95.2–99.3 | 96.9 | $59.24 | 13.4s | 4,490 | 96% |
One generation of movement
Each vendor's current model against the one it replaced, matched by price tier. Whether an upgrade is worth the migration is a different question from who leads. On this suite the answer is not uniformly yes.
Current generation against the one before it
One row per pair. The hollow dot is the model that was replaced, the filled one is its successor, and the number beside it is the move in points. Sorted by how much the upgrade bought.
Five findings worth acting on, and one caveat.
Single-step rule application is solved.
The top 9 models sit above 95 on the full suite, and on the easier families the leaders are separated by less than their own confidence intervals. That is a capability statement, not a ranking. For a claim adjudicated against one accumulator, a policy applied to one clean record, or a measure applied to one member, the model is no longer the constraint. The retrieval, the plumbing and the review path around it are.
Depth still separates the field.
The 144 tasks marked hard at authoring time are long claim chains across a household, coordination of benefits, contradictory plan documents, and near-miss pairs. On those, the spread widens from 42.2 points to 44.9. Read the hard column if the queue you are automating is not the easy half of the work.
Price and capability have come apart.
Gemini 3.7 Flash scores within five points of the leader at $6.44 per thousand tasks against the leader's $51.54, a factor of 8. At payer volumes that ratio decides whether a queue gets automated at all. That is why every ranking on this site carries a price.
Over-refusal is the larger compliance problem.
Across the field, models did something a plan is not permitted to do on 0.8% of the tasks they should have declined, and refused work a plan must carry out on 8.7% of the tasks they should have completed. The second number is the one that kills a deployment quietly. A refusal reads as caution in a pilot and as a useless tool in production. Reporting only the first number rewards a model that refuses everything.
Code-set recall is the weakest link, and retrieval fixes it.
On code tasks with the governing rule supplied in the prompt the field averages 97.0. On tasks where the model must answer from memory it averages 94.5. That gap is the argument for putting a code-set lookup in front of a model before pointing it at coding work, and here it is measured rather than assumed.
This is one run of a focused suite.
210 tasks, 3 samples each where the budget allowed and one where it did not. Confidence intervals are reported, and models whose intervals overlap should be read as close rather than ranked. Tasks were written and reviewed internally; external clinical and compliance review is planned. Synthetic material is cleaner than a real intake queue, so read these scores as an optimistic estimate. The methodology pagelists the known limitations.
10 families, 210 tasks
Ordered by how much the family separates the field. The ones at the top are where the choice of model still matters.
Tasks were run up to 3 times per model. The leaderboard averages a model's attempts. “Right every try” under Operations is the stricter number, and it carries the attempt count beside it because models did not all finish the same number of passes. Rows marked ×1 were run once and sit slightly higher than they would after resampling. Read those against the interval, not the point. 880 transport failures are excluded from scoring rather than counted as zeros.
Full methodology, including what this benchmark does not measure, is on the methodology page. Every response from the 28 ranked models, with its own reasoning where the vendor exposes it, is in transcripts. The 19,923 figure above covers the whole run, including the reasoning-tier variants, which are scored but not ranked.