Bevaya Benchmark: Insurance-Trained AI Beats Frontier Models on Loss Runs, Insurance's Hardest Documents
PR Newswire
NEW YORK, Sept. 23, 2026
Across 346 real loss runs, every leading general-purpose AI model scored between 78% and 85% field accuracy, from the most expensive configurations available to the least. Bevaya's insurance-trained InsurGPT™ loss run model scored 93.1%. The difference is the training, not a newer or larger model.
NEW YORK, Sept. 23, 2026 /PRNewswire/ -- Bevaya, the AI Agent platform built for insurance, today announced benchmark results comparing insurance-trained and general-purpose AI on the industry's hardest documents, loss runs. Bevaya's InsurGPT™ loss run model proved both more accurate and faster than the strongest general-purpose model tested, the combination insurers need to move submissions and claims through without a person opening every file.
Demand for AI in insurance has risen, and the pressure has shifted with it: pilots forgave errors, production does not. Insurers now need AI that is accurate, trusted, and governable, and regulators increasingly expect them to show how an automated decision was reached. That standard is what has brought insurers to Bevaya, whose AI Agents read, analyze, and recommend across underwriting, claims, and policy servicing, returning verified data a professional can act on. Bevaya has delivered more than 120 production deployments, including at three of the top five U.S. property and casualty carriers. The benchmark isolates the model behind that work and measures it against current general-purpose AI under identical conditions. Production adds verification and review on top of it.
"We built InsurGPT on hundreds of millions of non-public insurance documents, labeled by insurance practitioners rather than crowdsourced workers, and this benchmark shows why that matters," said Chaz Perera, co-founder and CEO of Bevaya. "General-purpose AI keeps getting better at general work. On specialized insurance work, each new generation is not moving the needle, and a general model on its own is not workable in production. You have to train on the documents themselves and on the judgment behind them: how an underwriter reads a loss run, where the expertise actually sits. None of that is in any public dataset."
- The most accurate model in the test. Bevaya's loss run model led every accuracy measure, scoring 93.1% field accuracy against 85.3% for the strongest general-purpose model. Every general-purpose model landed between 78% and 85%, whatever its generation or price. None broke 86%.
- Fewer errors, more documents ready to use. Fewer than half as many wrong fields means nearly twice as many loss runs came back needing no correction, because a loss run is usable only when every value on it is correct. Bevaya's model also read each document in less than half the time, which for an underwriting desk is turnaround on a submission.
- Strongest where a mistake costs most. Claim numbers, policy numbers, and claimant IDs tie a loss to the right policy and the right file. On these identifiers, Bevaya's model scored 86.6% against 72.8%, the widest gap of any category of field in the benchmark. On fields that depend on carrier-specific vocabulary, such as claim type and claim status, general-purpose models fell as low as 59% while Bevaya's model scored 82% and 89%.
In production, Bevaya adds a second model that checks every answer against the source document and a confidence score on every field. Anything uncertain goes to the insurer's own staff for review through Bevaya's patented human-in-the-loop technology, and every action is recorded to a full audit trail. That is what carries accuracy from 93.1% in the benchmark to the 98%+ Bevaya delivers in production. For loss runs specifically, more than 90% are completed with no human touch.
Bevaya selected loss runs for the benchmark because they are among the most difficult documents in insurance to read. Each carrier formats its own differently, and total incurred, the figure underwriters rely on most, is often absent and must be calculated from other columns. These conventions cannot be inferred from the page, and general-purpose models have no public source from which to learn them. InsurGPT, Bevaya's ensemble of specialist AI models, is trained on more than 300 million non-public insurance documents labeled by what Bevaya calls model tutors: insurance practitioners on staff whose sole job is creating training data. The loss run model alone required seven months of their annotation.
"An underwriter does not receive 93% of a loss run. They receive the whole document, or they receive a task," said Ratish Dalvi, SVP, AI and engineering at Bevaya. "Only a fully correct document goes straight through, and ours came back with nothing to fix nearly twice as often as the strongest frontier model. The story of economic history is a story that tends toward specialization, and this benchmark is what that looks like in insurance. Bevaya Labs published every model, every field, every line of business, and every document length, so anyone can check the work."
Full results, the method, and the Bevaya Labs research are at www.bevaya.ai/benchmarks.
About Bevaya
Bevaya is the AI agent platform built for insurance. Bevaya's AI agents read, analyze, and recommend across underwriting, claims, and policy servicing, including triage and clearance, rating, coverage analysis, and next-step recommendations, with 98%+ accuracy. Powered by InsurGPT™, an ensemble of specialized AI models trained on 300M+ insurance documents, Bevaya brings the precision insurance work demands. Across 120+ production deployments at the industry's largest carriers, brokers, and TPAs, Bevaya's AI agents deliver 3-4× capacity gains and measurable ROI from day one. Visit www.bevaya.ai
Media Contact: media@bevaya.ai
View original content to download multimedia:https://www.prnewswire.com/news-releases/bevaya-benchmark-insurance-trained-ai-beats-frontier-models-on-loss-runs-insurances-hardest-documents-302886931.html
SOURCE Bevaya
