AI Billing Tools Added $942 Million in Hospital Costs, Insurers Say

Hospitals' use of AI tools to submit insurance claims added $942 million in health spending over two years, according to an analysis from the Blue Cross Blue Shield Association. TechCrunch
The Association found a sharp rise in patients recorded as having complex conditions. It said there was a gap between the billing codes and the treatment given, with no evidence that the care itself had changed. TechCrunch
The analysis was based on Blue Cross Blue Shield claims data and was published on the Association's official website on Sept. 22. Blue Cross Blue Shield Association Coverage of the release on Sept. 24 described the result as more intense coded care adding $942 million over two years compared with 2023. Reuters
That comparison focuses on commercial hospital claims. A Reuters article on March 12, 2026, titled 'US insurers and hospitals turn to new AI for age-old battle over charges vs payments,' described an analysis of that claims set. Reuters The March account is useful context. Automation is not limited to one side of the billing process.
Coverage of the current dispute identifies Luke Chalker as senior vice president of the Blue Cross Blue Shield Association and Shiv Rao as founder of AI startup Abridge. TechCrunch
The broader context here will be familiar to anyone who has followed hospital billing software. Clinical documentation integrity, or teams that check doctors' notes are complete, computer-assisted coding, or older software that suggests billing codes, and payer-side claim review, or insurers checking those claims, have existed together for years. Large language models, the AI systems behind chatbots, lower the cost of pulling possible codes from free-text notes. The incentives are straightforward. Completeness pays.
In my view, the technical question is narrower than whether AI coding works. For people who build or use these systems, the issue is provenance and auditability, meaning whether each code can be traced and checked. A model can pick up a condition a doctor mentioned in passing, apply the correct ICD logic, the standard rules for labeling diagnoses, and attach it to a claim with no change in tests ordered, length of stay, or mix of procedures. That is the gap the Association alleges. Payers see higher intensity in the codes. They do not see it in the care used.
Looking at what this means for system design, logging will matter as much as accuracy. Which sentence in the notes triggered a code. What confidence level was required. Whether a human coder or clinician approved it. Whether the same visit would have been coded differently without the AI suggestion. Those details decide whether the tool becomes reliable automation or a steady risk of upcoding, or billing for more complex care than was given.
Worth flagging for technology leaders, the long arc still points toward better records. Structured, searchable notes allow for utilization review, or checks on what care was needed, safety work, and long-term research that paper charts never supported. The near-term task is to align what the AI is asked to optimize. Capturing codes alone is incomplete. Capturing codes tied to clinical evidence, with a clear trace back to the note, is the version both hospitals and insurers can work with.


