How Machine-Learned Utilization Review Became the Most Contested Technology in American Health Care — and the Legal, Legislative, and Market Counter-Reaction of 2025–2026
Between 2023 and 2026, the American health-care system entered its first full-scale confrontation with algorithmic administration. What began as an efficiency program — the deployment of machine-learning models and natural-language-processing engines to clear rising volumes of prior authorization (PA) requests and adjudicate claims — mutated into a systemic governance crisis. Investigative reporting by ProPublica and Reuters documented insurers whose physicians signed off on tens of thousands of algorithmically generated denials at a pace of seconds per case,[1][2] while STAT News revealed a UnitedHealth model alleged to override treating physicians and terminate post-acute care for Medicare Advantage enrollees despite a known error rate approaching 90 percent.[3][4]
The empirical record assembled by federal agencies confirms the structural problem. In 2024, Medicare Advantage (MA) insurers issued nearly 53 million PA determinations and denied 4.1 million of them;[5] of the small fraction of denials that were appealed (11.5 percent), more than eight in ten were overturned.[5] A 2026 HHS Office of Inspector General (OIG) report found that MA organizations reversed 95 percent of appealed skilled-nursing-facility denials — a overturn rate the OIG interpreted as evidence that the initial denials were clinically indefensible.[8] The data imply that the denial itself, not the review, is the product: an automated friction event that rationes care by exploiting appeal fatigue.
The counter-reaction has been swift and bipartisan. California's SB 1120 (effective January 1, 2025) established the template — a licensed physician, not an algorithm, must make the final medical-necessity determination[12] — and during the 2025–2026 sessions Arizona, Maryland, Nebraska, Texas, Alabama, Indiana, Utah, Washington, and Georgia followed with statutes prohibiting AI from serving as the sole basis of an adverse determination.[11][12] At the federal level, CMS's Interoperability and Prior Authorization Final Rule (CMS-0057-F) compressed decision timelines and, beginning March 31, 2026, forces payers to publish approval, denial, and appeal metrics that will for the first time expose the black box to quantitative sunlight.[7]
This report dissects the technical architecture of algorithmic utilization review, quantifies its clinical and economic externalities, analyzes the litigation and legislative response, and models the emerging payer–provider AI arms race. Its central conclusion is that "human oversight" mandates will only be as effective as their audit and enforcement architecture: without enforceable definitions of meaningful human review, the 1.2-second physician signature of the PXDX era will simply re-emerge in compliant paperwork.
Prior authorization is a relic of 1960s utilization management that survived because it worked as a cost-containment friction. By requiring provider submission before service delivery, payers insert a checkpoint at which medical necessity can be verified against plan criteria. The checkpoint, however, is labor-intensive: the American Medical Association's 2025 physician survey found that physicians complete roughly 40 PA requests per week, consuming on the order of 13 hours of physician and staff time weekly,[9] with 94 percent of respondents reporting that the burden contributes to burnout.[10] For payers, the same labor intensity represents a rising administrative cost line in a margin-compressed environment — particularly in Medicare Advantage, where capitated payments make every avoided or delayed service a direct margin contribution.
Three forces converged around 2018–2020 to make automation inevitable. First, PA volumes grew relentlessly: MA plans now process an average of 1.7 PA requests per enrollee per year — roughly 85 times the rate in traditional Medicare.[5] Second, advances in natural language processing made it technically feasible to ingest unstructured clinical documentation — the raw material of utilization review — at scale. Third, the Medicare Advantage business model rewarded denial velocity: under capitation, a denied post-acute stay that is never appealed is pure margin, and behavioral economics guaranteed that most beneficiaries would never appeal.[5]
The result was a migration from rules-based claims editing (deterministic checks of coding and eligibility) to probabilistic, machine-learned prediction of medical necessity. This migration is qualitatively different. A rules engine denies a claim because a code is invalid; a predictive model denies a patient because their record statistically resembles records previously denied. The epistemic shift — from verification to prediction — is precisely what renders the system a "black box": the adverse determination is no longer derivable from published plan criteria alone, but from proprietary weights trained on proprietary data, shielded as trade secrets, and in many deployments invisible to the treating physician, the beneficiary, and even the plan's own regulators.
Physicians perceive the consequences directly. In the AMA's December 2025 survey, 74 percent of physicians said denial rates have increased over recent years, 60 percent said they are concerned AI is worsening the trend, and only 24 percent believed payer reviews are consistently conducted by qualified clinicians.[10] Just one in three physicians trusts the integrity of the prior authorization process at all.[9] This trust deficit is the political fuel behind the legislative wave analyzed in Section 5.
Although vendors differ, deployed utilization-review AI shares a common reference architecture. An ingestion layer extracts structured claims data (ICD-10, CPT, revenue codes) and applies optical character recognition and NLP to unstructured inputs — physician notes, discharge summaries, imaging reports. A feature-engineering layer maps these inputs to plan coverage criteria, often encoded as clinical pathways. A decision layer — ranging from gradient-boosted classifiers to large language models — produces a recommendation: approve, deny, or flag for human review. Crucially, in the deployments at issue in litigation, the "human review" stage was not a safeguard but a ratification mechanism: physicians were presented with pre-populated denial queues and batch-signature interfaces.
Figure 1 · Reference Architecture of Algorithmic Prior Authorization
This architecture exhibits four recurrent failure modes, each documented in litigation or federal oversight.
Failure mode 1 — Training-data misalignment. A model trained on historical payment decisions learns payer behavior, not clinical necessity. If historical denials were themselves erroneous, the model industrializes the error. The nH Predict complaint alleges precisely this: the model was tuned to flag post-acute care for termination at scale, and its outputs were treated as presumptively correct even when treating physicians certified continued medical necessity.[4][19]
Failure mode 2 — Group-data substitution. Predictive denial inherently relies on cohort statistics. Applied to an individual, it denies coverage because the patient resembles a group, violating the foundational insurance-law principle — now codified in CMS's 2024 rule and a dozen state statutes — that medical-necessity determinations must be based on the individual's clinical circumstances.[6][12]
Failure mode 3 — Automation bias and rubber-stamp review. When a human signs an algorithmic queue, cognitive science predicts deference to the machine, and operational incentives (claims per hour) reinforce it. Cigna's PXDX system is the canonical example: physicians spent an average of 1.2 seconds per claim and one doctor could sign off on as many as 60,000 denials per month.[1][16] The "human in the loop" was a compliance fig leaf.
Failure mode 4 — Opaque criteria and un-auditable outcomes. Because model weights and training corpora are trade secrets, regulators cannot verify fairness, and plaintiffs cannot prove discrimination without court-ordered discovery. This is why the 2025–2026 statutes pair the human-review mandate with audit and inspection rights — Texas and Maryland explicitly authorize the insurance commissioner to examine the algorithm itself.[11][12]
The most authoritative public data come from CMS's Medicare Advantage prior-authorization metrics, analyzed by KFF.[5] Three patterns matter for this report.
Pattern 1 — Denial intensity is rising and concentrated. The MA-wide denial rate moved from 5.8 percent (2021) to 7.4 percent (2022), dipped to 6.4 percent (2023), and reached 7.7 percent in 2024 — 4.1 million denied requests.[5] But the aggregate conceals concentration: in 2024, UnitedHealth Group plans denied 12.8 percent of requests, roughly three times Elevance Health's 4.2 percent.[5] The 2022 insurer-level data show a similar spread, with CVS Health at 13 percent and Kaiser Permanente at 10.4 percent against Elevance's 4.2 percent.[13] Denial intensity is thus a governance choice, not an epidemiological constant.
Figure 2 · Medicare Advantage Prior Authorization Denial Rate, 2021–2024 (line chart)
Pattern 2 — Algorithmic deployment tracks denial spikes. The clearest natural experiment is UnitedHealth's post-acute care denial trajectory during the nH Predict deployment window: 10.0 percent (2020), 16.3 percent (2021), and 22.7 percent (2022) — a more-than-doubling in two years, contemporaneous with the algorithm's rollout and with internal pressure on clinical staff to follow its outputs.[19][3]
Figure 3 · UnitedHealth Post-Acute Care PA Denial Rate During nH Predict Deployment (line chart)
Pattern 3 — The appeal desert proves initial denials function as friction, not judgment. Only 11.5 percent of 2024 MA denials were appealed, yet 80.7 percent of appeals succeeded.[5] The OIG's 2026 skilled-nursing-facility analysis pushed the overturn rate to 95 percent and concluded the figure "indicates that some initial denials were erroneous."[8] An earlier OIG study found 13 percent of PA denials outright contradicted Medicare coverage rules.[5] Economically, this is a rationing mechanism with a low expected cost: if 88.5 percent of denials never meet resistance, even an 80 percent loss rate on the remaining 11.5 percent still suppresses utilization across the whole portfolio. The rational provider response — documented in AMA surveys, where 79 percent of physicians report patients abandoning treatment due to PA — is exactly what the model prices in.[9]
Figure 4 · The Appeal Desert: 2024 MA Denial Funnel (analytical bar chart)
Figure 5 · PA Denial Intensity by Parent Insurer, Medicare Advantage 2022 (bar chart)
| Metric | Value | Interpretation |
|---|---|---|
| Total PA determinations | ~52.9 million | ≈1.7 requests per enrollee; ~85× traditional Medicare intensity[5] |
| Fully/partially denied requests | 4.1 million (7.7%) | Highest rate in the 2021–2024 series[5] |
| Share of denials appealed | 11.5% | "Appeal desert": friction rationing[5] |
| Appeal overturn rate | 80.7% | Initial denials frequently clinically indefensible[5] |
| SNF denial overturn rate (OIG 2026) | 95% | Near-total reversal where review actually occurs[8] |
| Highest / lowest payer denial rate (2024) | 12.8% / 4.2% | UnitedHealth vs. Elevance; 3× spread[5] |
| Physicians reporting rising denial rates (AMA 2025) | 74% | 60% attribute worsening trend to AI[10] |
The class action Hobbs v. UnitedHealthcare, filed in the District of Minnesota in November 2023, is the paradigmatic AI-denial lawsuit. The complaint alleges UnitedHealth deployed an algorithm — "nH Predict" — to identify Medicare Advantage enrollees for termination of post-acute skilled nursing and rehabilitation coverage, that the model operated with a 90 percent error rate when measured against subsequent appeals outcomes, and that its outputs systematically overrode the treating physicians' certifications of medical necessity.[4][15] Corroborating the pleading, STAT News reported that UnitedHealth pressured clinical staff to follow the algorithm to cut off rehab care, and that when federal administrative law judges heard appeals of these denials, roughly 90 percent were reversed.[3][47] Procedurally, the case has matured into the most important discovery vehicle in the field: in 2025 a federal judge ordered UnitedHealth to produce broad discovery on the algorithm's operation, while a separate ruling dismissed five of seven counts but allowed the core claims to proceed — a split that signals both the viability and the doctrinal difficulty of AI-denial theories.[1][14] A parallel ProPublica investigation documented a related UnitedHealth algorithm for mental-health coverage that had been deemed illegal in three states by the end of 2021.[18]
If nH Predict illustrates algorithmic prediction gone adversarial, Cigna's PXDX system illustrates algorithmic process gone absurd. Reuters found that Cigna's physicians denied more than 300,000 payment requests in two months of 2022 through PXDX, spending an average of 1.2 seconds per claim, with individual doctors signing off on up to 60,000 denials per month.[2][16] ProPublica's follow-up showed the commercial logic with unusual clarity: when Cigna added a specific nerve test to the PXDX denial list, executives estimated the move would reject more than 17,800 claims per year — the algorithm as a revenue instrument, calibrated line-item by line-item.[1] Federal litigation followed: a California class action accuses Cigna of mass algorithmic denial without individualized review, and in 2024–2025 a federal judge advanced the class claims over Cigna's motion to dismiss, holding that plaintiffs plausibly alleged a benefits-denial process administered by automation rather by the plan's fiduciary judgment.[9][16]
| Case / Forum | System at Issue | Core Allegation | Status (as of mid-2026) |
|---|---|---|---|
| Hobbs v. UnitedHealthcare (D. Minn.) | nH Predict | AI with ~90% error rate terminated post-acute care, overriding treating physicians; ~90% ALJ reversal on appeal[4] | Core counts survived dismissal; broad algorithmic discovery ordered (2025)[14] |
| Cigna PXDX class action (N.D. Cal.) | PXDX | Mass denial of 300k+ claims in two months; 1.2-second physician "review"[2] | Class claims over automated benefits denial advanced[16] |
| ERISA algorithmic-denial actions (multiple) | Various medical-necessity models | AI-driven denials breach plan fiduciary duties; participants entitled to individualized review | Courts permitting selected counts to proceed; doctrine forming[16] |
| State enforcement actions (mental-health algorithm) | UnitedHealth behavioral-health model | Systemic denial of mental-health coverage contrary to parity law | Deemed illegal in three states by end of 2021[18] |
Together these cases establish litigation as the de facto transparency regime for algorithmic utilization review: because vendors will not voluntarily disclose model cards, and regulators lacked audit authority until the 2025–2026 statutes, discovery is the only public audit. Every motion to compel in Hobbs is therefore a precedent in the law of algorithmic accountability.
The federal floor was set in January 2024, when CMS's final rule permitted MA plans to use AI for coverage determinations but required that medical-necessity decisions be based on each individual's circumstances and reviewed by a physician or appropriate health-care professional.[12] The CMS-0057-F Interoperability and Prior Authorization Final Rule then built the transparency architecture: standard PA decisions within 7 calendar days (from January 2026), expedited decisions within 72 hours, and — critically — public posting of approval, denial, and appeal metrics, with the first disclosures due March 31, 2026.[7] Meanwhile, in a notable policy irony, the CMMI WISeR model (launched January 1, 2026) actively tests AI-driven prior authorization in traditional Medicare across six states for services vulnerable to fraud — signaling that Washington's posture is not anti-AI but pro-auditability.[5]
The states, however, are where the human-oversight mandate became hard law. California's SB 1120, effective January 1, 2025, prohibits AI from being the sole means of denying, delaying, or modifying care and reserves the final medical-necessity determination to a licensed physician competent in the relevant clinical issues.[12] The 2025 session added four states — Arizona (HB 2175), Maryland (HB 820), Nebraska (LB 77), and Texas (SB 815)[12] — and the 2026 session added at least six more: Alabama (SB 63), Indiana (HB 1271, uniquely extending the rule to AI downcoding of claims), Utah (SB 319), Washington (SB 5395), Maryland again (HB 1563, imposing quarterly adverse-determination reporting with AI flagging), and Georgia (SB 544).[11] As of early April 2026, at least 25 additional states have issued AI-in-insurance guidance under the NAIC model bulletin, creating a de facto national supervisory expectation even without new statutes.[6]
Figure 6 · Cumulative State Statutes Governing AI in Utilization Review, 2023–2026 (bar chart)
| State | Citation | Effective | Core Requirement |
|---|---|---|---|
| Colorado | DOI AI Rules | 2023 / 2025 | First-in-nation AI-in-insurance framework; governance, risk management, and anti-discrimination duties for health insurers (Oct 2025)[12] |
| California | SB 1120 | Jan 1, 2025 | Final medical-necessity determination only by licensed physician; AI cannot be sole basis for denial; audit and periodic-validation duties[12] |
| Arizona | HB 2175 | 2025 | Provider-independent review before insurer denial; sole reliance on algorithmic output prohibited[12] |
| Maryland | HB 820 / HB 1563 | 2025 / 2026 | AI must incorporate individual clinical history; 2026 law adds quarterly adverse-determination reporting with AI flagging and commissioner investigation triggers[11][12] |
| Nebraska | LB 77 | 2025 | AI output cannot be sole basis of medical-necessity denial; mandatory disclosure of AI use to providers, enrollees, and public website[12][17] |
| Texas | SB 815 | 2025 | AI barred from adverse medical-necessity determinations; permitted only for admin support/fraud detection; commissioner audit authority[12] |
| Alabama | SB 63 | Oct 1, 2026 | Mirrors CMS guardrails; annual certification that AI does not rely on group datasets or discriminate[11] |
| Indiana | HB 1271 | Jul 1, 2026 | Prohibits AI-only claim downcoding; also prohibits provider AI claim submission without human review[11] |
| Utah | SB 319 | Jan 1, 2027 | Disclosure of AI use; adverse determinations by independent medical judgment, not dictated by AI[11] |
| Washington | SB 5395 | 2026 | Sole reliance on AI for denial/delay/limitation prohibited; only licensed professionals may issue adverse determinations; AI-denial counting reported to commissioner[11][20] |
| Georgia | SB 544 | Jan 1, 2027 | Permits AI automation but bars adverse determinations without licensed-provider review and approval[11] |
| Illinois | Utilization-review law | In force | Only a "clinical peer" may make an adverse determination; "algorithmic automated process" cannot be sole decision-maker[6] |
Read together, the statutes converge on a doctrinal consensus: AI may approve; only humans may deny. Approval automation is broadly permitted — Washington and Georgia explicitly bless it — because false approvals cost the payer money and thus self-correct, whereas false denials externalize cost onto sick patients who rarely appeal. The legislative judgment is therefore asymmetric by design. The unresolved question is enforcement: a "human review" requirement without a minimum-review standard recreates PXDX under a compliant header. Maryland's quarterly reporting and Washington's AI-denial counting are the first statutory attempts to measure the substance of review, and they will likely become the model for the next legislative generation.[11]
The deployment of payer-side denial AI has triggered a rational, mirror-image response on the provider side, and the interaction is best modeled as an arms race with four observable strategies.
Provider strategy 1 — Counter-automation of documentation. Health systems and revenue-cycle vendors are deploying NLP to pre-emptively align clinical documentation with payer criteria, and — following Indiana's symmetric regulation — to police their own AI-assisted claim submission.[11] Large provider systems now run "denial-prediction" models that estimate the probability of a PA denial before submission and route high-risk requests to specialist coders. The effect is to shift the contest from the clinical record to the feature space of the payer's model.
Provider strategy 2 — Industrialized appeals. Because the overturn rate is 80–95 percent, appeals are an arbitrage with positive expected value once the marginal cost of filing falls. AI-drafted appeal letters, auto-populated from the electronic record, are collapsing that cost. Payers understand this: the economic defense of the denial friction model depends on appeal costs staying high, so provider-side automation directly attacks the payer's margin mechanism.
Payer counter-strategy — Criteria migration and administrative re-rating. As statutes freeze the denial decision with a human, payers migrate control upstream: tightening the written clinical criteria themselves, expanding PA lists (subject to state disclosure laws), and — as Indiana's downcoding provision anticipates — shifting from binary denial to payment-level reduction, which is legally less regulated and clinically less visible.[11]
Payer counter-strategy — Compliance theater risk. The cheapest response to a human-review mandate is to keep the algorithm and add a signature. The 1.2-second precedent proves the operational feasibility; only audit rights (Texas, Maryland), periodic-validation duties (California), and overturn-rate triggers (federal legislative proposals to penalize plans whose denials are frequently reversed[5]) make theater detectable.
| Actor | Move | Counter-Move | Equilibrium Effect |
|---|---|---|---|
| Payer | AI batch denial (PXDX-era) | Provider appeal mills; media exposure; litigation | Denial friction monetized until reputational/legal cost exceeds margin[1] |
| Payer | Predictive denial (nH Predict-era) | Class actions; discovery into model; OIG audits | 90% overturn rates become evidence of systematic error[4][8] |
| State | Human-oversight mandates (2025–26) | Payer criteria migration; downcoding; signature theater | Regulation shifts venue of control, not necessarily volume[11] |
| Provider | Documentation AI + auto-appeals | Payer model retraining on adversarial docs | Symmetric automation; contest moves to model governance[11] |
| Federal | Metrics transparency (CMS-0057-F) | Metrics gaming; denominator management | Public dashboards create naming-and-shaming enforcement[7] |
The deeper structural insight is that both sides are converging on the same technology — large language models over clinical text — with opposite objective functions. This symmetry has a silver lining: provider-side AI is the only economically viable counterweight to payer-side AI at population scale. A solo practice cannot out-appeal a Fortune 5 insurer; a health-system NLP pipeline can. The policy risk is that smaller, un-augmented providers — safety-net clinics, rural hospitals — become the collateral damage of a machine-vs-machine contest they cannot afford to enter, widening the very health-equity gaps the CMS health-equity analysis (currently unenforced[5]) was designed to monitor.
Figure 7 · Clinician Sentiment on Algorithmic Utilization Review (AMA, Dec 2025 — bar chart)
For legislators: define "human review" operationally — minimum review time is unenforceable, but batch-signature prohibition, mandatory reviewer credentialing (the "clinical peer" standard), and per-reviewer denial-volume reporting are measurable and would have detected PXDX instantly. Attach overturn-rate triggers: federal proposals to penalize plans whose appealed denials are frequently reversed convert the OIG's 95 percent statistic into a price.[5][6]
For regulators: exercise the new audit powers. Texas's commissioner-inspection authority and Maryland's investigative trigger are only as potent as their first enforcement action. Regulators should also demand model cards and validation studies as condition of market access, mirroring Colorado's governance framework, and publish the CMS-0057-F metrics in comparable, machine-readable form so that the Figure 2 and Figure 5 comparisons become permanent public infrastructure.[7][12]
For payers: the litigation docket is now a cost center larger than the administrative savings. Model governance — independent validation against individual-circumstance criteria, drift monitoring, denominator discipline, and a genuine clinical escalation path — is cheaper than discovery. The voluntary 2025 payer pledge to route clinical PA denials to medical professionals is a recognition of this; its credibility depends on whether the March 2026 metric disclosures show overturn rates falling.[6][7]
For providers and digital-health entrepreneurs: the white space is audit-grade automation for the underserved middle market — appeal automation, PA-preparation, and denial-prediction as-a-service for practices too small to build in-house. The regulatory tailwind is real: every statute that mandates individualized review increases the unit economics of documentation quality, and every disclosure mandate creates demand for compliance tooling on the payer side as well.[11]
Outlook to 2027: expect (1) consolidation of the "AI may approve; humans must deny" norm into federal legislation; (2) the first commissioner enforcement actions under the 2026 audit statutes; (3) migration of payer control from denial to downcoding and criteria design, provoking a second legislative wave (Indiana's downcoding ban is the sentinel); and (4) symmetric provider-side AI becoming standard revenue-cycle infrastructure. The black box will not be opened; it will be counterbalanced — and the quality of that balance, not the elegance of the models, will determine whether algorithmic administration in health care becomes a productivity story or a rationing scandal.