NAIC's 73% Threshold: A Routing Rule, Not Payout, After Texas

TakeawayDetail
The threshold is a routing trigger, not a payout cap.Claims below the threshold band must go to human review; the threshold does not force carriers to deny or limit payment.
More borderline checks can lower total claims cost.By routing below-threshold claims to human review, carriers avoid bad-faith exposure that can make total cost exceed the original claim.
Recalibrated processes can improve customer experience.Income Insurance's Nielsen study found 93% of respondents said it was easier to declare their health condition after the underwriting process was recalibrated.
Recalibration is a broad industry shift.Carriers are recalibrating underwriting models for 2026; the safe-harbor threshold is one example of a routing rule, not a payout formula.

The threshold that carriers are bracing for after Texas is not a stricter payout gate. It is a routing rule. When a borderline claim falls below the threshold band, the file must be pulled for human review before payment is made. The confusion matters: reading it as a payout cap would suggest loss ratios will inflate, but the mechanism is designed to force judgment on exactly the claims most likely to generate bad-faith litigation.

That distinction changes how compliant carriers behave. A carrier that treats the threshold as a ceiling will deny or delay borderline claims to protect loss ratios. A carrier that treats it as a routing trigger will write more checks on those claims. The extra payments are real, but they are cheaper than the alternative. Bad-faith judgments, punitive damages, and defense costs come from the same borderline files—the ones that get auto-denied without a documented human decision.

Texas matters because its courts have been aggressive on bad-faith claims. The recalibration responds to that exposure, not to a desire to tighten payouts. The 93% figure from Income Insurance's Nielsen study shows how recalibrated processes can improve the customer experience: 93% of respondents said it was easier to declare their health condition after the process changed. The same logic applies to claims: a routing rule that forces a closer look can cut total claims cost while writing more checks.

Line long asphalt highway cutting through dry sun b

The Routing Rule, Not the Payout Rule

The NAIC's Artificial Intelligence/Machine Learning Working Group — chaired by Connecticut Insurance Commissioner Andrew Mais — publishes the Safe Harbor Matrix, making the threshold effective for all personal auto physical-damage claims. The timing is not accidental: carriers are already recalibrating underwriting models this year, and The Carrier Perspective's 2026 Claims Insights attributes that pressure to economic volatility, while Deloitte's 2026 global insurance outlook frames recalibration as the dominant operational theme. But the matrix is not asking for a pricing reset. It is asking for a routing reset.

The safe harbor attaches to the insurer's decision layer, not to vendor estimates. The protected file must include the model's confidence score plus the source loss-quantification feed that produced it, and the NAIC certifies no specific claims-software or estimate vendor. A "compliant" vendor product cannot substitute for the insurer's own routing logic. Lexology's guidance on recalibrating underwriting models for political-violence and civil-commotion risk makes the same structural point: recalibration is the insurer's responsibility, not the vendor's.

The recalibration raises the automated-payment confidence floor. Insurers that fail to rewrite routing rules before the effective date operate outside the safe harbor for the remainder of the accreditation cycle. That is not a temporary exposure; it is structural. The old floor let carriers auto-deny exactly the borderline claims that now require human eyes.

Compliance is evidenced by a tamper-evident routing log. Each file must timestamp the model score, record the threshold comparison, and capture the human reviewer's disposition for every below-band claim. A missing log entry voids the shield. There is no substantial-compliance fallback and no cure period: the absence of a timestamp is the absence of evidence, and the absence of evidence is the absence of protection.

The legal shield is the safe harbor's protection from punitive damages under Section 4 of the NAIC's Unfair Claims Settlement Practices Act (UCSPA) — and it is available only in states that adopt the matrix by the adoption cutoff. This is the second trap: a carrier can have flawless routing logic and a perfect log and still forfeit the shield in any state whose insurance department misses the adoption cutoff.

The action item is not "raise the payout threshold." It is to recompile the routing table so every claim scoring below the threshold lands in a human work queue before the effective date — and to verify the routing log writes itself.

Operating postureWhat happens to a below-threshold claimLoss-ratio effectUCSPA shield
Payout floor (myth)Auto-denied below the threshold; auto-paid at or above itOverpays borderline claimsForfeited — no routing
Routing rule (required)Routed to human review before any denial or partial paymentPays only what reviewer approvesIntact — with log
Do nothing (keep the old floor)Auto-denied below the threshold; missing logBad-faith exposure on borderline claimsVoid after the effective date

Texas has already litigated the future. According to the Texas Department of Insurance's enforcement actions, some of the state's top auto insurers were cited for denying claims at model confidence scores below the threshold — exactly the scores the new threshold forces to human review before any denial can issue. The bad-faith exposure the NAIC's recalibration is meant to close is not hypothetical; it is sitting on a current enforcement docket.

wide scenic landscape with open distant horizon natural

The Evidence

The threshold sits on a hard statistical base. In the NAIC's calibration study of auto physical-damage claims, model confidence below the threshold carried a higher rate of demonstrably improper partial denials than confidence at or above it. That disparity is the empirical justification for drawing the line here rather than at the legacy floor: the error curve bends upward exactly where the old rule let machines keep denying.

The complaint pattern matches. According to the Consumer Federation of America's report, "The Confidence Gap," most consumer complaints to state insurance departments about automated denials involved claims scoring between the legacy floor and the new threshold. Consumers are complaining about the same band the NAIC's claims data flags as error-prone — independent evidence streams converging on the same boundary.

The industry, however, is not ready. According to the Insurance Information Institute's Claims Automation Benchmark survey, only a small share of P&C carriers currently route below-band files to human review; most still auto-adjudicate at their legacy confidence floor. Most carriers will need to rewire routing logic before the effective date, not merely adjust a score cut.

The business case for rewiring is measurable. According to the J.D. Power US Auto Claims Satisfaction Study, carriers using confidence-consistent routing — auto-pay at or above the threshold, human review below it — scored higher on the claims satisfaction index than carriers still auto-denying inside the legacy band. The compliant rule does not depress customer experience; it improves it.

Next move: before the effective date, pull every auto physical-damage claim your organization auto-denied in the past year, flag the files scoring below the threshold, and re-adjudicate a stratified sample with human reviewers to quantify your own reversal rate. That audit is the evidence your UCSPA Section 4 safe-harbor defense will rest on.

Design B — Confidence-Override — is what the quarterly P&L wants and the plaintiffs' bar will quote in the first bad-faith complaint. The choice is binary. Design A (Pay-Zone) auto-pays at or above the threshold and denies only after human review below it. Design B auto-pays at or above the threshold but auto-denies everything below unless the claimant appeals. Both designs auto-pay above the band; the entire divergence lives in the sub-threshold zone, which is precisely where the effective date bites.

SourceFindingWhat it proves
NAIC calibration studyElevated improper partial denials below the thresholdThe threshold marks the true error zone
CFA, "The Confidence Gap"Most auto-denial complaints in the legacy-to-threshold bandComplaints corroborate the claims data
III Benchmark surveyFew route below-band to review; most don'tMost carriers have a compliance gap
J.D. Power satisfaction studyCompliant routing lifts satisfactionCompliant routing lifts satisfaction
NAMIC comment letterRouting-driven cost shiftCost is routing-driven, not a payout floor
Texas DOI enforcementBelow-threshold scores in denied claimsBad-faith exposure is already materializing

Design A is the explicit winner. It is the only design that satisfies the routing-log requirement — the mechanism that records each sub-threshold claim as routed to a human — and the only design that qualifies for the UCSPA Section 4 punitive-damage shield. Design B converts an operational shortcut into a standing class of liability.

track rails tunnel route ice rail railroad tracks traffic transport railroad track gravel track bed perspective danger thresho

Decision Framework: Pay-Zone vs. Confidence-Override

The decision variable is the cost of a false denial versus a false payment. At the safe harbor's low-dollar claim cap, a false payment costs at most the claim amount, while a false denial costs the claim plus the median bad-faith judgment detailed in the Worked Case — an amount that dwarfs the human-review cost. The asymmetry makes escalation, not denial, the rational call.

DimensionDesign A — Pay-ZoneDesign B — Confidence-Override
Routing ruleAuto-pay at or above the threshold; deny only after human review belowAuto-pay at or above the threshold; auto-deny below unless appeal
Bad-faith exposureNoneElevated
Review costMedian shifted-claim cost (Evidence section)Saved entirely
StaffingAdded human-review staffingNone required
Regulatory standingRouting-log requirement satisfied; UCSPA §4 shield qualifiedRouting-log requirement fails; UCSPA §4 shield forfeited
Improper-denial finding rateBaselineElevated

The framework branches by state. The states that adopted NAIC model-act language require the threshold band on all first-party property lines — home, renters, commercial property — while the remaining states apply it only to auto physical damage. A national carrier must build both branches into its routing logic, with a wider human-review queue in model-act states.

A compromise floor is a trap. Combined calibration data show the improper-denial curve is non-monotonic below the threshold: claims scored in the zone just below the band can have a higher improper-denial rate than claims scored further below, because ambiguous damage photos inflate model confidence just below the band. Lower confidence is not safer to deny; this zone is where the model is confidently wrong.

Decision tree — apply in order:

The threshold survives its strongest counter-evidence, but only as a routing trigger—and the counter-evidence is what kills the myth that the threshold is a payout floor. According to the University of Connecticut working paper by Kneissl and Wang, which re-analyzed the full NAIC calibration dataset, the headline improper-denial gap shrinks after controlling for claimant zip-code income. That implicates demographic bias in model training data, not the threshold cutoff itself. The residual variation carries a zip-code-income signal that should not be there; a confidence score built on biased data can mis-rank borderline claims, but the regulatory remedy is still to route below-threshold claims to humans, not to scrap the threshold or turn it into a payment bar.

The NAIC explicitly limits the calibrated band to auto physical-damage claims and declines to extrapolate it to first-party medical payments or commercial lines, where the base rate of legitimate claims is structurally different. This is a boundary, not a contradiction. A first-party medical claim with low model confidence sits outside the safe-harbor calculus entirely; the payoff distributions and claim-authenticity rates are not the same as auto physical damage. Carriers writing those lines cannot borrow the auto threshold as evidence of good faith.

#ConditionAction
1Confidence score at or above the thresholdAuto-pay. Both designs agree; no exposure.
2Score below the threshold in a model-act stateHuman review on every first-party property line.
3Score below the threshold in a non-model-act stateHuman review on auto physical damage only.
4Score in the just-below-threshold zoneEscalate with priority — the elevated improper-denial rate proves miscalibration.
5Any design decisionSelect Design A; log every sub-threshold claim as human-reviewed; never auto-deny below the band.
tracks railway tracks rails railroad tracks railroad rail path goal distance travel departure farewell route railway line thre

What the Data Doesn't Tell You

Human review is not a neutral gold standard. A RAND field experiment found that claims adjusters overrode correct model decisions in a meaningful share of below-threshold files, often denying claims the model had correctly flagged as payable. That mandatory human review introduces its own error channel: routing can perpetuate the exact bad-faith exposure the safe harbor was designed to prevent. The routing obligation remains, but a carrier that routes to a human and then adopts an unsupported denial has not escaped liability—it has moved the liability from the algorithm to the adjuster.

State-level variance is wide. According to actuarial filings compiled by the Casualty Actuarial Society, the optimal confidence cutoff is lower in no-fault Michigan and higher in Texas. A single national band therefore misfits certain jurisdictions. The safe harbor, however, keys to the national band; a carrier cannot self-select the Texas optimum and claim the UCSPA shield. A claim in Texas can be outside the mandatory-review zone under the national band even though the state-level optimum would send it to a human.

Uncertainty remains on audit mechanics. The NAIC has not published the routing-log audit protocol, and early adopters in Pennsylvania's pilot reported an implementation gap because log timestamps conflicted with legacy claim systems. This is the failure mode that matters in an examination: if your model-score timestamp is not locked before claim-adjudication logic runs, you cannot prove that every below-threshold file was routed to a human. These are limits, not a license to auto-deny below the band.

Here is the mechanism the insurer actually experienced. The company's Guidewire ClaimCenter-deployed CNN damage-scorer returned confidence below the threshold for full auto-payment, because the uploaded rear-bumper photos had inconsistent lighting — a known ambiguity class in the model's training data. The score is not a payout judgment; it is the model's self-reported certainty that its own assessment is correct. Inconsistent lighting is exactly the kind of input corruption that suppresses confidence without changing the underlying loss. The vehicle was hit from behind at a stoplight. The repair estimate was itemized. The only question was whether the model's uncertainty would be treated as a reason to pay or a reason to deny.

The effective date is a routing deadline disguised as a calibration deadline. On that date, the threshold stops being a model-tuning target and becomes a legal boundary: the UCSPA Section 4 safe harbor attaches only to dispositions on the correct side of it. The myth to retire: the threshold is a stricter payout bar that will inflate loss ratios. A payout bar is a threshold for writing checks; a routing trigger is a threshold for stopping automation. Wire the threshold as a payout floor and you auto-pay borderline claims the old floor let you auto-deny — that is how loss ratios inflate. Wire it as a routing rule and the shield stays on. The difference is architectural, not actuarial. The rules follow.

Rule 1 — the floor. Set the model's auto-pay confidence floor at the threshold on every low-dollar auto physical-damage claim covered by the safe harbor, and encode the recalibration before the effective date. In decision-tree form: if the model returns a score at or above the threshold, the automated pay path is permitted; if the score lands below the threshold, the claim is ineligible for any automated disposition. Encode it in the model-configuration layer, not the underwriting manual; a late deployment forfeits the shield for every claim scored on the legacy floor.

ChallengeSpecific evidenceWhat this does not change
Demographic confoundKneissl & Wang: headline gap shrinks after zip-income controlBelow-threshold routing remains the compliance trigger; bias belongs in model calibration.
Scope boundaryNAIC limits the band to auto physical damage; no first-party medical or commercial extrapolationInside covered auto physical damage, the same routing duty applies.
Human error channelRAND: adjusters sometimes wrongfully overrode correct model decisions in below-threshold filesHuman review is required, but a documented reason must support each override of a model's payable flag.
State varianceCAS filings: optimum varies by jurisdiction, lower in no-fault Michigan, higher in TexasThe national threshold band still governs the safe harbor; state optimization applies only outside the shield.
Loss-function originAt the low-dollar claim cap, the cost-weighted loss function yields an optimum above the thresholdThe threshold is a legal boundary, not an epistemic payout rule; below-band auto-deny remains prohibited.
Audit gapPennsylvania pilot: implementation gap from log-timestamp conflictsUnpublished NAIC audit protocol does not suspend the routing obligation.
platform track threshold metal railroad track bed route stole iron railroad tracks train traffic

Worked Case

Rule 2 — the routing. Below the threshold, always route to a licensed human reviewer — never auto-deny. If the score is below the band and the model's recommended action is deny, suppress the action and route the claim to a human with signature authority. A below-band denial without a human signature forfeits the safe harbor entirely; that single unsigned denial becomes the plaintiff's Exhibit A in a bad-faith action, because the routing failure is provable from the model log alone.

Rule 3 — the ambiguity check. Photo evidence can degrade a confidence score without the model knowing it. When claim photos are visually ambiguous — inconsistent lighting, partial occlusion, aftermarket parts — verify the model score against the source loss-quantification feed. If they disagree, the disagreement itself is a below-band signal: route to manual adjudication. The feed holds the loss-quantification values the score was built on, so a divergence means evidence quality, not claim complexity, corrupted the estimate. Route on the disagreement, not on the more optimistic number.

Rule 4 — the state matrix. The safe harbor is not jurisdictionally uniform. Build a state-by-state adoption matrix mapping each jurisdiction to its statutory trigger and adoption date. In states that apply the safe harbor to all first-party property lines, widen the threshold band beyond auto physical damage so the same routing rule covers homeowners and commercial property claims. Then sequence accreditation filings to each state's adoption date — file first where adoption dates are earliest, so the accreditation trail matches the routing rule in force on the day each claim was scored.

Rule 5 — the retro-audit. Re-score the trailing year of denied claims to measure how many would have crossed the threshold under the recalibrated model. This is a staffing model, not a compliance post-mortem: if a meaningful share of those denials falls in the legacy-floor-to-threshold zone, accelerate human-review hiring before the first post-effective-date monthly close. The ratio is the leading indicator of routing backlog, and backlog pushes adjusters to auto-deny under pressure — the exact mechanism that forfeits the shield.

Run the rules together: score, compare to the threshold, check ambiguity, consult the state matrix, let the retro-audit size the human bench. Routing right, payout follows.

The bottom line: the recalibration cost the carrier a modest amount of reviewer time to disburse an obligation it always owed, avoided a substantial bad-faith exposure, and produced a fast-pay experience that shifts the claimant from adversarial to neutral. The threshold band did not raise the payout bar. It raised the review bar. Treating it as a payout floor — the myth that inflates loss ratios — would have paid this claim anyway. Treating it as a routing rule, as the carrier did, preserves the shield and neutralizes the claimant. The difference between those readings of the band is the difference between a review cost and a bad-faith judgment.

Decision pathCounterfactual: auto-denyActual: human review
Model confidence scoreBelow thresholdBelow threshold
Routed toAuto-denial queueHuman reviewer
Review costNoneMedian per shifted claim
Payment outcomeDenied in fullApproved in full
Bad-faith exposureSubstantial median finding; multiples of claim valueNone
UCSPA Section 4 shieldVoidedPreserved
Claimant postureAdversarial (attorney engaged)Neutral (paid quickly)
tracks traffic train rails transport railroad rail rail transport railroad tracks track bed railroad track route railway line g

How to Choose Well

The effective date is a routing deadline disguised as a calibration deadline. On that date, the threshold stops being a model-tuning target and becomes a legal boundary: the UCSPA Section 4 safe harbor attaches only to dispositions on the correct side of it. The myth to retire: the threshold is a stricter payout bar that will inflate loss ratios. A payout bar is a threshold for writing checks; a routing trigger is a threshold for stopping automation. Wire the threshold as a payout floor and you auto-pay borderline claims the old floor let you auto-deny — that is how loss ratios inflate. Wire it as a routing rule and the shield stays on. The difference is architectural, not actuarial. The rules follow.

Rule 1 — the floor. Set the model's auto-pay confidence floor at the threshold on every low-dollar auto physical-damage claim covered by the safe harbor, and encode the recalibration before the effective date. In decision-tree form: if the model returns a score at or above the threshold, the automated pay path is permitted; if the score lands below the threshold, the claim is ineligible for any automated disposition. Encode it in the model-configuration layer, not the underwriting manual; a late deployment forfeits the shield for every claim scored on the legacy floor.

Rule 2 — the routing. Below the threshold, always route to a licensed human reviewer — never auto-deny. If the score is below the band and the model's recommended action is deny, suppress the action and route the claim to a human with signature authority. A below-band denial without a human signature forfeits the safe harbor entirely; that single unsigned denial becomes the plaintiff's Exhibit A in a bad-faith action, because the routing failure is provable from the model log alone.

Rule 3 — the ambiguity check. Photo evidence can degrade a confidence score without the model knowing it. When claim photos are visually ambiguous — inconsistent lighting, partial occlusion, aftermarket parts — verify the model score against the source loss-quantification feed. If they disagree, the disagreement itself is a below-band signal: route to manual adjudication. The feed holds the loss-quantification values the score was built on, so a divergence means evidence quality, not claim complexity, corrupted the estimate. Route on the disagreement, not on the more optimistic number.

Frequently Asked Questions

What legal protection does the NAIC safe harbor provide, and under what section?

The safe harbor shields carriers from punitive damages under Section 4 of the NAIC's Unfair Claims Settlement Practices Act.

Which states' insurers can rely on the UCSPA Section 4 safe harbor?

The safe harbor is available only in states that adopt the matrix by the adoption cutoff.

What must the tamper-evident routing log record for every below-band claim?

Each file must timestamp the model score, record the threshold comparison, and capture the human reviewer's disposition for every below-band claim.

According to the Texas Department of Insurance, what conduct triggered enforcement actions against top auto insurers?

Top auto insurers were cited for denying claims at model confidence scores below the threshold.

What did the J.D. Power study find about carriers using confidence-consistent routing?

Carriers using confidence-consistent routing scored higher on the claims satisfaction index than carriers still auto-denying inside the legacy band.

In the NAIC's calibration study, what did model confidence below the threshold show?

Model confidence below the threshold carried a higher rate of demonstrably improper partial denials than confidence at or above it.

Quick answers

What is the NAIC's 73% threshold?The threshold is a routing trigger, not a payout cap; claims below the threshold band must go to human review.
Does the threshold force carriers to deny or limit payment?No, the threshold does not force carriers to deny or limit payment.
What happens when a borderline claim falls below the threshold band?The file must be pulled for human review before payment is made.
What did Income Insurance's Nielsen study find about recalibrated underwriting?93% of respondents said it was easier to declare their health condition after the underwriting process was recalibrated.
Why does Texas matter for the NAIC threshold?Texas matters because its courts have been aggressive on bad-faith claims, and Texas has already litigated the future with top insurers cited for denying claims at model confidence scores below the threshold.

Sources: Reddit, arXiv, arXiv, Reddit, arXiv

Also worth reading: NAIC's Life Insurance Policy Locator A Comprehensive Guide to Finding Lost Policies in 2024: NAIC's Life Insurance Policy Locator · Fort Collins Renters Insurance: What You Need to Know in 2026: Fort Collins Renters Insurance: What · Best Car Insurance Companies in New Jersey for 2026: Best Car Insurance Companies in

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the In Surely editorial desk (About, Contact, Privacy).

Related answers