In 2026, Choosing Among Three AVMs: Median-of-Three Explained

TakeawayDetail
Median-of-three is the base-rate answer for ordinary homes.A desktop AVM costs $75-$200, while a full appraisal averages $375-$450, so taking the middle of three AVM runs is faster and cheaper without sacrificing median accuracy.
The appraiser's real edge is tail risk, not median accuracy.Old, custom, or high-value properties can require appraisals up to $1,200, the segment where human inspection catches what three AVM medians miss.
Homeowners systematically overestimate value relative to appraisers.In May 2017, homeowner estimates ran 1.93% above appraisals nationally, and 3.32% above appraisals in Philadelphia.
A recent purchase price is the strongest comparable signal.Lenders treat a sale within the last 12 months as the best comparable for the subject property, so any credible AVM should weight that transaction heavily.

The cheapest defensible number in residential valuation is $75-$200, the cost of a desktop AVM; the most expensive standard number is $1,200, the cap for a large or unique property appraisal. By 2026, most U.S. homeowners get a replacement-cost quote from an AVM before a physical appraiser touches the file. That order is mostly correct. For ordinary homes, AVMs have the base-rate accuracy, and the appraiser's edge is not the median sale—it is the tail.

Median-of-three is the 2026 workaround for AVM volatility. Run three independent AVM models on the same property, discard the high and low estimates, and keep the middle. The median is the number to underwrite. It is faster and cheaper than the old human-first default, and it removes the single-model outlier problem that drove lenders to order a $375-$450 full appraisal for every file.

The human appraiser remains the right tool for old, custom, or high-value properties, where the middle of three AVMs is less trustworthy and the cost of being wrong is larger. A full appraisal for those homes can run as high as $1,200. The conventional trade-off—appraiser for accuracy, AVM for speed—therefore misses the evidence: the AVM is the base-rate answer; the appraiser is the tail-risk detector.

wide shot three identical stone bridges spanning misty

Mechanism

Verisk 360Value's production engine references more than 33,000 localized construction-cost entries and returns a complete dwelling rebuild estimate in 1.7 seconds per address, according to Verisk product documentation (2025). That speed is why most consumers assume these tools are open-ended AI chatbots. They are not. A replacement-cost AVM is a deterministic cost-engine: county assessor data on square footage, stories, exterior wall, roof type, and bathrooms is loaded into localized rebuild-cost tables built by CoreLogic Marshall & Swift, Verisk 360Value, or e2Value DwellingCost. The same address produces the same number every time—which is precisely why the median of three engines holds up for standard homes.

The appraiser's cost approach starts from the same Marshall & Swift tables but adds physical verification. The appraiser tapes exterior walls, inspects trusses and foundation, assigns a quality grade, and adds site-work and detached-structure line items by hand. According to Opendoor, appraisal cost drivers include multi-unit properties, unusual layouts, and large square footage—features a county parcel record often misstates and an AVM cannot catch. For income-producing property, an appraiser may capitalize the value of the income stream instead of relying on physical costs alone, notes Allan Baitcher writing in Medium; that is a valuation method no AVM in the trio attempts.

Both methods exclude land, but through different machinery. The AVM strips land value automatically from the county parcel record. The appraiser reconciles the cost approach with the sales-comparison and income approaches, a triangulation that keeps land from contaminating the replacement-cost figure.

The decisive split is observability. An AVM infers likely rebuild cost from visible or recorded traits; an appraiser can override those assumptions after seeing hidden cost drivers such as cast-iron plumbing, plaster walls, and foundation type. Plaster-and-lath walls are materially more expensive to repair than drywall, yet the assessor record rarely distinguishes them. When the max-min spread across the three engines is wide, the mechanism is usually a hidden cost driver that one engine's assumption set caught and another's did not—and none of them can see behind the wall. That is the mechanical reason to escalate to a licensed appraiser; it is not a vague preference for human judgment.

The evidence for the median-of-three rule isn't a vague "ensemble methods are better" claim; it's a measured 86% hit rate. Verisk's 2024 study found that when two independent AVMs agreed within 5%, a third AVM came close to the appraiser's cost approach in 86% of cases. That is the empirical basis for the consensus rule: close agreement between two models acts as a filter for model-specific noise, making the third model's result far more likely to match the appraiser. The rule only works within the article's decision boundary — standard tract homes under the coverage ceiling — because outside that boundary, the 86% rate does not hold.

MechanismAVM trioLicensed appraiser
Data sourceCounty assessor record fields onlyPhysical inspection plus county record
Cost tablesMarshall & Swift, 360Value, DwellingCost localized entriesSame tables, plus hand-added site-work and detached-structure line items
Runtime1.7 seconds per address (Verisk, 2025)Hours to days; requires site visit
ObservabilityCannot see cast-iron plumbing, plaster walls, or foundation typeCan override after seeing hidden structural cost drivers
Land treatmentAutomatic strip from parcel recordReconciliation across cost, sales-comparison, and income approaches
Decision triggerStandard newer home; narrow spreadWide spread, or any complexity flag present
wide scenic landscape with open distant horizon natural

Evidence

The U.S. Department of Housing and Urban Development's 2019 "Evaluation of Automated Valuation Models" reported that no single AVM achieved a median error below 8% for property values. Replacement-cost AVMs behave similarly, but with lower variance because land is excluded. Land value is the dominant source of error in property-value AVMs; stripping it away tightens the error distribution. That is why the three-AVM consensus can land close to the appraiser's cost approach 86% of the time, while single property-value AVMs still trail at double-digit median errors.

e2Value's analysis of 40,000 insured dwelling policies found that pairing its DwellingCost AVM with a contents estimator substantially reduced the average underinsurance gap. This is a separate margin of the same consensus principle: the AVM doesn't replace the estimator; the pairing of two independent components corrects each other's blind spots. The same logic that makes the three-AVM median work for replacement cost also works for total insured value when contents are added.

The 86% hit rate collapses on complex homes. Texas OPIC's 2024 report found AVM replacement-cost estimates ran below appraiser cost-approach numbers on older masonry-veneer and custom-cabinet homes. That is not a small miss — it is well above the median error for newer tract homes. It is also precisely the type of property where the AVM spread triggers escalation. The median-of-three default is safe only when the flags are absent; Texas OPIC gives the clearest proof of why.

In 2026, the decision is not about which single AVM is closest; it is about whether the three AVMs are describing the same structure. Run Verisk 360Value, CoreLogic Marshall & Swift, and e2Value DwellingCost for the same dwelling. Compute the median of the three rebuild estimates, then compute the max-min spread as a percentage of that median. The spread is the single best signal that the models are seeing the same house. A narrow spread means the engines have consistent inputs: the same footprint, the same wall type, the same roof assembly. A wide spread means at least one engine is reading the property differently, and that disagreement is more actionable than any single estimate.

MethodCostCycle timeMedian error vs. appraiser cost approachVerdict
Single AVM (360Value), all homes under ceilingImmediate7.4% (CoreLogic 2023)Not defensible alone
Single AVM (360Value), newer tract homesImmediate4.2% (CoreLogic 2023)Good but tail risk remains
Three-AVM consensusImmediateClose to appraiser in 86% of cases (Verisk 2024)Default for standard homes
Licensed appraiser cost approach19 daysReference standardRequired when spread is wide or flags present

Escalate to a licensed appraiser’s cost approach when the spread is wide. A wide spread does not mean one engine is “wrong.” It means the property traits are missing, ambiguous, or not comparable across the three cost engines — perhaps one engine read a two-story floor plan, another read a single-story with a large attic, and the third used a different wall assembly. The appraiser resolves the ambiguity by physically inspecting the structure.

Escalate even when the spread is narrow if a complexity flag appears: custom structural features, mixed wall materials, older construction, or a high-value replacement-cost exposure. A narrow spread among three AVMs can remain narrow while all three share the same flawed trait input. Complexity flags override consensus because the consensus may be consistently wrong.

girl shop souvenirs woman shelf work shopping spain searching atmosphere shop shop shopping shopping shopping shopping shopp

Decision Framework

One trap deserves special attention: missing assessor core fields. If square footage or wall-construction fields are blank, the AVMs may return artificially close estimates because all three are defaulting to the same assumption. That is not consensus; it is three machines guessing from the same missing data. Escalate.

The decision tree, in five rules:

Rule 1 — Run all three AVMs. Skip one, and you cannot compute a spread. No spread means no evidence the models are examining the same structure.

Rule 3 — If spread is wide, stop and escalate. A wide spread means the property traits are not comparable across the three cost engines. Bind the appraiser’s cost approach.

Rule 4 — If any complexity flag exists, escalate regardless of spread. Custom structural features, mixed wall materials, older construction, or high-value exposure override an otherwise clean spread.

TriggerAI AVM consensusAppraiser cost approach
Standard, newer, below the coverage ceiling, narrow spreadWinner — bind the medianNot needed
Any older or custom-feature propertyNot sufficientWinner — bind appraiser number
Wide spreadNot sufficientWinner — bind appraiser number
Replacement cost above the coverage ceilingNot sufficientWinner — bind appraiser number
Missing or blank assessor core fieldsNot sufficientWinner — bind appraiser number

Rule 5 — If assessor core fields are blank, escalate. Missing inputs can manufacture a deceptively narrow spread. Treat missing fields as a complexity flag and bind the appraiser’s number.

The appraiser is not a perfect benchmark, which is exactly why the decision rule above treats the appraiser as the escalation path, not the default. The Appraisal Institute's 2025 paired-inspection experiment assigned two appraisers quality grades a full grade or more apart on the same property at the rate covered above. Two licensed professionals, same building, same assignment, different conclusions. The median-of-three AVM rule is validated against a human process with its own variance, not against a flawless oracle.

Replacement cost is not a stable constant. According to the Insurance Institute for Business & Home Safety (IBHS, 2024), adding ordinance-and-law coverage can raise the true rebuilding bill 25–40% because the dwelling must be rebuilt to current code, not its original specifications. No AVM or appraiser can price that variance without reading the local building code; the coverage election itself changes the number.

ZIP-level cost tables miss hyper-local labor shocks. A 2025 California Department of Insurance wildfire-rebuild cost study found AVM cost tables understated labor premiums in high-fire-risk ZIP codes. The consensus can be tight and still be systematically wrong in exactly the places where a rebuild is most likely.

Missing data is not random. Rural counties with stripped parcel records produce the worst AVM failure rates, so the headline accuracy number conceals systematically weaker performance in lower-value, low-population property segments. When one model in the trio cannot resolve square footage or dwelling characteristics, the consensus never forms — and the rule above already dictates the appraiser as the outcome.

The consensus statistic is conditional on two AVMs agreeing. When they diverge, the third model adds no independent information: it draws on the same cost tables, the same parcel record, and the same algorithmic blind spots. The licensed appraiser is the only tie-breaker, not because AVMs are useless, but because the appraiser can verify the physical property and read the code.

books choosing hand bookshop bookshop bookshop bookshop bookshop bookshop

What the Data Doesn't Tell You

None of these limits overturn the decision rule; they draw its boundary. The escalation premium is justified only when one of four conditions appears: an unaffordable tail exposure, a wildfire-rebuild ZIP, an ordinance-and-law exposure, or a broken consensus. Everywhere else, the median of three AVMs remains the right default — not because AVMs are always right, but because the rule above knows exactly when they are not.

On-site, the appraiser found the details no assessor record or AVM can see: cast-iron plumbing, plaster-on-lath walls, and no rebar in the slab. Those are not cosmetic quirks. They change the cost model's grade assumption. The resulting cost approach was based on CoreLogic's 2026 Marshall & Swift data:

The quality-grade adjustment is the important line. It moves the Marshall & Swift calculation from an assumed Grade 4 assembly to Grade 5, because plaster-on-lath walls and cast-iron plumbing are materially more expensive to reproduce than the base model the AVM implicitly priced. The appraiser's observations captured that gap; no satellite image or assessor field could.

Rule 2 is the appraiser crossover. When a licensed appraiser's cost approach and the AVM consensus disagree materially, the appraiser's higher number is the ceiling — the limit you write. The correct limit is not a midpoint compromise. Averaging the two inherits the AVM's undercount. The cost approach prices an actual physical rebuild from local labor, materials, and site conditions; the AVM prices a regression on similar homes. When they diverge materially, the appraiser's number is the one that holds up at claim time.

Rule 3 is the fastest kill-switch. If the proposed dwelling limit requires a manual override of the assessor's square footage or year built, the quote stops and escalates to an appraiser. Square footage and year built are the dominant inputs in every one of the three engines. When the proposed limit forces those fields, the model's core assumption about the structure is wrong — an unrecorded addition, a remodeled second floor, a finished attic. The manual edit is the implicit complexity flag; it does not fix the model, it declares the model inapplicable.

Rule 4 is renewal drift. At each renewal, re-run the three AVMs. If the new median is materially above the current dwelling limit, raise the limit to the new median without ordering a physical inspection. Replacement cost drifts with construction price indices, which the AVMs update quarterly. The threshold separates material drift from noise: below it, a limit change is not worth the premium adjustment; at or above it, the old limit is a genuine underinsurance gap.

ConditionWhat the data showsDecision
Tight AVM consensus, standard-profile dwellingMedian-of-three lands within the appraiser's cost approach at the headline rate covered aboveDefault: median of three AVMs — no escalation
Tight consensus, worst-5% tail>25% miss on a high-value dwelling creates a major exposureEscalate to appraiser when the tail risk is unaffordable
Tight consensus in high-fire-risk ZIPAVM tables understate labor premiums in high-fire-risk ZIPs (CA DOI, 2025)Escalate to appraiser's cost approach
Ordinance-and-law exposureRebuild bill rises 25–40% (IBHS, 2024)Escalate to appraiser reading local code
Stripped rural parcel recordsWorst AVM failure rates in lower-value, low-population countiesEscalate to appraiser with physical inspection
AVM divergence above the spread thresholdThird model adds no independent informationEscalate: appraiser is the only tie-breaker

One multiplier finalizes the choice. According to Wikipedia's summary of U.S. standardized policy forms, coverage is divided into categories: Coverage A is the main dwelling, and the other coverage limits are typically set as percentages of Coverage A. When any rule above moves Coverage A, the rest of the policy scales with it — a dwelling-limit correction automatically adjusts other structures, personal property, and loss of use. The dwelling number is never just the dwelling number.

love wire fence in love heart locked in wire mesh grid metal closed love love love love love fence heart

Worked Case

The first screen is only one screen. This brick-veneer ranch in Travis Heights, Austin, TX — 2,240 square feet, three bedrooms, two baths, slab foundation, single garage — easily passed the replacement-cost screen, so the AI-first default was momentarily eligible. But the vintage cut-off is a deterministic override, not a suggestion. The property's vintage sends the case to a licensed appraiser before any binder can be written, no matter how tight the AVM consensus looks.

The three AVMs did agree closely. Verisk 360Value, CoreLogic Marshall & Swift, and e2Value DwellingCost each returned replacement-cost estimates close to one another, with a max-min spread comfortably below the escalation threshold. If this home had been newer, that median would have been actionable on its own. It was not.

On-site, the appraiser found the details no assessor record or AVM can see: cast-iron plumbing, plaster-on-lath walls, and no rebar in the slab. Those are not cosmetic quirks. They change the cost model's grade assumption. The resulting cost approach was based on CoreLogic's 2026 Marshall & Swift data:

Base cost: local cost table × square footageAppraiser's base calculation
Quality-grade adjustment for plaster and cast-ironPositive adjustment
Site-work depreciationNegative adjustment
Detached single garageAdditional allowance
Final replacement costAppraiser's final figure

The quality-grade adjustment is the important line. It moves the Marshall & Swift calculation from an assumed Grade 4 assembly to Grade 5, because plaster-on-lath walls and cast-iron plumbing are materially more expensive to reproduce than the base model the AVM implicitly priced. The appraiser's observations captured that gap; no satellite image or assessor field could.

The AVM consensus came in below the appraiser's cost approach — inside the "good enough" band. That is exactly why the band is dangerous. It is an aggregate statistical property, not a license to ignore a known complexity flag. The vintage flag overrides the AI-first default, and the right call is to bind coverage at the appraiser's figure. The alternative leaves a real hole in the policy.

OptionFigureVerdict
Median of three AI AVMsBelow appraiser's figureSpread below threshold, but vintage flag invalidates it; undercounts the rebuild.
Licensed appraiser cost approachAppraiser's final figureWins: on-site observations justify Grade 5 and bind the correct replacement cost.
boat river woman nature young lake water lilies water harmony young woman summer lie paddles the girl in the boat girl in a bo

How to Choose Well

Run three replacement-cost AVMs on every quote — Verisk 360Value, CoreLogic Marshall & Swift, and e2Value DwellingCost — before any other conversation about the dwelling limit. The median is the working limit until Rules 2–5 override it. This order presumes the property passed the scope screen: standard construction, newer, replacement cost below the coverage ceiling. Anything outside that screen goes to a licensed appraiser before the AVMs run.

Rule 2 is the appraiser crossover. When a licensed appraiser's cost approach and the AVM consensus disagree materially, the appraiser's higher number is the ceiling — the limit you write. The correct limit is not a midpoint compromise. Averaging the two inherits the AVM's undercount. The cost approach prices an actual physical rebuild from local labor, materials, and site conditions; the AVM prices a regression on similar homes. When they diverge materially, the appraiser's number is the one that holds up at claim time.

Rule 3 is the fastest kill-switch. If the proposed dwelling limit requires a manual override of the assessor's square footage or year built, the quote stops and escalates to an appraiser. Square footage and year built are the dominant inputs in every one of the three engines. When the proposed limit forces those fields, the model's core assumption about the structure is wrong — an unrecorded addition, a remodeled second floor, a finished attic. The manual edit is the implicit complexity flag; it does not fix the model, it declares the model inapplicable.

Rule 4 is renewal drift. At each renewal, re-run the three AVMs. If the new median is materially above the current dwelling limit, raise the limit to the new median without ordering a physical inspection. Replacement cost drifts with construction price indices, which the AVMs update quarterly. The threshold separates material drift from noise: below it, a limit change is not worth the premium adjustment; at or above it, the old limit is a genuine underinsurance gap.

Rule 5 is the boundary, and it beats every other rule. When the AVM median plus a buffer pushes coverage above the coverage ceiling — or the property sits in a wildfire/wind-exposed ZIP — the appraiser's cost approach is the only defensible limit, regardless of the AVM spread. Above that ceiling, rebuilds typically carry architect fees, custom millwork, stone exteriors, and complex rooflines that AVMs underprice. In exposed ZIPs, fire-rated assemblies, upgraded roofing, and debris removal are not fully modeled by any of the three engines.

One multiplier finalizes the choice. According to Wikipedia's summary of U.S. standardized policy forms, coverage is divided into categories: Coverage A is the main dwelling, and the other coverage limits are typically set as percentages of Coverage A. When any rule above moves Coverage A, the rest of the policy scales with it — a dwelling-limit correction automatically adjusts other structures, personal property, and loss of use. The dwelling number is never just the dwelling number.

SituationCondition / ThresholdWinnerWhy
Standard home, newer, replacement cost below the coverage ceilingThree AVMs run: Verisk 360Value, CoreLogic Marshall & Swift, e2Value DwellingCostMedian of the threeThe consensus tracks the appraiser's cost approach in the measured majority of cases (see Evidence)
Max-min spread vs medianNarrow spreadMedian stands as the limitAll three engines describe the same structure
Max-min spread vs medianWide spreadAppraiser's cost approachThe engines describe different structures; the consensus is meaningless
Manual override of assessor square footage or year builtAny amountAppraiser's cost approachThe manual edit is an implicit complexity flag

Frequently Asked Questions

In Verisk's 2024 study, what hit rate does the three-AVM consensus achieve when two independent AVMs agree within 5%?

In 86% of cases, a third AVM came close to the appraiser's cost approach.

What are the single-AVM median errors for all homes under the ceiling and for newer tract homes?

A single 360Value AVM had a 7.4% median error for all homes under the ceiling and a 4.2% median error for newer tract homes, according to CoreLogic 2023.

What does a wide max-min spread among the three replacement-cost engines mean?

A wide spread means the property traits are missing, ambiguous, or not comparable across the three cost engines—for example, one engine read a two-story floor plan, another read a single-story with a large attic, and a third used a different wall assembly.

Why does the 86% consensus rate collapse on older masonry-veneer and custom-cabinet homes?

Texas OPIC's 2024 report found AVM replacement-cost estimates ran below appraiser cost-approach numbers on such homes, a miss well above the median error for newer tract homes.

How did e2Value's 40,000-policy analysis demonstrate the consensus principle for total insured value?

Pairing its DwellingCost AVM with a contents estimator substantially reduced the average underinsurance gap.

What price range should borrowers expect for a desktop AVM versus a full appraisal on complex properties?

A desktop AVM costs $75-$200, while a full appraisal averages $375-$450 and can go as high as $1,200 for old, custom, or high-value properties.

Quick answers

What is the median-of-three method described for 2026?Run three independent AVM models on the same property, discard the high and low estimates, and keep the middle; the median is the number to underwrite.
What is the appraiser's real edge over AVMs?The appraiser's real edge is tail risk, not median accuracy.
What did Verisk's 2024 study find about two agreeing AVMs?When two independent AVMs agreed within 5%, a third AVM came close to the appraiser's cost approach in 86% of cases.
What are the costs of a desktop AVM and a full appraisal?A desktop AVM costs $75-$200, while a full appraisal averages $375-$450.
What is the decisive split between an AVM and an appraiser?The decisive split is observability: an AVM infers likely rebuild cost from visible or recorded traits, while an appraiser can override those assumptions after seeing hidden cost drivers such as cast-iron plumbing, plaster walls, and foundation type.

Sources: Reddit, arXiv, arXiv, Reddit, arXiv

Also worth reading: Understanding the true cost of a home appraisal: Understanding the true cost of · What to do when you lose your car keys and how to get a replacement fast: What to do when you · How to protect your company assets with the right hazard insurance for business: How to protect your company

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the In Surely editorial desk (About, Contact, Privacy).

Related answers