Skip to content
Daily AI Intel

AI in Insurance · Regulation & Fairness in Insurance AI

What is proxy discrimination and why does it matter for insurance AI

Proxy discrimination occurs when a seemingly neutral factor in an insurance AI model closely correlates with a protected characteristic like race, producing discriminatory outcomes even without directly using that characteristic — a significant concern since sophisticated models can find many such subtle correlations.

Key takeaways

  • Proxy discrimination happens when a neutral-seeming factor closely correlates with a protected characteristic, producing indirect discriminatory effects.
  • This can occur even when a model never directly uses a protected characteristic like race as an explicit input.
  • Sophisticated AI models can identify and rely on many subtle correlations that simpler traditional models might never have specifically captured.
  • Regulators increasingly require specific testing for proxy discrimination given how easily it can arise in complex, data-rich AI models.

Discrimination Without Directly Using a Protected Characteristic

Proxy discrimination occurs when a seemingly neutral factor used in an insurance AI model closely correlates with a protected characteristic — such as race, gender, or national origin — effectively producing discriminatory outcomes even though the model never directly uses that protected characteristic itself as an explicit input.

How Proxy Discrimination Actually Happens

Many data points that seem entirely unrelated to a protected characteristic on their surface can, in practice, correlate closely with that characteristic due to broader societal patterns — certain geographic areas, for example, might have a demographic composition strongly correlated with race or ethnicity, meaning a model using detailed location-based factors could inadvertently produce pricing or underwriting outcomes that closely track racial demographics, even without race ever being directly used as an input.

Why This Is a Particularly Significant Concern for Sophisticated AI Models

More sophisticated AI models, capable of identifying and using subtle patterns across very large, complex datasets, can potentially identify and rely on many more of these indirect correlations than a human actuary working with a simpler, more traditional model might ever specifically notice or intentionally build into a pricing formula, making proxy discrimination a genuinely more significant, harder-to-detect risk as models become more sophisticated and data-rich.

Why the Absence of a Directly Used Protected Characteristic Doesn’t Resolve the Concern

A common but mistaken assumption is that simply avoiding the direct use of protected characteristics as explicit model inputs is sufficient to avoid discriminatory outcomes — proxy discrimination specifically demonstrates why this assumption is incomplete, since a model can still produce genuinely discriminatory effects through indirect correlations, even while never directly referencing the protected characteristic itself.

Most legal frameworks addressing discrimination, including in insurance specifically, recognize that facially neutral practices producing discriminatory effects can still constitute illegal discrimination under disparate impact theory, meaning the technical absence of a directly used protected characteristic doesn’t automatically make an outcome legally acceptable if proxy discrimination is genuinely occurring.

Why This Has Prompted Specific Regulatory Testing Requirements

Given this genuine, documented risk, a growing number of insurance regulators have introduced or proposed specific requirements for insurers to test AI models specifically for proxy discrimination effects, going beyond simply confirming that protected characteristics aren’t directly used, to more thoroughly examine whether the model’s actual outcomes still correlate problematically with protected characteristics through indirect pathways.

Bottom Line

Proxy discrimination occurs when a seemingly neutral factor in an insurance AI model closely correlates with a protected characteristic, producing indirect discriminatory effects even without directly using that characteristic as an input — a particularly significant concern for sophisticated AI models capable of identifying many subtle correlations across large datasets, which is why regulators increasingly require specific proxy discrimination testing beyond simply confirming protected characteristics aren’t directly used.

Frequently asked questions

Can you give a hypothetical example of how proxy discrimination might occur?

A commonly cited hypothetical example involves geographic location data — if certain neighborhoods are strongly correlated with a specific racial or ethnic composition, using detailed location-based pricing factors could indirectly produce discriminatory pricing outcomes correlated with race, even though race itself was never directly used as a factor in the model.

Is proxy discrimination illegal even though the protected characteristic isn't directly used?

Generally yes — most legal frameworks addressing discrimination recognize that facially neutral practices producing discriminatory effects can still constitute illegal discrimination, meaning the absence of a directly used protected characteristic doesn't automatically make an outcome legally acceptable if proxy discrimination is occurring.

Sources

  1. [1]State insurance regulation resources — National Association of Insurance Commissioners
  2. [2]AI Risk Management Framework — National Institute of Standards and Technology
ET

Written by Editorial Team

Last updated July 29, 2026

Get one well-sourced answer a week

No spam. Unsubscribe anytime.