AI credit fair-lending review guide

How should lenders evaluate AI credit decisions for fair-lending risk?

A useful review does more than calculate a disparity metric. It defines the decision population, tests outcomes across documented cohorts, investigates the context behind material findings, checks whether stated decision reasons are supported, and preserves a record a qualified reviewer can reproduce.

Avarent guide · reviewed 2026-08-13

Direct answer

An AI credit-decision review should answer five questions: What decision was made? Which population and comparison groups were evaluated? Where did outcomes diverge? What data and candidate reasons help explain the finding? What evidence and human action were recorded afterward?

DecisionWhich workflow, outcome, product, and period are in scope?
PopulationWhich applicants or accounts are included, excluded, and compared?
OutcomeWhich rates, gaps, thresholds, and uncertainty checks were calculated?
ContextWhich inputs, policy choices, and alternative explanations require review?
EvidenceCan a qualified reviewer reproduce the finding and see the next action?

What should be measured?

Start with the actual decision outcome, not a generic model score. Depending on the workflow, useful measures can include favorable-outcome rates, adverse impact ratio, statistical parity difference, pricing or term differences, exception patterns, and the stability of results across time windows. Every measure should retain its population definition, comparison group, denominator, exclusions, configuration, and interpretation limits.

A threshold crossing is a screening signal. It does not identify cause or establish a legal conclusion by itself. Sample size, missing or inferred demographic attributes, reference-group selection, policy context, and statistical uncertainty can materially change the interpretation.

See Avarent's formulas and interpretation limits.

What belongs in a reviewable evidence packet?

  • The evaluation question, workflow, outcome, population, comparison group, and time window.
  • The data fields used, excluded, derived, missing, or inferred.
  • The metric definitions, thresholds, configuration, software version, and calculation output.
  • Material findings with supporting records and alternative explanations to investigate.
  • Candidate adverse-action reasons and the fields supporting or contradicting them.
  • Reviewer identity, review date, disposition, next action, and unresolved limitations.
  • An export manifest that lets another reviewer locate and reproduce the artifacts.

Inspect Avarent's four-page synthetic example.

A practical review sequence

  1. 01

    Define one decision question

    Name the workflow, outcome, population, comparison, time window, and decision the review must support.

  2. 02

    Reproduce a baseline

    Confirm that the supplied inputs and agreed method reproduce a known measure within an accepted tolerance.

  3. 03

    Investigate material findings

    Review population design, sample size, inputs, policy context, reason codes, and plausible alternative explanations.

  4. 04

    Record the human decision

    Document what was concluded, what remains uncertain, who owns the next action, and when the finding will be revisited.

Who should participate?

Compliance and fair lendingDefine the policy question, review implications, and own escalation.
Model riskChallenge methodology, data, configuration, validation boundaries, and reproducibility.
Lending and productExplain the workflow, policy intent, exceptions, and operational context.
Data and engineeringVerify source fields, transformations, versions, lineage, and access.
Legal counselProvide qualified legal interpretation when the facts require it.
Security and procurementReview data handling, provider risk, contracting, retention, deletion, and exit.

Primary sources

This guide is an operational framework, not legal advice. Review the underlying requirements and guidance directly: