Four institutions' 2024 fair-lending attestations, scored under one shared CREST seed so their disparity intervals are directly comparable — without any institution sharing loan-level data. Synthetic demonstration.
Disparity = White approval rate − Black approval rate (pp). "Actionable" = 95% CI excludes zero. Click a row to open its report.
| Institution | Geography | W−B disparity | 95% CI | Status | Coverage |
|---|---|---|---|---|---|
A
Institution A — Large Urban (NYC Metro)
DEMOBANKA0000000000A1 |
NYC metro | +11.5pp | +8.1 … +15.0 | Actionable | 100% |
B
Institution B — Small Rural (Upstate NY)
DEMOBANKB0000000000B2 |
Upstate NY | +9.0pp | +4.0 … +13.9 | Actionable | 100% |
C
Institution C — Mid Suburban (Mixed NY)
DEMOBANKC0000000000C3 |
Mixed NY | +0.9pp | -4.8 … +6.7 | Indeterminate | 100% |
D
Institution D — Mid Suburban (Mixed NY)
DEMOBANKD0000000000D4 |
Mixed NY | +13.4pp | -1.5 … +28.4 | Indeterminate | 100% |
A real, modest gap with enough minority sample that the de-biased, cluster-robust 95% CI excludes zero. An examiner should investigate.
The largest raw (argmax-proxy) gap here, but STRATA's calibrated posterior is uncertain on those records and the de-biased 95% CI includes zero: a proxy artifact, not reliable evidence. Do not act on it.
2 estimate(s) are statistically actionable (95% CI excludes zero):
| Comparison | Outcome | Point (pp) | 95% CI (pp) | N ref/cmp | Actionability |
|---|---|---|---|---|---|
| White vs. Black | Approval rate | +11.5 | +8.1 … +15.0 | 1636 / 1075 | Actionable |
| White vs. Hispanic | Approval rate | +6.5 | +3.3 … +9.6 | 1636 / 1254 | Actionable |
| White vs. Asian | Approval rate | +1.9 | -1.8 … +5.5 | 1636 / 727 | Indeterminate |
1 estimate(s) are statistically actionable (95% CI excludes zero):
| Comparison | Outcome | Point (pp) | 95% CI (pp) | N ref/cmp | Actionability |
|---|---|---|---|---|---|
| White vs. Black | Approval rate | +9.0 | +4.0 … +13.9 | 1759 / 401 | Actionable |
| White vs. Hispanic | Approval rate | +2.9 | -3.4 … +9.1 | 1759 / 201 | Indeterminate |
| White vs. Asian | Approval rate | +0.7 | -7.1 … +8.4 | 1759 / 130 | Indeterminate |
No estimate is statistically actionable at the 95% level: every interval below includes zero.
| Comparison | Outcome | Point (pp) | 95% CI (pp) | N ref/cmp | Actionability |
|---|---|---|---|---|---|
| White vs. Black | Approval rate | +0.9 | -4.8 … +6.7 | 1785 / 276 | Indeterminate |
| White vs. Hispanic | Approval rate | -0.5 | -4.6 … +3.6 | 1785 / 617 | Indeterminate |
| White vs. Asian | Approval rate | +3.5 | -2.9 … +9.8 | 1785 / 197 | Indeterminate |
1 estimate(s) are statistically actionable (95% CI excludes zero):
| Comparison | Outcome | Point (pp) | 95% CI (pp) | N ref/cmp | Actionability |
|---|---|---|---|---|---|
| White vs. Black | Approval rate | +13.4 | -1.5 … +28.4 | 1453 / 875 | Indeterminate |
| White vs. Hispanic | Approval rate | +22.8 | +5.2 … +40.4 | 1453 / 393 | Actionable |
| White vs. Asian | Approval rate | +2.1 | -5.5 … +9.7 | 1453 / 146 | Indeterminate |
The real deployed STRATA (M8-BGT) model scored 20 fictional name+address probes, to confirm it behaves sensibly. Rows with iu ≥ 0.30 (high uncertainty) are highlighted.
| Name (fictional) | Neighborhood | wht | blk | his | asn | oth | Dominant | iu | Tier |
|---|---|---|---|---|---|---|---|---|---|
| Sean Sullivan | Upper East Side | 0.94 | 0.02 | 0.02 | 0.01 | 0.01 | white | 0.07 | BG |
| Emily Anderson | Upper East Side | 0.94 | 0.01 | 0.03 | 0.01 | 0.02 | white | 0.07 | BG |
| Gregory Meyer | Riverdale | 0.93 | 0.01 | 0.02 | 0.00 | 0.03 | white | 0.07 | BG |
| Deshawn Washington | Harlem | 0.00 | 0.95 | 0.02 | 0.00 | 0.02 | black | 0.05 | BG |
| Andre Jefferson | Bed-Stuy | 0.01 | 0.95 | 0.01 | 0.00 | 0.04 | black | 0.05 | BG |
| Keisha Booker | South Bronx | 0.00 | 0.97 | 0.02 | 0.00 | 0.01 | black | 0.03 | BG |
| Maria Garcia | Washington Hts | 0.07 | 0.00 | 0.91 | 0.00 | 0.01 | hispanic | 0.10 | BG |
| Carlos Rodriguez | Jackson Hts | 0.07 | 0.01 | 0.91 | 0.01 | 0.01 | hispanic | 0.11 | ZIP |
| Luis Morales | Bronx | 0.03 | 0.01 | 0.95 | 0.00 | 0.00 | hispanic | 0.05 | BG |
| Wei Chen | Flushing | 0.00 | 0.00 | 0.00 | 1.00 | 0.00 | asian | 0.00 | ZIP |
| Minh Nguyen | Sunset Park | 0.00 | 0.00 | 0.00 | 1.00 | 0.00 | asian | 0.00 | BG |
| Priya Patel | Jackson Hts | 0.00 | 0.00 | 0.00 | 0.99 | 0.01 | asian | 0.01 | ZIP |
| Jisoo Kim | Flushing | 0.00 | 0.00 | 0.00 | 1.00 | 0.00 | asian | 0.00 | ZIP |
| David Smith | Park Slope | 0.84 | 0.10 | 0.01 | 0.01 | 0.04 | white | 0.18 | BG |
| Jennifer Obrien | Upper West Side | 0.91 | 0.00 | 0.04 | 0.01 | 0.03 | white | 0.10 | BG |
| Wei Chen | Harlem | 0.01 | 0.00 | 0.01 | 0.90 | 0.08 | asian | 0.11 | BG |
| Deshawn Washington | Upper East Side | 0.01 | 0.94 | 0.02 | 0.00 | 0.02 | black | 0.05 | BG |
| Maria Garcia | Riverdale | 0.11 | 0.01 | 0.85 | 0.00 | 0.03 | hispanic | 0.15 | BG |
| John Smith | Washington Hts | 0.58 | 0.31 | 0.07 | 0.01 | 0.03 | white | 0.48 | BG |
| Aisha Okafor | Bed-Stuy | 0.00 | 0.95 | 0.01 | 0.00 | 0.04 | black | 0.06 | BG |
STRATA proxies race/ethnicity on records without self-reported demographics, from applicant name + property address. It returns a five-class probability vector, a dominant race, an index of uncertainty (iu), and an address lookup-quality tier (block-group+tract, tract, ZCTA, or name-only). Raw geographic identifiers never leave the STRATA container — only anonymized cohort tokens are emitted.
CREST estimates each disparity with a de-biased, Fisher-consistent weighted method-of-moments estimator that inverts the proxy's misclassification using the calibrated posterior P(true race | proxy) — correcting the attenuation (regression dilution) that biases a plain plug-in of the proxy toward zero. The 95% interval is a tract-cluster-robust interval anchored to the property census tract, so it reflects proxy + finite-sample + within-tract correlation and tests for a systematic disparity rather than a chance in-sample gap. It is closed-form and deterministic, keyed to a published seed, so results are reproducible and comparable across institutions.
Reading the result. "Actionable" means the 95% CI excludes zero — a statistically stable finding. "Indeterminate" means the CI includes zero — the point estimate is not reliable and should not be acted on alone. A large gap can be indeterminate (a proxy artifact); a small gap can be actionable.
Screening vs. confirmatory analysis. These disparity estimates are marginal (unconditional) statistics — the correct tool for identifying which institutions warrant closer examination, analogous to a raw HMDA disparity ratio. The estimator assumes lending outcomes are independent of the proxy conditional on true race (Y ⊥ proxy | true race). Geography can challenge this: neighborhood factors that affect lending (property values, local income, LTV distributions) are correlated with race through census tract, and the proxy encodes the tract. Sensitivity analysis shows that under race-correlated geography confounding (the realistic case in segregated markets), the marginal estimator may overstate the true race effect by approximately 1–3.5 percentage points depending on the strength of the neighborhood–race correlation. This does not disqualify the screening statistic — every standard HMDA disparity analysis faces the same issue — but it means these estimates flag where to look, not the final adjudicated effect size. For confirmatory analysis controlling for legitimate underwriting factors (LTV, property value, loan purpose), use CREST's covariate-adjusted mode, which recovers the within-geography race effect stably across confounding levels.
Method reference: arXiv:2504.21259.
All data is synthetic. No real lender, applicant, name, or address appears; LEIs are anonymized. The records are assembled from real, auditable public components so they read plausibly:
| Component | Source | Role |
|---|---|---|
| Base LAR records | CFPB hmda-test-files 2024 (U.S. public domain / CC0) |
Schema-valid 2024 filing fields, valid code distributions, and real FFIEC 2024 census geography (tract/county/income). Contains no applicant name or usable street address — those are not in HMDA. |
| Applicant race | U.S. Census 2020 P.L. 94-171, table P2 | Each record's race is drawn from the real racial composition of its census tract (5,411 NY tracts), so names match the geography. |
| Name & property address | Synthesized (aequumAI) | Drawn to be consistent with the record's tract demographics. |
| Embedded disparities | Designed (aequumAI) | Approval/denial rates by race, and Institution D's geo-proxy artifact, create the four review scenarios. |
| STRATA / CREST | aequumAI · arXiv:2504.21259 | Proxy scoring and geographic-uncertainty confidence intervals. |
Scope: geography is New York only in this example (the vendored Census tract table covers NY's 5,411 tracts). The disparities are illustrative, designed to exercise the actionable-vs-indeterminate distinction — not empirical findings.