1990 U.S. Census · 20,640 block groups

California Housing Atlas

Explore every district in the California housing dataset, see which variables actually move price, estimate a district's value with a random forest running in the page, and check how far that model can be trusted.

Districts
20,640
block groups, 600–3,000 residents each
Median home value
$179,700
IQR $119,600 – $264,725
Median household income
$35,300
dataset stores this as 3.53 ($10k units)
Best model error
±$50,462
random forest, held-out test RMSE
Explore

Where the money is

Each dot is one district, placed at its longitude and latitude and shaded by median home value. Filter the population down and the map answers questions the summary statistics cannot — like how much of California's high-value housing is within sight of water.

Median value $15k $500k Hover a district for detail · click to load it into the estimator
Districts shown
20,640
Median value
$179,700
Median income
$35,300
Inland share
31.7%
0.00
15.00
$0k
$500k
1
52
Ocean proximity

Home values were truncated at $500,001 when the census was published, so 992 districts (4.8%) sit on that ceiling as a hard line of dots. Any model trained on this data systematically under-predicts genuinely expensive housing.

Correlate

What actually predicts price

Pearson correlation across all 20,640 districts. Pick any cell to plot that pair — the scatter usually shows why a strong-looking coefficient is or isn't trustworthy.

negative none positive

Median income is the single strongest linear predictor of home value (r = 0.688). The flat line of points along the top of that scatter is the $500,001 cap.

Predict

District value estimator

A 24-tree random forest, trained on 16,512 districts and exported into this page, runs on every keystroke. Nothing is sent anywhere — the trees are evaluated in your browser.

$35,300
29 yr
5.4
20.3%
2.8
Location
Predicted median home value
$0
Comparable districts in the real data

The forest never saw the test districts during training, and its held-out error is ±$51,871. Read any single estimate as a range, not a number — and expect it to be low for anything genuinely above half a million dollars.

Validate

Model bench

Three candidates on identical features and splits. Root mean squared error, in dollars, so lower is better — and the gap between training and held-out error is where overfitting shows itself.

5-fold CV error (bars show ±1 SD) Held-out test error
Show the numbers as a table
ModelCV RMSESDTest RMSETrain RMSE
Where the forest looks
Prediction error on the 4,128 held-out districts

The distribution is centred but left-skewed: the long tail of large negative errors is the model under-predicting expensive coastal districts, exactly as the $500,001 cap predicts it would.