Where the money is
Each dot is one district, placed at its longitude and latitude and shaded by median home value. Filter the population down and the map answers questions the summary statistics cannot — like how much of California's high-value housing is within sight of water.
Home values were truncated at $500,001 when the census was published, so 992 districts (4.8%) sit on that ceiling as a hard line of dots. Any model trained on this data systematically under-predicts genuinely expensive housing.
What actually predicts price
Pearson correlation across all 20,640 districts. Pick any cell to plot that pair — the scatter usually shows why a strong-looking coefficient is or isn't trustworthy.
Median income is the single strongest linear predictor of home value (r = 0.688). The flat line of points along the top of that scatter is the $500,001 cap.
District value estimator
A 24-tree random forest, trained on 16,512 districts and exported into this page, runs on every keystroke. Nothing is sent anywhere — the trees are evaluated in your browser.
The forest never saw the test districts during training, and its held-out error is ±$51,871. Read any single estimate as a range, not a number — and expect it to be low for anything genuinely above half a million dollars.
Model bench
Three candidates on identical features and splits. Root mean squared error, in dollars, so lower is better — and the gap between training and held-out error is where overfitting shows itself.
Show the numbers as a table
| Model | CV RMSE | SD | Test RMSE | Train RMSE |
|---|
The distribution is centred but left-skewed: the long tail of large negative errors is the model under-predicting expensive coastal districts, exactly as the $500,001 cap predicts it would.