Regression · Data Quality · Interactive Application

Property Value Analysis
& Price Estimator

How well can property characteristics predict historical sale prices for properties excluded from model fitting? I built an analysis and interactive application to answer that question—and examine where the estimates fall short.

Tools & Skills Python pandas scikit-learn Linear Regression Ridge Property-Grouped Splits Streamlit

Historical holdout evaluation

I selected a linear-regression pipeline using validation results, refitted it on training plus validation data, and evaluated it on 33,555 transactions representing 28,657 properties excluded from fitting.

0.8520 Dollar R²
$70,113 Mean Absolute Error
$126,742 Root Mean Squared Error
61.42% Lower MAE than Median Reference

The model reduced average absolute error compared with a simple reference that predicted the development data's median sale price for every holdout transaction. Both were evaluated on the same transactions.

Historical evaluation boundary
These results describe selected Cook County transactions from 2013–2019. They do not establish accuracy for current market values or future sales. The labeled source had already been explored before the rebuilt modeling workflow, so this is an internal historical evaluation.

Preparing records without inventing corrections

The labeled source contained 204,792 transaction records. I removed 13 exact duplicates and selected transactions marked Pure Market Filter == 1, leaving 167,375 records for the rebuilt analysis.

204,792 Source Transactions
13 Exact Duplicates Removed
167,375 Selected Transactions
60 Retained Source Fields Reviewed

Conflicting room counts

In 36 records, extracted bedrooms exceeded total rooms. I marked both counts as missing because the available evidence did not establish which value was correct.

One unverified bathroom count

I also marked one unverified bathroom count as missing. I retained the transaction rather than inventing a replacement value or removing the entire record.

Every selected description matched the extraction phrase. That confirmed extraction coverage, not the accuracy of every recorded count. Missing counts were handled through median imputation fitted inside the modeling pipeline.

Source questions remain open
Dataset attribution, version, redistribution terms, the population-filter construction, age reference date, room-count conventions, and geographic code mappings remain incompletely verified.

From property records to an interactive estimate

I organized the work into three analysis notebooks, followed by shared Python preparation and inference modules and a deployed application.

01 · pandas / Jupyter Review Data Quality Check duplicates, selected population, distributions, and source-field roles.
02 · Python / pandas Prepare Features Transform area, extract room counts, and construct a town–neighborhood key.
03 · scikit-learn Separate Properties Keep each property's transactions together across training, validation, and holdout.
04 · scikit-learn Compare Models Compare the reference, linear regression, and Ridge using validation results.
05 · Python / joblib Evaluate & Save Report holdout errors, inspect price bands, and save the evaluated pipeline.
06 · Python / Streamlit Verify & Deploy Match notebook and application preparation, validate inputs, and check hosted predictions.

Keeping the same property out of multiple splits

A property can appear in more than one transaction. Splitting individual rows could therefore place transactions for the same property in both fitting and evaluation data. I grouped records by property identifier before splitting.

Property-separated evaluation splits
Split Transactions Properties Purpose
Training 100,333 85,971 Fit candidate pipelines
Validation 33,487 28,657 Select the configuration
Holdout 33,555 28,657 Evaluate the selected configuration

No properties overlapped across these splits. After selection, I refitted the chosen pipeline on training plus validation data. The holdout was excluded from fitting, and the configuration was retained after evaluation.

What this split measures
It evaluates historical transactions for properties excluded from fitting within the 2013–2019 period. It is not a chronological test on future years.

Combining physical characteristics and location

The selected model uses seven prepared features: natural-log building area, natural-log land area, age, total rooms, bedrooms, bathrooms, and a combined town–neighborhood key.

Inputs Property Characteristics Area, age, rooms, location
Preparation Model Features Impute, scale, encode
Prediction Historical Sale Price Predict log price, then exponentiate

Why transform area?

Property areas can span a wide range. A natural-log transformation compresses that range, allowing the model to represent proportional differences in area rather than treating each extra square foot as having the same effect on log price.

xbuilding = ln(Building Square Feet)

Land area uses the same transformation. Source values were checked for positivity and finiteness before taking logarithms.

How does linear regression predict?

The model combines prepared features using learned coefficients. It predicts the natural logarithm of sale price, rather than sale price directly.

ẑ = β0 + ∑ βjxj

Here, ẑ is predicted log price, β values are fitted coefficients, and x values are pipeline outputs. Encoding location creates multiple indicator columns from the single geographic feature.

Keeping preparation inside the pipeline

Numeric imputation and scaling, categorical encoding, and regression were fitted within the training pipeline during model comparison. This prevents validation and holdout data from determining those fitted preparation steps.

Why compare Ridge?

Ordinary linear regression minimizes squared log-price residuals. Ridge adds a penalty on large coefficients, which can help stabilize a model when predictors carry overlapping information.

Ridge objective = ∑(zi − ẑi)² + α ∑βj²

zi = ln(observed sale price). α controls the penalty strength. Ridge was a comparison candidate; the selected configuration used ordinary linear regression.

Converting predictions back to dollars

Predicted Price = exp(ẑ)

Exponentiation returns the estimate to the dollar scale. Because exponentiation is nonlinear, directly converting a log-price prediction does not automatically produce an unbiased estimate of mean dollar price. This workflow does not apply a mean-price bias correction.

The fitted coefficients describe predictive associations. They do not establish that changing a property feature would cause a particular change in sale price.

What do the reported errors mean?

Dollar metrics describe errors after converting the predictions back to prices. Log metrics describe errors on the scale used to train the regression.

MAE: average absolute dollar error

MAE treats overestimates and underestimates as positive error amounts, then averages them. The holdout MAE was $70,112.72.

MAE = (1/n) ∑ |yi − ŷi|

y is observed price, ŷ is predicted price, and n is the number of evaluated transactions.

RMSE: greater weight on larger errors

RMSE squares each error before averaging. Large misses therefore influence it more strongly. The holdout dollar RMSE was $126,741.65.

RMSE = √[(1/n) ∑(yi − ŷi)²]

RMSE is expressed in dollars when calculated from dollar prices.

Log RMSE: error on the training scale

Log RMSE was 0.4326. Log residuals reflect multiplicative differences between observed and predicted prices. For example, a prediction twice the observed price has a log residual of ln(2).

Log RMSE = √[(1/n) ∑(ln(yi) − ẑi)²]

A log RMSE of 0.4326 is neither a dollar amount nor a 43.26% error rate.

R²: squared error relative to a mean reference

R² compares the model's squared errors with those from predicting the evaluation set's mean. A value of 1 indicates perfect predictions; 0 matches that reference, and negative values perform worse.

R² = 1 − [∑(yi − ŷi)² / ∑(yi − ȳ)²]

Dollar R² was 0.8520. Log R² was 0.7904, calculated using log prices instead. R² is not a percentage of predictions that are correct.

Two different references
R² uses the evaluation set's mean as its mathematical reference. The separately reported 61.42% MAE improvement compares the model with a development-fitted median-price predictor. These are different comparisons.
Recorded holdout results
Metric Result
Dollar MAE $70,112.72
Dollar RMSE $126,741.65
Dollar R² 0.8520
Log RMSE 0.4326
Log R² 0.7904
Mean signed error −$16,054.83
Median-price reference MAE $181,730.97

Mean signed error is predicted price minus observed price, averaged across transactions. The negative overall value indicates average underestimation. Positive and negative errors can cancel, so signed error should be read alongside MAE and RMSE.

None of these aggregate metrics provides an uncertainty interval or an error guarantee for an individual property.

Overall performance hides different error patterns

Higher-price transactions had larger dollar errors. Lower-price transactions had larger relative errors. Looking at price bands makes these differences visible.

Average absolute dollar error by actual sale-price band

Bar lengths start at zero and are proportional to dollar MAE. Exact results and transaction counts appear below.

Holdout errors by actual sale price
Actual Price Band Transactions Dollar MAE Mean Signed Error Median APE
Under $100,000 6,636 $35,004.55 +$27,978.79 55.53%
$100,000–$249,999 12,774 $49,369.09 −$5,714.50 23.93%
$250,000–$499,999 9,378 $65,838.01 −$20,111.43 15.40%
$500,000–$999,999 3,606 $130,951.57 −$51,874.02 15.72%
$1,000,000 or more 1,161 $344,583.28 −$237,491.33 17.52%

Median absolute percentage error, or median APE, is the median of each transaction's absolute error divided by its observed sale price, expressed as a percentage. It describes relative error rather than dollar error.

Reading the result

The model overestimated transactions below $100,000 by $27,978.79 on average. For transactions of $1 million or more, it underestimated by $237,491.33 on average.

The lowest-price band had a median absolute percentage error of 55.53%, compared with 17.52% in the highest band. A smaller dollar error can therefore still be substantial relative to a property's observed price.

Interpretation boundary
These bands use actual sale prices, which are known during evaluation. They are not available for assigning an unknown property to an error band before its sale. These patterns also do not establish demographic fairness or explain the causes of the errors.

Preserving the evaluated workflow in the estimator

I moved the reviewed preparation logic into shared Python modules so the notebook and application use consistent transformations. The application loads the saved pipeline, prepares the seven features, and converts its log-price prediction to dollars.

What visitors can enter

  • Building area and land area.
  • Recorded property age.
  • Total rooms, bedrooms, and bathrooms.
  • Town and neighborhood codes.

Unknown room counts can use the pipeline's fitted median imputation. The application rejects inconsistent inputs such as bedrooms exceeding total rooms.

What I verified

  • Application preparation matched the selected notebook features.
  • Inference predictions matched the reviewed workflow across 167,375 selected transactions.
  • Reloading the saved model preserved predictions.
  • A hosted example matched its local result: both displayed $248,328.

Location dropdowns are limited to combinations observed during model fitting. Verified geographic names remain unavailable, so the application uses source codes. Implementation consistency does not extend the model's evaluated accuracy beyond its historical scope.

Try a historical estimate

Explore how the fitted model responds to property characteristics. The output is a demonstration in the historical data context, not a verified current-market appraisal. The model does not take a sale year as an input, so it does not produce a year-specific estimate.

Try Estimator ↗

What this project demonstrated

This project connected data-quality decisions, regression modeling, property-separated evaluation, and a working application. The selected model outperformed the median-price reference on the historical holdout, while price-band analysis showed why overall metrics alone are insufficient.

What I Learned Make decisions traceable Document uncertain counts, protect split boundaries, and verify that application predictions preserve the reviewed workflow.
Limitation Historical accuracy has a boundary Source definitions remain unresolved. Current-market and future-year performance have not been established.
Next Improvement Verify inputs, then test later sales Confirm provenance and prediction-time definitions, then evaluate chronologically and develop uncertainty estimates.

The model achieved a dollar R² of 0.8520 and reduced MAE by 61.42% against the median-price reference. Its usefulness still depends on the population, historical context, input definitions, and substantial differences in error across price bands.