🏠 Airbnb Price Prediction & Price Classification
Complete Data Science Workflow – EDA, Feature Engineering, Regression, and Classification
This project presents a full end-to-end data science pipeline applied to Airbnb listing data.
The goal was to build models that:
Predict continuous log-price (Regression)
Classify listings into three price tiers: Low / Medium / High (Classification)
The workflow includes data cleaning, exploratory data analysis, extensive feature engineering, model training, evaluation, interpretation, and final model selection.
📌 Table of Contents
1.Project Overview
2.Dataset & Objective
3.Data Cleaning
4.Exploratory Data Analysis (EDA)
5.Feature Engineering
6.Regression Modeling
7.Price Classification Modeling
8.Model Comparison & Winners
9.Key Insights & Lessons Learned
🔍 1. Project Overview
This assignment demonstrates practical experience in:
log_price was transformed into 3 balanced classes using quantiles.
Class distribution remained roughly uniform in both train and test sets.
Three models were trained:
Logistic Regression
Decision Tree Classifier
Random Forest Classifier
✔ Results:
image
Findings:
Logistic Regression showed the best overall performance
It made fewer severe errors and rarely confused low-priced listings with high-priced ones
Decision Tree overfit and performed poorly
Random Forest was strong but slightly less balanced than Logistic Regression
🏆 Classification Winner:
Logistic Regression
🧠 8. Key Insights & Takeaways
Feature Engineering mattered more than model selection
After engineering features, R² jumped from 0.38 → 0.73.
Spatial features were the most powerful
cluster_id and distance_to_centroid were heavily used by the Random Forest.
Regression and Classification answer different business questions
Regression → “How much will this listing cost?”
Classification → “Is this listing cheap, medium, or expensive?”
Balanced classes allowed fair evaluation
Quantile binning ensured no class dominated the dataset.
Pipelines ensured clean, reproducible workflows
🛠️ 9. Tools & Technologies Used
Python
Scikit-Learn
Pandas
NumPy
Seaborn & Matplotlib
KMeans Clustering
Machine Learning Pipelines
OneHotEncoder & StandardScaler
🎉 Final Notes
This project demonstrates a complete ML workflow, from raw data to insights and deployment-ready models.
Both Regression and Classification models were evaluated, and the impact of Feature Engineering was clearly visible in the results.