Vehicle Price Prediction and Market Analysis
This project analyzes the Australian Vehicle Prices dataset (from Kaggle) and builds multiple models to:
Understand the key factors influencing car prices
Predict vehicle prices using regression models
Segment vehicles into meaningful clusters
Classify vehicles into Standard vs Premium categories
The project includes extensive EDA, feature engineering, regression modeling, clustering, and classification.
- Dataset Overview
The dataset includes 16,774 vehicle listings from the Australian market, containing attributes such as:
Brand, Model
Year
Kilometres driven
Engine size
Fuel consumption
Transmission
Drive type
Doors, Seats
Body type
Color
Price (target variable)
Research Question:
What are the key factors that influence vehicle prices in the Australian market, and how accurately can we predict a car’s price?
- Data Cleaning & Feature Engineering
Steps Performed:
Removed listings with missing Price
Filled missing categorical values with "Unknown"
Converted text fields such as "4 doors" → numeric values
Converted "5 seats" → numeric values
Created new features:
Car_Age = 2024 – Year
Kilometres_Clean (outlier-adjusted)
Efficiency = Engine_Liters / FuelConsumption
Is_Luxury based on known luxury brands (e.g., BMW, Mercedes, Porsche)
- Exploratory Data Analysis (EDA)
Below are the included graphs as Markdown placeholders.
3.1 Price Distribution
images:distribution_of_vehicle_prices.peg
3.2 Kilometres Distribution
distribution of kilometres .peg
3.3 Engine Size Distribution
distribution of engine size
3.4 Fuel Consumption Distribution
distribution of fuel consumption
3.5 Price vs Kilometres
price_vs_kilometres.peg
3.6 Price vs Year
price_vs_year.peg
3.7 Price Distribution by Brand
Price by Brand
3.8 Price by Drive Type
Price by DriveType
3.9 Price by Doors
Price by Doors
3.10 Price by Seats
Price by Seats
3.11 Price by Body Type
Price by Body Type
3.12 Price by Transmission Type
Price by Transmission
3.13 Correlation Heatmap
Correlation Heatmap
Key Insights from EDA:
Newer vehicles and low-kilometre vehicles are significantly more expensive.
Engine size and number of cylinders strongly correlate with price.
Luxury brands (Porsche, Lamborghini, Ferrari, McLaren) show extreme right-skewed distributions.
SUVs, Coupes, and Premium categories have much higher median prices.
Fuel efficiency has a mild negative correlation with price (larger engines consume more fuel and tend to be pricier).
- Regression Modeling
Three models were trained:
Models:
Linear Regression
Random Forest Regressor
XGBoost Regressor
4.1 Actual vs Predicted Prices
actual vs predicted vehicle prices
4.2 Residuals Plot
residuals plot
4.3 Feature Importance (Linear Regression)
feature importance
Key Regression Findings:
Linear Regression underfits high-priced vehicles due to non-linearity.
Random Forest and XGBoost perform significantly better with nonlinear interactions.
Cylinders, Engine Size, Year, and Kilometres are the top predictors of price.
- K-Means Clustering (Unsupervised Learning)
K=3 was selected based on interpretability.
5.1 PCA Visualization of Clusters
Clustering PCA
Cluster Interpretation:
Cluster Characteristics
0 Older, lower engine size, lower price
1 Mid-range vehicles, balanced features
2 Newer, powerful engines, premium pricing
6. Advanced Modeling: Random Forest & XGBoost
6.1 Performance Comparison
model preformance comparison
R² Scores:
Random Forest: ~0.89
XGBoost: ~0.83
Linear Regression: ~0.67
6.2 Feature Importance (Random Forest)
top 15 features
Important Features:
Car Age
Kilometres
Cylinders
Engine Liters
Efficiency
Luxury Brand Indicator
Drive Type (AWD / FWD)
- Classification Model: Standard vs Premium Vehicles
Created binary label:
0 = Standard vehicle
1 = Premium vehicle
7.1 Class Distribution
standard vs premium
7.2 Confusion Matrices
Decision Tree:
decision tree
KNN:
knn
Random Forest Classifier:
random forest
Classification Findings:
Random Forest achieved the highest accuracy and best balance across classes.
Premium vehicles are harder to classify due to smaller sample size (class imbalance).
Important features for classification included:
Cylinders, Engine Size, Luxury Brand Indicator, Drive Type, Cluster ID.
- Conclusions & Insights
Market Insights:
Price is most strongly driven by engine size, cylinders, year, kilometres, and luxury brand.
SUVs and Coupes show consistently higher price ranges.
Luxury brands dominate the top of the market with extreme price outliers.
Modeling Insights:
Simple linear models cannot capture the nonlinear structure of car prices.
Random Forest provides the best regression performance (R² ≈ 0.89).
Clustering reveals meaningful market segments.
Classification models perform well but are limited by class imbalance.