Skip to content

Latest commit

 

History

26 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RideIQ — Uber Product Analytics

Simulating a Senior Product Analyst role at Uber — analyzing 200,000+ rides across New York City to diagnose 3 critical business problems using SQL, Python, and Power BI.

Python SQL PowerBI Status


The Problem

Uber's New York operations show a concerning pattern — 8 in 10 riders never return after their first month. Meanwhile, a significant portion of booked rides never complete, and surge pricing may be driving away demand rather than maximizing revenue.

This project investigates all three problems end-to-end — from raw data to business recommendations.


Business Problems

# Problem What We Found
1 Rider retention drop D7: 34%, D30: 19% — 8 in 10 riders lost within a month
2 Cancellation funnel leakage 6 AM has lowest ride completion rate (81%)
3 Surge pricing impact All hours fall in $10–$15 avg fare range (Medium Surge)

Key Findings

  • Peak demand — Friday 7 PM peaks at 1,999 rides/hour
  • Retention crisis — Only 19% of riders still active after 30 days
  • Cancellation pattern — Early morning hours (6 AM) show highest cancellation rates
  • Fare distribution — Average fare $11.36, median $8.50 — majority of rides are short trips
  • Rider segments — 3 distinct segments identified: Low, Mid, and High value riders

Business Recommendations

Problem Recommendation
Retention Target D7 churned riders with discount voucher via push notification
Cancellation Investigate early morning supply-demand gap — increase driver incentives at 6 AM
Surge Run A/B test — cap surge at 1.5x in 2 cities for 30 days, measure demand vs revenue

Dashboard Preview

Uber Analytics Dashboard


Tech Stack

Tool Purpose
Python (Pandas, Matplotlib, Seaborn, Scikit-learn) Cleaning, EDA, A/B testing, segmentation
SQL (PostgreSQL) Funnel analysis, cohort retention, LTV, surge revenue
Power BI (DAX) Executive dashboard — 3 pages, 5 DAX measures

Project Structure

rideiq-uber-analytics/
├── data/
│   ├── raw/                  ← Original Kaggle dataset
│   └── cleaned/              ← Processed dataset (195,065 rows)
├── notebooks/
│   ├── 01_data_cleaning.ipynb
│   ├── 02_eda.ipynb
│   ├── 03_ab_test.ipynb
│   └── 04_segmentation.ipynb
├── sql/
│   ├── 01_funnel.sql
│   ├── 02_retention_cohort.sql
│   ├── 03_ltv.sql
│   ├── 04_cancellation.sql
│   ├── 05_supply_demand.sql
│   └── 06_surge_revenue.sql
├── dashboard/
│   └── uber_dashboard.pbix
└── problem_statement.md

Notebooks

Notebook Description
01_data_cleaning.ipynb Cleaned 200K rows — removed 4,935 invalid records, fixed data types, engineered time features
02_eda.ipynb 8+ visualizations — hourly demand, fare distribution, day-of-week patterns, cohort heatmap
03_ab_test.ipynb Simulated surge cap A/B test — Z-test, p-value: 0.46
04_segmentation.ipynb KMeans clustering (k=3) — Low, Mid, High value rider segments

SQL Queries

File Analysis
01_funnel.sql Ride completion funnel by hour
02_retention_cohort.sql Monthly cohort retention using window functions
03_ltv.sql Rider lifetime value by passenger segment
04_cancellation.sql Cancellation rate analysis by hour
05_supply_demand.sql Peak demand hours by day of week
06_surge_revenue.sql Revenue and surge level classification by hour

Dataset

Property Value
Source Uber Fares Dataset — Kaggle
Raw rows 200,000
Cleaned rows 195,065
Columns 11 (after feature engineering)
Location New York City
Period 2009–2015

Built by Aditya Sharma · LinkedIn · GitHub

About

Uber product analytics — 200K NYC rides analyzed using SQL, Python & Power BI to solve retention, cancellation, and surge pricing problems.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages