EcoAir Insight is an AI-powered air quality intelligence platform designed to analyze, forecast, and visualize air pollution data across India. This document serves as a comprehensive technical overview of the system's architecture, ML pipeline, backend services, and frontend engineering decisions.
EcoAir Insight adopts a decoupled architecture separating the machine learning pipeline, the RESTful backend, and the interactive frontend.
- Frontend: React SPA (Single Page Application) built with Vite, utilizing Leaflet for geospatial mapping and Framer Motion for fluid UX.
- Backend: Python FastAPI service orchestrating data retrieval, spatial queries, and external AI integrations.
- Database: SQLite (
ecoair.db), utilizing SQLAlchemy ORM. Serves as the central data store for monitoring stations, historic AQI, and precomputed ML predictions. - ML Pipeline: A standalone Python batch processing system using
scikit-learnandjoblibfor parallelized model training and inference.
- User Interaction: User clicks on the Leaflet map or searches for a location (managed by
MapPage.jsx). - Geospatial Resolution: The frontend sends the
(lat, lon)to the FastAPI backend (GET /analysis). - Orchestration:
find_nearest_station_with_datacalculates the Haversine distance to locate the closest active monitoring station in SQLite.get_current_aqiretrieves the latest pollution metrics.get_predictionqueries the 5-year precomputed forecast.generate_ai_insightscalls an LLM to generate contextual analysis based on the specific pollutant mix.
- Presentation: Data is returned to the client and rendered in a glassmorphic React
InfoPanelwith interactive Chart.js visualizations.
Rather than running inference synchronously on the backend, the ML pipeline operates as an offline batch process (train_model.py) that precomputes forecasts. This architectural decision ensures the API remains highly responsive.
The core predictive model is the HistGradientBoostingRegressor from sklearn.ensemble. It was chosen for its native support for missing values and fast execution on large datasets.
- Hyperparameters:
max_depth=6,learning_rate=0.05,max_iter=200,random_state=42. - Parallelization: Training is distributed across available CPU cores using
joblib.Parallelwith thelokybackend. A separate model is trained for each individual monitoring station.
For each station, historic data is aggregated by month. The time-series nature of the data is captured through engineered features:
- Cyclical Temporal Features:
Month_Index,Month,sin_month, andcos_month(capturing seasonal pollution patterns). - Autoregressive Features:
lag_1,lag_2, andlag_3(the AQI of the previous 3 months). - Rolling Statistics:
rolling_mean_3androlling_std_3to capture short-term trends and volatility.
The pipeline generates a 60-month (5-year) forecast using an autoregressive loop.
- The model predicts the next month's AQI.
- The predicted AQI is appended to the feature set, simulating
lag_1for the subsequent step. - Rolling means and temporal features are recalculated for
t+1, and the cycle repeats 60 times. - Confidence Intervals: An upper bound and a lower bound (
max(0, val - std_dev)) are computed using the standard deviation of the forecasted trajectory.
Raw data arrives as multi-tab Excel files. The script normalizes column headers, filters out non-data sheets, enforces date and aqi validity, and extracts a canonical set of stations (stations.csv) and cleaned metrics (cleaned_data.csv).
The backend (backend/app/main.py) acts as the high-performance orchestration layer.
The primary endpoint powering the application is:
GET /analysis:- Query Parameters:
lat(float),lon(float) - Response Shape: Returns a heavily nested JSON payload including
stationmetadata,is_fallbackflag,distance_km,current(PM2.5, PM10, etc.),prediction(forecast array),health(risk level and advice), andai_insights(LLM-generated markdown).
- Query Parameters:
To simplify deployment, the backend utilizes SQLite. A notable engineering decision is the Background Migration Thread:
- During FastAPI startup (via the
@asynccontextmanager async def lifespan), the app checks if theStationtable is empty. - If empty, it spawns a daemon thread
threading.Thread(target=run_migration, daemon=True)to executemigrate_to_db.py. - This ensures the FastAPI server starts immediately and is ready to accept requests without blocking, returning a graceful
"Database is currently being populated"error to early API callers until the CSV data is fully loaded into SQLite.
The frontend is engineered for a premium, highly responsive user experience.
State is managed locally in top-level components (e.g., MapPage.jsx) and passed down via props.
- The
handleClickfunction translates Leaflet map click events (yieldinglatlng) into asynchronous Axios calls to the/analysisbackend endpoint. - To prevent UI jank, a dedicated
loadingstate toggles a<Loader />component while the backend aggregates data and resolves the LLM prompt.
- Glassmorphism & Fluid Animation:
framer-motionhandles the entrance and exit of the<InfoPanel />. The panel uses a glassmorphic design (backdropFilter: "blur(16px)") to overlay data cleanly above the interactive map. - Graceful Fallbacks: If the user clicks an area without immediate coverage, the backend returns
is_fallback=Truewith thedistance_kmto the nearest station. The UI renders a specific warning state to inform the user of the spatial approximation. - Data Visualization: Forecast data is routed into the
<PredictionChart />component, which manages Chart.js configuration for rendering the predicted AQI alongside its confidence intervals.
- Offline ML Inference over Real-time: Serving a scikit-learn model in the request-response cycle can introduce latency spikes, especially when forecasting 60 steps autoregressively. By precomputing predictions and storing them in SQLite alongside station data, the
/analysisendpoint guarantees low-latency O(1) reads for predictions. - SQLite for Production?: Given the read-heavy nature of the application (monitoring stations don't change frequently, and historical/predicted data is batch-updated), SQLite is highly efficient. The background migration thread makes cold-start deployments (e.g., on Render or Vercel) seamless.
- FastAPI: Selected over Flask/Django for its native async capabilities and automatic Pydantic validation. The
@asynccontextmanagermade the background database initialization trivial to implement.
- Spatial Interpolation: Currently, the system falls back to the nearest station. Implementing Inverse Distance Weighting (IDW) or Kriging could provide estimated AQI for unmonitored coordinates based on a web of nearby stations.
- Database Migration: Moving from SQLite to PostgreSQL + PostGIS would allow for native bounding-box and spatial queries (
ST_DWithin), optimizing thefind_nearest_stationlogic which currently relies on application-level Haversine distance calculations. - Model CI/CD: Integrating Airflow or Prefect to automatically trigger
train_model.pyas new monthly data arrives, seamlessly updating thepredictionstable.