CAN WE PREDICT AIR POLLUTION? A Data-Driven Time Series Analysis of AQI
By Shubham Pawar | Data Science Honours
CAN WE PREDICT AIR POLLUTION? A Data-Driven Time Series Analysis of AQI
By Shubham Pawar | Data Science Honours
Shah & Anchor Kutchhi Engineering College, Mumbai
1. The Invisible Threat & The Predictive Gap
Mumbai winters frequently bring unpredicted, severe haze, contributing to the 4.2 million global deaths caused annually by air pollution. Currently, cities operate on a reactive gap — closing schools and issuing advisories only after pollution peaks. This project aims to bridge that gap by analyzing historical time-series data to build a forecasting model capable of predicting AQI spikes 3 to 5 days in advance. By extracting recurring seasonal and event-driven patterns, we can shift from reactive emergency scrambling to proactive, scheduled management.

[embed]
2. Understanding AQI & Preparing the Data
AQI scales from 0 (Good) to 500 (Hazardous), tracking critical pollutants like PM2.5, PM10, NO2, and SO2. Fluctuations are heavily driven by human activity (stubble burning, Diwali) and weather (temperature inversions trapping ground-level air). Our dataset comprises 3–4 years of daily time-series AQI readings from a public CPCB monitoring station. Because time-series algorithms fail with missing dates, aggressive preprocessing was applied: short gaps were forward-filled, long gaps used 7-day rolling averages, and statistical outliers lacking corroborating pollutant spikes were neutralized as sensor errors.

[embed]
3. Statistical Profiling & Exploratory Data Analysis (EDA)
Due to the right-skewed nature of AQI data (where a few hazardous days pull the average up), the median is a more honest baseline than the mean. Multi-year line plots instantly reveal a massive winter peak and a monsoon trough. However, monthly box plots expose a hidden truth: while July and August have the lowest overall AQI, they exhibit high variance because rain instantly clears the air, but dry gaps allow rapid pollution buildup. Additionally, weekday AQI consistently runs 8 to 12 points higher than weekends due to commercial traffic.

4. Proving Patterns: Hypothesis Testing & Time-Series Mechanics
Visual patterns require statistical proof. A Welch’s T-Test confirmed that winter pollution is significantly worse than summer (p < 0.001) — a structural reality, not chance. More alarmingly, a Mann-Kendall Trend Test proved a slow, structural year-over-year upward trend in baseline pollution (p = 0.03). To model this, the data was decomposed into Trend, Seasonality, and Residuals (noise/spikes). Because lags 1–7 and 365 showed strong autocorrelation (today predicts tomorrow, and this day last year predicts today), applying a 30-day moving average perfectly filtered daily noise to reveal the true underlying seasonal shape.

5. Forecasting Models & Core Data Findings
Using Seasonal Naive forecasting (mirroring last year’s exact dates) and Exponential Smoothing (Alpha = 0.3 to minimize Mean Absolute Error), we successfully projected 30-day forward trajectories. The models and data generated five undeniable findings:
- Winter AQI mathematically doubles summer AQI.
- The foundational baseline is worsening every year.
- Diwali triggers a 40–60% spike that clears in 3–5 days.
- Weekdays are consistently worse than weekends.
- Asymmetric Recovery: The air gets dirty in 2 days but takes 10 to 14 days to clear.
[embed]
6. Visualizing for Real-World Action
Charts translate data into policy. Governments must stop reacting to peaks and instead implement fixed October-January industrial/traffic restrictions and subsidize stubble-clearing tech before the season begins. Because air takes 14 days to clear, single-day interventions are useless. For individuals, October to January is the high-risk window for stocking meds and running purifiers, while the July-August monsoon season is the only reliable window for heavy outdoor physical activity.

7. Conclusion, Limitations, & The Path Forward
We can predict air pollution usefully, if not perfectly. Even basic time-series methods combined with rigorous testing extract actionable intelligence from public datasets. However, this study is limited by its reliance on a single monitoring station and the exclusion of live weather variables, which widens uncertainty beyond a 14-day forecast window. The future of this research lies in incorporating ARIMAX models for exogenous weather variables, transitioning to live CPCB API feeds, and deploying LSTM deep learning to capture complex, non-linear temporal patterns.
[embed]
메타데이터
- post_id
- f757da61fa87
- slug
- can-we-predict-air-pollution-a-data-driven-time-series-analysis-of-aqi-f757da61fa87
- url
- https://medium.com/@theshubham1412/can-we-predict-air-pollution-a-data-driven-time-series-analysis-of-aqi-f757da61fa87
- canonical_url
- https://medium.com/@theshubham1412/can-we-predict-air-pollution-a-data-driven-time-series-analysis-of-aqi-f757da61fa87
- author_url
- https://medium.com/@theshubham1412
- status
- ok
- fetched_at
- 2026-06-17 08:20:12