Production-grade analytical system for UIDAI Data Hackathon 2026. Engineered Mann-Kendall trend tests and ARIMA forecasting to detect declining Aadhaar updates across 36 states/UTs. Processed 1.2M+ records and identified 26 states with significant declining trends.
CSV → Dask Chunked Pipeline → Validation → ARIMA / Mann-Kendall → Streamlit Dashboard
DECISION
Dask over Pandas for data processing
1.2M+ records caused Pandas OOM crashes on 16GB RAM. Dask chunked evaluation processes data in memory-efficient batches.
Alternatives: Modin — rejected for immature API compatibility. Spark — overkill for 1.2M records, complex setup.
Tradeoffs: Dask adds complexity (lazy evaluation, debugging difficulty). Worth it for the memory efficiency gain.
DECISION
ARIMA for forecasting
Statistical rigor with interpretable results — critical for government stakeholders who need to explain methodology
Alternatives: Prophet — rejected for black-box seasonality that couldn't be explained to non-technical stakeholders
Tradeoffs: ARIMA requires manual parameter tuning. auto_arima mitigated this but extended compute to 8 minutes.
DECISION
Streamlit for dashboard
Fastest path from analysis to interactive dashboard — Python-native, no frontend code needed
Alternatives: Plotly Dash — more flexible but requires React knowledge. Power BI — not programmatically reproducible.
Tradeoffs: Streamlit less customizable than Dash. For this use case, speed of delivery outweighed design flexibility.
Pandas OOM crash at 800K records
Rewrote pipeline using Dask for chunked, lazy evaluation — processing 100K rows at a time
Always profile memory usage before assuming in-memory processing will work at scale
Date formats varied across 5+ regional conventions
Built a date parser that tries pandas' date parser, falls back to custom format detection per state
Government CSVs have no data standards — always budget 20% of pipeline time for data cleaning
Systematic missing data in specific districts
Flagged and excluded districts with >30% missing data rather than imputing — imputation would have hidden the reporting problem
Not all missing data is random — systematic gaps indicate reporting infrastructure failures, not statistical noise
Unit: pytest for each pipeline stage (validation, cleaning, analysis). Integration: full pipeline test on 100K record subset (2 min runtime). Regression: weekly scheduled run to catch data drift.
Plotly dashboards deployed to Streamlit Cloud via GitHub. Pipeline runs locally on demand — could be migrated to Airflow for scheduled execution.
Manual review of dashboard for data anomalies. No automated monitoring — next step would add Great Expectations for data quality checks.
Package the pipeline as a reusable CLI tool with configurable parameters. Other hackathon participants could benefit from the same Dask pipeline pattern for government datasets.