ZUBER
Back to Projects

Data Science

UIDAI Analytics

PythonStreamlitARIMAPlotlyDaskPandas

Production-grade analytical system for UIDAI Data Hackathon 2026. Engineered Mann-Kendall trend tests and ARIMA forecasting to detect declining Aadhaar updates across 36 states/UTs. Processed 1.2M+ records and identified 26 states with significant declining trends.

Key Metrics

  • 1.2M+ records processed
  • 36 states/UTs analyzed in 12 seconds
  • 15-minute full pipeline runtime on 4 vCPUs

Architecture

CSV → Dask Chunked Pipeline → Validation → ARIMA / Mann-Kendall → Streamlit Dashboard

Engineering Decisions

DECISION

Dask over Pandas for data processing

1.2M+ records caused Pandas OOM crashes on 16GB RAM. Dask chunked evaluation processes data in memory-efficient batches.

Alternatives: Modin — rejected for immature API compatibility. Spark — overkill for 1.2M records, complex setup.

Tradeoffs: Dask adds complexity (lazy evaluation, debugging difficulty). Worth it for the memory efficiency gain.

DECISION

ARIMA for forecasting

Statistical rigor with interpretable results — critical for government stakeholders who need to explain methodology

Alternatives: Prophet — rejected for black-box seasonality that couldn't be explained to non-technical stakeholders

Tradeoffs: ARIMA requires manual parameter tuning. auto_arima mitigated this but extended compute to 8 minutes.

DECISION

Streamlit for dashboard

Fastest path from analysis to interactive dashboard — Python-native, no frontend code needed

Alternatives: Plotly Dash — more flexible but requires React knowledge. Power BI — not programmatically reproducible.

Tradeoffs: Streamlit less customizable than Dash. For this use case, speed of delivery outweighed design flexibility.

Engineering Challenges

PROBLEM· High difficulty

Pandas OOM crash at 800K records

Rewrote pipeline using Dask for chunked, lazy evaluation — processing 100K rows at a time

Always profile memory usage before assuming in-memory processing will work at scale

PROBLEM· Medium difficulty

Date formats varied across 5+ regional conventions

Built a date parser that tries pandas' date parser, falls back to custom format detection per state

Government CSVs have no data standards — always budget 20% of pipeline time for data cleaning

PROBLEM· Medium difficulty

Systematic missing data in specific districts

Flagged and excluded districts with >30% missing data rather than imputing — imputation would have hidden the reporting problem

Not all missing data is random — systematic gaps indicate reporting infrastructure failures, not statistical noise

Testing

Unit: pytest for each pipeline stage (validation, cleaning, analysis). Integration: full pipeline test on 100K record subset (2 min runtime). Regression: weekly scheduled run to catch data drift.

Deployment

Plotly dashboards deployed to Streamlit Cloud via GitHub. Pipeline runs locally on demand — could be migrated to Airflow for scheduled execution.

Monitoring

Manual review of dashboard for data anomalies. No automated monitoring — next step would add Great Expectations for data quality checks.

Future Redesign

Package the pipeline as a reusable CLI tool with configurable parameters. Other hackathon participants could benefit from the same Dask pipeline pattern for government datasets.