AI-powered legal aid platform that turns hours of legal consultation into a single conversation for Indian citizens. Built a multilingual voice pipeline using Whisper-large-v3 supporting Hindi/English/Hinglish, integrated Groq's Llama 3.3 70B for instant legal counsel. Full-stack architecture with React 18, FastAPI, and ChromaDB. Simplifies FIR filing as an AI "Station House Officer".
User → React Frontend → FastAPI → LangChain → Groq Llama 3.3 / ChromaDB / Whisper
DECISION
FastAPI over Django Rest Framework
Native async support for concurrent WebSocket connections and audio stream processing
Alternatives: DRF — rejected for WSGI-based blocking IO that would bottleneck audio pipelines
Tradeoffs: FastAPI has fewer built-in features (admin panel, ORM). Compensated with SQLAlchemy async + custom admin.
DECISION
Groq Llama 3.3 over OpenAI GPT-4
Sub-100ms inference latency for real-time legal counsel; critical for conversational UX
Alternatives: GPT-4 — rejected for 2-3s latency that made real-time conversation impossible
Tradeoffs: Smaller context window (8K vs 128K) and less consistent structured output. Designed prompt templates to work within context limits.
DECISION
Whisper-large-v3 for multilingual STT
Best-in-class Hindi/English code-switching support with 90%+ WER on legal domain
Alternatives: Google STT — rejected for poor Hindi legal terminology recognition
Tradeoffs: Larger model size increases first-inference latency. Solved with warm-start inference and connection pooling.
Legal terminology in Hindi not recognized by Whisper
Built a post-processing pipeline mapping Whisper output through a legal term dictionary before LLM injection
Domain-specific ASR requires a custom vocabulary layer — no off-the-shelf model handles legal Hindi adequately
LLM hallucination on legal citations
Implemented retrieval-augmented generation with ChromaDB — every citation must be grounded in a retrieved legal document
For accuracy-critical domains, RAG is not optional — it is the only acceptable architecture
End-to-end latency exceeded 4s on first build
Parallelized audio capture with Whisper inference, added response streaming, implemented prompt caching
Perceived latency matters more than actual latency — streaming made 3s feel instant to users
Unit: pytest for FastAPI endpoints (30+ tests across REST and WebSocket routes). Integration: ChromaDB query accuracy evaluated against 200 legal question-answer pairs. E2E: latency benchmarking via custom timing middleware.
Frontend: Vite build deployed to Vercel. Backend: Docker Compose with FastAPI + Uvicorn behind Nginx reverse proxy. CI/CD via GitHub Actions for automated build and deploy.
Prometheus metrics for endpoint latency (p50/p95/p99), WebSocket connection count, and LLM response times. Grafana dashboard for real-time observability.
Replace LangChain with direct LLM API calls for simple retrieval tasks — LangChain overhead is unnecessary for single-model queries. Only use it for multi-agent orchestration.