ZUBER

Engineering Journal

Technical decisions, architecture notes, and lessons learned.

· Nyay Sahayak

Why FastAPI over Django for AI Backends

FastAPI's native async, Pydantic validation, and automatic OpenAPI docs made it the clear choice over Django Rest Framework for latency-sensitive AI applications.

For Nyay Sahayak, the backend needed to handle concurrent WebSocket connections for real-time legal chat, process audio streams from Whisper, and serve a REST API — all under 200ms latency targets. Django REST Framework was evaluated but rejected because: - WSGI-based by default (async support is bolted on, not native) - Heavier footprint for a focused API service - More boilerplate for request/response validation FastAPI won on: native async (Uvicorn), Pydantic v2 for zero-overhead validation, automatic OpenAPI docs that the frontend team consumed directly, and dependency injection for clean separation of concerns. The tradeoff: fewer built-in features (admin panel, ORM). We compensated with SQLAlchemy async for the ORM layer and a custom admin dashboard.

ArchitectureBackendPython
· Nyay Sahayak

LangChain — When Orchestration Justifies the Abstraction

LangChain's abstraction cost was worth it for multi-LLM routing and chain composition, but it added debugging complexity that required custom tooling.

For Nyay Sahayak's legal AI pipeline, the choice was between calling LLMs directly via HTTP or using LangChain as the orchestration layer. Direct calls were simpler but lacked: - Built-in chain composition for multi-step legal reasoning - Standardized prompt template management - Easy swapping between LLM providers - Built-in retry and fallback logic LangChain added value through: RunnableSequence for composing legal reasoning chains, ChatPromptTemplate for versioned prompts, and CallbackHandler for tracing. The cost: debugging LangChain internals is painful. Stack traces from chained operations are hard to read. We built a custom logging middleware around the callback system to get visibility. For simple single-LLM tasks, LangChain is overkill. For multi-step AI workflows with provider fallbacks, it earns its complexity.

ArchitectureAIOrchestration
· Kumbh Plus

PostgreSQL > MongoDB for Civic Tech Data

PostgreSQL's relational integrity, JSON support, and geospatial extensions outperformed MongoDB for structured civic registration data with complex relationships.

For Kumbh Plus, the data model involves: person records with multiple aliases, relationships between missing/found reports, location data across 80M+ attendees, and temporal tracking of report status changes. MongoDB was initially attractive for schema flexibility — person records could have varying fields. But the data integrity requirements (a person can't have two different ages in different reports) demanded relational constraints. PostgreSQL won on: foreign key constraints for referential integrity, JSONB for flexible person attributes, PostGIS for geospatial queries across the Kumbh Mela ground, and CTEs for recursive relationship matching. The tradeoff: schema migrations require careful planning. We used Alembic for versioned migrations with rollback support. For unstructured content like legal documents in Nyay Sahayak, PostgreSQL's JSONB + full-text search handled both structured metadata and unstructured content — no MongoDB needed anywhere.

ArchitectureDatabaseData Engineering
· Kumbh Plus

PWA Over Native for Mass-Event Applications

A PWA with Service Workers and IndexedDB was the right call for Kumbh Mela — instant access, no install friction, and full offline capability.

For Kumbh Plus, reaching 80M+ attendees across 4,000+ hectares meant: - Users won't install an app for a temporary event - Internet connectivity is unreliable in tent cities - Many users have low-end Android devices - Distribution through app stores is too slow for emergency deployment A PWA solved all of these: shareable via SMS/QR code in seconds, Service Workers for offline-first data entry, IndexedDB for local storage with background sync, and a responsive design that works on any screen size. The limitations: no push notifications on iOS, limited access to device hardware. For this use case — data entry and lookup — neither was needed. Native apps would have been over-engineered. A PWA with a well-designed offline sync protocol was the minimum-viable architecture that maximized reach.

ArchitectureFrontendMobile
· CodeThon CLI

Building a Multi-LLM Router Without the Bloat

CodeThon CLI's multi-LLM provider system started as a simple abstraction over HTTP clients and grew into a pluggable provider architecture without over-engineering.

CodeThon CLI needed to support 7 LLM providers: OpenAI, Anthropic, NVIDIA, DeepSeek, Together AI, Ollama (local), and LM Studio (local). The naive approach: 7 separate API client implementations with duplicated request/response handling. The over-engineered approach: a PluginProviderFactoryAbstractVisitor pattern (no). The right approach: a Provider interface with unified input/output schemas, a registry pattern for provider discovery, and each provider implementing exactly: authenticate(), generate(), and stream(). Key decisions: - Streaming is the default, not an option — all providers support it - Error handling is provider-specific but returns a unified error type - Rate limiting is handled at the router level, not per-provider - Local providers (Ollama, LM Studio) share the same connection pattern The router itself is a simple weighted distribution: prefer fastest provider first, fall back on failure, track latency per provider for adaptive routing. No message queues, no complex orchestration — just an interface + a registry + a fallback chain.

ArchitectureCLIAI Infrastructure
· Nyay Sahayak

Building a Multilingual Voice Pipeline for Indian Languages

Whisper-large-v3 + Groq's Llama 3.3 delivered sub-3-second end-to-end latency for Hindi/English/Hinglish legal voice queries — but accent handling required significant prompt engineering.

Nyay Sahayak's voice feature faced three challenges: 1. Code-switching — users mix Hindi and English mid-sentence 2. Accent variation — Indian English and regional Hindi accents 3. Legal terminology — domain-specific words that Whisper doesn't know Whisper-large-v3 handled code-switching well out of the box. The bigger challenge was legal terminology in Hindi — words like "FIR" (First Information Report), "IPC section" (Indian Penal Code). Solution: post-processing pipeline that maps Whisper output through a legal term dictionary before passing to the LLM. For example, "FIR" was sometimes transcribed as "fear" — context-based correction fixed this. The latency breakdown: audio capture (500ms) → Whisper inference (800ms) → LLM generation (1.2s) → response (200ms) = ~2.7s end-to-end. Acceptable for a conversation, but too slow for real-time voice chat. The bottleneck is the LLM, not Whisper. Future improvement: speculative decoding for LLM generation to cut 300-400ms off response time.

AIVoiceNLPFull Stack
· UIDAI Analytics

Processing 1.2M Aadhaar Records — Pipeline Design Lessons

Building a data pipeline for 1.2M+ government records taught me that assumptions about data quality are always wrong — defensive parsing saved the project.

The UIDAI Analytics system processes Aadhaar enrollment/update data across 36 states and UTs. The raw data came in CSV files with inconsistent formatting, missing values, and encoding issues. Initial approach: load everything into Pandas, clean in memory. Crashed at 800K records (16GB RAM was not enough). Solution: chunked processing with Dask, streaming data through validation → cleaning → transformation → aggregation stages. Key lessons: - Date formats varied across 5+ regional formats. Standardize early or fail late. - Missing values weren't random — specific districts had systematic missing data. Statistical imputation would have been wrong. We flagged and excluded instead. - The Mann-Kendall test for trend detection is computationally intensive on 1.2M records. Vectorized implementations in SciPy handled 36 states in 12 seconds. - ARIMA forecasting with auto_arima found optimal parameters per state in parallel. Total runtime: 8 minutes for all 36 states. The pipeline now runs in 15 minutes on 4 vCPUs and 8GB RAM — from raw CSVs to published dashboards.

Data EngineeringPythonPerformance
· Nyay Sahayak

Designing a RAG Pipeline with ChromaDB for Legal Documents

ChromaDB's local-first design made it the right choice for legal document retrieval — but chunking strategy and embedding model selection required more iteration than expected.

Nyay Sahayak needs to retrieve relevant legal documents (IPC sections, CrPC, landmark judgments) to ground LLM responses in verified law. The RAG pipeline went through three iterations. Iteration 1: Naive chunking. Split documents by page, embed with text-embedding-3-small, top-5 retrieval. Result: 60% relevance. Chunks too large, context window wasted on irrelevant content. Iteration 2: Semantic chunking. Split by section headers and paragraphs, used multilingual-e5-small for Hindi compatibility. Top-5 retrieval hit 82%. Better, but legal cross-references ("see section 302") were missed because the linked content was in a different chunk. Iteration 3: Hybrid chunking + hierarchical retrieval. Primary chunks by section boundary (500-800 tokens), linked to parent document. Cross-reference edges stored as metadata. Retrieval strategy: dense embedding search + keyword BM25 fallback + cross-reference traversal. Final pipeline: query → embed (15ms) → ANN search in ChromaDB (5ms) → cross-reference expansion (10ms) → re-rank with cross-encoder (50ms) = 80ms total retrieval time. 94% relevance on legal domain test set. Key decisions: - ChromaDB over Pinecone: self-hosted, no API costs, works offline. Tradeoff: no built-in hybrid search — implemented it manually. - e5-small over Ada-002: better Hindi legal terminology recall. Tradeoff: self-hosting the model adds infrastructure complexity. - Re-ranking is essential: without a cross-encoder stage, the top-5 from pure embedding search included too many false positives on legal nuance.

ArchitectureAIDatabasePython
· Nyay Sahayak

FastAPI WebSocket Patterns for Real-Time Legal Chat

WebSocket connection management, audio streaming, and graceful disconnection handling in FastAPI required careful state machine design to prevent ghost connections.

Nyay Sahayak's legal chat uses WebSockets for real-time bidirectional communication. The client streams audio chunks via WebSocket, receives streaming LLM responses, and maintains session state. Architecture pattern: each WebSocket connection maps to a ConversationSession object with: - AudioBuffer (accumulates chunks until silence detected) - TranscriptionStream (Whisper inference on buffer flush) - ResponseStream (SSE-style token streaming back through the WebSocket) - SessionState (conversation history, legal context, user profile) The state machine: Connecting → Ready → Listening → Processing → Responding → Ready (loop). Each state has a timeout. If the client doesn't transition within 30s, the session is garbage collected. Key gotchas: - FastAPI's WebSocket disconnect detection is not instantaneous. A client that force-closes the tab leaves a ghost connection for up to 60s. Solution: heartbeat ping/pong every 10s with 3-strike timeout. - Audio streaming over WebSocket requires careful buffer sizing. 5-second chunks give the best latency/accuracy tradeoff for Whisper. - Streaming LLM responses back through WebSocket is trivial — until you need to handle cancellation (user interrupts the AI mid-response). Solution: asyncio cancellation tokens per response stream. Lesson: WebSocket code is harder to test than REST. Every state transition needs both success and failure test cases.

BackendPythonFull StackPerformance
· Nyay Sahayak

Docker Compose for Local AI Development Stacks

Docker Compose with a FastAPI backend, ChromaDB, Ollama for local LLMs, and PostgreSQL gave us a reproducible AI dev environment that matched production.

For Nyay Sahayak and Kumbh Plus, the development stack includes: FastAPI server, ChromaDB vector store, PostgreSQL database, Redis cache, and optional Ollama for local LLM inference. Compose structure: - api: FastAPI with hot-reload (bind mount) - chromadb: persistent volume for embeddings - postgres: init scripts for schema + seed data - redis: ephemeral cache (no persistence needed) - ollama: GPU passthrough for local inference (optional, conditional profile) Key patterns that worked: - Healthcheck dependencies prevent race conditions. Postgres must be accepting connections before FastAPI starts. ChromaDB must respond to /api/v1/health before the API tries to create collections. - Profile-based service selection: `docker compose --profile local-llm up` adds Ollama. Default profile skips it for faster startup. - .env file for secrets with .env.example committed to repo. No hardcoded API keys anywhere. - Volume mounts for ChromaDB persistence: losing embeddings on container restart is painful. - CPU limits on AI services: ChromeDB and Ollama can consume all available resources. Setting `deploy.resources.limits.cpus` prevents laptop fans from taking off. The tradeoff: Docker adds ~500ms to request latency for proxied services vs direct host execution. For development, the reproducibility is worth it. For production, services run natively with Docker only for CI/CD consistency.

DevOpsBackendInfrastructure
· Nyay Sahayak

Multi-Agent Orchestration with LangGraph State Machines

LangGraph's state graph model beats sequential chains for multi-step legal assistance — each agent handles one task and passes verified results to the next.

Nyay Sahayak's legal assistance pipeline involves multiple steps: understand the query → retrieve relevant laws → draft a response → verify citations → format for the user. A linear chain with LangChain's RunnableSequence fails when a step produces unexpected output. Example: retrieval returns no relevant documents for an unusual legal question. In a chain, the LLM would hallucinate an answer. In a graph, the agent can branch to a "request clarification" path. LangGraph solution: - Agent 1 (Classifier): determines query type (filing, advice, rights inquiry) - Agent 2 (Retriever): fetches relevant legal documents from ChromaDB - Agent 3 (Drafter): composes response based on retrieved documents - Agent 4 (Verifier): checks citations against source documents - Agent 5 (Formatter): presents in user-friendly format The state graph connects these with conditional edges. If Verifier detects a hallucinated citation, the graph loops back to Drafter with a correction instruction. If Retriever finds zero results, the graph routes to a "human escalation" node instead of letting Drafter proceed. Key insight: the state object passed between nodes is the most critical design decision. Ours includes: conversation_id, query, retrieval_results[], draft_response, verification_results[], error_log[], and routing_metadata. Every node can read the full state but only writes to its designated fields. Cost: LangGraph adds ~200ms overhead per step vs direct function calls. For a 5-agent pipeline, that's ~1s of orchestration overhead. Worth it for the safety guarantees.

AIArchitectureOrchestration
· CodeThon CLI

TypeScript over Python for Developer Tool CLIs

For CodeThon CLI, TypeScript's npm distribution, native async, and rich terminal ecosystem made it the better choice than Python for a cross-platform CLI.

CodeThon CLI needed to be: installable with one command, cross-platform (macOS, Linux, Windows), and capable of streaming real-time output from 7 LLM providers. Python was the initial choice. Click for CLI framework, httpx for async HTTP, rich for terminal output. But three problems emerged: 1. Python distribution is painful. "pip install codethon" fails if Python isn't installed, or the wrong Python version is installed, or pip isn't on PATH. npx codethon just works. 2. Windows support in Python CLIs is a second-class experience. Rich library has Windows rendering issues. Event loops (asyncio) behave differently. 3. Streaming output from subprocesses (which CodeThon uses for code generation) is simpler in Node.js — native streams compose naturally. TypeScript advantage: - `npx codethon` — zero-install execution via npm. Game-changer for adoption. - Commander.js + Inquirer + Chalk + Ora = mature, well-tested CLI ecosystem - Native event loop handles concurrent provider requests without asyncio boilerplate - Single binary distribution via pkg or bun build for offline use The tradeoff: Python has better ML/AI libraries (LangChain Python is more mature than LangChain JS). For CodeThon — which calls LLMs via HTTP API — this doesn't matter. The CLI is a client, not an ML pipeline. Decision rule: if your tool calls LLMs via API, TypeScript is fine. If your tool trains or runs models locally, Python is required.

CLITypeScriptArchitectureDeveloper Tools
· Parishtha

E-commerce Data Modeling with MongoDB for Parishtha

MongoDB's document model worked well for Parishtha's product catalog — but order management revealed the relational gaps that required careful schema design.

For Parishtha, an e-commerce platform for spiritual and wellness products, MongoDB was chosen for flexible product schemas (different product categories have different attributes). Product catalog model: embedded documents for variants (size, color, material) and dynamic fields via MongoDB's schema flexibility. Query patterns: category browse (filtered), product search (text index), and product detail (single document lookup). What worked: - Product catalog with embedded variants: one document per product, variants as sub-documents. Single query returns all options. - Dynamic attributes: incense products have "burn time" and "scent notes" fields. Apparel has "size" and "fabric". No schema migration needed. - Text indexes on product names and descriptions work well for search. What required careful design: - Order management: Orders reference products by ID but need to snapshot the product state at time of purchase (price, variant). Embedded order items within a user's order document keeps reads fast but makes analytics queries harder. - Inventory tracking: MongoDB transactions (multi-document) were needed for atomic decrement operations. Not as natural as PostgreSQL's row-level locking. - Cart persistence: Using MongoDB TTL indexes for abandoned cart cleanup works well. 7-day expiry with a last-updated timestamp. Lesson: MongoDB is excellent for the "product catalog" use case. For the "transactional e-commerce" use case, a relational database for orders + inventory + payments is cleaner. Splitting the data model across MongoDB (catalog) and PostgreSQL (orders) would have been the ideal architecture.

DatabaseBackendArchitecture
· Kumbh Plus

Offline-First Sync Protocol for PWA Applications

Kumbh Plus's offline sync protocol uses timestamp vectors and conflict markers — not CRDTs — for pragmatic conflict resolution in a humanitarian context.

Kumbh Plus's core requirement: volunteers in tent cities with unreliable internet must be able to submit missing person reports, search existing reports, and receive updates — all offline. The sync protocol design: - Service Worker intercepts all API calls. When online, passes through. When offline, caches the request in IndexedDB with a pending flag. - On connectivity restore, the SW replays pending requests in order. If the server responds with a conflict (stale data), the client shows a merge UI. - Timestamp vectors (last-written timestamp per record) determine write order. Last-write-wins for non-critical fields. Manual merge for critical fields (person identity data). - Background sync API triggers replay when connection is detected. The user doesn't need to refresh. Key design decisions: - IndexedDB schema mirrors the API schema. Local reads are instant — no network latency for search queries. - Service Worker cache-first strategy for static assets. Network-first for API calls with IndexedDB fallback. - Sync queue with exponential backoff for failed requests. Max 5 retries before marking as failed. - Conflict detection: both client and server maintain a version counter per record. If versions diverge, the most recent change to each field wins (field-level, not document-level). What I would do differently: add a sync status indicator in the UI. Users in our pilot didn't know if their report was submitted or pending sync. A green/amber/red indicator would have reduced duplicate submissions significantly. CRDTs would have been academically elegant but operationally unnecessary. Field-level LWW with manual merge for critical data was the pragmatic choice that worked.

FrontendMobileArchitectureFull Stack
· CodeThon CLI

Building Systematic AI Evaluation Pipelines

A structured eval framework with accuracy, hallucination, latency, and cost metrics — tested across 7 LLM providers — revealed that faster models don't always mean better user experience.

When CodeThon CLI supports 7 LLM providers, the natural question is: which one should it use for which task? We built an evaluation pipeline to answer this systematically. The eval framework: - Test dataset: 100 prompts across 4 categories (code generation, explanation, debugging, architecture design) - Metrics per response: accuracy (semantic similarity to reference), hallucination rate (factual claims not in the prompt), latency (TTFT and total generation), cost per token - Automated grader: a separate LLM (Claude) evaluates each response against reference answers and flags hallucinations Key findings: - For code generation: DeepSeek was 40% cheaper than GPT-4 with 94% of the accuracy. Within 5 percentage points for less than half the cost. - For explanation tasks: Claude outperformed every provider on accuracy (96%) but had the highest cost. Worth it for user-facing explanations. - Groq's Llama 3.3 had sub-100ms TTFT but higher hallucination rates (12% vs 5% for Claude). Great for streaming, but needs verification layer. - All providers hallucinated on library versions. None got "what is the latest version of X" consistently right. Infrastructure: the eval pipeline runs as a scheduled GitHub Action weekly. Results go to a JSON file that the CLI's router downloads at startup for adaptive provider selection. The eval system runs on GitHub Actions with Ollama for local model tests and API calls for cloud providers. Total runtime: ~45 minutes for 700 evaluations (100 prompts × 7 providers). Cost: ~$3 per run in API fees. Lesson: without systematic evaluation, you're guessing. With it, you can make data-driven decisions about provider selection, routing, and cost optimization.

AIInfrastructurePythonTesting
· Nyay Sahayak

Prompt Engineering Patterns for Structured Legal Output

Legal AI requires structured output with zero hallucination tolerance — chain-of-thought + schema-enforced generation + citation grounding achieves this.

Nyay Sahayak's core challenge: an LLM generating legal advice that sounds authoritative but is wrong could cause real harm. We needed structured, verifiable output. The prompt architecture (3 layers): 1. System prompt: role definition ("you are a legal assistant for Indian law"), behavior constraints ("only cite sections you have retrieved"), output format (JSON schema) 2. Context injection: retrieved legal documents, conversation history, user's state (for location-specific laws) 3. User message: the actual query Structured output via Pydantic schema: ```python class LegalResponse(BaseModel): summary: str # Plain-language answer citations: list[Citation] # All cited sections must be in context disclaimers: list[str] # Required disclaimers next_steps: list[str] # Actionable steps for the user ``` The LLM must produce valid JSON matching this schema. FastAPI's response_model validates it server-side. If validation fails, the response is rejected and the LLM retries once. Chain-of-thought prompting for reasoning: "First, identify the relevant IPC/BNS sections for this query. Second, explain how each applies to the user's situation. Third, cite the specific section number and text. Fourth, structure the response in the required JSON format." The COT step is critical — without it, the model jumps to conclusions and cites irrelevant sections. Citation grounding enforcement: a post-processing step verifies that every citation in the response exists in the retrieved context. If a citation is not grounded, the response is flagged for human review, not shown to the user. This system caught 34% of initial model outputs that contained ungrounded citations. The end user never sees them — they get a "request under review" message instead. Safety over speed.

AINLPPython