Engineering
Decisions.
Database-First over API-First
Building a wrapper around the Gmail API introduces massive latency (often 1-2 seconds per thread fetch). By syncing metadata to PostgreSQL in the background and tracking the historyId, the frontend UI feels native and instantaneous. The tradeoff is storage cost and database sync complexity, which is heavily outweighed by the UX improvement.
Gemini 2.0 Flash vs GPT-4
Email processing is a high-volume, low-latency requirement. GPT-4o was too slow and expensive for bulk thread processing. Gemini 2.0 Flash offers structured JSON enforcement and processes 10,000 tokens in under a second, bringing the cost down to roughly ~₹0.25 per email without sacrificing extraction quality.
Single-Pass Execution
Many "AI Assistants" use ReAct loops or multi-agent frameworks. This introduces catastrophic hallucination risk and unbounded costs. We explicitly chose a functional, pipeline-based approach: The LLM acts as a pure parser returning JSON. Traditional backend functions (FastAPI) handle all the routing, persistence, and state.
Strict Contracts (DTOs)
We define exact Pydantic models for ThreadIntelV1. If the AI deviates from the schema, the backend simply rejects or retries it at the integration boundary, ensuring the database is never corrupted with malformed output.