# AskAILab - [User Stats](https://authorsnote.askailab.online/user-stats.md): How I turn Tripadvisor's public traffic numbers into concrete capacity math for model serving and user profiling. - [FastAPI load performance](https://authorsnote.askailab.online/fast-api-load-performance.md): What a 1-core, 4 GB FastAPI service actually serves per second, what sets that ceiling, and how to measure it without fooling yourself. - [Customer Segmentation](https://authorsnote.askailab.online/customer-segmentation.md): How I group users into segments: the signals, the feature engineering, how I pick and evaluate the clustering method, and how segments get computed offline but served online. - [Calculating time to run a pipeline](https://authorsnote.askailab.online/calculating-time-to-run-a-pipleine.md): How I estimate end-to-end runtime and throughput for a multi-stage data/AI pipeline, and what changes when stages run in parallel. - [Tag clustering algorithm](https://authorsnote.askailab.online/tag-clustering-algorithm.md): How I cluster tags — why they behave nothing like numeric feature vectors, and how hierarchical, k-means, and graph-based methods each behave on tag data. - [Estimating daily load and LLM cost for a travel chatbot](https://authorsnote.askailab.online/user-usage.md): How I estimate Tripadvisor's AI trip-planning chatbot daily token load and LLM API cost from public traffic data, including prompt caching, RAG context sizing, and a sensitivity analysis. - [RAGs](https://authorsnote.askailab.online/rags.md): How I explain retrieval-augmented generation from first principles: the pipeline stage by stage, the chunking and embedding decisions, index choice, how to score retrieval quality, and the failure modes that bite in production. - [Hybrid Logical Clocks](https://authorsnote.askailab.online/hybrid-logical-clocks.md): How one 64-bit integer gives you both wall-clock time and causal ordering, with the failure modes it exists to prevent and a runnable Python implementation. - [Vector clocks for CRDTs](https://authorsnote.askailab.online/vector-clocks-for-crdt.md): How vector clocks establish happens-before and expose write conflicts, why their size is a problem in real clusters, and what dots and interval tree stamps trade away. - [Sequential Database Scan](https://authorsnote.askailab.online/sequential-database-scan.md): How PostgreSQL decides between a sequential scan and an index scan under LIMIT, what each plan shape costs per page and per tuple, and the mitigations that actually hold. - [Optimizing Redis with Lua scripts](https://authorsnote.askailab.online/optimizing-redia-via-lua-script.md): How Redis Lua scripting buys atomicity and fewer round trips, what EVALSHA saves on the wire, and why one slow script is a cluster-wide latency event. - [Caching LLM Chats to Quickly Answer User Queries Without RAG](https://authorsnote.askailab.online/caching-llm-chats-to-quick-answer-user-query-without-rag.md): How prefix caching, context-cache IDs, and semantic answer caches let you skip the retrieval round trip, with the hit-rate math and the invalidation rules that make it safe. - [What to Expect in FDE GenAI Roles](https://authorsnote.askailab.online/what-to-expect-in-fde-gen-ai-roles.md): What a Forward Deployment Engineer working in Gen-AI actually does, what the interview loop tests, how the teaching and knowledge rounds are scored, and how to prepare for both. - [Sample Resume](https://authorsnote.askailab.online/sample-resume.md): A Gen-AI and Forward Deployment Engineer resume template in resume shape, followed by commentary that explains every choice using this resume's own lines as the worked examples. - [Why Need RAG if Gemini Can Answer](https://authorsnote.askailab.online/why-need-rag-if-gemini-can-answer.md): Long context and world knowledge do not remove the need for retrieval: the private-data, freshness, provenance, cost and latency limits, the counter-case where RAG is genuinely unnecessary, and the rule I use to choose. - [Latest Terms in GenAI - 30 Sep](https://authorsnote.askailab.online/latest-terms-in-gen-ai-30-sep.md): A dated snapshot of the Gen-AI vocabulary that mattered on 30 September, with what each term does and why it matters in production. - [Scaling a Vector Database](https://authorsnote.askailab.online/scalablae-vector-database.md): How a vector store actually grows to billions of embeddings: ANN index choice, the recall/latency/memory triangle, quantization arithmetic, sharding and replication, and the hot-partition failure mode. - [Case Study: Building High-Capacity Vector Databases with Open-Source Techniques](https://authorsnote.askailab.online/case-stude.md): A worked case study on taking a vector store to billions of embeddings with open-source components: sizing, architecture, latency budget, recall strategy, and the tradeoffs that decide it. - [Data Size Estimation](https://authorsnote.askailab.online/data-size-estimation.md): The unit ladder I use to size any Gen-AI workload, with worked arithmetic for a book corpus, the whole of English Wikipedia, and a billion-review retrieval index. - [Complete RAG Tutorial](https://authorsnote.askailab.online/complete-rag-tutorial.md): An end-to-end RAG pipeline you can run right now: chunk, embed, index, retrieve, prompt, generate, evaluate, with verified stdlib-only output and the production path for each stage. - [Interview Experience](https://authorsnote.askailab.online/interivew-data.md): Notes from a senior backend loop — a FastAPI rate-limiter race, an atomic Redis sliding-window log, a Postgres index regression, and an active-active CRDT editor — with the answers I gave and the ones I'd give now. This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.