Writing on AI systems, MLOps, LLMs, edge deployment, and the occasional deep dive into what actually works in production.
After shipping NOVAR, a document RAG system backed by FastAPI, ChromaDB and Gemini, there are real gaps between what the tutorials show you and what actually holds up under real usage. This is a writeup of the decisions that mattered: chunking strategy, retrieval tuning, streaming SSE responses, and session isolation.
Most MLOps content assumes you have a full platform team. This is about getting a model from notebook to monitored production endpoint as a solo engineer, using Docker, a small FastAPI wrapper, and GitHub Actions — nothing more.
"Edge AI" gets used to mean everything from a quantized model on a Raspberry Pi to running inference in a browser. Here's what it means in the context of MiraiQ, the real constraints it imposes, and the architecture decisions that follow from them.
Implementing the Cox proportional hazards loss in TensorFlow for the breast cancer survival prediction project exposed a few sharp edges — tied event times, how to handle censored samples in batched training, and numerical stability. Notes on how each was resolved.