Data and ML Engineer
StreamingoMay, 2025 - Present1 yr 4 months
Built and deployed production RAG pipelines using LangChain, enabling enterprise document retrieval, context-aware reasoning, and tool-assisted question answering. Fine-tuned open-source LLMs (Qwen Family) using PEFT with LoRA and QLoRA, enabling efficient on-premise inference while significantly reducing GPU memory requirements. Designed and implemented a production multi-agent AI platform using Google ADK and LangGraph with stateful orchestration, conditional routing, tool calling, and memory for customer support and enterprise documentation workflows. Implemented systematic LLM evaluation framework tracking correctness, latency, and hallucination control metrics across agentic workflows, enabling rapid iteration cycles. Improved model accuracy through spatial and temporal analytics, identifying performance gaps and implementing post-processing fixes. Built end-to-end MLOps pipelines for multiple models, handling training, inference, and deployment in production. Tuned CUDA settings for YOLO and MTO models, improving inference speed and GPU efficiency. Developed reusable Python SDKs and internal libraries to standardize model training, inference, and deployment across AI teams. Restructured inference workflow to run projects in parallel, cutting daily processing time by 5 hours. Set up CI/CD pipelines with Bitbucket and Jenkins to automate builds, tests, and deployments. Led migration from service-based setup to cloud-native architecture using Helm and Kubernetes, improving scalability. Consolidated all GPU machines into a single Kubernetes cluster, removing manual GPU management and improving resource usage. Managed GCP infrastructure (GKE, GCS, networking), optimizing cloud operations and reducing annual costs by 10-12%.