profile-pic

Varshith B

Senior AI Engineer with 5+ years building end-to-end AI systems — from data curation, model training and fine-tuning to inference optimization and production deployment.

Currently at Livepeer working on real-time video diffusion models and AI infrastructure. Previously at Sprinklr (4+ years) building production ML at scale — LLM agents, NLP, computer vision, and ML infrastructure serving enterprise clients.

I go beyond black-box usage — I understand the full stack from model architecture to GPU serving and like to build systems that are fast, reliable, and cost-efficient.

Areas: LLMs, AI Agents, Generative AI, Diffusion Models, NLP, Computer Vision, Model Training & Fine-Tuning, Inference Optimization, TensorRT, PyTorch, Transformers, MLOps, Distributed Systems

Open to Senior AI/ML Engineer, Forward Deployed Engineer, Applied AI Scientist, or Founding Engineer roles.

  • Role

    Senior AI Engineer

  • Years of Experience

    5.8 years

  • Professional Portfolio

    View here

Skillsets

  • MLC
  • Elasticsearch
  • TensorRT
  • Redis
  • PyTorch
  • Python
  • Prometheus
  • ONNX
  • MQTT
  • MongoDB
  • C
  • LangChain
  • Kubernetes
  • Kafka
  • Helm
  • gRPC
  • Grafana
  • Docker
  • C++

Professional Summary

5.8Years
  • Oct, 2024 - Present1 yr 9 months

    Senior AI Engineer

    Livepeer
  • Aug, 2020 - Oct, 20244 yr 2 months

    Lead ML Engineer

    Sprinklr
  • May, 2019 - Jul, 2019 2 months

    Summer Internship

    American Express

Work History

5.8Years

Senior AI Engineer

Livepeer
Oct, 2024 - Present1 yr 9 months
    Rewrote the LongLive real-time video generation model for full torch.compile compatibility with zero graph breaks, achieving 55-60% inference speedup on H100 and 40% on RTX 5090. Also resolved FP8 quantization incompatibilities in the compiled path, unifying FP8, VACE conditioning, and LoRA into a single compiled inference pipeline. Accelerated StreamV2V by 6x to 50 FPS on RTX 4090 (from 8 FPS) by engineering full TensorRT support overcame static graph limitations by externalizing the KV-cache and patching the diffusers library, unblocking real-time workflows for tools like TouchDesigner. Integrated TensorRT support for SDXL-Turbo into the StreamDiffusion framework to enable high resolution real-time generation. Optimized tensor movement between CPU and GPU, setting a new performance standard of 21 FPS for 1024x1024 image generation on a 4090 GPU. Trained and released TemporalNet2 models for SD-Turbo and SDXL architectures, solving critical video consistency issues (flickering). These models are now used by the daydreamlive to enforce temporal coherence in real-time generative streams. Engineered ControlNet support for the StreamDiffusion framework, enabling real-time, controllable image generation with a minimal latency overhead of only 5 ms per ControlNet on the SD-Turbo model. Furthermore, integrated quantization support, increasing frame rates by 30% through optimized model precision. As part of the core infrastructure team, contributed to optimizing Livepeers AI inference infrastructure, helping to identify and resolve critical bugs. This teamwork resulted in a more robust system, achieving a 95% startup success rate and reducing average startup time to 8 seconds. Implemented audio support for Comfystream, enabling real-time audio processing pipelines alongside video workflows. Designed and built a concurrent processing system that handles multiple input streams simultaneously, with features including audio-only, video-only, and combined audio-video modes. Optimized the queue system to efficiently manage different data rates between audio and video streams, resulting in a more responsive user experience.

Lead ML Engineer

Sprinklr
Aug, 2020 - Oct, 20244 yr 2 months
    Machine Learning Smart Compose: Developed a highly-efficient Real-time Next Word Suggestion Service utilizing an Encoder-Decoder Model and MQTT. Abandoning causal generation, the service utilizes conditional probability on possible candidates acquired from the Trie to ensure Compliance, Personalization, and Latency constraints and thereby enabling the service to serve requests under 100 milliseconds. Smart Reply: Played a critical role in the comprehensive lifecycle management, from deployment to maintenance of Context-aware Response Generation System for customer care agents, ensuring optimal scalability, performance and cost efficiency. The system was built on top of a single, generalized GPT2 model, which allowed for the generation of personalized responses appropriate for different partners using special tokens. Logo Detection: Engineered a high-performing logo detection service built on top of YOLOv5 that exceeds the capabilities of Google Cloud Vision(GCV). To avoid re-training the network for each new logo, the service is split into presence and classification steps, enabling rapid integration of new client logos from weeks to 30 minutes with minimal samples. The service detects over 5000 logos currently. Smart Grammar: Developed a Real-Time Grammar Correction Service utilizing a T5 model trained on data generated from wikiedits. The service is designed to detect, highlight, and rectify spelling, punctuation, grammar, and word choice errors made by customer care agents. Subsequently integrated the service into quality management score to evaluate agents, garnering adoption by 80+ clients. FAQ Discovery: Built a system that identifies frequently asked questions (FAQs) from previous customer care conversations using gpt-4 and gpt-3.5, guiding customers towards these FAQs, alleviating the need for direct interaction with customer service representatives. Executive Insights: Built a RAG-based insights engine that analyzes social, listening, and benchmarking data to generate executive-level recommendations. The pipeline clusters enriched content into themes, compiles questions of interest, and uses LLM agents (ReAct framework via Autogen) to produce insight cards covering content strategy, engagement patterns, and thought leadership guidance. Multi-task NLP: Trained a multi-task NLP encoder model that facilitates simultaneous processing of multiple similar NLP tasks. Trained the encoder using MLM first on the domain data to reduce perplexity and then fine-tuned it on downstream tasks. This helped replace multiple single task models, improving latency and reducing resource consumption by 60%. Voice Agent Assist: Developed a Real-Time Agent Assist for Voice, providing comprehensive voice analysis like emotion, loudness, pace of speech, participation ratio to help nudge the customer care agent to provide superior customer service, incorporating a critical Voice Message Compliance feature for real-time detection and masking of inappropriate language using a novel strategy to make already existing wav2vec2 models pseudo auto-regressive which inherently arent, adding an overhead of just 400ms. LLM Inference: Worked on optimizing LLM inference throughput/latency and training time, facilitating a cost-effective transition towards in-house LLMs that are 10x more economical than GPT 3.5 and 4 using TensorRT, MLC, Speculative/Lookahead Decoding, vLLM. Additionally, engaged in open-source contributions to improve functionalities within the MLC framework. MultiPool: Engineered an aio-multiprocess based framework to enhance GPU utilization, necessitating only minor modifications in the form of decorator additions to existing pipelines, boosting throughput across all GPU services by 4 times. Transformer Optimization: Developed an optimization framework for generative models used for conversational products incorporating TensorRT and ONNX accelarations, significantly improving latency by 3 times and reducing the operational costs by 50%.

Summer Internship

American Express
May, 2019 - Jul, 2019 2 months
    Developed a robust and efficient API using Spring Boot, facilitating seamless data synchronization between a collection of dynamic web pages and an internal database to eliminate manual effort, ensuring data consistency and integrity.

Major Projects

1Projects

Slooo

    A Xonsh-based fault-injection framework incorporating shell functionalities and high-level language for evaluating distributed systems fault tolerance. Includes utilities for workload generation and resource usage monitoring. Created APIs for controlled fault injection and uncovered bugs in RethinkDB's re-election mechanism.

Education

  • Bachelors of Technology in Computer Science and Engineering

    Indian Institute of Technology Madras (2020)
  • XII, Telangana State Board in MPC

    Narayana Junior College (2016)
  • X, Telangana State Board

    KKRs Gowtham Concept School (2014)