profile-pic

Amit kumar singh

Lead Software Engineer with over 10 years of experience in product development and artificial intelligence across analytics, telecom, and after-sales domains. Adept at evaluating business needs and implementing comprehensive strategies to deliver products that enhance revenue and drive growth.
  • Role

    Lead Software Engineer

  • Years of Experience

    9.3 years

Skillsets

  • Kubernetes
  • Celery/redis
  • Ai / ml stack
  • Terraform
  • SQL
  • REST
  • JavaScript
  • Hibernate
  • Apache Kafka
  • AWS - 9 Years
  • gRPC
  • GCP
  • Docker
  • Spring Boot
  • Python
  • Java
  • Java - 11 Years

Professional Summary

9.3Years
  • Apr, 2025 - Present1 yr 4 months

    Lead Software Engineer

    Visa
  • Apr, 2021 - Mar, 20253 yr 11 months

    Lead Data Engineer / ML Ops Engineer

    FICO
  • Dec, 2019 - Apr, 20211 yr 4 months

    Senior Data / ML Ops Engineer

    Syncron India
  • Feb, 2017 - Dec, 20192 yr 10 months

    Software Engineer II

    Accenture Solutions

Applications & Tools Known

  • icon-tool

    gRPC

  • icon-tool

    REST APIs

Work History

9.3Years

Lead Software Engineer

Visa
Apr, 2025 - Present1 yr 4 months
    Designed low-latency auth and token pipelines handling 8,000 TPS at peak infra patterns directly transferable to high-concurrency LLM inference gateways. Introduced Kafka-based event sourcing for identity events, strengthening auditability and system resilience. Delivered sub-100ms auth paths across distributed microservices, reducing checkout failures by 1%. Drove cross-team architecture standards for gRPC service contracts adopted across 4+ payment teams; optimized Spring Boot services improving P95 stability by 5%. Led merchant onboarding initiative increasing adoption by 12% within 90 days, expanding CyberSource transaction volume.

Lead Data Engineer / ML Ops Engineer

FICO
Apr, 2021 - Mar, 20253 yr 11 months
    Architected event sourcing for large-scale IAM migrations ensuring real-time processing and data consistency; reduced IAM pipeline latency by 40% using Kafka Streams and gRPC. Built and operated ML model deployment and monitoring infrastructure on Kubernetes (EKS/GKE) with Terraform IaC the same orchestration patterns foundational to production LLM serving. Delivered CI/CD pipelines for multi-cloud deployments improving release velocity by 30%; migrated legacy services to cloud-native workloads, enhancing scalability by 10%. Improved test coverage from 60% to 92%, cutting production defects by 68%; reduced MTTR by 60% through automated rollback and targeted health checks.

Senior Data / ML Ops Engineer

Syncron India
Dec, 2019 - Apr, 20211 yr 4 months
    Built an end-to-end MLOps framework for deploying, monitoring, and rolling back ML models using Docker, Kubernetes, and Airflow established canary release workflows that reduced deployment risk by 6% and shortened lead time by 22 days. Architected observability pipelines (Prometheus/Grafana) that reduced ML model downtime by 50% and tracked inference accuracy, achieving >95% accuracy in production. Optimized service communication via gRPC and HTTP/2, improving response times by 35%.

Software Engineer II

Accenture Solutions
Feb, 2017 - Dec, 20192 yr 10 months
    Transitioned telecom systems from legacy monoliths to Spring Boot microservices; developed AdTech ML components improving audience targeting accuracy by 6%. Implemented Redis caching strategies and established unit/integration testing frameworks, raising coverage and enforcing quality gates.

Major Projects

2Projects

Lex-India Compound AI Legal Reasoning Engine

    Architected an autonomous, hallucination-free Agentic RAG system for Indian jurisprudence enforcing strict legal hierarchy (Constitution Supreme Court High Court) via a 13-step deterministic Python state machine no LangChain, no framework lock-in. Engineered a multi-tier hybrid search pipeline (Dense Vectors, Sparse BM25 with morphological stemming, Reciprocal Rank Fusion, and Cross-Encoder reranking) to achieve ultra-high precision on complex legal texts. Pushed jurisdiction authority logic into Vector DB metadata (SC weight=1.0, HC=0.7, District=0.2), making legal precedence algorithmic rather than LLM prompt-dependent. Implemented a kill-switch abstention engine with domain-specific rules detecting repealed statutes (IPCBNS), wrong jurisdictions, and non-binding rulings ensuring zero hallucination liability in production. Optimized to run fully locally on RTX 5060 Ti (16GB VRAM) using vLLM with AWQ 4-bit quantization; zero data sent to external APIs.

InstructifyAI Privacy-First Enterprise LLM Platform

    Built a local, zero-data-leakage AI inference platform for regulated industries (Finance, Insurance, Healthcare) with resource-aware infrastructure that auto-detects hardware and falls back between GPU (HuggingFace) and CPU (llama.cpp) backends. Engineered a document-scoped inference gateway with dynamic LoRA/QLoRA/DoRA adapter loading at runtime, preventing generic model hallucinations on proprietary internal policy documents. Designed multi-teacher ensemble generation strategies (Vote / Concat / Rerank) to improve deterministic accuracy for audit-heavy compliance workflows. Built a human-in-the-loop curation pipeline via Label Studio webhooks with full data provenance, versioned taxonomies, and audit logs. Deployed on a distributed microservices stack: FastAPI + Celery + Redis + PostgreSQL + MinIO, containerized via Docker/Podman with single-command make targets for GPU and CPU variants.

Education

  • Bachelor of Technology, Computer Science

    Biju Patnaik University of Technology (2013)