profile-pic

Supreeth Gowda

Build and maintain data solutions for enterprise-grade applications using various AWS services and data engineering tools.

Building and maintaining configuration driven Data ingestion framework and Data Orchestration framework using AWS Services such as Lambda, DynamoDB, MWAA and Python

Building and Maintaining the Modern Data Stack built using Fivetran, Airflow, Snowflake (Enterprise Data Platform), Data Build Tool (DBT) and PowerBi

Hackathons involving Amazon Bedrock/Chat-GPT for image recognition and web scrapping data analysis.


  • Role

    Data Engineering Consultant

  • Years of Experience

    19.1 years

Skillsets

  • Parquet
  • Excel
  • Ganglia
  • GenAI
  • Github
  • Governance
  • HP LoadRunner
  • Java
  • Lake Formation
  • Monitoring
  • MySQL
  • Docker
  • performance
  • PySpark
  • RapidMiner
  • Rbac
  • Redshift
  • Resource access manager
  • SiteScope
  • Spark
  • Weka
  • Amazon AppFlow
  • Terraform - 5 Years
  • Apache Airflow
  • AWS
  • dbt
  • EMR
  • Fivetran
  • Power BI
  • Snowflake
  • Machine Learning
  • Alerting
  • Python - 8 Years
  • Ansible
  • Apache Hudi
  • Apache NiFi
  • Athena
  • C
  • cost optimization
  • CSV
  • Data Architecture
  • Delta Lake

Professional Summary

19.1Years
  • Jan, 2023 - Present3 yr 6 months

    Data Engineering Consultant

    Mindera
  • Nov, 2021 - Present4 yr 8 months

    Independent Data Engineering Consultant (Part-time)

    Encore Data Intelligence Services
  • Jun, 2021 - Nov, 2021 5 months

    Data Engineer II

    Amazon Web Services
  • Aug, 2012 - May, 20141 yr 9 months

    Research Assistant

    Worcester Polytechnic Institute
  • Jun, 2014 - Jul, 20184 yr 1 month

    Software Test Engineer, Data Services

    Sonos
  • Jul, 2018 - May, 20212 yr 10 months

    Big Data Engineer

    Meredith Corporation
  • Mar, 2010 - Jan, 20121 yr 10 months

    Performance Tester

    CGI
  • Senior Software Development Engineer in Test

    Meredith Corporation

Applications & Tools Known

  • icon-tool

    AWS Lambda

  • icon-tool

    DynamoDB

  • icon-tool

    Fivetran

  • icon-tool

    Snowflake

  • icon-tool

    PowerBi

  • icon-tool

    Python

  • icon-tool

    Apache Airflow

  • icon-tool

    Amazon Athena

  • icon-tool

    Delta Lake

  • icon-tool

    Redshift

  • icon-tool

    Spark

  • icon-tool

    Glue

  • icon-tool

    Jupyter Notebook

  • icon-tool

    Postman

  • icon-tool

    MySQL

  • icon-tool

    Ansible

  • icon-tool

    Docker

  • icon-tool

    GitHub

  • icon-tool

    Pycharm

  • icon-tool

    Jenkins

  • icon-tool

    Perforce

  • icon-tool

    Eclipse

  • icon-tool

    AWS

  • icon-tool

    Lambda

  • icon-tool

    Airflow

  • icon-tool

    Amazon Bedrock

  • icon-tool

    Amazon EMR

  • icon-tool

    AWS Glue

  • icon-tool

    Jupyter Notebook

  • icon-tool

    GitHub

  • icon-tool

    PowerBI

  • icon-tool

    AWS services

  • icon-tool

    Terraform

  • icon-tool

    PagerDuty

  • icon-tool

    S3

  • icon-tool

    Redshift

  • icon-tool

    Apache Spark

  • icon-tool

    Zeppelin

  • icon-tool

    Splunk

  • icon-tool

    AWS CLI

  • icon-tool

    RapidMiner

  • icon-tool

    Weka

Work History

19.1Years

Data Engineering Consultant

Mindera
Jan, 2023 - Present3 yr 6 months
    Led the production implementation of a Snowflake-to-Salesforce integration using Amazon AppFlow, defining the integration architecture, authentication strategy, and deployment approach while coordinating Salesforce, DevOps, QA, and Data Engineering teams to deliver a secure, production-ready solution. Designed and evolved a modern data platform using Snowflake, Airflow (MWAA), dbt, Fivetran, Power BI, and AWS, building reusable configuration-driven ingestion and orchestration frameworks supporting scalable batch and streaming data pipelines. Administered the Snowflake platform during its adoption across the engineering team, implementing RBAC-based access control, governance, and workload management to support secure and scalable data operations. Built Snowflake cost and utilization dashboards to identify expensive queries and warehouse usage patterns, enabling data-driven optimization that reduced critical pipeline runtimes by nearly 50% while lowering unnecessary compute costs. Automated infrastructure deployments using Terraform and developed centralized error logging, monitoring, and alerting frameworks, improving platform reliability and operational visibility. Developed custom Python utilities to parse and validate CSV and Excel files before ingestion into Snowflake, improving data quality and ingestion reliability.

Independent Data Engineering Consultant (Part-time)

Encore Data Intelligence Services
Nov, 2021 - Present4 yr 8 months
    Providing independent data engineering consulting for a healthcare provider, designing and maintaining a data aggregation solution that consolidates disparate datasets into standardized reporting formats for business stakeholders. Built Python-based data processing and reporting workflows to automate data consolidation, improve consistency, and support ongoing business reporting requirements.

Data Engineer II

Amazon Web Services
Jun, 2021 - Nov, 2021 5 months
    Supported a centralized enterprise data warehouse managing 45-50 PB of data, providing production support and technical guidance for data integration across multiple engineering teams. Managed and optimized thousands of ETL workflows supporting the migration from Amazon Redshift to a modern Data Lake architecture while ensuring high platform reliability. Built Apache Spark on EMR pipelines to offload and transform Redshift data into Parquet, enabling Redshift Spectrum and a scalable Data Lakehouse architecture. Contributed to an enterprise AWS Glue Catalog sharing platform using AWS Glue, Lake Formation, Resource Access Manager (RAM), and Redshift Spectrum, reducing data duplication and improving cross-team data sharing.

Big Data Engineer

Meredith Corporation
Jul, 2018 - May, 20212 yr 10 months
    Designed and built an automated CCPA data deletion framework that removed customer data across downstream systems within minutes, reducing a 45-day regulatory SLA to near real-time execution. Evaluated Apache Hudi, Delta Lake, Snowflake, and Athena CTAS through proof-of-concepts before selecting the optimal production implementation. Designed and implemented an enterprise AWS Data Lake using AWS Glue, Athena, EMR, Glue Data Catalog, and PySpark for large-scale data ingestion and transformation. Built and orchestrated scalable data processing workflows using Apache Airflow, ensuring reliable and timely data availability.

Software Test Engineer, Data Services

Sonos
Jun, 2014 - Jul, 20184 yr 1 month
    Built a PySpark validation framework performing byte-level data comparison across distributed systems, uncovering data inconsistencies that traditional sample-based testing failed to detect. Performed root-cause analysis through automated error log investigation, identifying defects in upstream processing logic before production deployment. Designed and implemented data pipelines using Python, PySpark, MySQL, AWS, Apache NiFi, Docker, Ansible, Ganglia, and GitHub. Developed and presented engineering innovations at internal hackathons involving automation, voice integration, testing tools, and PySpark-based solutions.

Research Assistant

Worcester Polytechnic Institute
Aug, 2012 - May, 20141 yr 9 months
    Research Assistant (Data Mining and Software Development) - co-advisor was the President of the International Educational Data Mining Society. Developed Java-based automation using RapidMiner and Weka APIs to build machine learning detectors and engineer complex features from educational datasets. Built feature extraction pipelines from log files and relational databases to identify student learning behaviors including boredom, confusion, frustration, and engagement within Intelligent Tutoring Systems (ASSISTments).

Performance Tester

CGI
Mar, 2010 - Jan, 20121 yr 10 months
    Participated in performance testing strategy, requirements analysis, and test plan development for enterprise web applications using HP LoadRunner. Developed automated performance test scripts using HTTP/HTML, Ajax, C, and Java, incorporating reusable components and robust error handling. Configured monitoring with SiteScope, executed load, stress, spike, and endurance testing, analyzed system bottlenecks, and collaborated with development teams to improve application performance.

Senior Software Development Engineer in Test

Meredith Corporation
    Built automated validation frameworks using AWS Glue, EMR, Spark, and Python to verify large-scale ETL pipelines across data warehouses, databases, and Amazon S3.

Achievements

  • Pat on the back - recognizing and appreciating the commitment at work at CGI Group Inc.
  • Corona - recognizing the 'most successful team' for Quarter - FY 11 at CGI Group Inc.
  • Pat on the back
  • Most successful team
  • Corona

Education

  • Master of Science (MS), Learning Sciences and Technologies

    Worcester Polytechnic Institute (2014)
  • Bachelor of Engineering (BE), Computer Science

    Visvesvaraya Technological University (2009)

Certifications

  • Certified programmer for the java platform (scjp) 5.0

  • Certified programmer for the java 2 platform (scjp) 5.0