profile-pic

MOHAMED THANVEER

Data Engineer with nearly 4 years of experience in designing and implementing scalable data platforms for both on-prem and cloud solutions. Hands-on experience in designing and developing complex data pipelines using Databricks, Azure Data Factory/Azure Synapse to achieve business and data requirements. Good experience in working on cloud platforms such as AWS and Azure using technologies like Databricks, Azure Synapse, Azure Function App, and Snowflake. Detailed experience in handling issues in projects, proactive client communication and project management, determined leadership, and dynamic client relations.

  • Role

    Data Engineer

  • Years of Experience

    4.10 years

Skillsets

  • Github
  • IICS
  • Azure DataBricks
  • SQL
  • Spark SQL
  • SnowSQL
  • PostgreSQL
  • Snowflake
  • Oracle SQL
  • Microsoft SQL
  • Jenkins
  • GitLab
  • Python - 03 Years
  • Databricks
  • Azure Synapse Analytics
  • Azure Function App
  • Azure DevOps
  • AWS
  • Apache NiFi
  • Azure Data Factory - 2 Years
  • Azure - 2 Years
  • PySpark
  • Apache Spark

Professional Summary

4.10Years
  • Apr, 2024 - Present2 yr 4 months

    Data Engineer

    Mastek
  • Aug, 2021 - Apr, 20242 yr 8 months

    Data Engineer

    Avasoft

Applications & Tools Known

  • icon-tool

    Databricks

  • icon-tool

    Azure Synapse Analytics

  • icon-tool

    Azure Data Factory

  • icon-tool

    Snowflake

  • icon-tool

    AWS S3

  • icon-tool

    GitLab

  • icon-tool

    MS SQL

  • icon-tool

    Azure DevOps

Work History

4.10Years

Data Engineer

Mastek
Apr, 2024 - Present2 yr 4 months
    Fidelity Investments (Finance Client) | Nov 24 Till Date Data Engineer Technologies: Apache NiFi, Snowflake, AWS Lambda, Postgres, GitHub, Datadog --- Designed and implemented a reusable file ingestion framework enabling customers to upload data into S3, which automatically trigger AWS Lambda to transfer files to a staged bucket for Snowpipe ingestion into Snowflake, replacing legacy EC2 + SnowSQL workflows and reducing infrastructure costs and manual overhead. Developed a Snowflake wrapper stored procedure to automate data loading from external stages, streamlining the pipeline and achieving a ~60% reduction in onboarding time for new file-based data sources. Onboarded new ServiceNow tables into Aurora PostgreSQL by performing schema mapping, integration design, and data modeling using Apache Kafka plugins, enabling near real-time access for downstream analytics. Additionally, leveraged Apache NiFi to build scalable, low-code data pipelines for ingesting ServiceNow data into Snowflake. Designed and developed a Python-based automation framework for creating Snowflake tables, views, metadata, inventory, and audit objects. Integrated automated unit testing, validation report, and logging features, reducing development and testing effort by ~75% while improving onboarding speed, code quality, and deployment consistency. Onboarded 20+ tables and objects from MS SQL Server and Oracle databases into Snowflake using Apache NiFi, and orchestrated workflow through Control-M. Performed a Proof of Concept for migrating Apache NiFi authentication from PingFederate to Microsoft Entra ID, validating authentication and functionality of the NIFI flow. Designed and implemented Fact and Dimension data models based on ServiceNow data to support reporting and analytics requirements and improve dashboard query performance. Partnered with platform and data teams to streamline Snowflake environments across DEV/SIT/UAT/PRD, resulting in the cleanup of 500+ objects (tables, views, procedures, stages, and functions), improving system maintainability and performance. Performed EC2 rehydration activities across DEV, UAT, and PROD environments, including server restoration, configuration validation, and application readiness checks, ensuring successful environment recovery with minimal downtime and no impact to business operations. New Balance (Manufacturing Client) | Apr 24 Oct 24 Data Engineer Technologies: Azure Blob Storage, Snowflake, ADF, GitHub Designed and implemented an event-driven ADF pipeline with a dynamic, single Snowflake stored procedure to facilitate the seamless ingestion of data from over 150 files across 15 regions. Data is ingested directly from Azure Blob Storage into Snowflake staging tables, triggered immediately upon file arrival using the COPY INTO command. Configured the stored procedure to automatically trigger the corresponding curated procedure for each region once all staging tables for that region are refreshed, ensuring data consistency and efficient processing. Implemented an alerting mechanism within ADF that sends notifications to a Teams channel in case of any pipeline failure, ensuring prompt issue resolution. Developed a robust exception handling framework and created a monitoring table to track and manage ETL-related activities, ensuring transparency and operational control. Recreated existing MS SQL Server views within Snowflake, maintaining continuity and optimizing performance. Implemented various data quality frameworks tailored to the clients specific requirements, ensuring high data integrity throughout the migration process.

Data Engineer

Avasoft
Aug, 2021 - Apr, 20242 yr 8 months
    Danella (Construction Client) | Dec 23 Mar 24 Data Engineer Technologies: AWS S3, Databricks, Spark-SQL, PySpark, GitLab --- The client tasked us with creating an Operational Summary Report for their business, as they currently lack a data warehouse. In response, our proposal involves the development of a Data Lakehouse using databricks to consolidate data from diverse sources such as Samsara and TrackVia. Engaged with business stakeholders, gaining understanding of the entity within their business to derive actionable metrics for each entity. After a comprehensive analysis, designed a data model comprising both dimension and fact tables tailored to accommodate the defined metrics. Collaborated with the Cloud Infra team to provision and configure a Databricks workspace with unity catalog enabled for efficient data management and governance. Designed and Developed extraction notebooks to pull data from Samsara and TrackVia API endpoints. Additionally, created a silver transformation notebook framework to streamline the development phase, reducing the team's effort. Developed databricks notebooks to address business-critical requirements, including the comparison of daily vehicle movements on TrackVia and Samsara, as well as the analysis of daily, week-to-date, month-to-date, and year-to-date purchase orders made by each employee. Implemented Slowly Changing Dimension (SCD) type 2 methodology for critical dimensional tables in accordance with client requirements and data quality framework to check data integrity and data consistency across the Data Lakehouse. The notebooks were scheduled using databricks workflow for 1 hour once during business hours. By implementing the Operational Summary Report, Danella successfully boosted employee engagement with Samsara and TrackVia application by 31% within a month. Goddard Systems (Education Client) | May 23 Nov 23 Data Engineer Technologies: MS SQL, Azure Function App, Azure Storage Queue, Azure Data Factory, GitHub Leveraging SQL Server, the client desires an event-driven data mart solution for Master Data Management, integrating data from various sources and implementing data governance rules to data. Collaborated with business stakeholders to identify different master data objects within their organization and designed a data model through analysis and obtained approval from the Director for the proposed design. Established SQL Server replication between the source MS SQL server databases and the MDM's data mart. Additionally, created triggers and stored procedures to enable event-based data flow from the sources to the destination data mart. Utilized Azure Data Factory (ADF) to integrate with the SOCI, facilitating the extraction of social media information related to schools. Subsequently, load this data into the MDM's data mart using a stored procedure. Developed an Azure Function app to systematically purge lead records based on lead hygiene rules defined by business stakeholders, leading to a 19% reduction in licensing costs with the external marketing vendor. Additionally, use the Azure Function app to send an event message to Azure Storage Queue whenever there are changes in the master data objects. Worked closely with the integration team to build consumable views based on requirements. These views are used by the team to transmit data to marketing vendors like MERGE for running marketing campaigns. Nestle (Supply chain management client) | Jul 22 Apr 23 Data Engineer Technologies: Azure Databricks, Azure Data Lake Storage Gen2, Spark-SQL, PySpark, MS SQL, Azure DevOps Client requires a comprehensive Data Lakehouse solution built on Azure Databricks to consolidate and analyze various data sources including structured and unstructured data. Designed overall cloud data platform by utilizing databricks. Data was imported into delta lake as CSV files from source systems such as Exceedra application and Drive. Developed databricks notebooks to address critical requirements, such as calculating metrics for seasonal marketing product counts and metrics for the number of products on shelves, also implemented incremental data processing through the medallion architecture. --- Worked with the analytics team to gather the requirements and build a logic to design and implement views that satisfy the sales and marketing needs. Data cleansing rules and validation logic were implemented to remove duplicates and invalid customers, resulting in the provision of unique customers and the removal of inactive customers, additionally implemented SCD type 2 for business-critical dimensional tables based on the requirement from the client. Implemented the Data Quality framework with a metadata approach where emails were sent to the respective stakeholders when the data quality checks failed using the Azure Logic app. Rent-A-Center (Retail client) | Oct 21 Jun 22 Data Engineer Technologies: AWS S3, Databricks, Snowflake, Spark-SQL, PySpark, Oracle SQL, GitLab The client requirement was to migrate from an on-premises data warehouse to a cloud-based modern data platform. The legacy ecosystem relied on Oracle and Informatica PowerCenter for data extraction, transformation, and ETL processing. As part of the modernization initiative, the platform was migrated to AWS, with Databricks serving as the ETL and Data Lakehouse platform, Snowflake as the enterprise data warehouse, and Informatica Intelligent Cloud Services (IICS) utilized for cloud-based data ingestion and orchestration where applicable. During the initial migration, historical data was extracted from Oracle and loaded into the AWS S3 staging zone using AWS DMS. Developed Databricks notebooks to process the staged data and load it into Snowflake. Re-engineered existing Informatica PowerCenter mappings and Oracle PL/SQL procedures into scalable Databricks ETL pipelines using PySpark and Spark SQL, improving maintainability and supporting complex business transformations. Worked with IICS to orchestrate cloud data ingestion workflows, monitor pipeline executions, and integrate source data into the AWS data lake for downstream processing. Designed and developed Databricks notebooks to ingest JSON payloads from Amazon S3, flatten nested structures, and load data into the Raw layer. Implemented data quality validations and standardization in the Confirmed layer, followed by business transformations in the Curated layer to build analytics-ready datasets. Developed incremental Delta Lake to Snowflake loading pipelines, enabling efficient Change Data Capture (CDC) processing and optimized warehouse loading.

Achievements

  • Boosted employee engagement with Samsara and TrackVia application by 31% within month
  • Lead to a 19% reduction in licensing costs with the external marketing vendor

Testimonial

Goddard Systems

Sid (Data and Analytics Director)

I want to recognize the efforts that Thanveer has put in for

SSDB since he started with Goddard in June/July of this year.

He has been very diligent and thorough in documenting all data sources, integrations and touch points since he has joined.

He has been very instrumental in implementing the SSDB objects in time giving us the necessary traction when needed and testing the different flows from all scenarios making sure it meets our performance criteria.

He goes above and beyond to make sure our environments, git are in sync and brings up any discrepancies to the table for discussion.

He does the data analysis, creating reports for us to understand the anomalies as they come up

Lastly but not least, he is ready to stretch when needed, putting in long hours regardless of whether he is in office or home

Thank you Thanveer, your efforts are truly appreciated and very commendable!

Major Projects

4Projects

Operational Summary Report for Danella

Dec, 2023 - Mar, 2024 3 months
    Developed a Data Lakehouse for operational reports consolidating data. Designed models and implemented SCD Type 2 for accurate metrics tracking.

Master Data Management at Goddard Systems

May, 2023 - Nov, 2023 6 months
    Event-driven data mart solution for integrating and governing master data using SQL Server and Azure resources.

Nestle Data Lakehouse

Jul, 2022 - Apr, 2023 9 months
    Built a scalable Data Lakehouse using Databricks for consolidated marketing and supply chain analytics.

Data Warehouse Migration for Rent-A-Center

Oct, 2021 - Jun, 2022 8 months
    Migrated on-prem ETL pipelines to Databricks and Snowflake. Implemented modernized transformation workflows.

Education

  • B.E Computer Science and Engineering

    AMS College of Engineering (2021)

Certifications

  • Databricks certified data engineer associate

  • Snowpro core certification

Interests

  • Football
  • Games
  • Long Rides