profile-pic

Somebuddha Paul

Having 8 Years of Total IT experience with special emphasis in design, development, architecture, administration and implementation of data management and governance applications. Cloud and data processing expert with multiple projects (~6 nos) for Regeneron, IBM, Intellus Learning(part of McMillan) and internship in startup CleverLogik. Seeking to draw on proven software development and engineering skills to increase and improve applications.

  • Role

    Back End Developer

  • Years of Experience

    8.5 years

Skillsets

  • Web Crawling
  • Web Crawling
  • NIFI
  • Kubernetes
  • Dremio
  • Cloud
  • BigPanda
  • AWS
  • Docker - 2 Years
  • Databricks - 2 Years
  • Containers - 3 Years
  • Relational Database - 6 Years
  • API - 4 Years
  • Rest APIs - 4 Years
  • SQL - 6 Years
  • Web Development
  • Airflow
  • Spark - 4 Years
  • Python - 8 Years
  • Privacera
  • NIFI
  • MySQL
  • Mongo DB
  • Kubernetes - 3 Years
  • Jira
  • Image Processing
  • Github
  • Dremio
  • Django
  • Cloud Computing
  • AWS - 6 Years

Professional Summary

8.5Years
  • May, 2022 - Present4 yr 2 months

    Technical Lead

    Visionet Systems
  • May, 2020 - May, 20222 yr

    Senior Software Engineer

    IBM
  • Nov, 2017 - Mar, 20202 yr 4 months

    Software Engineer

    Intellus Learning INC
  • Nov, 2015 - Oct, 20171 yr 11 months

    Software Developer

    Cleverlogik Technologies PVT LTD

Applications & Tools Known

  • icon-tool

    Amazon EKS

  • icon-tool

    S3

  • icon-tool

    EMR

  • icon-tool

    Athena

  • icon-tool

    EC2

  • icon-tool

    SNS

  • icon-tool

    SQS

  • icon-tool

    IAM

  • icon-tool

    Dremio

  • icon-tool

    Airflow

  • icon-tool

    Jenkins

  • icon-tool

    Sharepoint

  • icon-tool

    AWS Athena

  • icon-tool

    Hive

  • icon-tool

    Redshift

  • icon-tool

    Kubernetes

  • icon-tool

    Jira

  • icon-tool

    GitHub

  • icon-tool

    AWS

  • icon-tool

    Django

  • icon-tool

    MySQL

  • icon-tool

    MongoDB

  • icon-tool

    Pandas

  • icon-tool

    Scrapy

  • icon-tool

    BeautifulSoup

  • icon-tool

    Selenium

  • icon-tool

    OpenCV

  • icon-tool

    Celery

Work History

8.5Years

Technical Lead

Visionet Systems
May, 2022 - Present4 yr 2 months
    Experience of design and development of Data Pipelines using Amazon Web Services such as Amazon EKS, MWAA, S3, Lambda functions, EMR, Athena, EC2, SNS, SQS, IAM. Developed Spark applications using PySpark - SQL for data extraction, transformation and aggregation from multiple file formats for analysing & transforming the data to uncover insights into the customer usage patterns. Experience in using SAAS products like Dremio, Privacera, Acryl for Connected datalake platform. ETL using Databricks Expertise in designing and implementing data pipelines using Dremio Perform requirement analysis and produce the technical design specifications. PII tag identification program on multiple datalakes. Create policies for multiple business units and individuals. Product development in Cloud platforms Good understanding of Spark Architecture including Spark Core, Spark SQL, Data Frames, Spark Streaming, Driver Node, Worker Node, Stages, Executors and Tasks. Build Data pipeline using Nifi flow for Ingestion Build Post ingestion data pipeline process using PySpark with Airflow and Jenkins Build connection between SFTP server to extract data from business side Build connection with Sharepoint to extract data and store the data after processing Data quality verification using AWS Athena and Hive and Redshift

Senior Software Engineer

IBM
May, 2020 - May, 20222 yr
    Senior developer in Compute side Practice on Kubernetes Creating , managing, deploying clusters Perform operations on node group(Worker nodes) Perform Billing, Regression and Migration on Kubernetes Cluster and worker nodes Automation testing on Zonal Compute Perform Billing, Regression and Migration on Cloud resources hostOS, kube, genctl packages deployment on clusters Modified Automation codes using Pyhton 3 Deep knowledge on IBM Cloud Platform Perform cloud operations in IBM cloud Create, modify and delete cloud resources manually and automated Comfortable with Jira, Github Orchestrated efficient large-scale software deployments. Worked on new compute VM features like local disk, placement groups

Software Engineer

Intellus Learning INC
Nov, 2017 - Mar, 20202 yr 4 months
    Developing REST API using Django framework. Building the application from scratch. Creating, managing AWS EKS clusters and Node groups Perform cloud operations in AWS(EKS, Lambda, EC2, S3, RDS). Define modules based on requirements Identify ways to improve data reliability, efficiency and quality. Run Application as containers Working with Pandas, DataFrame, MySQL, MongoDB Working with celery script, crontab, Jenkins Image Processing with Python 3(PIL, pytesseract and OpenCV, paramico) Image to text processing with Tesseract and other text processing APIs Web crawling using Scrapy, BeautifulSoup, Selenium Comfortable with Jira, GitHub Familiar with data structure

Software Developer

Cleverlogik Technologies PVT LTD
Nov, 2015 - Oct, 20171 yr 11 months
    Data acquisition using Python Scrapy, Selenium, Beautiful Soup, MongoDB Identify ways to improve data reliability, efficiency and quality Machine Learning with Rare technologies Perform cloud operations in AWS. Developing API using Django framework. Practicing Jira, Github regular basis

Major Projects

8Projects

Connected Data Lake (CDL)

May, 2022 - Present4 yr 2 months
    Creation of access management layer in Privacera, Acryl. Responsible for EKS platform and performed Create, Modify, Upgrade EKS and node groups. Managed the setup having new/existing PODS and provided rollback, restore, replicas, backup and restart. Enabled access management and control using Privecera policies for resource control, specifying audits, discovery scan and other features. Create or modify IAM policies on EKS and Node Groups for accessing other services like S3, Airflow. Senior developer on tag policy creation for UDA platform based on Discovery Scan data PII tag identification on multiple data lakes using Privacera Discovery. Time and memory utilization on discovery program by configuring Privacera YAML files using Kubernetes and Privacera manager deployments. Kafka enablement for Privacera Dev and Prod environment. S3 Tag based policy creation by Discovery generated tags VDS creation, giving access for different BUs on vds.

Data pipeline

May, 2022 - Present4 yr 2 months
    Build Data pipeline using Nifi flow for Ingestion Perform Post ingestion data pipeline process using PySpark with Airflow and Jenkins. Responsible for data processing with ETL process from SFTP to S3 for Commercial business unit. Build connection with Sharepoint to extract data and store the data after processing. Data verification with AWS Athena (MySQL) , Hive and Redshift.

INSIGHTS

    Created APIs for 'Insights' which is a reporting API for instructor who can log in and see activities (Like attendance, student profile) of all his students and course details (like how many people like the course, students who are enrolled, who are all attending etc). Created APIs where Institute Admin can see all and their usage within institute level.

iClicker

    Getting images of different questions from Amazon S3 and convert image into text using different OCR APIs (mainly Tesseract). Identify questions and answers/ multiple choices using data analysis tools and stored in database. Managed all the processes with AWS Lambda. Build the process in AWS serverless architecture (AWS Lambda) Fetch images from S3 buckets Convert the image to text and remove unwanted data in AWS Lambda Process and structure the data and store into AWS RDS (MY SQL) Send Notification using SNS and queue-based system using SQS

Intellus

    Web scraping books, pdf, epub using Scrapy, beautifulsoup and selenium. Get information and normalize the data and store into database

EPAM(Merck project)

    Assembler and PQA Build Automation using Robot Framework Verify and test two different product using Veeva APIs Assembler to create documents using multiple documents and components Migrating documents from one veeva env to another using Python and veeva apis PQA to verify product quality Identifies bugs and report to developer team to fix

Regeneron Pharmaceuticals

    Data pipeline Build Data pipeline using Nifi flow for Ingestion Perform Post ingestion data pipeline process using PySpark with Airflow and Jenkins Responsable for data processing with ETL process from SFTP to S3 for Commercial business unit Build connection with Sharepoint to extract data and store the data after processing Data verification with AWS Athena (MySQL) , Hive and Redshift

IBM Cloud

May, 2020 - May, 20222 yr
    Isolation at CPU level by making sure Regression suite is executed and comes clean on various test conditions. Established AMD CPU processor. Convert or modify automation framework with object-oriented approach.

Education

  • Master of Computer Applications

    Christ University (2016)
  • Bachelor of Computer Applications

    Pailan School of international studies (2013)

Certifications

  • Aws cloud practitioner

AI-interview Questions & Answers

Hello. I'm Zumbhuda Paul with 8 years of experience in cloud and data engineering platforms. I'm an expert in building data pipelines using ETLs, using PySpark and different AWS services, independent and reusable Python projects, and with using Django, AWS, optimizing existing Python projects. Currently, I'm working in Vision as a technical lead, developing my skills, like how to handle teams, distribution of work, planning for roadmaps for my team. I contributed to 2 different projects here, the DID pipeline and connected to the platform building in DID. Developing an ETL pipeline using NiFi, AWS services like Airflow, EMR, S3, RDB with device database like Hive Athena. In CDL, I developed a connected data platform, building a platform for managing access to all users, by creating policies for groups and data lakes with Kubernetes and AWS services like EKS, Node Group, S3, IAM, and SaaS products like Privacera, Dremio. Previously, I was in IBM as a senior software engineer for 2 years and was developing IBM Cloud as a product. I was responsible for developing new VM features and managing compute cloud services like clusters, routers, network, and groups using Python and Kubernetes. Here, I learned about cloud services in-depth. Before IBM, I was in a company named Intellus Learning as a software engineer and worked on different products. One was iClicker and another one was Insights. The Insights application is a reporting application for instructors and students. We developed this product from scratch using Python, Django, AWS services like EC2, and DBs like MongoDB and MySQL. I created an image processing pipeline to convert images to meaningful text using Python and AWS services like Lambda, SNS, SQS, S3 API Gateways, and MySQL. So, here, I learned about building independent and usable Python projects.

Developing a company on how to ensure the data quality and integrity through the pipeline. When developing a complex ETL process, ensuring data quality and integrity throughout the pipeline involves several key practices. Data validation, for example, involves implementing data validation checks at each stage of data processing to ensure data conforms to expected formats, ranges, and business rules. Data profiling also involves analyzing source data to understand its structure, content, and quality, as well as handling anomalous, missing values, and duplicates before loading. Error handling, in particular, requires designing a robust error handling and logging mechanism to capture errors, enabling quick identification and resolution. Automated testing is also essential, as it allows for the development of automated testing capabilities to validate data transformation, consistency, and correctness at various pipeline stages, track data lineage, and ensure transparency and traceability. Enforcing consistent standards, such as naming conventions, data types, and code standards across the ETL process, can also maintain uniformity and reduce errors. Other key practices include monitoring and alerts, backup and recovery, and incremental loading. By integrating these practices, you can ensure data quality and integrity throughout your data processing and deliver accurate data to downstream applications.

Planning for data disaster recovery in Snowflake involves several key steps. One is understanding requirements. Another one is data replication, backups, and snapshots, automated failures, then access controls, testing, documentation, and continuous monitoring. Understanding requirements is defined by the recovery time objective (RTO) and recovery point objective (RPO) based on business needs. Data applications utilize Snowflake in multiple regions, requiring us to replicate data across regions. We should take backups and snapshots on a regular basis. We can schedule automated failures and configure the process to quickly switch back to the previous replica in the event of a disaster. Access controls ensure that proper security measures are in place for the replica data, which is unnecessary during a disaster. Testing is regulated throughout the disaster recovery plan simulation process. Documentation should include everything, including roles, decision plans, responsibilities, and detailed recording steps. And continuous monitoring helps improve the recovery process by detecting early stages when the recovery process is prompted.

To deploy a Spark application in AWS, ensuring scalability and cost efficiency, we would use Amazon EMR. First, create an EMR cluster and select Spark as the application. Choose a mix of on-demand and spot instances to optimize costs. Next, enable auto-scaling to dynamically adjust the number of instances based on the workload. For cost-effectiveness, specify multiple instance types and pricing options to achieve a flexible and cost-effective configuration. For data storage, use AWS S3, which provides data consistency. To execute a Spark job, use an EMR step or automate the service using the console. For monitoring and logging, use CloudWatch to set alarms and track cluster performance. To troubleshoot issues, enable logging to S3. Finally, configure the cluster to auto-terminate after the job is completed, which will help save costs.

How can you use the skill to enhance the security and reliability of data processing system? How can you use Linux system administration skills to enhance the security and reliability of data processing system. Using Linux system administration skills, you can enhance the security and reliability of data processing system through following practices. The one is user and permission management. They implement strict user access control management and group policies. Use principles at least the principle of least privilege to restrict access to sensitive data and use system resources, a firewall, and network security. Configure a firewall, use a secure communication protocol like SSH, SSL, TLS to encrypt data transmission. System updates and patching: apply regular system patching to the Linux operating system and install software when you install software. Monitoring and logging: check the monitoring to promote system performance. You can use software to monitor with promoters to track performance, predict anomalies, configure logging like system logs or audit logs. You can use this to do system logging. Back up and recovery systems: we can use this.

What techniques would you use to handle schema evolution in Snowflake data warehousing without causing downstream for customers to handle schema evolution in Snowflake without causing downtime for customers, you can use the following techniques. 0-copy cloning. Create 0-copy clones of the database or schema to make changes without affecting the live environment. Time travel. Utilize Snowflake's time travel feature to revert to previous states of the data if needed, ensuring data integrity. And during schema change, apply schema changes in an online, non-disruptive manner by adding new columns or tables without dropping existing ones, using values and maintaining backwards compatibility with the existing schema while transitioning to the new schema. Transactional DDL. Leverage Snowflake's support for transactional DDL to ensure schema changes are automated and do not leave the database in an inconsistent state. State rollout implements schema changes in stages, gradually rolling out updates and validating each step to minimize impact.

In summing up the following pattern, what is intended to check if an element is in the list and the element itself. We don't need to run the loop, which is an input list. There is no element called "input chat." One thing is a wrong value. There is an input list. There is nothing called "input chat." And the second thing is you can directly convert the list to a set. Automatically, duplicates will be deleted. We don't need to run a loop. And we have to return the set. Examine the following code is needed to check if all elements are unique. Now, after getting the set, you have to convert the set back to a list and then check if the input list and the output list are equal or not. If they are equal, then return true. Otherwise, return false.

There is no need for a group by here. We can order by, we can do.

To handle increment loading in an ETL pipeline that ingests data into Snowflake from multiple changing data sources, we can follow this strategy: 1. Implement CDC mechanism, for example, a log-based CDC tool, to capture changes, inserts, updates, and deletes from the data source, and load the captured data into a staging area in Snowflake. 2. Dedicatedly staging tables for each data source to temporarily store incremental data, and ensure each source data has a timestamp column to track the latest modification time, facilitating the identification of new and updated records. 3. The ETL process extracts the incremental data based on the timestamp or serial mechanism, applies necessary transformation, data cleaning, and duplication, and then merges the incremental data into the target Snowflake table using a merge statement to handle inserts, updates, and deletes. 4. Schedule a job to run regularly in intervals, hourly or daily, to process and load the incremental data. 5. Implement error handling and monitoring, and include data validation checks to ensure the accuracy of the loaded data. By following this strategy, you can implement efficient incremental loading in your ETL pipeline, handling changes from multiple data sources through staging and transformation.

To transition to an architecture that supports both real-time and batch processing, and optimizes for Python-based computation, modernization, and shared components, there are several steps to consider. First, for real-time data ingestion, use tools like Apache Kafka or Amazon Kinesis to handle similar data streams. For batch storage, utilize AWS S3 or Hadoop HDFS. For real-time data processing, employ Apache Flink or Apache Spark Streaming for Python-based computation. For batch processing, apply PySpark for Python-based computation. Data storage can utilize real-time data and batch data together in a scalable and flexible manner using Snowflake. For orchestration, Apache Airflow can be used to schedule and manage both real-time and batch computation. Additionally, ETL pipelines can be leveraged. In monitoring and logging, the same tools can be used for both real-time and batch processing.

There is a system for managing cross-region data applications within AWS. Another is the entire processing system.