profile-pic

Venkatesh Mamillapalli

I am a Software Developer with more than 12 years of experience. I have worked with several kinds of systems. My principal area is Java and its frameworks, Talend. I am constant in contact with Java, J2EE, ETL, Unix, Springboot, Microservices, Webservices(SOAP/REST), Jenkins, Maven, ANT, SQL, Postgres, MongoDB, MySQL, Swing, OSGI, HTML, CSS, JavaScript, Angular, GIT, JUnit, Mockito, Grafana, Docker, Kubernetes, UCD.


I am experienced working with legacy and new systems, always using the best practices and improving the codes that I maintain, being cautious by testing and guaranteeing all the functionalities.


The IT industry always grows and change so fast and because of that, every day is a good day to learn new technology and don't stay behind. I am very willing and really like to share knowledge with my colleagues, my greater satisfaction is when i can help or teach someone.

  • Role

    Technical Architect

  • Years of Experience

    11 years

Skillsets

  • Solr
  • MySQL
  • Oracle
  • PostgreSQL
  • Rally
  • Redshift
  • REST
  • Shell Scripting
  • Snowflake
  • SOAP
  • Mssql
  • Springboot
  • SVN
  • Swing
  • Talend
  • Teradata
  • Unix
  • WSDL
  • XML
  • XSD
  • Hive
  • Python
  • SQL
  • Apache CXF
  • AWS
  • Azure
  • Azure Data Factory
  • Databricks
  • Git
  • Groovy
  • Java
  • Jenkins
  • Jira
  • JMS
  • JSON
  • JUnit
  • Maven
  • Microservices
  • MongoDB

Professional Summary

11Years
  • Dec, 2021 - Mar, 20242 yr 3 months

    Technical Architect

    Wipro
  • Jul, 2016 - Nov, 20215 yr 4 months

    Technical Lead

    Capgemini
  • Oct, 2015 - Jun, 2016 8 months

    Senior Software Engineer

    CGI
  • Jun, 2012 - Sep, 20153 yr 3 months

    System Engineer

    TCS

Applications & Tools Known

  • icon-tool

    AWS

  • icon-tool

    Jenkins

  • icon-tool

    Maven

  • icon-tool

    Bitbucket

  • icon-tool

    Azure Data Factory

  • icon-tool

    Databricks

Work History

11Years

Technical Architect

Wipro
Dec, 2021 - Mar, 20242 yr 3 months
    ETL Developer. Created a detailed business analysis, outlining problems, opportunities and solutions for a business. Migrated talend ESB services to camel. Fixed production issues for talend ESB services. Supported Order management services and worked on enhancements. Developed camel services using springboot. Software development using Agile-model. Responsible for Requirement management and clarification with Customer.

Technical Lead

Capgemini
Jul, 2016 - Nov, 20215 yr 4 months
    ETL Developer. Created a detailed business analysis, outlining problems, opportunities and solutions for a business. Created microservices using Talend ESB using different camel components. Using java code and routines downloaded files from Azure blob and then Load data from flat files to STG and then to LDR then applying business rules then load it to master tables. Design and Implement ETL to data load from Source to target databases and for Fact and Slowly Changing Dimensions (SCD) Type 2 to capture the changes. Involved in designing jobs using Talend Components, create complex mappings developing database connections. Job run, execution and debugging with error handling. Implemented java code and routines to download the files from Azure blob and then Load data from flat files. Developed RESTful web services using spring boot based on the client requirement. Created Talend jobs to get data from RESTful services and loading into AWS S3 bucket. Implemented microservices. Integrated java code inside Talend studio by using components like tJavaRow, tJava, tJavaFlex and Routines. Software development using Agile-model. Responsible for Requirement management and clarification with Customer (exchange via Jira).

Senior Software Engineer

CGI
Oct, 2015 - Jun, 2016 8 months
    ETL Developer. Main area of work is data integration, data extracting from different DBs and flat files using Talend Open studio for data integration enterprise and retrieve data files to different types of flat files. Created microservices using Talend ESB using different camel components. Using java code and routines we are using to download the files from Azure blob and then Load data from flat files to STG and then to LDR then applying business rules then load it to master tables. Involved in designing jobs using Talend Components, create complex mappings developing database connections. Job run, execution and debugging with error handling. Implemented java code and routines to download the files from Azure blob and then Load data from flat files. Software development using Agile-model. Responsible for Requirement management and clarification with Customer (exchange via Jira).

System Engineer

TCS
Jun, 2012 - Sep, 20153 yr 3 months
    ETL Developer. First we move all the data from different data bases to MSSQL staging tables using Talend. Depending upon the requirement we create job to load data in to master warehouse. Main area of work is data integration, data extracting from different DBs and flat files using Talend Open studio for data integration. It depends upon client requirements. Developed REST and SOAP web services using spring framework to get the data from different host systems to MyAviva portal using JBOSS. Responsible for CI/CD process. Responsible for Requirement management and clarification with Customer. Software development using Agile-model.

Major Projects

5Projects

Lowes

Dec, 2021 - Present4 yr 9 months
    Involved in migrating Talend ESB services to camel and fixing production issues for Talend ESB services. Support for Order management services and enhancements.

MetLife

Jul, 2019 - Nov, 20212 yr 4 months
    Designed ETL processes for data loading into databases, implementing microservices, and developing RESTful web services.

ICENTER

Jul, 2016 - Jun, 20192 yr 11 months
    Data integration and extraction, creating microservices, managing data transactions and loading processes.

AVIVA

Jun, 2012 - Oct, 20153 yr 4 months
    Moving data to MSSQL staging tables and creating jobs to load data into master warehouse depending on requirements.

Daimler

Nov, 2015 - Jun, 2015
    Developing workflows, updating client data, and ensuring quality data integration as per specifications.

Education

  • Bachelor of Engineering (EEE)

    G. Pullareddy Engineering College (2012)
  • Class XII (Higher Secondary Certificate Examination)

    Nalanda Junior College (2008)
  • Class X (Secondary School Certificate Examination)

    SSVM High School (2006)

A Few Screening Questions Before You Begin

Would you like to understand more about my background by giving a brief introduction of myself? Yeah, my name is Venkatesh. I have been working in the IT industry for the past 11 plus years. I've worked with various technologies. I have experience in talent acquisition for the past 11 years. I have experience in AWS and have exposure to Azure. I also have experience in Java, specifically working with advanced Java. I work with different tools in talent, including talent BI, talent big data, talent cloud, and talent ESP. These experiences have given me a better understanding of REST APIs and how to read and expose services using APIs in talent. I also have a good understanding of CAML, as I successfully migrated talent jobs to CAML-based code. In addition, I have experience in DevOps technologies, including Docker and Kubernetes. I have exposure to repository systems like GitHub, Stash, and Bitbucket. I work on various operating systems, including Linux and Windows. I have a good amount of experience in real SQL and SQL queries, as well as working with databases like SQL and Snowflake. Most recently, I worked on SQL clusters and fixed bugs in the production environment. I've also worked in L1 and L2 support, where I gained a deeper understanding of technical bugs. I'm able to fix those bugs quickly, often within half an hour or an hour. If additional code changes are required, we raise a change request. If there are small configuration changes, we raise a service request. I've worked in various environments, including support, production, development, and maintenance projects. I've also worked in the bank, banking and finance, retail, and insurance domains. I'm a quick learner and can pick up new technologies easily, without any issues. I've worked with Python, C++, and other technologies, and have successfully delivered projects despite initial unfamiliarity with the technologies. I also have on-site experience.

So what would be my approach to the rendering speed and stream processing system, if you can talk a little bit about that. See, I have worked on AWS and Kafka related stuff, I used it in Talend ESP because we have some Kafka components. These Kafka components are there where we used all the Kafka related things in Talend, and this stream processing everything I did in Talend only because AWS requires all those setups and everything in AWS, that's why I did it in Talend ESP, where it has more freedom to do and also how the partitioning is done and how the processing is done, all these things I did in Talend using the Kafka components and also I used some AWS components in Talend, like I used AWS Glue and all these data retrieval parts in AWS, and I connected some Spark related stuff using the Talend ESP, that is somewhat we have done, and for stream processing like as it is, I mean it's a continuous process, right. We have to use the Talend ESP only because Talend ESP has a routing system where you have 24/7 routes where you can retrieve any kind of data, and if it's a file-based system or if it's a Kafka related thing, all these things you can retrieve the data at any point of time, if there is a queue mechanism also, then it will be used in Talend ESP is more useful in that case, that's why you can use the Talend ESP for any of the streaming process if it is like a 24/7 where you are getting the data, then you don't have any problem with that is one good thing about the Talend ESP compared to Talend Data Integration because Talend Data Integration you don't have, I mean this routing mechanism is there in the Talend ESP, and also like where you have hierarchical mapper where you can do the transmission from XML to JSON to XML, all these things are available in the hierarchical mapper, and also you have different conditions where you can set there in the Talend ESP, and when you connect to a straight up ESI, then definitely, if you use the Kafka components and Spark related stuff, all these things you can implement it in the Talend ESP, that is possible in Talend using those Kafka related components and also Spark related components, which you can route in Talend ESP, and AWS components you have, I mean those components where you can utilize that also to connect to AWS and to retrieve the data.

How to integrate a python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database? Python based machine learning model with immediate python on database?

How to design a talent job to not only process the data transfer but also handle error logging and recovery. First, when using a talent job, we have pre-job, main job, and post-job. When using a talent USB, we have a catch mechanism to handle exceptions, similar to try-catch blocks in Java. We can use try-catch blocks to get exceptions and log them as desired. Once an exception occurs or an error is encountered, we want to log it, such as a 500 internal server error or a 404 not found error. We can have a logging table at the end of the job to store these exceptions. Whenever an exception comes, we can catch predefined errors and throw warnings. We can also kill the job at any point in time. These features are possible in talent. Additionally, we can use a tlog catcher to catch all exceptions and errors, and store them in any of the tables. When the first job completes without errors, we can check the table to see if the data is loaded properly, whether it's a full load, incremental load, or any other type of load. The next day, when running the same job, we can check the previous job's status in the logging table. If it's a full load, we can check if the load was completed successfully. If it's a partial load, we can check if it's an incremental load. We can use this information to run the job based on the previous job's status. Context is important in a talent job, and we can specify the context using tags or tmc. This allows us to specify from where to where the talent job should run, such as running the job for only one day or two days. To achieve this, we need to have try-catch blocks, log catches, and warnings in the correct place in the job, so that it's easy to log and get the status of the job in talent.

I leveraged EC2 support instances within my database by utilizing their scalability and flexibility to process large datasets. However, I didn't work with EC2 in Pakistan because I was in a talent developer position where I focused mainly on database and components, and I didn't get much chance to work on EC2. I think that's where my lack of experience in data processing workflows and EC2 technologies like AWS and Lambda functions is evident. On ES3, I've worked for many years, but I didn't get a chance to work on EC2 till now, although I'm very eager to do so. Data processing flows, when you create a talent job and data processing flows, are similar to what I explained to talent ESP. Any file system, like the file component, allows you to get file data at any point in time, irrespective of the talent DA jobs that are scheduled. Once the file is there, it will be processed. If you see the C file, it focuses on the job, which will deploy to any other system where you can have data processing for that particular instance. Whenever you put the file there, you can easily get the file. I think in the same way, EC2 also has different instances where you can see the file size and how the data is being processed. All these things can be done using EC2 instances, and it depends on the size of the data. If the size is more, the processing will be more, and cost optimization will also be affected. The processing time will be more, and the cost will also be higher. However, if the data size is less, the processing will be much faster, and the cost will be very less. The EC2 instance we're using can have a very low amount of cost when the size is less. When the size is more, the EC2 instance will take time to process the entire workload, and different systems will be required for validation and loading data, which will take more time and cost.

You think you can have one log table where all the talent jobs got executed so that you can check whether this particular job was executed successfully before running any of the jobs. We can have one free job concept where it will check that particular log table and that particular log table to see whether the talent job was executed successfully or not. The only thing is you need to pass that particular talent job or else whenever you run it can take the talent job, I mean talent job name also, based on the talent job name, it will check that particular log table whether the talent job was executed successfully or not. If it is not executed successfully, then we can have different conditions, such as if you got a warning based on the error, if it is like having data errors or this kind of thing. We can use the send email component also where you can send the mail at the end of any talent job which is executed successfully and also if there is any error in the talent job. We can have a mailing mechanism where it will check whether the job is successful or not, whether it has any errors, right. All these things can be logged in the same way and at the same time, it can send the email or it can just call the talent job, but before the talent job, it can check the talent job execution whether it is successful or not, and after the talent job, we can update the same table that this talent job was successful at some point of time and what is the duration of the talent job, when it was started, when it was ended, all these things. We can have one log table so that like before running the job, we can check the talent job execution like the previous one or how it is done, and after the execution, how much time it took, all these things. If it is like any timeout or whenever you are running the APA jobs, like suppose if there is any timeout errors, we can have like sleeps, sleep additions and from there like we can get the logs and how much time that talent job, whether it is successfully updated or maybe like it failed with any of the timeout errors, all these things like we can have that kind of thing, but that is what I told right, you can use the pre-job and you can use the post-job, pre-job execution and post-job, it will update the log table where then the job is started and when the job is ended, all these things it can capture, I mean it can capture that entire thing, I mean even for if you are using for AWS in different regions, that is also like we can have like one, I mean database where like it will check the executions, I mean you should have like some place, some database where it can capture all the data, the pre-job and post-job where you can utilize that can be possible using this.

The job involves transferring data from an RDS instance to a STU, how does it work? What steps do you take to debug this primitive issue and ensure that it is successful? When there is no clear pattern, first we need to see whether the STU connection is successful or not. Because when the AWS connection is successful, then we can have a different kind of approach. First and foremost thing is to check the STU connection from the talent server, if it is correct or not. Then, as mentioned, we check the AWS connection, then it will go to RDS input. Whether the RDS input is getting the data or not, we need to see whether you logged it or not. If you have not logged in, then we need to put some logging there and see what kind of data is coming from that RDS input. From there, you need to see the Tmap, how the mapping is done. If there is any null pointer exception, if there is a class cast exception, if there is any runtime exception, we need to see it in the Tmap. But the thing is, we need to make sure that all the connections and also the data which is coming from the jobs, all those things should be having, I mean we need to print that and see the data. Also, when any connection issue comes, immediately we think, if you use any other SME main component, we need to have multiple, if there is any connection failure, we need to wait for some sleep and also after that we can have multiple connection options. It's like for every one minute you can try a connection like that, you can try three times. So that if the connection is established, then you can go for this. Like you can have AWS connection and after that you can go for this RDS input, Tmap, all these things. But make sure that you use the TLABRO to understand what is happening. When the connection is established, definitely there is an issue with the data. But if it is a data issue, definitely I think it will be there in the backlogs. You can have a runtime server where you can see the logs there and you can check. Like if the log is enabled with the info, then you can activate the debug mode so that in which component the exact error comes, then you can easily find it out if you are running the job in a debug mode. That is entirely up to the tag configuration. You need to ask if any of the person is handling these tags or not, then we can enable that in debug mode. All these jobs, you can enable that in debug mode and see at which point this error comes and which component this error comes. So we can enable that in the debug mode.

Not selecting the disk options. Because in the team app, you have, in the front, left zone, options for this and these options. First, whenever you are doing the joinings, whether the joinings are correct or not, you need to see. Like, whenever maybe if it is a cost approach, cost match, unique match, all these things, we need to see. And also, whether you want to reload at each home because when you have, like, 2 inputs, one of the input is keep on changing it, then you can go for this redirecting. So otherwise, you can have the same option. Like, you can just select the data, and then you can use this disk option. There, when you select the display, definitely, auto memory errors will not come. Mostly, it will not come. And if you are still getting auto memory errors, then what you need to do is, you need to increase the size of XMS and XMS. Like, we'll be having options where you can increase the talent process size also so that you won't get the out of memory errors. But most thing is that whenever any of the input or output comes in, make sure that you need to use the proper joints and proper conditions. Like, you can use the proper filters, whenever you want to have any filter kind of thing. I mean, in my case, I prefer that if you are using 1 to 1 mapping, don't use the t map because t map is somewhat large compared to t Java. You can use the t Java row component for 1 to 1 mappings. Other than that, if you are having any of the join condition, then you can use the t join. Suppose if you don't want to have a unique match, all these kind of conditions, use t map instead. And if you are specifically using the t map, then you can have expression filters. You can have different joins and also you can have the different matches. And whenever you have to catch, I mean, and catch this filter whenever any of the filter work from it. Whenever any of the expression filter is dealing with data also, you can have it. All these kind of things, we can do it in a team. But the main thing is that if they are using the team app, we need to make sure that we can have this option, but this option when it goes selected. And also, make sure that if it is increasing more of the thing, then make to increase the XMS and the XMS size of the JVM and that talent process which are the best options which we can do to the d map. So that's it. It will not have any failures.

This number tool I used to track for this tracking purpose because it was useful. We can use some tool there. Yeah. There's tracking that can be tracked. I mean, every server is there as well. The tracking can be done. I think both of these you can try. But ladies are through the process, you know. Definitely, we do have all the servers done, also the beginning of the API down or the beginning of the process.

When we're running the sprint priority site, we have the data available. So, any other part of the tables, it's easy to go, then we can use temporary tables. Otherwise, you can have the CPUs where you can have this CTS. It sounds like adding bulk queries. Like, suppose if you're doing so many joins, like, there are joins, you should not put them in the inner files. You can have those files in the city, and that city will be proper. You won't only have to get the queries on again and again. You can have this or maybe issue of scoring some data already. Temporary data structures where you can have the data. And also, you can even partition tables also. Yeah. That is also cool. If you're doing partitioning, then yeah, that is also wonderful. But the thing is, the bigger guys, such as SQL and all these things. Right? Then I use CTEs also, this temporary thing. Sorry. Temporary things. Then I think that time, we can have those kinds of things. I think based on that, and I think we can get all the. And also, like, what's the suggestion that when you're running any of the queries, then no need to put all those things in the queries because, and whenever you are running those queries using another tool. Then there's a condition for while all these groups are supported by. Then now this connection, we can use it to use your timing. That we can do any kind of large-scale query. So let me know. We can just use temporary tables where we can get the data and we can just reuse it so that, like, you know, you don't need to run the queries again.

What strategies can you utilize for reducing the time it takes to various competitors? The first thing is when you're having the power added, you need to get it from the data. You should not hit the database in my product also, they have used the view, and the view running time is very much not effective. Sometimes the view takes a lot of time to run so that loading the particular dataset takes a lot of time. That's why always we need to use the table in the back-end configuration where the data is instead of putting it from the loop. That is the best option which I suggest for the power delay as we already loaded the data. We can just take it from the table and put it in a code. And also, what other things are done, any of the conditions are acting that we can implement in a view instead of writing everything through our way, all the calculations and everything, then we can have all those things in the view. What about our calculations, what are the things we are looking for? As it is only a select query, right, we are not updating or we are not doing anything with the view, then you can do any other transformation things that you want to show other than a table, then those things you can put in a view, then that you can separate it from Power BI. That can be possible. But if you don't have, like I mean, so many queries, then I suggest that all the data printing up a table that you will be running it for any group. That data, we can load it so that it will be easy because it just takes the data and puts it in the body. Like, you can have different kinds of things. You can have tables, you can have charts. All these things, you can have enough of it. Look. My version is how to get the data from the teams instead of calling the view. If it is very big SQL queries to load the dataset itself, it will take a lot of time because in Power BI, we have different options to get the data get the dataset. That we can do it, and it will reduce the time. The loading time to load the dashboard. Right, Power BI dashboard. That will take very less amount of time. Yeah. That can be possible. I think that strategy, I think it'll be useful so that we can have no data will be loaded with.