SaaS Engineer (Cloud)
Temenos IndiaOct, 2021 - Apr, 20253 yr 6 months
Delivered uninterrupted L2 cloud support for Temenos T24 banking suite across 3.5 years, supporting mission-critical operations for global banking clients. Delivered L2 application support ensuring high availability and seamless user experience for global banking clients. Monitored application health using Grafana and SolarWinds dashboards, tracking key SLIs (latency, error rate, availability) against defined SLOs. Troubleshot P1 issues by analyzing server logs, application metrics, and distributed traces; generated detailed RCA reports for stakeholders. Led and managed P1/P2 incident bridge calls, coordinating with cross-functional teams to drive quick resolution and minimize business impact. Acted as primary incident coordinator, ensuring timely stakeholder communication and SLA adherence during critical outages. Tracked SLIs (response time, uptime, error rates) and ensured services remained within defined SLO and SLA boundaries. Performed root cause analysis (RCA) for production incidents and implemented preventive measures to reduce recurrence. Utilized Squadcast for alerting and incident escalation, ensuring rapid response to system issues. Managed incident, service request, and change tickets using ServiceNow and Jira within defined SLAs. Analyzed impact of system changes and identified opportunities to enhance customer experience and self-service capabilities. Created and maintained SOPs, work instructions, and test scenarios to ensure consistency in service delivery. Acted as single point of contact for users, coordinating with internal teams to resolve issues efficiently. Informally led a team of junior SaaS engineers, conducting knowledge-transfer sessions on T24 platform troubleshooting, log analysis, and SLO tracking. Reviewed and validated tickets raised by junior team members, providing guidance on prioritization, RCA documentation, and escalation paths. Onboarded new team members by walking them through SOPs, runbooks, and monitoring dashboards, reducing ramp-up time significantly. Monitored ticket queues and performed regular follow-ups to ensure closure and customer satisfaction.