Senior Site Reliability Engineer

Chennai, IndiaFull-timePosted Jul 10, 2026

This role is for one of the Weekday's clients

Salary range: Rs 3000000 - Rs 4500000 (ie INR 30-45 LPA)

Min Experience: 7+ years

Location: Chennai

JobType: full-time

The Senior SRE is responsible for deployment, updates, and operational support for environments hosting our leading client’s cloud-based solutions. This role ensures operational excellence, a seamless client experience, and continuous improvement across infrastructure and delivery processes. The ideal candidate combines strong technical capabilities with the ability to lead delivery through influence and hands-on engineering expertise.

Requirements

Key Responsibilities

  • Manage deployments, upgrades, maintenance, and operational support for cloud environments.
  • Ensure high availability, scalability, performance, and reliability of production systems.
  • Define, monitor, and improve SLAs, SLOs, and SLIs.
  • Drive automation initiatives and Infrastructure as Code (IaC) adoption.
  • Perform Root Cause Analysis (RCA) and implement preventive actions.
  • Optimize cloud infrastructure, operational efficiency, and costs.
  • Enhance monitoring, observability, security, and deployment processes.
  • Collaborate with Engineering, Project Management, Customer Success, and cross-functional teams to deliver reliable services.

Required Skills

  • Strong hands-on experience with AWS  cloud platforms.
  • Expertise in Kubernetes for container orchestration and cluster management.
  • Experience with Terraform for Infrastructure as Code (IaC).
  • Proficiency in Ansible for configuration management and automation.
  • Hands-on experience with Helm for Kubernetes application deployments.
  • Experience managing MariaDB and MongoDB databases in production environments.
  • Strong understanding of CI/CD pipelines, deployment automation, and DevOps practices.
  • Experience with monitoring, observability, logging, and alerting tools (e.g., Prometheus, Grafana, ELK, CloudWatch, Azure Monitor).
  • Good knowledge of Linux administration, networking, DNS, load balancing, and cloud security best practices.
  • Experience troubleshooting production environments, conducting Root Cause Analysis (RCA), and improving platform reliability.
  • Understanding of SRE principles, including SLAs, SLOs, and SLIs.
  • Scripting experience using Bash, Python, or Shell for automation.
  • Excellent problem-solving, communication, and stakeholder management skills.

Preferred Experience

  • Experience working in Site Reliability Engineering, DevOps, or Cloud Operations roles.
  • Experience managing large-scale, production cloud environments.
  • Ability to thrive in a fast-paced, customer-focused environment.
  • Strong analytical mindset with a proactive approach to continuous improvement.

Must-have skills

AWS, Kubernetes, Site Reliability Engineering

Good-to-have skills

Helm Charts, IaC, monitoring

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free