Senior Manager - IT Systems Engineering

Supreme IT Park, India · Tower A, IndiaRegularPosted Jul 20, 2026

We are seeking a talented individual to join our GIS at Marsh. This role will be based in Mumbai. This is a hybrid role that has a requirement of working at least three days a week in the office.

Senior Manager – Site Reliability Engineering (SRE) / Reliability Operations

Role Overview

We are looking for a Senior Manager – SRE / Reliability Operations to drive reliability engineering practices across critical services and platforms. In this role, you will establish and mature SLI/SLO/SLA and error budget governance, lead operational excellence (ITIL-aligned incident/problem/change), and partner closely with engineering, architecture, and DevOps teams to reduce toil through automation and improve resilience, observability, and release reliability.

We will count on you to:

  • SRE fitness: SLI/SLO/SLA & error budget management
    • Define, implement, and continuously improve SLIs and SLOs for services/products in collaboration with product and engineering teams
    • Align SLAs with business expectations and ensure operational commitments are measurable and reportable
    • Create and manage error budgets, including:
      • Error budget policies and burn-rate thresholds
      • Release gating recommendations based on error budget status
      • Executive reporting on reliability posture and trade-offs
    • Build SLO dashboards and reliability scorecards for leadership and stakeholders
    • Conduct periodic SLO reviews and drive corrective actions when reliability trends degrade
  • Software engineering, automation, and toil reduction
    • Identify operational toil and repetitive manual work; design and deliver automation to reduce or eliminate it
    • Develop tooling/scripts/services using standard languages (e.g., Python, Go, Java, PowerShell, or similar) aligned to the team’s stack
    • Implement self-healing patterns (automated remediation, auto-rollback, auto-scaling, safe retries)
    • Standardize operational runbooks and embed automation into runbooks wherever possible
    • Improve operational efficiency through event-driven workflows and platform capabilities
  • System design: reliability, resiliency, and stability
    • Partner with architecture and engineering teams to design systems for:
      • High availability, fault tolerance, and graceful degradation
      • Scalability and performance under load
      • Robust dependency management and timeout/retry/circuit-breaker strategies
    • Lead or contribute to resiliency reviews, failure-mode analysis, and capacity planning
    • Design and facilitate resiliency testing (dependency failure simulations, controlled fault injection where appropriate)
    • Define and validate RTO/RPO targets and align with disaster recovery strategies
  • Cloud adoption & reliability enablement
    • Support cloud adoption/migration initiatives (AWS/Azure/GCP as applicable), ensuring reliability-by-design
    • Establish operational patterns for cloud-native services including:
      • Auto-scaling, health checks, multi-AZ/region strategies
      • Infrastructure-as-Code (IaC) practices in partnership with platform/DevOps teams
    • Contribute to secure, compliant, and standardized cloud operations (guardrails, baselines, tagging/monitoring standards)
  • Operational management (ITIL-aligned)
    • Own and/or drive improvements in:
      • Incident Management: triage, mitigation, communications, major incident handling, post-incident reviews
      • Problem Management: root cause analysis, trend analysis, corrective/preventive actions, known error database hygiene
      • Change Management: change risk assessment, change validation, release readiness, change success metrics
    • Facilitate and standardize post-incident reviews (PIRs) focusing on blameless learning, actionable follow-ups, and prevention
    • Improve operational governance and documentation quality (runbooks, SOPs, service ownership, on-call readiness)
  • Observability & monitoring (Datadog)
    • Build and maintain observability practices using Datadog, including metrics, logs, traces (APM), synthetics, and RUM (as applicable)
    • Establish dashboard standards for service health, SLO compliance, and operational KPIs
    • Design alert strategies focused on actionable alerts, noise reduction, correlation, and routing
    • Implement monitoring for golden signals (latency, traffic, errors, saturation) and service-specific signals
    • Tune alerts based on incident learnings and error budget burn rates; reduce false positives
  • DevOps & CI/CD reliability
    • Collaborate with DevOps teams to improve CI/CD pipelines for reliability and speed, including:
      • Automated testing strategy integration (unit/integration/smoke)
      • Deployment strategies (blue/green, canary, rolling, feature flags)
      • Automated rollback and deployment verification
    • Improve release reliability and change failure rate through quality gates and operational readiness checks
    • Promote “shift-left” reliability: testing, observability instrumentation, and operational requirements embedded early in the SDLC

Deliverables / Outcomes (What success looks like)

  • Defined SLIs/SLOs and error budgets for critical services with leadership-ready reporting
  • Measurable reduction in toil and faster operational response through automation
  • Improved reliability KPIs (availability, MTTR, incident volume/severity, change failure rate)
  • Mature Datadog dashboards and alerting with reduced noise and improved signal quality
  • Standardized incident/problem/change practices with consistent PIR execution and follow-through
  • Increased platform resiliency validated via reviews and targeted resiliency testing

What you need to have:

  • 7+ years (or appropriate level) experience in SRE / Production Operations / Platform Engineering / DevOps supporting enterprise applications
  • Strong knowledge of:
    • SLA/SLO/SLI concepts and practical error budget implementation
    • Incident response, troubleshooting, and root cause analysis
    • System reliability patterns (timeouts, retries, backoff, circuit breakers, bulkheads)
  • Proven capability to build automation using one or more programming/scripting languages (e.g., Python/Go/Java/PowerShell)
  • Hands-on experience with Datadog (dashboards, monitors, APM/tracing/logs correlation)
  • Working knowledge of CI/CD pipelines and modern DevOps practices
  • Understanding of cloud fundamentals and cloud operations (AWS/Azure/GCP), including reliability considerations

What makes you stand out?

  • Experience implementing SLO programs at scale across multiple teams/services
  • Experience with Kubernetes and container platforms (or managed equivalents)
  • Infrastructure-as-Code experience (Terraform/CloudFormation/Bicep)
  • Familiarity with ServiceNow (or similar) for ITSM processes and reporting
  • Experience with chaos testing or resiliency validation exercises
  • Strong security and compliance awareness in production operations (least privilege, auditability)

Behavioural Competencies

  • Strong stakeholder management and ability to explain reliability trade-offs to leadership
  • Structured problem-solving and data-driven decision making
  • Ability to operate calmly during major incidents and lead technical bridges
  • Ownership mindset and continuous improvement orientation
  • Strong documentation discipline and ability to drive standardization

Why join our team:

  • We help you be your best through professional development opportunities, interesting work and supportive leaders.
  • We foster a vibrant and inclusive culture where you can work with talented colleagues to create new solutions and have impact for colleagues, clients and communities.
  • Our scale enables us to provide a range of career opportunities, as well as benefits and rewards to enhance your well-being.

Marsh (NYSE: MRSH) is a global leader in risk, reinsurance and capital, people and investments, and management consulting, advising clients in 130 countries. With annual revenue of over $24 billion and more than 90,000 colleagues, Marsh helps build the confidence to thrive through the power of perspective. For more information, visit corporate.marsh.com, or follow us on LinkedIn and X.

Marsh is committed to embracing a diverse, inclusive and flexible work environment. We aim to attract and retain the best people and embrace diversity of age, background, caste, disability, ethnic origin, family duties, gender orientation or expression, gender reassignment, marital status, nationality, parental status, personal or social status, political affiliation, race, religion and beliefs, sex/gender, sexual orientation or expression, skin color, or any other characteristic protected by applicable law.

Marsh is committed to hybrid work, which includes the flexibility of working remotely and the collaboration, connections and professional development benefits of working together in the office. All Marsh colleagues are expected to be in their local office or working onsite with clients at least three days per week. Office-based teams will identify at least one "anchor day" per week on which their full team will be together in person.


Marsh (NYSE: MRSH) is a global leader in risk, reinsurance and capital, people and investments, and management consulting, advising clients in 130 countries. With annual revenue of over $27 billion and more than 95,000 colleagues, Marsh helps build the confidence to thrive through the power of perspective. For more information, visit corporate.marsh.com, or follow us on LinkedIn and X.

Marsh is committed to embracing a diverse, inclusive and flexible work environment. We aim to attract and retain the best people and embrace diversity of age, background, caste, disability, ethnic origin, family duties, gender orientation or expression, gender reassignment, marital status, nationality, parental status, personal or social status, political affiliation, race, religion and beliefs, sex/gender, sexual orientation or expression, skin color, or any other characteristic protected by applicable law.

Marsh is committed to hybrid work, which includes the flexibility of working remotely and the collaboration, connections and professional development benefits of working together in the office. All Marsh colleagues are expected to be in their local office or working onsite with clients at least three days per week. Office-based teams will identify at least one “anchor day” per week on which their full team will be together in person.

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free