Site Reliability Engineer
Responsible at the advanced level for writing code and the team’s technical requirements gathering. Independently completes work following banking technology standards and contributes to the overall stability and resiliency of banking technology within the Software Development Lifecycle (SDLC).
This role will help establish and mature Site Reliability Engineering practices across critical banking technology services. The position will focus on improving reliability, observability, incident response, release safety, disaster recovery readiness, automation, and continuous improvement. The ideal candidate will partner closely with application development, infrastructure, operations, and product teams to define standards, implement resilient engineering patterns, reduce operational toil, and support highly available customer-facing services.
Primary Responsibilities
- Work independently and within the boundaries of the approved Software Development Lifecycle (SDLC) to process, design, and develop applications that solve business needs and minimize risk to the Bank by writing clean, secure, and resilient code.
- Author organized, efficient, and secure source code at an advanced level in a minimum of one programming language, applying appropriate data structures, algorithms, and engineering practices to solve business problems.
- Contribute to the design, implementation, and adoption of Site Reliability Engineering standards, practices, and governance across critical technology services.
- Define, implement, and mature Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) for key customer-facing services.
- Establish consistent SLI calculation methodologies and data sources across observability platforms, logs, metrics, dashboards, synthetic monitoring, and related monitoring tools.
- Partner with development teams to implement baseline resiliency patterns, including timeouts, retries, circuit breakers, autoscaling, fault tolerance, and rollback-safe release practices.
- Improve observability maturity by enhancing signal quality, reducing alert noise, increasing actionability, and expanding proactive alerting and synthetic monitoring capabilities.
- Standardize dashboards and reporting to provide end-to-end visibility into service health, customer experience, availability, latency, error rates, throughput, and operational performance.
- Support native Azure monitoring integration and ensure critical workloads have appropriate monitoring, alerting, resiliency, and capacity controls in place.
- Standardize and enhance incident response playbooks, severity models, escalation paths, and recovery procedures to improve consistency and reduce Mean Time to Recovery (MTTR).
- Introduce and support automated incident response capabilities, including event-driven recovery, runbooks, and self-healing approaches where appropriate.
- Contribute to release engineering practices, including Blue-Green and Canary deployments, reversible releases, rollback-safe changes, change risk reduction, and safer production deployments.
- Support disaster recovery validation, failover testing, and controlled chaos engineering exercises to improve resilience and validate recovery capabilities.
- Help scale SLO-driven engineering practices, including error budgets, reliability ownership, and service health accountability within product and engineering teams.
- Participate in or help mature an SRE Community of Practice by contributing to standards, decision pathways, knowledge sharing, and enterprise adoption of reliability practices.
- Support advanced observability capabilities, including trend analysis, anomaly detection, predictive insights, and customer journey observability.
- Institutionalize continuous improvement practices such as blameless postmortems, systemic issue tracking, recurring issue elimination, and operational efficiency improvements.
- Optimize infrastructure and application performance through autoscaling, right-sizing, automation, and reduction of manual operational toil.
- Regularly review pull requests, provide feedback, and execute on the change management of the request.
- Utilize source code management tools to manage and deploy code/applications, ensure compliance with SDLC policies, and support merge conflict resolution.
- Independently analyze and critique technical and business requirements to ensure completeness, accuracy, feasibility, and alignment to resiliency and reliability goals.
- Collaborate, document, and communicate technical implementation details clearly and concisely with business, technical, product, infrastructure, and operational stakeholders.
- Conduct code reviews, providing constructive feedback on code quality, maintainability, security, reliability, and performance improvements.
- Contribute to conversations with business or technical stakeholders and teams regarding application architecture, service design, production readiness, resiliency, and supportability.
- Understand and adhere to the Company’s risk and regulatory standards, policies, and controls in accordance with the Company’s Risk Appetite. Identify risk-related issues requiring escalation to management.
- Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
- Maintain M&T internal control standards, including timely implementation of internal and external audit points together with any issues raised by external regulators, as applicable.
- Complete other related duties as assigned.
Supervisory/Managerial Responsibilities
No supervisory responsibilities.
Education and Experience Required
- Associate’s degree and a minimum of 5 years’ systems analysis and/ or application development work experience or Bachelor's degree and a minimum of 3 years’ systems analysis and/ or application development work experience. In lieu of degree, a combined minimum of 7 years’ education and/or relevant work experience, including a minimum of 3 years’ systems analysis and/or application development work experience
- Advanced proficiency in a minimum of one relevant programming language.
- Experience developing, supporting, or operating applications within a formal Software Development Lifecycle.
- Experience with application reliability, production support, observability, monitoring, incident response, or operational stability practices.
- Experience working with source code management tools and deployment processes.
- Experience analyzing technical and business requirements and translating them into reliable, secure, and scalable technology solutions.
Education and Experience Preferred
- Experience with Site Reliability Engineering practices, including SLOs, SLIs, SLAs, error budgets, incident management, and reliability governance.
- Experience with observability and monitoring tools such as Dynatrace, Azure Monitor, application logs, metrics platforms, dashboards, alerting tools, or synthetic monitoring.
- Experience with Azure cloud services, cloud-native monitoring, autoscaling, and resilient application design.
- Experience implementing resiliency patterns such as retries, timeouts, circuit breakers, graceful degradation, and fault-tolerant design.
- Experience with release engineering practices, including Blue-Green deployments, Canary releases, rollback strategies, and change risk reduction.
- Experience with automated runbooks, event-driven recovery, self-healing systems, or other automation used to reduce operational toil.
- Experience with disaster recovery planning, failover testing, resiliency validation, or chaos engineering.
- Experience contributing to incident response playbooks, severity models, escalation standards, postmortems, and continuous improvement practices.
- Experience partnering with application development, infrastructure, operations, product, and business teams to improve service reliability and customer experience.
- Advanced analytical skills specific to application development, production stability, and service health.
- Experience working in a team environment.
- Ability to work autonomously.
- Ability to multitask on complex projects.
- Strong organizational skills.
- Strong time management skills.
- Proficient verbal and written communication skills.