Senior Manager, Site Reliability Engineering
OPOWER
Reston, VA$122k–$264kPosted Jul 22, 2026
Skip to main content.
sitemap
Profile
Sign Out
View More Jobs
Senior Manager, Site Reliability Engineering
Reston, VA, United StatesAustin, TX, United States
Be the First to Apply
Job Identification
340641
Job Category
Product and Research
Posting Date
07/21/2026, 02:29 PM
Role
People Manager
Job Type
Regular Employee
Does this position require a security clearance?
Yes
Years
10+ years
Applicants
Less than 10 applicants
Additional Info
Visa / work permit sponsorship is not available for this position
Applicants are required to read, write, and speak the following languages
English
Job Description
Capacity Ingestion and Management:
- Supports team members designing and architecting infrastructure and/or service, sharing guidance on practices and terms for reliability and functionality.
- Supervises team members and provides direction to ensure accurate forecasting of demands for infrastructure and response to capacity needs, ensuring systems have sufficient resources to handle current and future workloads and identifying resource gaps.
- Maintains a collaborative relationship with the software development team to develop infrastructures, ensuring features are reliable and scalable according to deployment requirements.
- Implements expectations for identifying opportunities for prototyping and manages prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding) to explore novel approaches.
Incident and Service Lifecycle Management:
- Monitors data collection, triage, technical analysis, and redirection, ensuring team members maintain and optimize operations and infrastructure reliability.
- Provides support to team members monitoring services, ensuring they maintain up-to-date knowledge of performance and document their condition.
- Leverages advanced knowledge to aid team members in performing incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
- Monitors comprehensive health and performance reporting and ensures team members take appropriate actions based on trends in data.
- Ensures team members adhere to procedures when performing provisioning to support infrastructure, applications, and services.
- Encourages team members to experiment with new approaches for and perform decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
Automation:
- Implements standards for identifying and recommending opportunities for automation and assesses potential benefits to enhance operational efficiency.
- Takes a proactive role in reviewing and offering feedback on design, automation tools, or scripts, acting as a leader during implementation.
- Shares strategies for conducting testing on automations to ensure they perform tasks correctly and produce expected results.
Technical Communication and Guidance:
- Reviews and provides feedback on release notes and ensures team members communicate comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams.
- Proactively anticipates and articulates the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations.
- Serves as a resource to team members on...