Welcome back!Don't have an account? Sign upContinue with GoogleContinue with LinkedInorEmailPasswordForgot password?LOG INSend a MessageTypeTitleDescriptionCancelSendFOR EMPLOYERSLOG INSIGN UPView all jobsNewSenior CloudOps EngineerLocationWashington, United StatesWork modeFully remoteExperience5+ yearsPosted2 days agoSkills & LanguagesMust-haverequiredMongodb5 Year(s)Redis5 Year(s)Ci/cd Automation5 Year(s)Google Cloud Platform - Gcp5 Year(s)Languages requiredEnglish
We are looking for a Senior CloudOps Engineer to design, build, and operate scalable, secure, and highly available cloud infrastructure on Google Cloud Platform (GCP) and Kubernetes. In this role, you will lead infrastructure automation, platform reliability, observability, and CI/CD initiatives while partnering closely with software engineering teams to deliver resilient production systems.
As a senior member of the Platform Engineering team, you will drive infrastructure best practices, improve developer productivity, and leverage AI-powered DevOps tooling to enhance operational efficiency. This role also participates in a rotating on-call schedule to support our production platform.
Key Responsibilities
Cloud Infrastructure
Design, deploy, and manage scalable, secure, and highly available infrastructure on Google Cloud Platform (GCP).
Architect and optimize GCP services including GKE, Compute Engine, VPC, IAM, Cloud Load Balancing, Cloud DNS, and Cloud Storage.
Continuously improve infrastructure reliability, scalability, performance, and cost efficiency.
Apply infrastructure security best practices, including least-privilege access, encryption, network segmentation, and compliance controls.
Partner with engineering teams to design cloud-native solutions and improve platform architecture.
Kubernetes & Platform Engineering
Design, deploy, and operate Kubernetes clusters running stateless and stateful workloads.
Manage Kubernetes networking, including Ingress Controllers, Service Mesh technologies, Virtual Services, and traffic management.
Build reusable deployment patterns using Helm, Terraform, and infrastructure automation tools.
Troubleshoot and optimize Kubernetes performance, networking, and cluster health.
Infrastructure Automation & CI/CD
Develop and maintain Infrastructure as Code (IaC) using Terraform and Ansible.
Build, maintain, and optimize CI/CD pipelines for automated application and infrastructure deployments.
Standardize deployment workflows and improve developer experience through automation.
Continuously improve release reliability, deployment speed, and operational efficiency.
Observability & Reliability
Design and maintain monitoring, logging, tracing, and alerting solutions using Prometheus, Grafana, and OpenTelemetry.
Define and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and platform health metrics.
Implement proactive alerting, capacity planning, and performance optimization.
Improve platform reliability through root cause analysis, incident reviews, and operational excellence.
Database & Platform Operations
Deploy, administer, and optimize MongoDB, Redis, and MySQL environments.
Ensure backup, disaster recovery, replication, and high-availability strategies are implemented.
Monitor database performance and optimize resource utilization.
Security & Identity
Implement secure infrastructure and cloud security best practices.
Manage Identity and Access Management (IAM), Role-Based Access Control (RBAC), and secrets management.
Implement Single Sign-On (SSO) and identity federation using OAuth 2.0, SAML, and OpenID Connect (OIDC).
Participate in security reviews, audits, and compliance initiatives.
AI-Powered DevOps
Integrate AI-assisted tools into infrastructure management, CI/CD pipelines, observability, and operational workflows.
Evaluate and implement AI agents, Model Context Protocol (MCP) servers, and emerging AI technologies to improve platform operations.
Apply...
Want jobs like this matched to you?
Swoopd scores fresh postings against your résumé so you only see the matches that matter.