System Reliability Engineer

Securiti
San José Province, Costa RicaPosted Jul 12, 2026
Skip to main content Search Find Jobs For Location Search Close Search Sign up for job alerts Go to Saved Jobs System Reliability Engineer System Reliability Engineer San José, Provincia de San José Saved Role Save role Apply now Function Corporate Function Team Enterprise Technology Role Type Permanent Work Location Remote Date Posted 07/17/2026 Explore this location Take a look at the map to discover what’s nearby. Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands. About the Role We are looking for a Site Reliability Engineer (SRE) to join our team and ensure the continuous, reliable operation of company services. This role involves proactive monitoring, incident response, and building resilient observability and escalation practices across our infrastructure. What You'll Do Ensure monitoring and uninterrupted operation of company services Write and maintain alerting rules and runbooks Perform triage of incoming incidents and initial diagnosis of issues Build and maintain escalation chains for incident response Perform technical incident resolution activities according to runbooks Participate in on-call rotations and post-incident reviews (RCA/postmortems) Continuously improve observability coverage and reduce alert noise/false positives Collaborate with development and infrastructure teams to identify reliability risks and implement preventive measures What You'll Bring Experience with observability tools (Grafana, ELK, VictoriaMetrics) Experience working with Linux Experience working with Kubernetes (k8s) Experience with AWS and Azure cloud platforms Ability to analyze incidents, identify root causes, and propose remediation steps Bonus Skills Experience with Infrastructure as Code (Terraform, Ansible, or similar) Scripting skills (Python, Bash) for automation of operational tasks Understanding of DevOps and CI/CD principles Experience with incident management tools (PagerDuty, Opsgenie, etc.) Effective communication skills and a collaborative approach to teamwork What You'll Get Comprehensive Health Coverage – Fully employer-paid medical, dental, and vision insurance for employees and eligible dependents, including virtual care services Wellbeing & Mental Health Support – Access to an Employee Assistance Program (EAP) that includes confidential therapy sessions, plus legal and financial counseling services Financial Protection Benefits – Company-provided life and disability insurance to help support employees and their families Paid Time Off & Global Recharge Days – Vacation time, statutory holidays, and additional company-wide VeeaMe Days dedicated to rest, wellbeing, and self-care Family-Friendly Leave Programs – Competitive maternity, paternity, adoption, and other leave benefits that support employees through important life moments Give Back to Your Community – Employees receive paid volunteer time each year through the Veeam Cares program Please note: The position is based in San Jose, Costa Rica. If the applicant is permanently located outside of Costa Rica, Veeam reserves the right to decline the application. All applications must be submitted in English. #LI-FT#LI-REMOTE Veeam Software is an equal...

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free