Site Reliability Engineer (SRE) On-Prem

TLV, IsraelPosted Jul 22, 2026

Every nation has data. Few can protect it. Fewer still can act on it.

Dream is the sovereign AI and national cyber-defense company for governments.

We help nations secure their most critical systems, connect fragmented information at a national scale, and turn their most sensitive data into decisions, all fully sovereign.

This is more than a job. It's a Dream job, where you'll work at a global scale alongside some of the best AI researchers, cyber operators, and government experts in the world.

The mission only works if the company behind it does. This role keeps Dream running at the scale our work demands. And our work demands a uniquely global scale.

The Dream Job

We are on an expedition to find an On-Premise Site Reliability Engineer (SRE)- someone who is passionate about building rock-solid, high-performance infrastructure and bringing order to complex environments. In this role, you will own the end-to-end reliability, automation, and deployment of DREAM’s platform across customer sites, working hands-on with cutting-edge AI, bare-metal, and hybrid cloud architectures.

You’ll collaborate closely with Product, R&D, and Architecture teams while serving as the ultimate technical authority for our customer deployments. From designing automated Ansible workflows and mastering Kubernetes to troubleshooting complex network topologies, you will eliminate toil, streamline cluster operations, and ensure every deployment is scalable, seamless, and mission-ready.

The Dream-Maker Responsibilities

  • Lead End-to-End On-Prem & Hybrid Deployments: Own the technical delivery and reliability of DREAM’s platform in close collaboration with Product, R&D, and customer technical teams. 
  • Architect, Execute & Improve K8s Deployments: Take a definitive hands-on role in deploying, configuring, operating, and continuously improving our platform using advanced, enterprise-grade Kubernetes architectures. 
  • Helm Chart Management: Design, modify, and manage Helm charts to package, version, and streamline complex application deployments across different environments. 
  • Drive Automation & Simplification: Design, implement, and maintain robust deployment automation using Ansible. You must have a passion for turning complex manual tasks into reliable, repeatable, single-click operations. 
  • Manage Infrastructure as Code: Utilize Git as the single source of truth to manage configurations, manifests, and automation playbooks, enforcing modern engineering best practices. 
  • Bridge On-Prem and Cloud: Leverage AWS resources (specifically EC2 and S3) for hybrid components, staging environments, or cloud-to-on-prem data flows. 
  • Technical Tier-3 Escalation: Serve as the ultimate technical authority for deployment, Linux networking, and Kubernetes orchestration issues. 
  • Continuous Improvement: Constantly refine our delivery pipelines, optimize bootstrap processes, and create rock-solid technical documentation. 


The Dream Skill Set

  • SRE / Delivery Mindset: 3–5 years of hands-on experience in enterprise infrastructure deployment, systems engineering, or an on-prem operational reliability role. 
  • Kubernetes & Helm Expert: Deep, production-grade experience with Kubernetes architecture, deployment, advanced troubleshooting, and CNI networking. Proven working experience creating, maintaining, and deploying applications using Helm charts. 
  • Ansible Mastery: Proven experience writing clean, scalable Ansible roles and playbooks for configuration management, automation, and infrastructure provisioning. 
  • Modern Workflows (Git & AWS): Solid working experience using Git for version control and collaborating on code/infrastructure. Practical experience provisioning and managing AWS resources (EC2 and S3). 
  • Core Systems & Linux: Strong Linux background (Ubuntu) with a deep understanding of system internals, containerized runtimes, and troubleshooting distributed applications. 
  • Solid Networking Knowledge: Hands-on experience with routing, firewalls, and switching topology (mainly Cisco) 
  • Storage Foundations: Working knowledge of storage protocols (iSCSI, SAN, local NVMe) and enterprise storage arrays (like DELL) interacting with Kubernetes Persistent Volumes. 
  • GPU & Accelerated Compute: Working knowledge of managing GPU-enabled Kubernetes nodes, including NVIDIA drivers/runtime and basic troubleshooting. 
  • Air-Gapped Deployments: Experience deploying and maintaining software in air-gapped or offline environments, including registry mirroring and artifact staging. 
  • Problem-Solver: Strong debugging and problem-solving skills in complex, distributed environments with an intense ownership and accountability mindset. 
  • Willingness to Travel: Ready to travel to customer sites for physical staging and deployments—at least 30%. 
  • Language: High-level English proficiency. 

Nice to have 

  • EU or any other additional citizenship. 
  • Experience with enterprise data components like MongoDB, PostgreSQL, Neo4j, or RabbitMQ. 
  • Working experience with project management tools such as Jira and Monday. 
  • Valid Israeli security clearance. 

 

Never Stop Dreaming...

If you think this role doesn't fully match your skills but are eager to grow and break glass ceilings, we’d love to hear from you! 

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free