Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About This Role
Crusoe is building a vertically integrated AI neocloud from the electron up. Because we own the entire stack—from clean energy generation to the GPUs running customer training jobs—we treat power as a control input, not a constraint. To support our massive scaling vector from 30,000 accelerators to 300,000 without a linear growth in headcount, we are building the Cloud Availability Platform Engineering (CAPE) organization. CAPE serves as the horizontal reliability spine beneath all of Crusoe Cloud. As the Founding Manager of Production Engineering in Tel Aviv, you will stand up our local presence from scratch, driving a cultural shift from reactive firefighting to software-defined engineering.
This is a hybrid leadership and technical role where you will hire, scale, and lead a local team of exceptional systems generalists while remaining deeply hands-on in the code and incident response. You will ensure that your team spends at least 30% of their bandwidth on strategic automation, tooling, and firmware optimization primitives to prevent the operational treadmill. If you want to bridge core software layers with physical infrastructure and write the playbook for an entire engineering site, this founding seat is for you. This is a full-time position located in Tel Aviv, Israel.
What You’ll Be Working On
Team Leadership & Founding Culture: Recruit, mentor, and establish a high-performing Production Engineering footprint in Tel Aviv, setting an uncompromising cultural standard for operational discipline and systems-first engineering.
Incident & On-Call Ownership: Partner with US and Dublin teams to run a follow-the-sun global on-call rotation, while championing a strict blameless post-mortem culture that targets systemic failures over human error.
Software-Defined Operations: Drive alert-reduction initiatives to improve fleet signal-to-noise ratios, automate routine manual workflows using modern runbook automation (e.g., Temporal), and build predictive monitoring to catch SEV1/SEV2 events before customers do.
Collaborative Governance: Act as the ultimate Production Gatekeeper across cross-functional compute, storage, networking, and platform teams, holding a strict line on Production Readiness Reviews and change control.
Strategic Reliability Engineering: Protect team bandwidth to ensure engineers spend at least 30% of their time on strategic automation, tooling, and firmware optimization primitives rather than drowning in incident response.
Physical-to-Digital Automation: Instill a software-first approach to physical problems, ensuring that any physical intervention occurring twice is successfully converted into a software-defined auto-remediation.
What You’ll Bring to the Team
Years of Infrastructure Experience: Minimum of 8+ years of experience working within infrastructure, SRE, or production engineering environments.
Engineering Leadership Track Record: Minimum of 2+ years of experience directly leading first-line engineering teams within a high-growth neocloud, hyperscaler, or large-scale distributed environment.
Non-Negotiable Coding Proficiency: Strong, hands-on software engineering fundamentals in Go, Python, C++, or a comparable systems language to build automation rather than scale through headcount.
Distributed Systems Depth: Expert-level command of Linux internals, container orchestration at scale, and root-cause analysis across complex physical-to-virtual boundaries.
Operational Execution Expertise: Proven track record of running tiered on-call models, establishing clear SLIs/SLOs and error budgets, and measurably reducing paging fatigue.
Bonus Points
AI Infrastructure Experience: Prior experience working at a neocloud or AI-infrastructure company operating massive GPU clusters.
High-Performance Fabric Exposure: Hands-on exposure to high-performance networks (such as InfiniBand or RoCEv2) or hardware internals (including BMC, firmware qualification, and attestation).
Accelerator Domain Knowledge: Deep familiarity with the failure modes of modern AI accelerators (NVIDIA or AMD platforms) and how they manifest from DCGM counters up to a customer's training run.
Scale Engineering: Experience managing step-change scaling milestones and building systems designed to absorb 10x fleet expansions.
Benefits:
Crusoe also offers a competitive benefits package designed to support financial security, health, and overall well-being. Our benefits are tailored to local market standards and include core offerings such as pension contributions and additional perks to support work-life balance.
Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.