Principal Operations Engineer, Reliability

California · San Francisco, CA · Texas · Washington · Austin, TX · New York, NY · Seattle, WAPosted Jul 17, 2026
Principal Operations Engineer, Reliability LocationAustin, TX; New York, NY; San Francisco, CA; Seattle, WAEmployment TypeFull timeLocation TypeOn-siteDepartmentOperationsCompensation$220K – $260KTo provide greater transparency to candidates, we share base pay ranges for all US-based job postings. Our compensation package includes base salary, equity for all full time roles, benefits, and, for applicable roles, commissions plans.This range reflects a good-faith estimate based on the applicable pay scale, the range previously established for this role, the compensation of individuals currently in equivalent positions, and/or the budgeted amount for the position. Actual compensation may vary based on experience, qualifications, and other factors.We welcome compensation discussions if this range doesn't meet your requirements. Outstanding candidates may be eligible for adjusted terms, plus meaningful equity that ensures you benefit directly from the company's long-term performance.Total compensation may also include equity in the form of restricted stock units. About FluidstackWe exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it. We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.We hire people who care deeply about this problem space. If that is you, please apply!How We OperateExtreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.Velocity. We drive everything forward as fast as possible.First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.The Data Center Operations TeamExamples of key problems the team is working onOperate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.Role ScopeOwn fleet reliability engineering: define availability targets, measure them honestly, and close the gap.Run root cause analysis on the fleet's worst incidents and drive corrective actions to done across every site.Build the failure data pipeline, facility and hardware both, that turns incident history into engineering priorities.Set the maintenance strategy (reliability-centered, condition-based) so the fleet spends effort where the failure data says to.What We're Looking ForThe below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.You've owned reliability for critical infrastructure and moved the availability number, not just reported it.You've led root cause analyses that found the real cause, not the convenient one.You work fluently with failure data: Weibull, Pareto, and FMEA are tools you actually use, not terms you know.You get corrective actions closed across teams you don't...

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free