Staff Software Engineer, Compute Architecture

Monolith AI
Livingston, NJ$188k–$275kPosted Jul 21, 2026
Skip to main content Staff Software Engineer, Compute Architecture CoreWeave Livingston, NJ Apply Join or sign in to find your next job Join to apply for the Staff Software Engineer, Compute Architecture role at CoreWeave Email or phone Password Show Forgot password? Sign in Sign in with Email or New to LinkedIn? Join now By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy. Staff Software Engineer, Compute Architecture CoreWeave Livingston, NJ 3 weeks ago Be among the first 25 applicants See who CoreWeave has hired for this role Apply Join or sign in to find your next job Join to apply for the Staff Software Engineer, Compute Architecture role at CoreWeave Email or phone Password Show Forgot password? Sign in Sign in with Email or New to LinkedIn? Join now By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy. Save Report this job CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.About The RoleAs a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers. The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems. This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.What You’ll DoDesign, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure.Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations.Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure.Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly.Translate incidents and hardware failure modes into software improvements that make the platform more resilient.Partner with...

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free