Staff Software Engineer, Compute Architecture
Monolith AI
Livingston, NJ$188k–$275kPosted Jul 21, 2026
Skip to main content
Staff Software Engineer, Compute Architecture
CoreWeave
Livingston, NJ
Apply
Join or sign in to find your next job
Join to apply for the Staff Software Engineer, Compute Architecture role at CoreWeave
Email or phone
Password
Show
Forgot password?
Sign in
Sign in with Email
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Staff Software Engineer, Compute Architecture
CoreWeave
Livingston, NJ
3 weeks ago
Be among the first 25 applicants
See who CoreWeave has hired for this role
Apply
Join or sign in to find your next job
Join to apply for the Staff Software Engineer, Compute Architecture role at CoreWeave
Email or phone
Password
Show
Forgot password?
Sign in
Sign in with Email
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Save
Report this job
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.About The RoleAs a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers. The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems. This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.What You’ll DoDesign, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure.Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations.Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure.Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly.Translate incidents and hardware failure modes into software improvements that make the platform more resilient.Partner with...