Member of Technical Staff - Distributed Systems
About Us
Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference.
As AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.
Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.
About the role
At Gimlet, we believe every hire changes the company.
As a an early-stage company, talent density matters more than headcount. The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.
The future of AI infrastructure will not be built on a single hardware platform. It will be built on systems capable of coordinating increasingly heterogeneous compute at unprecedented scale.
This role is an opportunity to help build that future. We are not optimizing for headcount, we are optimizing for talent density.
You will design and operate the distributed systems that schedule, route, and coordinate AI workloads across thousands of nodes and diverse hardware architectures.
What success looks like
In the first 12-18 months, you will help:
Build scheduling and orchestration systems that coordinate workloads across heterogeneous hardware
Improve reliability and fault tolerance for production AI infrastructure operating at scale
Create APIS and control planes that simplify deployment for customers running mission-critical workloads
Influence the architecture of a platform that will power the next generation of AI systems
You may be a good fit if
Strong software engineering fundamentals
Experience building or operating distributed systems in production environments
Comfort reasoning about concurrency, failure modes, and tradeoffs in large-scale systems
Strong candidates may also have
Experience with Kubernetes or Kubernetes-adjacent systems beyond basic usage
Experience designing service-oriented architectures using RPC or asynchronous messaging
Familiarity with scheduling, queues, or resource management systems
Experience building reliable APIs and operating systems under high load
Software development experience in languages commonly used for systems development (e.g., Go, C++, Python)
Why join now?
Gimlet is at the very beginning of its journey, and that's what makes this moment special. Most AI infrastructure companies are focused on deploying more compute. We are focused on making increasingly diverse compute work together, and that ambition touches every part of how we build and run this company.
As an early member of the team, you will have significant ownership over your work, partner directly with a small group of highly capable people, and help shape not just what we build, but how we scale the company.
We value people who are excited to work across domains, take ownership of meaningful problems, and help define what Gimlet becomes over the next several years.