Member of Technical Staff - ML Systems & Inference
About Us
Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference.
As AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.
Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.
About the role
Gimlet is seeking a Member of Technical Staff focused on ML Systems and Inference.
In this role, you will design and build the inference systems that execute full models end-to-end under real production constraints: across scheduling, memory management, and runtime performance. You will work at the intersection of model architecture, runtime behavior, and system performance to ensure inference is fast, predictable, and scalable.
This is not a traditional machine learning role. You won't be training models or tuning benchmarks in isolation. You'll be building the systems that determine how AI workloads are served, optimized, and executed in production.
This role is ideal for engineers who deeply understand how modern models execute in practice and who care about latency, throughput, and memory behavior across the full inference lifecycle.
What success looks like
In the first 12-18 months, you will help:
Build and optimize inference systems that improve latency, throughput, and efficiency for production AI workloads
Design execution strategies that intelligently balance batching, scheduling, concurrency, and resource utilization
Improve KV cache management, memory efficiency, and execution behavior across large-scale serving environments
Enable new model architectures and inference techniques to run efficiently in production
Partner with compiler, kernel, networking, and distributed systems engineers to drive end-to-end performance improvements
Influence the architecture of a platform that will help define how AI workloads are deployed over the next decade
You may be a good fit if
Strong software engineering fundamentals
Experience building or operating ML inference or model serving systems
Comfort reasoning about performance, memory usage, and system behavior under load
Strong candidates may also have
Experience with inference runtimes such as TensorRT-LLM, vLLM, or custom serving systems
Deep understanding of modern model architectures and attention mechanisms
Experience with batching, scheduling, and concurrency control in inference systems
Familiarity with KV cache management and memory placement strategies
Experience profiling and tuning latency- and throughput-critical systems
Software development experience in Python and C++
Why join now?
Gimlet is at the very beginning of its journey, and that's what makes this moment special. Most AI infrastructure companies are focused on deploying more compute. We are focused on making increasingly diverse compute work together, and that ambition touches every part of how we build and run this company.
As an early member of the team, you will have significant ownership over your work, partner directly with a small group of highly capable people, and help shape not just what we build, but how we scale the company.
We value people who are excited to work across domains, take ownership of meaningful problems, and help define what Gimlet becomes over the next several years.