Member of Technical Staff - ML Systems & Inference

San Francisco, CAFullTime$150k–$350kPosted Mar 10, 2026

About Us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference.

As AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.

Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.

We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.

About the role

Gimlet is seeking a Member of Technical Staff focused on ML Systems and Inference.

In this role, you will design and build the inference systems that execute full models end-to-end under real production constraints: across scheduling, memory management, and runtime performance. You will work at the intersection of model architecture, runtime behavior, and system performance to ensure inference is fast, predictable, and scalable.

This is not a traditional machine learning role. You won't be training models or tuning benchmarks in isolation. You'll be building the systems that determine how AI workloads are served, optimized, and executed in production.

This role is ideal for engineers who deeply understand how modern models execute in practice and who care about latency, throughput, and memory behavior across the full inference lifecycle.

What success looks like

In the first 12-18 months, you will help:

  • Build and optimize inference systems that improve latency, throughput, and efficiency for production AI workloads

  • Design execution strategies that intelligently balance batching, scheduling, concurrency, and resource utilization

  • Improve KV cache management, memory efficiency, and execution behavior across large-scale serving environments

  • Enable new model architectures and inference techniques to run efficiently in production

  • Partner with compiler, kernel, networking, and distributed systems engineers to drive end-to-end performance improvements

  • Influence the architecture of a platform that will help define how AI workloads are deployed over the next decade

You may be a good fit if

  • Strong software engineering fundamentals

  • Experience building or operating ML inference or model serving systems

  • Comfort reasoning about performance, memory usage, and system behavior under load

Strong candidates may also have

  • Experience with inference runtimes such as TensorRT-LLM, vLLM, or custom serving systems

  • Deep understanding of modern model architectures and attention mechanisms

  • Experience with batching, scheduling, and concurrency control in inference systems

  • Familiarity with KV cache management and memory placement strategies

  • Experience profiling and tuning latency- and throughput-critical systems

  • Software development experience in Python and C++

Why join now?

Gimlet is at the very beginning of its journey, and that's what makes this moment special. Most AI infrastructure companies are focused on deploying more compute. We are focused on making increasingly diverse compute work together, and that ambition touches every part of how we build and run this company.

As an early member of the team, you will have significant ownership over your work, partner directly with a small group of highly capable people, and help shape not just what we build, but how we scale the company.

We value people who are excited to work across domains, take ownership of meaningful problems, and help define what Gimlet becomes over the next several years.

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free