Senior Software Engineer - AI Agent Platform
Navan’s Cognition team builds and operates AI-powered travel and expense experiences used by customers around the world. We are developing intelligent agents that help users search, make decisions, and complete complex workflows through natural, personalized interactions.
Our team owns the AI platform end to end - from agent execution and LLM orchestration to state management, integrations, evaluation, observability, and production reliability.
We are looking for a Senior Software Engineer to own and evolve the runtime platform that powers Navan’s production AI agents.
This is a hands-on backend and platform engineering role for someone who combines deep TypeScript expertise, strong distributed-systems fundamentals, and exceptional production debugging skills with a practical understanding of LLMs and agentic systems.
You will work closely with engineers building AI agents. Your responsibility will be to provide the reliable runtime, infrastructure, abstractions, and observability they need to deliver new capabilities safely and quickly.
What you’ll do
- Own and evolve the production runtime responsible for executing and orchestrating AI agents.
- Design platform capabilities for agent execution, tool calling, streaming, state management, persistence, and long-running workflows.
- Build resilient integrations with multiple LLM providers and model-serving platforms.
- Design provider-routing and fallback strategies based on availability, latency, quality, and cost.
- Implement retries, timeouts, circuit breakers, rate-limit handling, idempotency, and graceful degradation.
- Ensure the platform remains available when external dependencies or infrastructure components experience outages.
- Build reliable mechanisms for loading, caching, versioning, and recovering agent configurations and artifacts.
- Create end-to-end observability for AI requests, including model, provider, agent, latency, token usage, cost, errors, retries, and fallback behavior.
- Define dashboards, alerts, SLOs, and runbooks for production AI workloads.
- Lead the investigation of complex production issues across application code, infrastructure, external providers, distributed state, and agent behavior.
- Improve platform scalability, concurrency, latency, and resource efficiency.
- Build reusable APIs and abstractions that allow agent developers to add capabilities without duplicating infrastructure logic.
- Strengthen platform quality through integration testing, load testing, failure injection, and dependency-outage simulations.
- Turn production incidents into architectural improvements, automated tests, monitoring, and operational safeguards.
- Collaborate with product, infrastructure, and engineering teams to translate customer and business requirements into platform capabilities.
- Mentor engineers and establish best practices for building and operating reliable production AI systems.
What we’re looking for
- 7+ years of professional software engineering experience, primarily in backend, platform, or distributed systems.
- Expert-level TypeScript and Node.js skills.
- Experience with NestJS or a comparable backend framework.
- Proven experience designing, building, and operating large production services.
- Strong understanding of distributed-systems patterns, including retries, backoff, idempotency, circuit breakers, caching, consistency, and failure recovery.
- A systematic debugging mindset and the ability to trace failures across multiple services and dependencies.
- Experience owning customer-facing systems where availability, latency, and correctness directly affect users.
- Strong experience with cloud infrastructure and managed services, preferably AWS.
- Experience with distributed caching and storage technologies such as Redis and S3.
- Hands-on experience with production observability: structured logs, metrics, tracing, dashboards, alerts, and SLOs.
- Experience participating in incident response and driving follow-up improvements.
- Strong API design, testing, and software architecture fundamentals.
- Excellent communication and collaboration skills across engineering, product, infrastructure, and AI teams.
- Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
AI platform experience
You should have hands-on experience with or a strong understanding of:
- LLM APIs, streaming, tool calling, and structured outputs.
- Agentic patterns such as reasoning loops, tool orchestration, and multi-step workflows.
- Tokens, context windows, rate limits, latency, and model-specific behavior.
- Multi-provider LLM integrations and the trade-offs between providers and models.
- How retries and fallbacks can affect response quality, latency, correctness, and cost.
- The observability required to understand a request across multiple LLM and tool calls.
- Techniques for controlling and optimizing LLM usage and cost.
Nice to have
- Experience building an AI gateway, agent runtime, inference platform, or workflow engine.
- Experience with OpenAI, Gemini/Vertex AI, AWS Bedrock, or similar platforms.
- Familiarity with model and provider routing based on quality, availability, latency, and cost.
- Experience with Kafka or other event-driven architectures.
- Familiarity with MCP or agent-to-agent communication protocols.
- Experience with Kubernetes and cloud-native infrastructure.
- Experience with Grafana, or similar tooling.
- Experience with chaos engineering, fault injection, or large-scale load testing.
- Experience building internal developer platforms or frameworks used by multiple engineering teams.
- Familiarity with conversational AI, RAG, memory systems, or generative user experiences.
What we offer
- The opportunity to shape a production AI platform used by customers worldwide.
- Complex engineering challenges at the intersection of distributed systems and generative AI.
- Direct influence over platform architecture, reliability, and technical direction.
- A collaborative environment with experienced engineers working across AI, product, backend, and infrastructure.
- Opportunities for professional growth, mentorship, and technical leadership.