Staff Engineer - Machine Learning
San Mateo, CAFull-time$211k–$261kPosted Jul 8, 2026
Google Chrome
Microsoft Edge
Apple Safari
Mozilla Firefox
Staff Engineer - Machine LearningFull-timeCompany DescriptionOrganizations everywhere struggle under the crushing costs and complexities of “solutions” that promise to simplify their lives. To create a better experience for their customers and employees. To help them grow. Software is a choice that can make or break a business. Create better or worse experiences. Propel or throttle growth. Business software has become a blocker instead of ways to get work done.There’s another option. Freshworks. With a fresh vision for how the world works.Freshworks Inc. builds uncomplicated service software that delivers exceptional employee and customer experiences. Our people-first approach to AI eliminates friction, helping businesses reduce complexity, lower cost-to-serve, and deliver faster, more human support through enterprise-grade yet easy-to-use CX and IT solutions. Nearly 75,000 companies, including Bridgestone, New Balance, Nucor, S&P Global, and Sony Music, trust Freshworks to power their Employee Experience (EX) and Customer Experience (CX) operations.Fresh vision. Real impact. Come build it with us.Job DescriptionWe are seeking a Machine Learning Staff Engineer to lead the development of core backend services powering our Agentic AI Platform. In this role, you will be the primary technical driver for building reasoning-driven agents, multi-agent orchestration, and outcome-based workflows.You will collaborate with the Principal AI Architect to bridge the gap between high-level design and scalable implementation. You will own the development of high-performance APIs and orchestration runtimes, integrating frameworks like LangChain, LangGraph, and LangSmith. This role is highly hands-on and requires a technical leader who can navigate a fast-paced environment while maintaining high engineering standards.Key ResponsibilitiesPlatform Backend DevelopmentLead Implementation: Act as the lead engineer for agent runtime orchestration services, focusing on reasoning, planning, and tool invocationAPI & SDK Design: Own the development of robust APIs and SDKs that enable internal and external teams to build and deploy agent workflowsState & Memory Management: Build and optimize stateful dialog management and memory services for multi-turn, context-aware agentsAgent Communication: Implement multi-agent communication (A2A) protocols and shared context systemsSystem Architecture & ExecutionMicroservices Leadership: Execute the transition to cloud-native, event-driven backend services using microservices or service mesh architecturesWorkflow Design: Build and manage task scheduling and long-running operations using tools like Temporal or AirflowPerformance Optimization: Optimize LLM orchestration for performance and cost through caching, batching, and token monitoringIntegrations & Data ServicesData Pipeline Ownership: Implement and maintain RAG pipelines and integrations with vector databases (e.g., Pinecone, Weaviate, FAISS)Telemetry Integration: Ensure all backend services support LangSmith-based evaluation and comprehensive observabilitySecurity & ReliabilityTenant Isolation: Implement RBAC, API authentication, and multi-tenant isolation logicEngineering Excellence: Set the standard for automated testing (unit, integration, load) and implement OpenTelemetry for full system visibilityMentorship: Provide technical guidance and code reviews for more junior backend engineers on the teamQualificationsRequired Qualifications:6+ years of backend engineering experience1-2+ years of hands-on experience in AI/ML platform developmentExpert Proficiency: Strong programming skills in Java and PythonFramework Expertise: Hands-on experience with LangChain, LangGraph, or LangSmithDistributed Systems: Deep understanding of event-driven architectures, microservices, and KubernetesCloud & Data: Solid experience with AWS, Terraform, and a mix of relational (PostgreSQL) and Vector...