Senior Director - AI Evaluation Platform

Santa Clara, CAFull-time$279k–$488kPosted Jul 8, 2026
Google Chrome Microsoft Edge Apple Safari Mozilla Firefox Senior Director - AI Evaluation PlatformFull-timeEmployee Type: RegularRegion: AMS - North America and CanadaWork Persona: Flexible or RemoteCompany DescriptionIt all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.Join us to put AI to work for people. Job DescriptionWhat you get to do in this role:The CoreAI team, part of the Platform and AI organization, is focused on building the evaluation infrastructure and frameworks that ensure the quality, reliability, and safety of AI capabilities across ServiceNow's platform. Our evaluation systems power the quality assurance behind ServiceNow's GenAI offerings, enabling our customers to trust the AI capabilities they depend on.As a Senior Director, AI Evaluation Platform, you will lead the teams responsible for building and scaling ServiceNow's evaluation platform — both internal-facing (used by research and engineering teams) and customer-facing (enabling customers to assess and monitor AI quality in their environments). Your leadership will be critical in establishing rigorous, scalable evaluation methodologies that raise the bar for AI quality across the industry.This is a fast-growing, high-impact role where you will:Lead and scale a team of talented researchers, applied researchers, and evaluation engineers building next-generation AI evaluation infrastructure.Define the strategy and roadmap for the evaluation platform, encompassing benchmark design, automated scoring pipelines, verifier systems, and quality monitoring frameworks for both internal and customer-facing use cases.Drive innovation in evaluation methodologies, including developing new approaches for agentic evaluation, multi-turn assessment, domain-specific benchmarking, and confidence calibration.Design and implement a comprehensive verifier stack spanning LLM-as-judge, rule-based, reference-based, and process verifiers to enable robust, multi-signal evaluation across diverse AI capabilities.Architect and deliver a unified evaluation platform that serves internal model development teams and external customers, with robust APIs, dashboards, and reporting capabilities.Collaborate with cross-functional teams across product, engineering, and research to translate business needs into evaluation requirements and deliver impactful quality assurance solutions.Stay ahead of the curve by researching and implementing the latest advancements in AI evaluation, including novel verifier architectures, safety and alignment evaluation, and emerging evaluation standards.QualificationsTo be successful in this role you have:Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.Extensive experience in leading and managing high-performing AI or platform engineering teams, with a focus on evaluation, quality assurance, or testing infrastructure at scale.Deep technical expertise in AI evaluation methodologies, including benchmarking frameworks, automated scoring systems, verifier design (LLM-as-judge, rule-based, reference-based, process verifiers), and statistical methods for assessing model quality, safety, and reliability.Proven track record of building and operating...

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free