AI Quality and Evaluation, Lead

Pune, IndiaPosted Jul 13, 2026
Google Chrome Microsoft Edge Apple Safari Mozilla Firefox AI Quality and Evaluation, LeadFull-timeTime Type: Full TimeDepartment: Product ManagementLocation: India - Pune - Adjunct 0fficeCompany DescriptionQAD is building a world-class SaaS company, and we are growing. We are looking for talented individuals who want to join us on our mission to help solve relevant real-world problems in manufacturing and the supply chain. This hybrid position requires candidates to be based in Pune with 3-4 days of in-office collaboration per week.Job DescriptionAbout the roleQAD's products increasingly include AI agents that make and execute recommendations inside customers' operations. These systems behave differently from traditional software — outputs vary, quality is judgment-based rather than binary, and the cost of getting it wrong matters. We're building the function that makes AI quality measurable, defensible, and continuously improved across our portfolio, and we're hiring the person to own it.You'll design and run the evaluation framework our AI products are measured against, define the criteria that gate every release, and monitor production performance so quality issues are caught early rather than after customers feel them. The role partners closely with product, engineering, and customer-facing teams, and reports to the Head of Product Operations.If you've worked on the quality and measurement side of LLM-based products and want a role where evaluation is genuinely load-bearing rather than an afterthought, this is that role.What you'll ownBuild and own a shared evaluation framework across the product organization: golden datasets, LLM-judge rubrics, and code-based checks. Measure not just response quality but the quality of the decisions the AI products produce — did the recommendation actually serve the customer outcome the product was built for?Own the technical criteria for stage-gate release reviews: define what evaluation evidence a product must produce to clear each gate, and each capability-tier progression. You don't chair the gates; the Head does. But a gate cannot pass without your evidence.Own decision auditability: every AI-driven recommendation must be logged with the context considered, the rationale, and the outcome — in a way that's faithful, end-to-end, and useful for both customer trust and continuous improvement of the product.Own production drift detection: monitor evaluation-score regression, human-override-rate increase, and exception-rate spikes. Treat these as leading indicators of customer issues and surface findings before they show up as support escalations. When a product regresses, you trigger a capability-tier review.Define blast-radius and rollback requirements with engineering, and gate releases on thresholds in CI.Partner with each product team's AI lead to translate "what good looks like" into rubrics — breaking quality into independent dimensions (correctness, constraint compliance, decision quality, latency, tone) rather than one blended score.Run a regular evaluation-review cadence across teams, surface regressions early, and build the organization's shared vocabulary for what product quality means in this domain.QualificationsWhat we're looking for5-8 years in product ops, AI/ML product, data, or quality engineering, with hands-on AI evaluation experience on production LLM systems.Strong fluency with agent architecture: traces and spans, tool calls, RAG/groundedness, LLM-as-judge, offline vs. online evaluations, drift monitoring.Practical technical skills — SQL plus Python or JavaScript; comfort with evaluation/observability platforms and Git.A discovery mindset toward evaluation: failure data and edge cases are the core of the job, not a postmortem activity.Bias toward decision quality over surface quality. We are not optimizing for hallucination rate alone — we're measuring whether the AI's recommendation actually served the customer's need.Additional InformationYour...

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free