QE for AI Engineer

Pune, IndiaPosted Jul 22, 2026

We are looking for someone who has validated AI systems in production - where outputs are non-deterministic, failure modes are ambiguous, and correctness is probabilistic. 

You will be responsible for evaluating, hardening, and governing GenAI and agentic systems before and after they go live. This role sits at the intersection of engineering, QA, and AI safety, ensuring that what we build works consistently, reliably, and safely - at scale.

You will partner closely with AI engineers, architects, and clients to define what “correct” means in an AI world — and prove it.

AI Quality Engineering & Evaluation

  • Define evaluation frameworks for LLMs, RAG pipelines, and multi-agent systems

  • Design test strategies for non-deterministic systems (semantic correctness, hallucination detection, consistency checks)

  • Build testing pipelines for LLM evaluation (offline + online evals)

  • Establish ground truth datasets, benchmarks, and scoring metrics (precision, recall, relevance, factuality)

 

Agent & Workflow Validation

  • Validate multi-agent orchestration flows.

  • Test failure modes: hallucination, incomplete reasoning.

  • Simulate edge cases and adversarial inputs.

  • Ensure robustness across multi-step workflows and chained reasoning tasks

     

RAG & Data Validation

  • Validate end-to-end RAG pipelines: 

    • Chunking quality

    • Embedding correctness

    • Retrieval relevance

    • Re-ranking effectiveness

  • Detect and quantify RAG failure points (retrieval gaps, stale data, hallucinations)

  • Ensure data lineage and traceability in AI responses

 

Guardrails, Safety & Governance

  • Ensure compliance with enterprise AI governance and auditability requirements

  • Validate explainability and traceability of AI outputs

 

Automation & Tooling

  • Automate prompt testing, regression testing, and response comparison
  • Integrate AI validation into CI/CD pipelines
  • Built or contributed to evaluation frameworks for LLM-based systems
  • Tested RAG pipelines end-to-end and identified failure points
  • Defined and executed non-deterministic test strategies
  • Automated LLM evaluation or prompt regression pipelines
  • Worked on agent-based or multi-step AI workflows
  • Debugged incorrect or hallucinated model outputs using structured methods
  • Established quality metrics where ground truth was unclear or evolving

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free