We are looking for someone who has validated AI systems in production - where outputs are non-deterministic failure modes are ambiguous and correctness is probabilistic.
You will be responsible for evaluating hardening and governing GenAI and agentic systems before and after they go live. This role sits at the intersection of engineering QA and AI safety ensuring that what we build works consistently reliably and safely - at scale.
You will partner closely with AI engineers architects and clients to define what correct means in an AI world and prove it.
AI Quality Engineering & Evaluation
Define evaluation frameworks for LLMs RAG pipelines and multi-agent systems
Design test strategies for non-deterministic systems (semantic correctness hallucination detection consistency checks)
Build testing pipelines for LLM evaluation (offline online evals)
Establish ground truth datasets benchmarks and scoring metrics (precision recall relevance factuality)
Agent & Workflow Validation
Validate multi-agent orchestration flows.
Test failure modes: hallucination incomplete reasoning.
Simulate edge cases and adversarial inputs.
Ensure robustness across multi-step workflows and chained reasoning tasks
RAG & Data Validation
Validate end-to-end RAG pipelines:
Chunking quality
Embedding correctness
Retrieval relevance
Re-ranking effectiveness
Detect and quantify RAG failure points (retrieval gaps stale data hallucinations)
Ensure data lineage and traceability in AI responses
Guardrails Safety & Governance
Ensure compliance with enterprise AI governance and auditability requirements
Validate explainability and traceability of AI outputs
Automation & Tooling
Required Experience:
IC
QE for AI Engineer • Pune, Maharashtra, India