Talent.com
HCL TechBee
AI Platform Reliability Engineer (PRE) /ArchitectHCL TechBee • Noida, India
AI Platform Reliability Engineer (PRE) /Architect

AI Platform Reliability Engineer (PRE) /Architect

HCL TechBee • Noida, India
26 days ago
Job description

Location - Noida, Bengaluru, Chennai, Hyderabad and Pune.

Role

An AI PRE Engineer (L4) acts as a Platform Reliability Architect and Production Readiness Leader responsible for defining, governing, and scaling enterprise-grade AI/ML platforms. This role ensures AI systems are production-ready, resilient, observable, secure, compliant, and cost-efficient at scale, while driving standardization, automation, and platform-wide governance across AI engineering, SRE, DevOps, and MLOps disciplines. The L4 PRE acts as a strategic bridge between platform engineering, AI teams, and business leadership, ensuring reliability and scalability of AI solutions.

Responsibilities

Architecture & Production Readiness Governance

  • Define and govern enterprise-wide production readiness frameworks across: Platform, data, model, application, and security layers
  • Establish organization-wide PRE standards, policies, and maturity models
  • Define and drive release certification governance (PRR frameworks) across all AI workloads
  • Own production readiness lifecycle from design → deployment → scale

SLO/SLI Strategy & Reliability Engineering

  • Define and govern: SLO/SLI/SLA frameworks for latency, availability, quality, safety, drift
  • Error budget strategy at platform and business levels
  • Establish reliability engineering models for AI systems
  • Align SLOs with: Customer experience, business KPIs, and cost targets

Reference Architecture & Platform Design

  • Define and publish: Enterprise AI reference architectures for: LLM applications, RAG pipelines, vector stores, agent frameworks
  • Batch, real-time, and streaming inference systems
  • Standardize: Platform architecture patterns across multi-cloud and hybrid environments
  • Drive platform abstraction and reusable AI services
  • Design and architect the NVIDIA Enterprise AI Software stack and deployment of GPU enabled Kubernetes clusters.

Deployment & Release Engineering Governance

  • Define enterprise standards for: Safe deployment strategies (canary, shadow, blue-green, A/B testing)
  • Prompt/model versioning, rollback, and release controls
  • Govern: Release pipelines, CI/CD, GitOps, and policy-as-code frameworks
  • Establish release risk assessment frameworks for AI workloads

Observability & AI Ops Strategy

  • Define enterprise observability strategy covering: Metrics, logs, traces, and AI-specific telemetry: Token usage, latency, model drift, GPU utilization, quality metrics
  • Standardize: Observability patterns for AI systems across platform
  • Drive adoption of: AIOps and predictive reliability models
  • Establish: SLO burn alerts, cost anomaly detection frameworks

Capacity Engineering & FinOps for AI

  • Own: Capacity planning and forecasting models for: Token usage
  • GPU/CPU scaling
  • Vector databases and caching layers
  • Define: Cost optimization strategies (FinOps for AI platforms)
  • Govern: Resource allocation, autoscaling policies, and utilization efficiency

Resilience Engineering & System Design

  • Define: Enterprise resilience patterns, including: Circuit breakers, fallbacks, retries, timeouts
  • Semantic caching and prompt optimization
  • Architect: High-availability AI systems with: Multi-region / multi-zone failover
  • Establish: DR/BCP strategies with defined RTO/RPO targets

AI Security, Risk & Compliance Governance

  • Define and govern: AI security frameworks: Secrets management
  • Private networking
  • Egress controls
  • Model and prompt security
  • Establish: Red-team validation and AI safety testing standards
  • Own: Compliance frameworks mapping (ISO 27001, SOC 2, GDPR/DPDP, HIPAA)
  • Ensure: Data privacy, auditability, and governance across AI platforms

Platform Engineering & Developer Enablement

  • Define: Enterprise delivery frameworks: CI/CD pipelines
  • SDKs, APIs, Helm charts, Terraform templates
  • Drive: Platform standardization and self-service enablement
  • Build: Reusable automation frameworks and PRE accelerators
  • Lead: Enablement programs (trainings, brown-bags, playbooks)

Cross-Functional Leadership & Stakeholder Engagement

  • Partner with: AI/ML teams, platform teams, DevOps, and business stakeholders
  • Act as: Technical advisor to leadership and customers
  • Validate: Architecture, NFRs (scale, latency, cost, safety) across programs
  • Drive: Enterprise-wide adoption of PRE practices

Incident Management & Reliability Operations

  • Own: Enterprise incident response frameworks for AI systems
  • Lead: Critical incident command and escalations
  • Govern: RCA, postmortems, and systemic reliability improvements
  • Establish: Continuous improvement and

Qualifications & Experience

  • Bachelor's degree in Computer Science / Engineering (mandatory)
  • Master's preferred (Distributed Systems / Cloud / AI Systems)
  • 12–18 years overall IT / platform engineering experience
  • 7–10 years in AI / platform / cloud engineering at scale
  • 5+ years in architecture / PRE / SRE strategy roles

Core Experience:

  • Enterprise-scale:
  • Production readiness frameworks, PRR governance
  • SLO/SLI/SLA design and reliability engineering
  • Deep expertise in:
  • Kubernetes platforms (EKS/GKE/AKS) at scale
  • GPU-aware AI infrastructure
  • Multi-cluster, multi-tenant environments
  • Strong experience in:
  • MLOps / LLMOps / AI production systems
  • Deployment strategies and release governance

Technology Experience:

  • Cloud: AWS / Azure / GCP (architect-level understanding)
  • Infra: Kubernetes, Docker, Terraform, GitOps
  • Observability: Prometheus, Grafana, OpenTelemetry
  • AI stack:
  • LLM deployment, RAG pipelines, vector DBs
  • Frameworks like LangChain, OpenAI
  • NVIDIA AI Enterprise (NVAIE)
  • Programming:
  • Python / Go / Java (platform engineering focus)

Certifications Required

  • NVIDIA Certified Professional – AI Infrastructure & Operations
  • NVIDIA DLI – Advanced AI Infrastructure / GPU platforms
  • Kubernetes Certifications (CKA / CKS – mandatory)
  • Cloud Architect Certifications (AWS/Azure/GCP – Professional level preferred)
  • DevOps / CI-CD certifications (preferred)
  • Linux Certifications (RHCE / LFCS – advanced preferred)


Skills Required
Docker, Kubernetes, Terraform, Gcp, Grafana, Prometheus, Go, Azure, Java, Aws, Python
Create a job alert for this search

AI Platform Reliability Engineer (PRE) /Architect • Noida, India

Similar jobs

Lead Engineer -AI & Knowledge Systems - Noida

NeoXamNoida, Uttar Pradesh, India

Lead Engineer - AI & Knowledge Systems.NeoXam is a leading B2B fintech software vendor serving buy-side investment managers and asset servicers globally.Our platform covers data management, reconci... Show more

Senior Software Engineer - AI Agent

Crimson EducationNew Delhi, Delhi, India

Crimson Education is the world’s leading college admissions consulting firm, with over 1,490 Ivy League offers and 2,410 to the US Top 15.We help ambitious students gain admission to the world’s to... Show more