Job Description :
Key Responsibilities :
- Own the platform infrastructure strategy for the AI capability programme across all six pods, covering CI/CD, environment provisioning, secrets management, observability, and reliability engineering.
- Design and maintain multi-environment infrastructure (development, staging, production) with appropriate isolation, promotion gates, and security controls for classified, internal, and public workloads.
- Define infrastructure-as-code standards (Terraform, Ansible, or equivalent) and drive their adoption across pods.
- Establish the observability stack metrics, logs, tracing, alerting covering both AI-model and application layers, in coordination with the MLOps Lead.
- Own incident response protocols and on-call rotation across the programme; conduct post-incident reviews and drive corrective actions.
- Define availability and reliability SLOs for each pod's production services; track error budgets and coordinate with Product Managers on trade-offs.
- Oversee infrastructure security hardening network policies, WAF operations, DDoS protection, VAPT remediation at infrastructure layer in coordination with the Security & Compliance Engineer.
- Manage relationships with NIC, IndiaAI Compute, MeghRaj, and other government infrastructure providers; coordinate capacity planning and reserved-versus-on-demand mixing.
- Ensure Data Residency, Data Storage, and Data Lifecycle Management compliance per the LoE (data within India, AES-256 at rest, TLS 1.3 in transit, sanitisation at project closure).
- Mentor DevOps/SRE Engineers and coordinate with MLOps Lead on shared platform responsibilities.
Technical Competencies :- Container & Orchestration : Docker, Kubernetes (production-grade), Helm, service mesh (Istio, Linkerd), operators; certified Kubernetes administrator credentials preferred.
- Infrastructure as Code : Terraform, Ansible, Pulumi, or equivalent; GitOps workflows (Argo CD, Flux).
- Cloud Platforms : AWS, Azure, GCP; government infrastructure (NIC, MeghRaj); familiarity with IndiaAI Compute service model.
- CI/CD : Jenkins, GitLab CI, GitHub Actions, Argo Workflows; artefact repositories (Nexus, Harbor).
- Observability : Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Jaeger; SLO/SLI methodology; incident management platforms.
- Infrastructure Security : Network security, secrets management (HashiCorp Vault, Infisical), certificate management, zero-trust patterns.
- Reliability Engineering : SRE principles, error budgets, capacity planning, chaos engineering basics, disaster recovery and business continuity.
- Scripting & Programming : Bash, Python, Go for tooling and automation.
- Compliance & Governance : MeitY Security Policy, ISO 27001, incident reporting timelines (24-hour breach reporting per LoE), audit trail management.
Experience :- 8+ years in infrastructure engineering, DevOps, or Site Reliability Engineering, with minimum 3 years in a lead or architect role.
- Demonstrated experience running production infrastructure for high-availability platforms at scale.
- Prior experience with government-hosted infrastructure (NIC, MeghRaj) or MeitY-empanelled CSPs preferred.
- Experience with on-premise, hybrid, and sovereign-cloud deployment models.
Educational Qualification :- B.Tech./B.E. in Computer Science, Information Technology, or related engineering discipline (Must have).
- M.Tech./M.S. in Computer Science, Information Technology, or related field desirable.
- Certifications in DevOps, SRE, or Cloud Infrastructure (AWS, Azure, GCP, or Kubernetes CKA/CKS) preferred.
- ITIL or equivalent service management certification desirable.
(ref:hirist.tech)