Talent.com
The Hartford
Lead Reliability EngineerThe Hartford • Hyderabad / Secunderabad, Telangana, India
Lead Reliability Engineer

Lead Reliability Engineer

The Hartford • Hyderabad / Secunderabad, Telangana, India
1 day ago
Job description

About the Role

We are looking for a Lead Reliability Engineer to own and evolve The Hartford's enterprise Software Delivery Framework (SDF) platform. You will be the technical authority on Jenkins (primary), Harness, Git / GitHub Enterprise, uDeploy, Nexus, and SonarQube — driving platform stability, service reliability, and cost-efficiency through strong operational engineering and governance. An AI-first mentality is a core expectation. Cross-team collaboration is a must — this role partners with engineering, security, observability, release management, and infrastructure teams to align standards and unblock delivery.

Key Responsibilities

  • Own the full SDF platform lifecycle: Jenkins, Harness, Git / GitHub Enterprise, uDeploy, Nexus, and SonarQube.
  • Ensure platform stability and availability across SDF tooling through proactive reliability engineering practices.
  • Define and enforce enterprise reliability standards, operational controls, and compliance guardrails.
  • Drive incident prevention and rapid recovery: observability, early risk detection, runbook maturity, and resilience testing.
  • Apply AI-first approaches for reliability diagnostics, intelligent alert triage, and auto-remediation opportunities.
  • Own end-to-end RCA for Sev1/Sev2 SDF incidents — from detection through corrective action and verified closure; publish stakeholder summaries.
  • Drive alert noise reduction across SDF tooling — enforce signal-to-noise standards and measurable alert quality gates.
  • Lead cost optimization initiatives across tooling and infrastructure without compromising reliability or developer experience.
  • Coach junior engineers — define escalation criteria, build triage playbooks, and close knowledge gaps.
  • Cross-team collaboration — partner with application engineering, observability, security, release management, and infrastructure teams to align standards and unblock delivery.
  • Lead end-user experience outcomes as the centralized SDF platform owner, ensuring every reliability and stability improvement delivers a simpler, faster, and more consistent experience across all SDF tools.
  • Build and maintain a self-onboarding framework so new teams can independently adopt SDF operational best practices.
  • Manage SDF infrastructure via Terraform and Ansible; automate provisioning, health checks, and self-healing.
  • Track and report platform reliability and efficiency metrics: availability, MTTR, incident recurrence, and cost savings.

Platform & Tools

Jenkins Harness Git / GitHub Enterprise uDeploy Nexus SonarQube Terraform Ansible Harness Python / Groovy / Bash REST APIs Rally

Requirements

  • Jenkins administration (required)
  • Harness CD pipelines
  • Git / GitHub Enterprise
  • uDeploy
  • Nexus Repository Manager
  • SonarQube
  • Platform reliability engineering and operational governance
  • Service stability, resilience, and observability practices
  • Cross-team collaboration (required)
  • End-to-end RCA ownership
  • Cost optimization and efficiency mindset
  • Coaching and mentoring junior engineers
  • Terraform / Ansible (IaC)
  • AI-first problem-solving approach
  • Regulated industry experience a plus


Skills Required
Jenkins, Git, Ansible, Groovy, Sonarqube, Bash, Terraform, Udeploy, Rest Apis, Rally, Harness, Nexus, Python

Create a job alert for this search

Lead Reliability Engineer • Hyderabad / Secunderabad, Telangana, India