This job offer is not available in your country.

Senior Site Reliability Engineer

ConfidentialChennai, India

8 days ago

Job description

Join us as we work to create a thriving ecosystem that delivers accessible, high-quality, and sustainable healthcare for all.

We are looking for a Senior Site Reliability Engineer to join our Service Operations, Site Reliability Engineering team within the Cloud Infrastructure Engineering division. This team is newly formed and is responsible for managing the fleet of systems owned by its sister teams in the Service Operations zone. We're looking for Site Reliability & Infrastructure Engineering experts to help us build out the tools and processes necessary for success. Come help us create a thriving ecosystem that delivers accessible, high-quality and sustainable healthcare for all!

The Team :

The Service Operations Site Reliability Engineering team is a newly formed branch of the Network Operations Center. The team sits within the Cloud Infrastructure Engineering (CIE) division which is responsible for delivering high quality and highly available SaaS infrastructure and internal tools. You will be working closely with our internal stakeholders from the R&D and Cloud Infrastructure organizations using modern toolsets to deliver resilient & scalable business solutions. Consistent with SRE practices, a key goal for this team is to measure & reduce toil, not only for ourselves but for the broader NOC & Service Operations organizations.

Job Responsibilities

. Provisioning and ongoing management of physical & virtual Linux machines using tools like Puppet, Ansible, and Terraform, to name a few
. Engage closely with sister teams to assume ownership of various system lifecycle tasks
. Automate away toil and / or create empowerment processes for transitioning high urgency work to the NOC's rapid response team
. Build automated monitoring & observability using tools such as Prometheus / AlertManager, iCinga, Grafana, etc.
. Participate in all Agile / scrum ceremonies including daily stand-ups, sprint planning, backlog grooming, etc.
. Participate in the team's on-call rotation (expected to begin late 2024, early 2025)
. Work closely with internal teams to integrate new monitoring & alerts into the NOC using Perl scripting to author custom parsing & mapping rules
. Develop metrics and observability dashboards which can be used to measure and track various success measures for the team & the business

Typical Qualifications

. 5+years of professional experience delivering SaaS solutions, preferably in a hybrid cloud environment

. Bachelor's or Master's degree in a Computer Science / Engineering program

. Proven experience using query languages to deliver observability solutions

. Proficiency working with one or more configuration management tools (Puppet, Chef, Ansible, etc.)

. Admin-level expertise with a Unix-based operating system

. Proven ops background using cloud-native best practices

. Proven proficiency with one or more scripting languages (Python, Ruby, Perl, Java, etc.)

. Proficiency working with Git & Atlassian suite or similar

. Proficiency working with containerized environments is a plus

. Experience creating technical documentation & standard operating procedures (SOPs)

Skills Required

Terraform, Java, Puppet, Grafana, Ansible, Linux, Icinga, Ruby, Prometheus, Python, Perl, Git

Create a job alert for this search

Senior Site Reliability Engineer • Chennai, India

Related jobs

Promoted

Senior Site Reliability Engineer

PoshmarkChennai, Tamil Nadu, India

We’re looking for an experienced Site Reliability Engineer to fill the mission-critical role of ensuring that our complex, web-scale systems are healthy, monitored, automated, and designed to scale...Show moreLast updated: 2 days ago

Promoted

Site Reliability Engineer

Zyoin GroupChennai, Tamil Nadu, India

Site Reliability Engineer (SRE).Chennai (Hybrid – 2 days in office).We are seeking a Site Reliability Engineer (SRE) responsible for leading reliability practices, ensuring scalable systems, and co...Show moreLast updated: 30+ days ago

Promoted

Site Reliability Engineer

Amicon Hub Serviceschennai, tamil nadu, in

Manage and scale production systems hosted on.Automate operational tasks using.Improve system reliability and reduce manual interventions through automation. Collaborate with development teams to en...Show moreLast updated: 4 days ago

Promoted

Site Reliability Engineer

ConfidentialChennai

The right candidate will put the customer first, understand their user stories and will identify ways to support them, ensuring stability and ability to scale and meeting the needs both of our cust...Show moreLast updated: 30+ days ago

Promoted

Site Reliability Engineer

XebiaChennai, IN

AWS Engineer with strong Python development and Chaos Engineering expertise.The ideal candidate will combine cloud engineering, DevOps, and chaos experimentation to improve reliability, fault toler...Show moreLast updated: 25 days ago

Promoted

Site Reliability Engineer

UplersChennai, IN

Uplers is hiring for one of the clients.SRE (Oracle Cloud Infrastructure).Remote | Mon–Fri | 10 : 30 AM – 7 : 30 PM IST.Use of personal device required. OCI cloud infrastructure using Terraform and GitL...Show moreLast updated: 23 days ago

Promoted

Reliability Engineer

Alp Consulting Ltd.Chennai, Tamil Nadu, India

Job Title : Reliability Engineer.Qualification : Diploma / BE (Mech.Experience of maintaining the Instruments, Valves, transmitters, Sensors, Control systems (DCS / PLC, SCADA), Analyzers and F &G system...Show moreLast updated: 30+ days ago

Promoted

Site Reliability Engineer - Chaos Management

Xebiachennai, tamil nadu, in

Promoted

Site Reliability Engineer

ConcordChennai, IN

Engineers (Individual Contributors).Strong SRE (Site Reliability Engineering).CI / CD, monitoring, automation, infrastructure as code, etc.Show moreLast updated: 17 days ago

Promoted

Senior Site Reliability Engineer- ELK Expert

iVedha Inc.Chennai, IN

Senior Site Reliability Engineer (SRE) – ELK Expert | Platform Engineering Practice.Must be available to work in the EST (US / Canada) Time Zone. Are you a Senior Site Reliability Engineer (SRE) with ...Show moreLast updated: 30+ days ago

Promoted

Poshmark - Senior Site Reliability Engineer - Cloud Infrastructure

POSHMARKChennai

Job Description : Were looking for an experienced Site Reliability Engineer to fill the mission-critical role of ensuring that our complex, web-scale systems ...Show moreLast updated: 17 days ago

Promoted

Senior Site Reliability Engineer

WSO2chennai, tamil nadu, in

Founded in 2005, WSO2 is the largest independent software vendor providing open-source API management, integration, and identity and access management (IAM) to thousands of enterprises in over 90 c...Show moreLast updated: 6 days ago

Promoted

Senior Site Reliability Engineer

Tata Consultancy ServicesChennai, Tamil Nadu, India

TCS is looking for Senior Site Reliability Engineer – AWS.Design, implement, and maintain scalable, secure, and highly available infrastructure on AWS. Develop and improve CI / CD pipelines, Infrastru...Show moreLast updated: 3 days ago

Promoted

Site Reliability Engineer 2

ConfidentialChennai

Work with team to plan, design and deploy new cloud technologies.Create, Maintain , and Enhance Automated Product Deployments. Develop, Modify, Support and maintain AWS based components through Infr...Show moreLast updated: 19 days ago

Promoted

Site Reliability Engineer - Cloud Platforms

LanceSoft, IncChennai

Role and Responsibilities : Reporting to Engineering, the Site Reliability Engineer will play a critical role in driving innovation and growth for the Banking Soluti...Show moreLast updated: 17 days ago

Promoted

Senior Site Reliability Engineer

Loyalytics AIChennai

Site Reliability / DevOps Engineer to be our first hire in this function, responsible for owning and scaling the reliability, observability, and infrastructure of our platform running entirely on M...Show moreLast updated: 13 days ago

Promoted

RELX - Site Reliability Engineer - IAC Terraform

REED ELSEVIER INDIA (a part of RELX India Pvt Ltd)Chennai

Job Description : - Lead initiatives to identify and eliminate manual, repetitive tasks through automation and tooling.Develop s...Show moreLast updated: 17 days ago

Promoted

Site Reliability Engineer

ElgebraChennai

Role Overview : We are seeking a highly experienced and technically proficient Site Reliability Engineer (SRE) to join our team in support of our c...Show moreLast updated: 2 days ago