Talent.com
This job offer is not available in your country.
Senior Site Reliability Engineer

Senior Site Reliability Engineer

ConfidentialChennai, India
8 days ago
Job description

Join us as we work to create a thriving ecosystem that delivers accessible, high-quality, and sustainable healthcare for all.

We are looking for a Senior Site Reliability Engineer to join our Service Operations, Site Reliability Engineering team within the Cloud Infrastructure Engineering division. This team is newly formed and is responsible for managing the fleet of systems owned by its sister teams in the Service Operations zone. We're looking for Site Reliability & Infrastructure Engineering experts to help us build out the tools and processes necessary for success. Come help us create a thriving ecosystem that delivers accessible, high-quality and sustainable healthcare for all!

The Team :

The Service Operations Site Reliability Engineering team is a newly formed branch of the Network Operations Center. The team sits within the Cloud Infrastructure Engineering (CIE) division which is responsible for delivering high quality and highly available SaaS infrastructure and internal tools. You will be working closely with our internal stakeholders from the R&D and Cloud Infrastructure organizations using modern toolsets to deliver resilient & scalable business solutions. Consistent with SRE practices, a key goal for this team is to measure & reduce toil, not only for ourselves but for the broader NOC & Service Operations organizations.

Job Responsibilities

  • . Provisioning and ongoing management of physical & virtual Linux machines using tools like Puppet, Ansible, and Terraform, to name a few
  • . Engage closely with sister teams to assume ownership of various system lifecycle tasks
  • . Automate away toil and / or create empowerment processes for transitioning high urgency work to the NOC's rapid response team
  • . Build automated monitoring & observability using tools such as Prometheus / AlertManager, iCinga, Grafana, etc.
  • . Participate in all Agile / scrum ceremonies including daily stand-ups, sprint planning, backlog grooming, etc.
  • . Participate in the team's on-call rotation (expected to begin late 2024, early 2025)
  • . Work closely with internal teams to integrate new monitoring & alerts into the NOC using Perl scripting to author custom parsing & mapping rules
  • . Develop metrics and observability dashboards which can be used to measure and track various success measures for the team & the business

Typical Qualifications

  • . 5+years of professional experience delivering SaaS solutions, preferably in a hybrid cloud environment
  • . Bachelor's or Master's degree in a Computer Science / Engineering program
  • . Proven experience using query languages to deliver observability solutions
  • . Proficiency working with one or more configuration management tools (Puppet, Chef, Ansible, etc.)
  • . Admin-level expertise with a Unix-based operating system
  • . Proven ops background using cloud-native best practices
  • . Proven proficiency with one or more scripting languages (Python, Ruby, Perl, Java, etc.)
  • . Proficiency working with Git & Atlassian suite or similar
  • . Proficiency working with containerized environments is a plus
  • . Experience creating technical documentation & standard operating procedures (SOPs)
  • Skills Required

    Terraform, Java, Puppet, Grafana, Ansible, Linux, Icinga, Ruby, Prometheus, Python, Perl, Git

    Create a job alert for this search

    Senior Site Reliability Engineer • Chennai, India

    Related jobs
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    PoshmarkChennai, Tamil Nadu, India
    We’re looking for an experienced Site Reliability Engineer to fill the mission-critical role of ensuring that our complex, web-scale systems are healthy, monitored, automated, and designed to scale...Show moreLast updated: 2 days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    Zyoin GroupChennai, Tamil Nadu, India
    Site Reliability Engineer (SRE).Chennai (Hybrid – 2 days in office).We are seeking a Site Reliability Engineer (SRE) responsible for leading reliability practices, ensuring scalable systems, and co...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    Amicon Hub Serviceschennai, tamil nadu, in
    Manage and scale production systems hosted on.Automate operational tasks using.Improve system reliability and reduce manual interventions through automation. Collaborate with development teams to en...Show moreLast updated: 4 days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    ConfidentialChennai
    The right candidate will put the customer first, understand their user stories and will identify ways to support them, ensuring stability and ability to scale and meeting the needs both of our cust...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    XebiaChennai, IN
    AWS Engineer with strong Python development and Chaos Engineering expertise.The ideal candidate will combine cloud engineering, DevOps, and chaos experimentation to improve reliability, fault toler...Show moreLast updated: 25 days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    UplersChennai, IN
    Uplers is hiring for one of the clients.SRE (Oracle Cloud Infrastructure).Remote | Mon–Fri | 10 : 30 AM – 7 : 30 PM IST.Use of personal device required. OCI cloud infrastructure using Terraform and GitL...Show moreLast updated: 23 days ago
    • Promoted
    Reliability Engineer

    Reliability Engineer

    Alp Consulting Ltd.Chennai, Tamil Nadu, India
    Job Title : Reliability Engineer.Qualification : Diploma / BE (Mech.Experience of maintaining the Instruments, Valves, transmitters, Sensors, Control systems (DCS / PLC, SCADA), Analyzers and F &G system...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer - Chaos Management

    Site Reliability Engineer - Chaos Management

    Xebiachennai, tamil nadu, in
    AWS Engineer with strong Python development and Chaos Engineering expertise.The ideal candidate will combine cloud engineering, DevOps, and chaos experimentation to improve reliability, fault toler...Show moreLast updated: 6 days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    ConcordChennai, IN
    Engineers (Individual Contributors).Strong SRE (Site Reliability Engineering).CI / CD, monitoring, automation, infrastructure as code, etc.Show moreLast updated: 17 days ago
    • Promoted
    Senior Site Reliability Engineer- ELK Expert

    Senior Site Reliability Engineer- ELK Expert

    iVedha Inc.Chennai, IN
    Senior Site Reliability Engineer (SRE) – ELK Expert | Platform Engineering Practice.Must be available to work in the EST (US / Canada) Time Zone. Are you a Senior Site Reliability Engineer (SRE) with ...Show moreLast updated: 30+ days ago
    • Promoted
    Poshmark - Senior Site Reliability Engineer - Cloud Infrastructure

    Poshmark - Senior Site Reliability Engineer - Cloud Infrastructure

    POSHMARKChennai
    Job Description : Were looking for an experienced Site Reliability Engineer to fill the mission-critical role of ensuring that our complex, web-scale systems ...Show moreLast updated: 17 days ago
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    WSO2chennai, tamil nadu, in
    Founded in 2005, WSO2 is the largest independent software vendor providing open-source API management, integration, and identity and access management (IAM) to thousands of enterprises in over 90 c...Show moreLast updated: 6 days ago
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    Tata Consultancy ServicesChennai, Tamil Nadu, India
    TCS is looking for Senior Site Reliability Engineer – AWS.Design, implement, and maintain scalable, secure, and highly available infrastructure on AWS. Develop and improve CI / CD pipelines, Infrastru...Show moreLast updated: 3 days ago
    • Promoted
    Site Reliability Engineer 2

    Site Reliability Engineer 2

    ConfidentialChennai
    Work with team to plan, design and deploy new cloud technologies.Create, Maintain , and Enhance Automated Product Deployments. Develop, Modify, Support and maintain AWS based components through Infr...Show moreLast updated: 19 days ago
    • Promoted
    Site Reliability Engineer - Cloud Platforms

    Site Reliability Engineer - Cloud Platforms

    LanceSoft, IncChennai
    Role and Responsibilities : Reporting to Engineering, the Site Reliability Engineer will play a critical role in driving innovation and growth for the Banking Soluti...Show moreLast updated: 17 days ago
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    Loyalytics AIChennai
    Site Reliability / DevOps Engineer to be our first hire in this function, responsible for owning and scaling the reliability, observability, and infrastructure of our platform running entirely on M...Show moreLast updated: 13 days ago
    • Promoted
    RELX - Site Reliability Engineer - IAC Terraform

    RELX - Site Reliability Engineer - IAC Terraform

    REED ELSEVIER INDIA (a part of RELX India Pvt Ltd)Chennai
    Job Description : - Lead initiatives to identify and eliminate manual, repetitive tasks through automation and tooling.Develop s...Show moreLast updated: 17 days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    ElgebraChennai
    Role Overview : We are seeking a highly experienced and technically proficient Site Reliability Engineer (SRE) to join our team in support of our c...Show moreLast updated: 2 days ago