Talent.com
This job offer is not available in your country.
Site Reliability Engineer

Site Reliability Engineer

XebiaHosur, Tamil Nadu, India
26 days ago
Job description

We are looking for a highly skilled AWS Engineer with strong Python development and Chaos Engineering expertise to design, build, and validate resilient, scalable, and automated cloud-native environments. The ideal candidate will combine cloud engineering, DevOps, and chaos experimentation to improve reliability, fault tolerance, and operational efficiency of critical systems.

Key Responsibilities

Cloud Engineering (AWS) :

Architect, implement, and manage secure, scalable, and cost-efficient AWS infrastructure (EC2, Lambda, EKS, S3, RDS, IAM, CloudFront, etc.).

Automate infrastructure provisioning and configuration using Terraform / CloudFormation and AWS SDKs.

Manage containerized workloads (Docker, Kubernetes, EKS).

Python Development :

Build automation scripts, deployment utilities, and infrastructure tooling using Python (Boto3, Flask, FastAPI, etc.) .

Develop custom monitoring / alerting integrations with APIs, SDKs, and third-party observability platforms.

Implement self-healing and resilience-focused automation scripts.

Chaos Engineering & Resiliency :

Design and execute chaos experiments (fault injection, latency, outages, resource failures) to validate system resilience.

Use tools like Gremlin, Litmus, Chaos Mesh, or AWS Fault Injection Simulator .

Partner with SRE and development teams to define SLIs, SLOs, and error budgets .

Document learnings from chaos tests and improve incident response & recovery playbooks.

DevOps & Observability :

Build and maintain CI / CD pipelines for automated deployments (Jenkins, GitHub Actions, GitLab CI, AWS CodePipeline).

Integrate observability frameworks (Prometheus, Grafana, ELK / EFK, CloudWatch, Datadog) for monitoring and tracing.

Ensure proactive alerting and real-time visibility into system health.

Security & Compliance :

Apply AWS security best practices for IAM, networking, and data protection.

Ensure compliance with internal and external regulatory frameworks (SOC2, ISO, GDPR, etc.).

Required Skills & Qualifications

6–10 years of experience in Cloud, DevOps, or SRE roles.

Strong hands-on expertise in AWS Cloud (certifications preferred : AWS DevOps Engineer / Solutions Architect).

Advanced Python development skills for automation and tooling (Boto3 a must).

Experience designing and running chaos experiments (Gremlin, AWS FIS, Litmus, Chaos Mesh, or custom Python-based fault injection).

Solid knowledge of IaC (Terraform / CloudFormation) .

Proficiency in containers & orchestration (Docker, Kubernetes, EKS) .

Strong background in monitoring, observability, and incident management .

Familiarity with DevOps toolchain (CI / CD, Git, Jenkins, GitLab, CodePipeline) .

Good understanding of resilient architectures, reliability principles, and disaster recovery .

Preferred Skills

Knowledge of Go / Shell scripting in addition to Python.

Experience with chaos testing in production-like environments .

Exposure to multi-cloud or hybrid-cloud environments .

Strong problem-solving and debugging skills.

What We Offer

Opportunity to lead cloud reliability & chaos engineering initiatives .

Culture focused on automation, resilience, and continuous improvement .

Growth opportunities through certifications, R&D projects, and leadership roles.

Create a job alert for this search

Site Reliability Engineer • Hosur, Tamil Nadu, India

Related jobs
Site Reliability Engineer

Site Reliability Engineer

AIONBengaluru, KA, IN
Quick Apply
AION is building the next generation of AI cloud platform by transforming the future of high-performance computing (HPC) through its decentralized AI cloud. Purpose-built for bare-metal performance,...Show moreLast updated: 30+ days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

ViewSonicBengaluru, Karnataka, India
Bachelor's degree in Computer Science, Engineering, or a related field.Site Reliability Engineer, DevOps Engineer, or similar, is preferred but not mandatory. Basic understanding of AWS solutions in...Show moreLast updated: 17 days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

ElgebraBangalore
Role Overview : We are seeking a highly experienced and technically proficient Site Reliability Engineer (SRE) to join our team in support of our c...Show moreLast updated: 4 days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

SynechronBengaluru, Karnataka, India
We have immediate opportunity for Senior Site Reliability Engineer.Senior Site Reliability Engineer.At Synechron, we believe in the power of digital to transform businesses for the better.Our globa...Show moreLast updated: 30+ days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

Core Minds Tech SOlutionsHosur
Job Description : - Engage with our product teams to understand requirements, design, and implement resilient and scalable infrastructure solutions&l...Show moreLast updated: 30+ days ago
  • Promoted
LSEG - Site Reliability Engineer

LSEG - Site Reliability Engineer

REFINITIV INDIA SHARED SERVICES PRIVATE LIMITEDBangalore
LSEG is a leading global financial markets infrastructure and data provider.Our purpose is driving financial stability, empowering economies and enabling customers to create sustainable growth.Our ...Show moreLast updated: 30+ days ago
Site Reliability Engineer

Site Reliability Engineer

Aqilea (formerly Soltia)Bangalore, Karnataka, India
Quick Apply
We are a consulting company with a bunch of technology-interested and happy people!.We love technology, we love design and we love quality. Our diversity makes us unique and creates an inclusive and...Show moreLast updated: 30+ days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

TavantBengaluru, Karnataka, India
With 25+ years of experience building innovative digital products and solutions, Tavant provides impactful results to its customers. It has been the frontrunner in driving digital innovation and tec...Show moreLast updated: 26 days ago
  • Promoted
  • New!
Site Reliability Engineer

Site Reliability Engineer

BayOne SolutionsBengaluru, Karnataka, India
Role : Site Reliability Engineer.The CXE Site Reliability Engineering (SRE) team manages the CI / CD pipelines and cloud infrastructure, ensuring seamless deployment, monitoring, and maintenance.Howev...Show moreLast updated: 16 hours ago
  • Promoted
Lead Site Reliability Engineer

Lead Site Reliability Engineer

Delta Air LinesBengaluru, India
Execute on the Incident, Change Management, Problem Management processes.Building and supporting reliable applications that meet development and maintenance requirements. Provide consultation and di...Show moreLast updated: 30+ days ago
  • Promoted
  • New!
Site Reliability Engineer

Site Reliability Engineer

ExasoftBangalore, IN
Responsibilities and Requirements : .Experience must be at least 10+ years in SRE.Multi Cloud, Hybrid Cloud – on Data center sites. Experience with multiple operating systems (.Operating Systems, Kern...Show moreLast updated: 16 hours ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

People Realm Recruitment Services Private LimitedBengaluru, Karnataka, India
Job Title- Site Reliability Engineer.Desired Years of Experience - 5 - 14 Years of Relevant Experience.A Career with a Leading Global Investment Management Firm’s Technology Team.Our client, a lead...Show moreLast updated: 20 days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

XebiaBengaluru, Karnataka, India
AWS DevOps Engineer with strong expertise in Observability and Site Reliability Engineering (SRE).The role requires hands-on experience with AWS services, Infrastructure as Code (IaC), CI / CD, monit...Show moreLast updated: 30+ days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

WhiteLotus Talent PartnersBengaluru, Karnataka, India
L0 and L1 Site Reliability Engineer (SRE) Support.Krutrim Cloud Site Reliability operations team and ensure the smooth functioning of our cloud infrastructure powered by. In this role, you will focu...Show moreLast updated: 30+ days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

Amicon Hub ServicesBengaluru, Karnataka, India
Manage and scale production systems hosted on.Automate operational tasks using.Improve system reliability and reduce manual interventions through automation. Collaborate with development teams to en...Show moreLast updated: 7 days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

ConcordBangalore, IN
Engineers (Individual Contributors).Strong SRE (Site Reliability Engineering).CI / CD, monitoring, automation, infrastructure as code, etc.Show moreLast updated: 18 days ago
  • Promoted
Senior Site Reliability Engineer

Senior Site Reliability Engineer

ViewSonicBengaluru, Karnataka, India
At ViewSonic Technologies, we’re passionate about building software that solves problems.We count on our site reliability engineers (SREs) to empower users with a rich feature set, high availabilit...Show moreLast updated: 30+ days ago
  • Promoted
Site Reliability Engineer

Site Reliability Engineer

UplersBengaluru, Karnataka, India
Uplers is hiring for one of the clients.Role Details : Position : SRE (Oracle Cloud Infrastructure) Type : 10-month contract (possible extension) Mode : Remote | Mon–Fri | 10 : 30 AM – 7 : 30 PM IST Pol...Show moreLast updated: 24 days ago
  • Promoted
  • New!
Site Reliability Engineer

Site Reliability Engineer

TrantorBengaluru, Karnataka, India
Job Title - Site Reliability Engineer Role- Contract (9 Months- Extendable) Exp- 5+ years Loc- Bangalore ( Hybrid) Notice- Immediate joiner only Duties : Responsible for maintaining and scaling pro...Show moreLast updated: 11 hours ago
  • Promoted
  • New!
Site Reliability Engineer

Site Reliability Engineer

ACL DigitalBengaluru, Karnataka, India
Service Management : Maintain application uptime / performance, manage system enhancements and defects, oversee daily operational activities, and ensure continuous improvement and adherence to ITIL be...Show moreLast updated: 11 hours ago