Talent.com
Sysmind
Site Reliability Engineer - Cloud InfrastructureSysmind • Pune
Site Reliability Engineer - Cloud Infrastructure

Site Reliability Engineer - Cloud Infrastructure

Sysmind • Pune
2 days ago
Job description

Job Summary :

We are seeking a proactive and experienced Site Reliability Engineer (SRE) with 46 years of hands-on experience in managing cloud-native infrastructure, ensuring high availability, and improving the reliability of enterprise applications. The ideal candidate should possess strong expertise in Kubernetes, Python, Linux, and modern observability tools such as Dynatrace and Prometheus. You will work closely with development, DevOps, and cloud teams to automate operations, optimize system performance, and build resilient production environments.

Key Responsibilities :

- Design, deploy, and manage highly available applications on Kubernetes clusters.

- Develop automation scripts and operational tools using Python.

- Administer, monitor, and troubleshoot Linux/Unix production servers.

- Implement and maintain monitoring, alerting, and observability solutions using Dynatrace, Prometheus, and related tools.

- Analyze production incidents, perform root cause analysis (RCA), and implement preventive measures.

- Monitor infrastructure performance, application health, and system capacity to ensure maximum uptime.

- Collaborate with DevOps and development teams to improve deployment strategies and operational efficiency.

- Support CI/CD pipelines and automate operational workflows wherever possible.

- Maintain system documentation, operational runbooks, and incident response procedures.

- Participate in on-call rotations and provide production support for critical enterprise applications.

- Drive continuous improvements in system reliability, scalability, security, and operational excellence.

Required Skills :

- 46 years of experience in Site Reliability Engineering (SRE), DevOps, or Platform Engineering.

- Strong hands-on expertise in Kubernetes administration and container orchestration.

- Proficiency in Python scripting for automation and operational tasks.

- Strong experience with Linux/Unix system administration and troubleshooting.

- Hands-on experience with monitoring and observability tools such as Dynatrace and Prometheus.

- Experience supporting production environments with high availability requirements.

- Knowledge of incident management, root cause analysis, and system performance tuning.

- Familiarity with CI/CD pipelines and DevOps methodologies.

- Strong analytical, troubleshooting, and communication skills.

Mandatory Skills :

- Kubernetes

- Site Reliability Engineering (SRE)

- Python

- Linux / Unix Administration

- Dynatrace

- Prometheus

- Monitoring & Observability

- Production Support

- Incident Management

- Automation & Scripting

(ref:hirist.tech)
Create a job alert for this search

Site Reliability Engineer - Cloud Infrastructure • Pune