Role : Site Reliability Engineer (SRE) Python & Kubernetes
Location : Bangalore
Experience : 4 to 6 Years
Role : Engineer
About the Role :
We are looking for a highly motivated Site Reliability Engineer (SRE) with strong expertise in Kubernetes, Python, Linux, and Monitoring Tools to ensure the reliability, scalability, and performance of enterprise applications. The ideal candidate will have hands-on experience in automation, container orchestration, system monitoring, and production support.
Key Responsibilities :
- Manage, monitor, and optimize applications running on Kubernetes clusters.
- Develop automation scripts and operational tools using Python.
- Monitor application and infrastructure health using Dynatrace, Prometheus, and other observability tools.
- Troubleshoot production issues, perform root cause analysis, and implement preventive solutions.
- Maintain and support Linux-based environments, ensuring system availability and performance.
- Collaborate with development, DevOps, and infrastructure teams to improve system reliability and deployment processes.
- Configure monitoring dashboards, alerts, and performance metrics.
- Participate in incident management, capacity planning, and continuous service improvements.
- Prepare technical documentation and follow operational best practices.
Required Skills :
- 46 years of experience in Site Reliability Engineering (SRE), DevOps, or Production Support.
- Strong hands-on experience with Kubernetes and containerized environments.
- Proficiency in Python scripting and automation.
- Strong knowledge of Linux/Unix administration and command-line operations.
- Experience with monitoring and observability tools such as Dynatrace and Prometheus.
- Excellent troubleshooting, debugging, and analytical skills.
- Understanding of system performance, reliability, and high availability concepts.
Preferred Skills :
- Exposure to Docker and container technologies.
- Knowledge of CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI.
- Familiarity with cloud platforms such as AWS, Azure, or GCP.
- Experience with Infrastructure as Code (Terraform or Ansible).
- Understanding of Agile and DevOps methodologies.
Why Join Us :
- Work on large-scale, mission-critical production systems.
- Build and manage cloud-native applications using Kubernetes.
- Gain hands-on experience with industry-leading monitoring and observability platforms.
- Collaborate with experienced SRE, DevOps, and Cloud Engineering teams.
- Accelerate your career in Site Reliability Engineering and cloud infrastructure.
(ref:hirist.tech)