Role Overview :
We are seeking a seasoned Site Reliability Engineer to join our high-performance infrastructure team. In this role, you will be the bridge between development and operations, ensuring our large-scale distributed systems remain resilient, scalable, and highly available. You will work closely with cross-functional engineering teams, product managers, and stakeholders to define service level objectives and implement automation that eliminates manual toil. Your contributions will directly impact the end-user experience by minimizing latency and downtime, ultimately driving the reliability standards that support our business growth across global markets.
Key Responsibilities :
- Architect and maintain robust Kubernetes clusters to ensure seamless container orchestration and high availability for mission-critical applications.
- Design and implement automated CI/CD pipelines to accelerate deployment cycles while maintaining rigorous quality and security standards.
- Conduct deep-dive incident analysis and post-mortems to identify root causes and implement preventative measures that enhance system stability.
- Optimize cloud infrastructure costs and performance by proactively monitoring resource utilization and scaling strategies.
- Collaborate with software developers to integrate reliability best practices into the application lifecycle, ensuring code is production-ready from day one.
Required Skillset :
- Demonstrated expertise in managing complex Kubernetes environments at scale, including troubleshooting networking, storage, and security configurations.
- Proven ability to write clean, maintainable code for infrastructure automation using languages such as Python, Go, or Bash.
- Strong proficiency in managing cloud-native monitoring and observability stacks to proactively detect and resolve system bottlenecks.
- Excellent communication skills with the ability to articulate technical risks and architectural decisions to both technical and non-technical stakeholders.
- A collaborative mindset that thrives in a fast-paced, hybrid work environment, balancing independent problem-solving with team-based project delivery.
- A Bachelor's or Master's degree in Computer Science or a related field, complemented by 6 to 10 years of hands-on experience in SRE or DevOps roles.
(ref:hirist.tech)