- Write clean, reusable, and efficient infrastructure code using tools such as Puppet, Python, Terraform, and Packer
- Collaborate closely with senior tech leads and engineering teams to support ongoing development and architectural discussions
- Troubleshoot and resolve complex infrastructure and system issues in collaboration with cross-functional teams
- Contribute to and maintain documentation related to infrastructure, processes, and standard operating procedures
- Ensure compliance with defined technical standards, policies, and operational best practices
- Learn and apply SRE (Site Reliability Engineering) principles to ensure system performance, reliability, and scalability
- Support product and development teams in designing, deploying, and scaling production data infrastructure, including:
- CloudSQL, Firestore, Pub/Sub, Kafka, Dataproc, Airflow
- Participate in migration projects, including transitioning legacy platforms to cloud-based infrastructure
- Use software development practices to enable self-service deployment of distributed systems and microservices
Skills Required
Devops, Puppet, Gcp, Site Reliability Engineering, Aws