Job Title : Senior ML Platform Engineer
Experience : 6 - 10 Years
Location : Bangalore
Job Summary :
We are looking for a highly skilled Senior ML Platform Engineer to design, build, and manage scalable machine learning platforms and cloud-native AI infrastructure. The ideal candidate will have strong expertise in backend engineering, ML platform operations, distributed systems, and cloud technologies, with hands-on experience in Azure ML, Kubernetes, Spark, and Airflow.
You will play a key role in enabling data scientists and ML engineers by developing reliable, scalable, and production-ready machine learning platforms while ensuring high availability, security, and operational excellence.
Key Responsibilities :
- Design, develop, and maintain scalable Machine Learning platforms and cloud-native AI infrastructure.
- Build and manage backend services and APIs using Python, Java, or Go.
- Design distributed systems and microservices with a focus on scalability, reliability, and fault tolerance.
- Develop, deploy, and manage Azure Machine Learning pipelines, endpoints, and model deployment workflows.
- Support and optimize Apache Spark, PySpark, and Apache Airflow environments.
- Containerize applications using Docker and orchestrate workloads using Kubernetes (AKS preferred).
- Build and maintain CI/CD pipelines for ML model deployment and platform services.
- Troubleshoot and resolve complex production issues across application, infrastructure, and data layers.
- Implement monitoring, logging, and observability using Grafana, Prometheus, and Loki.
- Develop event-driven solutions using Kafka or Azure Event Hub.
- Automate infrastructure provisioning using Terraform, ARM Templates, or other Infrastructure-as-Code tools.
- Collaborate with data scientists, ML engineers, DevOps teams, and business stakeholders to deliver production-ready AI solutions.
- Ensure platform security, governance, scalability, and operational best practices.
Required Skills :
- Strong backend development experience with Python, Java, or Go.
- Expertise in designing distributed systems, microservices, and scalable platform architectures.
- Hands-on experience with Azure Machine Learning services, pipelines, and deployment.
- Strong knowledge of Apache Spark, PySpark, and Apache Airflow.
- Experience with Docker and Kubernetes (AKS preferred).
- Strong troubleshooting and debugging skills across cloud infrastructure, applications, and data pipelines.
- Experience operating enterprise-scale Machine Learning or Data Platforms.
- Hands-on experience with observability tools including Grafana, Prometheus, and Loki.
- Knowledge of event-driven architectures using Kafka or Azure Event Hub.
- Experience with Terraform, ARM Templates, or other Infrastructure-as-Code tools.
- Familiarity with CI/CD pipelines and DevOps best practices.
(ref:hirist.tech)