Experience: 3 - 5 Years
Location: Bangalore
Mode of Work: Hybrid
Key Responsibilities:
- Build, deploy, and manage comprehensive MLOps and LLMOps pipelines on Azure.
- Design and oversee CI/CD pipelines for machine learning models and large language model workflows utilizing Harness or Azure DevOps.
- Streamline the promotion of models, prompts, and agent workflows between environments through automation.
- Establish approval gates, implement rollback mechanisms, and facilitate controlled release processes.
- Oversee the lifecycle of ML models and LLM-driven workflows, including their training, assessment, deployment, monitoring, and retraining.
- Administer Azure Machine Learning workspaces, computing resources, environments, model registries, and endpoints.
- Integrate LLM workflows and agent-centric architectures using LangGraph.
- Support the incorporation of Claude-based models, skills, and plugins into enterprise-level applications.
- Operationalize prompt versioning, orchestration strategies, and agent workflows in live production settings.
- Set up and govern Azure ML and Generative AI infrastructure via Terraform as Infrastructure as Code (IaC).
- Standardize development, test, and production environments for ML and GenAI workloads.
- Automate recurring platform operations and environment setup tasks.
- Deploy both real-time and batch inference solutions for machine learning and LLM use cases.
- Enable scalable Generative AI services that are reliable, high-performing, and cost-effective.
- Apply blue-green or canary deployment techniques where suitable.
- Implement monitoring practices to track model accuracy, data drift, prompt drift, and agent behavior.
- Ensure thorough logging, alerting systems, and observability for ML and GenAI platforms.
- Adhere to enterprise standards concerning security, access controls, and cost management.
- Collaborate closely with engineering, platform, and development teams.
- Guide teams toward adopting best practices in MLOps and LLMOps.
- Contribute to the creation and maintenance of reusable templates, frameworks, and documentation.
Required Skills And Qualification:
1. Core Skills:
- 3-5 years of experience in MLOps / ML Engineering / Cloud Engineering
- Proficient in designing and deploying end-to-end ML pipelines
- Terraform for Azure infrastructure automation
- Python for ML, automation, and GenAI workflows
2. Azure & Cloud:
- Azure Compute, Storage, Networking, and Identity
- Running ML & GenAI workloads at scale on Azure
- Supporting data pipelines for ML and LLM workloads
3. GenAI / LLMOps:
- Experience with LangGraph for LLM workflow and agent orchestration
- Hands-on exposure to Claude models, including skills/plugins integration
- Understanding of prompt management, agent execution, and orchestration patterns
4. DevOps & Operations:
- Monitoring, logging, and alerting practices
- Troubleshooting production ML/GenAI systems
- Cost-optimized design for compute-intensive AI workloads
Good To Have Skills:
- Experience with Responsible AI and enterprise governance
- Exposure to multi-agent architectures
- FinOps awareness for ML and GenAI workloads
- Experience supporting multiple teams via a shared AI platform
(ref:hirist.tech)