About the Role
We are looking for a hands-on Senior MLOps / LLMOps Engineer with 7–9 years of experience to help operationalize machine learning and Generative AI capabilities across enterprise products and platforms.
This role is focused on MLOps and LLMOps, including model lifecycle management, deployment automation, prompt and model governance, evaluation pipelines, observability, monitoring, controlled rollout, release evidence, and production support.
The role will support Alegeus' broader AI enablement and governed AI adoption model by helping ensure that AI/ML models, LLM-powered services, prompts, evaluation assets, and inference workflows are deployed through repeatable, traceable, secure, and auditable operational patterns.
This is not an AI Engineer or product feature engineering role. The candidate is not expected to primarily build AI-enabled product features, business APIs, or application workflows. Instead, this role will focus on creating the operational foundation that allows AI engineers, data scientists, foundational AI teams, and product engineering teams to deploy, monitor, evaluate, govern, and support AI/ML and LLM-powered capabilities in production.
This is also not a data platform engineering role. Enterprise data pipelines, data warehouse operations, and platform-level data engineering will be handled by the data platform team. This role will consume model artifacts, prompts, evaluation datasets, AI service configurations, model endpoints, and platform-provided data interfaces to establish reliable production operations for AI systems.
The ideal candidate should have strong experience with cloud engineering, DevOps, MLOps, AI/ML lifecycle tooling, Azure Machine Learning, MLflow, Azure OpenAI or related LLM platforms, CI/CD automation, monitoring, governance, and production support. A background in Python and/or .NET / C# engineering is preferred for automation, integration, and operational tooling.
Key Responsibilities
MLOps Engineering
- Build and maintain MLOps pipelines for model packaging, validation, registration, deployment, monitoring, rollback, and lifecycle management.
- Standardize model tracking, versioning, approval workflows, promotion gates, release evidence, and lifecycle governance using Azure Machine Learning, MLflow, or similar platforms.
- Support deployment of model endpoints, inference services, batch scoring jobs, and AI/ML service artifacts across environments.
- Create repeatable deployment patterns for moving models from experimentation to production.
- Implement CI/CD and release automation for AI/ML workloads using Azure DevOps, GitHub Actions, or similar tools.
LLMOps Engineering
- Build and maintain LLMOps practices for prompt lifecycle management, prompt versioning, prompt testing, deployment, rollback, and governance.
- Manage operational workflows for LLM-powered capabilities, including prompt assets, model configurations, system instructions, evaluation datasets, retrieval configurations, and AI service settings.
- Support evaluation and monitoring of LLM outputs for correctness, relevance, groundedness, completeness, hallucination risk, safety, latency, and cost.
- Implement prompt and model comparison workflows across environments and releases.
- Support safe rollout patterns for Generative AI capabilities using feature flags, approval gates, A/B testing, canary releases, and human-in-the-loop review where required.
- Track LLM-specific operational signals such as token usage, cost, latency, failed calls, prompt regressions, retrieval quality, and response quality trends.
- Help establish responsible AI controls including auditability, traceability, explainability, privacy, security, and compliance-aware operations.
Governed AI Operations & Auditability
- Help establish governed operational patterns for AI/ML and LLM systems so teams use shared deployment, monitoring, evaluation, and release paths instead of one-off or shadow processes.
- Ensure models, prompts, evaluation datasets, deployment artifacts, and AI configurations are versioned, attributable, and traceable across environments.
- Capture operational evidence for AI releases, including model version, prompt version, approver, deployment environment, test results, evaluation results, rollback plan, and production readiness status.
- Support separation-of-duties and human-in-the-loop controls for sensitive model releases, prompt changes, and production deployments.
- Work with Security, Governance, Risk, Privacy, Internal Audit, Architecture, and Platform teams to ensure AI operational controls are built into the release lifecycle.
- Contribute to reusable templates and standards for AI/ML and LLM release governance, production readiness, incident response, and audit evidence.
Evaluation, Monitoring & Observability
- Build and operate evaluation pipelines for both traditional ML models and LLM-powered systems.
- Implement monitoring for model performance, service health, latency, failures, cost, usage,
Encoded Definition of Done & Production Readiness
- Help define and operationalize an AI/ML Definition of Done for production releases.
- Ensure production releases include appropriate evidence for testing, evaluation, monitoring, security review, rollback, documentation, and supportability.
- Support automated quality gates for model validation, prompt regression testing, response quality checks, security checks, and deployment approval.
- Work with architecture, security, data science, and platform teams to define release criteria for different risk levels of AI/ML and LLM use cases.
- Support phase-gated rollout of AI/ML capabilities so production release is based on evidence, not intent.
- Help ensure regulated or sensitive AI workflows follow appropriate review, approval, vendor, privacy, and compliance controls before production use.
Cloud, DevOps & Platform Integration
- Work with Azure Machine Learning, Azure OpenAI, Azure AI services, Azure Storage, Azure .
Cross-Functional Collaboration
- Work closely with Data Scientists, AI Engineers, Product Engineering teams, Foundational AI teams, Platform teams, Security teams, Architects, Governance/Risk teams, and Product Managers.
- Provide operational guidance for moving AI/ML experiments and LLM-powered capabilities into production.
Required Qualifications
- 7–9 years of professional experience in MLOps, DevOps, ML engineering, cloud engineering, platform engineering, AI operations, or related roles.
- Strong hands-on experience deploying, monitoring, and supporting production-grade AI/ML services or cloud-native platforms.
- Practical experience with MLOps concepts such as model registry, experiment tracking, model packaging, deployment automation, model monitoring, rollback, and lifecycle governance.
- Experience with Azure Machine Learning, MLflow, model registries, or similar model lifecycle management tools.
- Familiarity with LLMOps concepts such as prompt versioning, prompt testing, LLM evaluation, AI observability, token/cost tracking, hallucination monitoring, and safe rollout.
- Experience with Azure OpenAI, OpenAI APIs, Azure AI services, or comparable LLM platforms.
- Programming or scripting experience in Python and/or C#/.NET for automation, tooling, and operational workflows.
- Experience with CI/CD automation using Azure DevOps, GitHub Actions, or similar tools.
- Working knowledge of Docker, cloud deployment patterns, environment promotion, configuration management, and production support.
- Strong understanding of logging, monitoring, alerting, dashboards, tracing, and operational metrics.
- Familiarity with secure release practices, access control, audit logging, approval workflows, and production governance.
- Ability to collaborate across data science, AI engineering, platform, product engineering, security, architecture, and governance teams.
Preferred Qualifications
- Experience operationalizing Generative AI systems involving LLMs, prompts, embeddings, RAG, chatbots, document extraction, summarization, classification, or workflow automation.
- Experience with Azure AI Document Intelligence, Azure AI Search, Azure App Services, Azure Functions, AKS, or related Azure services.
- Experience with traditional ML frameworks such as Scikit-learn, XGBoost, PyTorch, TensorFlow, or similar.
- Exposure to Semantic Kernel, LangChain, LlamaIndex, AutoGen, CrewAI, or similar orchestration frameworks from an operations and observability perspective.
- Experience with evaluation frameworks, golden datasets, regression testing, prompt/model comparison, or AI quality dashboards.
- Exposure to Terraform, Kubernetes, AKS, secrets management, RBAC, audit logging, and cloud security patterns.
- Experience with regulated enterprise environments such as healthcare, fintech, benefits administration, claims processing, or insurance.
- Familiarity with compliance-aware AI operations, data privacy, human-in-the-loop review, separation of duties, and responsible AI practices.
- Strong communication and collaboration skills, with the ability to work across global teams.
Skills Required
Docker, .NET, Azure Devops, Azure Machine Learning, Python