Job descriptionJob Description Aptean is seeking an experienced and hands-on Site Reliability Engineering Manager (Cloud Infrastructure & Operations) to lead a team of 15 engineers. You will be responsible for managing the infrastructure layer of our multi-tenant, cloud-hosted ERP products. This critical role encompasses platform reliability, product upgrades, cloud security, incident and preventive maintenance, disaster recovery, and compliance audits. You will also act as a stage-gate for all production deployments, ensuring release readiness, rollback capability, and platform stability. Principal Duties and Responsibilities - Cloud Infrastructure Oversight: Oversee provisioning, monitoring, and scaling of cloud environments (primarily Azure) for ERP products. Ensure optimal performance, cost control, and platform stability. - SaaS Product Operations: Own product environment availability (Dev, UAT, Prod), plan platform upgrades, apply security patches, and manage certificates and access. - Incident Management: Lead incident response for outages and degradation. Perform Root Cause Analysis (RCA), document learnings, and implement post-mortem action items. - Preventive Maintenance: Define and execute regular health checks, patching schedules, environment cleanups, and alert tuning. - Disaster Recovery Planning: Develop and test Disaster Recovery (DR) / Business Continuity Planning (BCP) plans. Ensure business continuity across all cloud-hosted environments. - Security & Compliance: Lead infrastructure-level compliance activities for SOC 2, ISO 27001, and secure deployment pipelines. Coordinate with infosec and audit teams. - Production Deployment Stage-Gate: Review and approve all deployment tickets. Validate readiness, rollback strategy, and impact analysis before production cutover. - Team Leadership: Lead, coach, and upskill a team of cloud and DevOps engineers. Foster a learning culture aligned with platform reliability and innovation. Qualifications - Education: Bachelor's degree (Required). Master's degree (Preferred). - B.E./B.Tech/MCA in Computer Science or equivalent. - Work Experience: - 10+ years of experience in Cloud Infrastructure / SaaS Operations. - 3+ years managing teams in a cloud product environment (preferably multi-tenant SaaS). - Certifications: ITIL or SRE certification preferred. Required Skills and Abilities - Strong hands-on knowledge of Azure (VMs, PaaS, Networking, Monitoring, Identity). - Experience with ERP platforms (SAP Cloud, Infor, Oracle Cloud, or custom-built ERP solutions). - Good grasp of DevOps practices, CI/CD pipelines, infrastructure as code (IaC). - Familiarity with SOC 2, ISO 27001, and data privacy compliance. Skills Matrix (Manager-Level & Team Needs) - Cloud Platform: Azure (App Services, VM, Networking, Storage, Defender) - Advanced - ERP Infra: Multi-tenant ERP hosting, Cloud DB tuning, PaaS scaling - Advanced - DevOps: CI/CD (Azure DevOps, GitHub Actions), Automation - Intermediate - IaC: Terraform / Bicep / ARM Templates - Intermediate - Monitoring & Logging: Azure Monitor, Application Insights, Log Analytics - Advanced - Incident Management: ITIL, On-call Runbooks, RCA Writing - Expert - Preventive Ops: Scheduled health checks, capacity management - Expert - Security & Access: IAM, Azure AD, Role-based Access, Secret Rotation - Advanced - Disaster Recovery: DR Drills, Geo-Redundancy, RTO/RPO - Advanced - Audit & Compliance: SOC 2, ISO 27001, Risk Registers - Advanced - Release Stage-Gate: Deployment approvals, Go/No-go criteria - Expert - Collaboration: Working with Product, Security, Dev teams - Expert - Tools: Azure DevOps, Jira, ServiceNow, Salesforce (case management) - Intermediate - Leadership: People development, Shift planning, Mentoring - Expert