Job Description This is a remote position.
SRE Observability Engineer (Offshore) Location: Offshore / Remote
Experience: 6+ Years
Employment Type: Contract / Full-TimeRole Overview
We are seeking an experienced SRE Observability Engineer to join our offshore team. This role is the backbone of our production reliability strategy — owning observability tooling, monitoring architecture, and incident response pipelines across a cloud-native AWS environment. The ideal candidate brings deep hands-on expertise with Dynatrace, Datadog, and AWS-native services, and can independently drive observability maturity from instrumentation to actionable insights.
Key Responsibilities - Design, implement, and maintain observability frameworks across distributed systems using Dynatrace and Datadog (APM, infrastructure monitoring, log management, synthetic monitoring, RUM)
- Define and enforce alerting standards, SLOs, SLIs, and error budgets for production environments
- Build and manage logging pipelines — ingestion, indexing, and retention strategies to maximize signal-to-noise for development and ops teams
- Develop dashboards, runbooks, and reports for uptime, latency, cost, and system health visibility
- Instrument AWS-native services (EC2, EKS, EMR, Lambda, DynamoDB, S3, Cognito) with APM agents and custom metrics
- Manage capacity planning and proactive anomaly detection to prevent incidents before impact
- Coordinate with DevOps, Platform, and Development teams on escalations, on-call rotations, and post-incident reviews
- Automate configuration management for monitoring agents and infrastructure using Ansible, Terraform, or CloudFormation
- Support CI/CD pipeline observability — tracing deployments, correlating releases with reliability metrics
- Mentor and train support personnel on observability tooling and best practices
- Evaluate and onboard new observability technologies; maintain awareness of APM industry trends
Requirements
Required Qualifications Experience 6+ years in IT with at least 3+ years in an SRE, Platform Engineering, or Observability-focused role
3+ years running production systems on AWS
Hands-on production experience with Dynatrace (OneAgent, Davis AI, dashboards, synthetic monitors, session replay)
Hands-on production experience with Datadog (APM, Log Management, Infrastructure, Monitors, SLOs)
Experience instrumenting containerized workloads — Docker, Kubernetes (EKS)
Technical Skills AWS services: EC2, EKS, EMR, Lambda, DynamoDB, S3, Cognito, CloudWatch
Infrastructure-as-Code: Terraform and/or CloudFormation
Configuration management: Ansible or equivalent
Scripting and automation: TypeScript, JavaScript, Python, or Bash
CI/CD pipelines and source control (Git, Jenkins, GitHub Actions, or equivalent)
Strong understanding of networking, DNS, load balancing, and distributed systems troubleshooting
Familiarity with Serverless architectures and event-driven patterns
Experience with Windows Server, IIS, and Linux-based environmentsObservability-Specific Skills
Designing SLO/SLI frameworks and error budget policies
Log structuring, indexing strategy, and cost-aware retention configuration
Distributed tracing (OpenTelemetry or vendor-native)
Synthetic monitoring and real-user monitoring (RUM) configuration
Alert fatigue reduction — tuning thresholds, suppression, and anomaly-based detection
Preferred Qualifications
Dynatrace Professional or Associate certification
Datadog Fundamentals certification
AWS Certified SysOps Administrator or AWS Certified DevOps Engineer
Experience with Couchbase or NoSQL monitoring
Familiarity with PagerDuty, OpsGenie, or similar incident management platforms
Experience in financial services, healthcare, or other regulated environments
Soft Skills Clear communicator — able to translate technical observability concepts to non-technical stakeholders
Proactive, ownership-driven mindset with bias toward automation over manual processes
Collaborative team player comfortable working across time zones in an offshore delivery model
Trusted and transparent — raises issues early, documents decisions, shares knowledge freely
Nice to Have
Experience with Splunk, Grafana, Prometheus, or New Relic alongside primary tooling
Exposure to FinOps practices and AWS Cost Explorer dashboards
Contributions to open-source observability projects or internal tooling
RequirementsRequired Qualifications Experience 6+ years in IT with at least 3+ years in an SRE, Platform Engineering, or Observability-focused role 3+ years running production systems on AWS Hands-on production experience with Dynatrace (OneAgent, Davis AI, dashboards, synthetic monitors, session replay) Hands-on production experience with Datadog (APM, Log Management, Infrastructure, Monitors, SLOs) Experience instrumenting containerized workloads — Docker, Kubernetes (EKS) Technical Skills AWS services: EC2, EKS, EMR, Lambda, DynamoDB, S3, Cognito, CloudWatch Infrastructure-as-Code: Terraform and/or CloudFormation Configuration management: Ansible or equivalent Scripting and automation: TypeScript, JavaScript, Python, or Bash CI/CD pipelines and source control (Git, Jenkins, GitHub Actions, or equivalent) Strong understanding of networking, DNS, load balancing, and distributed systems troubleshooting Familiarity with Serverless architectures and event-driven patterns Experience with Windows Server, IIS, and Linux-based environments Observability-Specific Skills Designing SLO/SLI frameworks and error budget policies Log structuring, indexing strategy, and cost-aware retention configuration Distributed tracing (OpenTelemetry or vendor-native) Synthetic monitoring and real-user monitoring (RUM) configuration Alert fatigue reduction — tuning thresholds, suppression, and anomaly-based detection Preferred Qualifications Dynatrace Professional or Associate certification Datadog Fundamentals certification AWS Certified SysOps Administrator or AWS Certified DevOps Engineer Experience with Couchbase or NoSQL monitoring Familiarity with PagerDuty, OpsGenie, or similar incident management platforms Experience in financial services, healthcare, or other regulated environments Soft Skills Clear communicator — able to translate technical observability concepts to non-technical stakeholders Proactive, ownership-driven mindset with bias toward automation over manual processes Collaborative team player comfortable working across time zones in an offshore delivery model Trusted and transparent — raises issues early, documents decisions, shares knowledge freely Nice to Have Experience with Splunk, Grafana, Prometheus, or New Relic alongside primary tooling Exposure to FinOps practices and AWS Cost Explorer dashboards Contributions to open-source observability projects or internal tooling