Skip to content
Skip to content
DevOps Jobs
ADT

Mid-Level Observability Engineer

ADT

Location
Onsite (Irving, TX)
Employment
Full-time
Level
Mid Level
Posted 1 day ago

About the Role

ADT is seeking a Mid-Level Observability Engineer to design and maintain monitoring solutions using Dynatrace across cloud-native and legacy environments. The role focuses on ensuring system reliability, defining SLOs, and automating alert remediation to support high availability and performance.

Skills

Dynatrace Observability DevOps Site Reliability Engineering Kubernetes Docker Python Bash Go PowerShell Cloud Computing CI/CD Service Level Objectives Infrastructure as Code OpenTelemetry Incident Management

Full job details

Summary: 

As our infrastructure and application landscape scales, ensuring high availability, performance, and reliability is critical. We are looking for a dedicated Observability Engineer to help us see clearly into our complex systems, predict issues before they impact customers, and empower our engineering teams with actionable insights.

 

As a Mid-Level Observability Engineer, you will be the driving force behind our application performance monitoring (APM) and infrastructure visibility. Acting as the resident Dynatrace subject matter expert, you will design, implement, and maintain monitoring solutions across our cloud-native and legacy environments. You will partner closely with DevOps, SRE, and software development teams to build a culture of proactive monitoring, define Service Level Objectives (SLOs), and automate alert remediation.

 

Duties and Responsibilities:

  • Dynatrace Administration: Own the deployment, configuration, and lifecycle management of Dynatrace OneAgent, ActiveGates, and integrations across all environments.
  • Instrumentation & Telemetry: Partner with development teams to instrument microservices, databases, and third-party applications to ensure full-stack visibility.
  • Dashboarding & Alerting: Create intuitive, role-based dashboards and configure intelligent, actionable alerts that reduce alert fatigue and accurately trigger on degraded user experiences.
  • Synthetic Monitoring: Design and maintain synthetic checks to simulate user behavior and monitor critical business transactions.
  • Root Cause Analysis: Leverage Dynatrace’s Davis AI to assist incident response teams in rapidly identifying and resolving bottlenecks and outages.
  • Automation & CI/CD: Integrate observability into the deployment pipeline (e.g., using Dynatrace Keptn or standard CI/CD tools) to evaluate performance metrics during the build and release phases.
  • SLI/SLO Management: Help define, measure, and report on Service Level Indicators and Service Level Objectives to hold the engineering organization accountable for reliability.

 

Requirements:

  • Experience: 3–5 years of experience in DevOps, Site Reliability Engineering (SRE), or Systems Administration, with at least 2 years heavily focused on observability and monitoring.
  • Dynatrace Expertise: Strong, hands-on experience deploying and managing Dynatrace (OneAgent, ActiveGate, Synthetics, Application Security, and Log Monitoring).
  • Cloud & Containerization: Working knowledge of modern cloud environments (GCP) and container orchestration platforms (Kubernetes, Docker).
  • Scripting & Automation: Proficiency in at least one scripting language (Python, Bash, Go, or PowerShell) to automate administrative tasks and API interactions.
  • Incident Management: Familiarity with integrating monitoring tools with ITSM and alerting platforms (ServiceNow, Jira).
  • Communication: Strong ability to explain complex technical issues to both engineering teams and non-technical stakeholders.

 

Preferred Skills (Nice to Haves):

  • Current Dynatrace Certifications (e.g., Dynatrace Associate or Professional Certification).
  • Experience configuring Infrastructure as Code (IaC) tools like Terraform or Ansible to automate agent deployments.
  • Familiarity with OpenTelemetry standards and how to ingest OTel data into Dynatrace.
  • Understanding of CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins) and automated performance gates.