Skip to content
Skip to content
DevOps Jobs
Workiy

Site Reliability Engineer

Workiy

Location
Onsite (Scottsdale, Arizona)
Employment
Contract
Level
Senior Level
Posted 2 weeks ago

About the Role

Workiy is seeking a Site Reliability Engineer to support large-scale, high-performance applications in a hybrid cloud and on-premises environment. The role focuses on building automation, managing transaction journeys, and implementing observability for real-time monitoring and incident resolution.

Skills

Site Reliability Engineering Kubernetes Cloud Infrastructure Automation Go Python Java Rust GCP Rancher Observability OTEL GraphQL Networking Protocols Database Management Distributed Tracing

Full job details

We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-performance applications running in a hybrid environment (on-premises and cloud). The ideal candidate will have strong experience in cloud infrastructure, Kubernetes, observability, automation, and production operations.

Requirements

  • Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
  • Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
  • Experience working with Programming languages such as Go, Python, Java, Rust etc.
  • Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any  time-series databases
  • Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
  • Experience maintaining containerized app in GKE/RKE/AKE environments.
  • Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
  • Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
  • Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
  • Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
  • Experience working with Programming languages such as Go, Python, Java, Rust etc.
  • Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any  time-series databases
  • Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
  • Experience maintaining containerized app in GKE/RKE/AKE environments.
  • Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
  • Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
  • Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.