Skip to content
DevOps Jobs
Origami Risk LLC

Site Reliability Engineer

Origami Risk LLC

Remote (Remote, Georgia) $100k - $120k/yr Full-time Senior Level 8 benefits + 5 perks
Posted 1 day ago

About the role

Origami Risk provides integrated SaaS solutions for risk, compliance, and insurance ecosystems. The Site Reliability Engineer improves system scalability, stability, and incident resolution through post-incident investigations and observability tool management.

Skills

Site Reliability Engineering Incident Management Observability Tools New Relic Data Dog SumoLogic JavaScript .NET SQL AWS Azure C# SaaS Operations CI/CD Infrastructure as Code Troubleshooting
Remote (Duluth, Georgia) Contract Senior Level 1 perks
Posted 1 day ago

About the role

Arctiq is a global technology services firm seeking a Site Reliability Engineer to support a client's vulnerability remediation program. The role focuses on triaging and remediating infrastructure and configuration vulnerabilities across cloud and on-premise environments using automated solutions.

Skills

AWS Kubernetes Vulnerability Management Infrastructure-as-Code Terraform Python Bash PowerShell CI/CD Wiz Linux Windows Server VMware Security Hardening Cloud Security Amazon EKS
Hybrid (Japan) Full-time Senior Level 2 benefits + 2 perks
Posted 1 day ago

About the role

Wayve is a leading developer of Embodied AI technology for autonomous driving. This role leads the Japan Fleet Reliability SRE team, ensuring the safety, availability, and reliability of autonomous vehicle operations while collaborating with global engineering teams.

Skills

Engineering management SRE Autonomous vehicles Incident management Observability Reliability engineering Team leadership Python Go Rust TypeScript Kubernetes Linux Cloud infrastructure System design Automation
Hybrid (Yokohama) Full-time Senior Level 2 perks
Posted 1 day ago

About the role

Wayve is building a global driving intelligence for autonomous vehicles. This role leads the Japan Fleet Reliability SRE team to ensure the safety, reliability, and availability of the autonomous vehicle fleet through engineering management and operational excellence.

Skills

Engineering management SRE Reliability engineering Incident management Observability Team leadership Autonomous vehicles Software engineering Automation Production systems Roadmap planning Performance management Technical strategy Stakeholder management Linux Kubernetes
Onsite (Tokyo) $8640k - $12960k/yr Mid Level 1 benefits + 2 perks
Posted 2 days ago

About the role

Indeed is the world's leading job site, helping people get jobs through AI and real-time data. This role partners with product teams to design, code, and maintain scalable, high-performance systems while driving modern reliability practices and operational excellence.

Skills

Site Reliability Engineering Cloud Infrastructure Kubernetes ArgoCD Terraform GitLab Datadog Java Kotlin Service Level Objectives Error Budgets Incident Response Observability Automation Generative AI System Architecture
Onsite (New York, NY) $250k - $500k/yr Full-time Senior Level 5 benefits + 2 perks
Posted 2 days ago

About the role

Polymarket is the world's largest prediction market platform, enabling users to trade on real-world event outcomes. This role supports the growth engineering team by building and maintaining the infrastructure layer for user acquisition, onboarding, and referral systems to ensure scalability and reliability.

Skills

Infrastructure engineering Platform engineering System reliability Data pipelines Growth engineering API management System architecture Production monitoring Incident management Scalability Performance optimization Software engineering Onboarding flows Referral systems User acquisition
Hybrid (Seattle, Washington) $320k - $485k/yr Full-time Senior Level 5 benefits + 3 perks
Posted 2 days ago

About the role

Anthropic is building reliable, interpretable, and steerable AI systems. This role on the Safeguards ML Infra team designs and operates production infrastructure to ensure safety classifiers are correctly deployed and verified across all platforms for every model launch.

Skills

Site Reliability Engineering Production Change Management Python Rust Cloud Platforms AWS GCP Incident Response Automation Configuration Management Canary Analysis Transformer Architectures LLM Inference System Architecture Operational Toil Reduction
Hybrid (Mountain View, California) $298k - $368k/yr Full-time Senior Level 3 benefits + 3 perks
Posted 2 days ago

About the role

Waymo is building the world's most trusted autonomous driver to improve mobility and save lives. As a Pipeline SRE Lead, you will drive reliability, observability, and incident management for critical release pipelines supporting a large-scale autonomous vehicle fleet.

Skills

Site Reliability Engineering C++ Distributed Systems Observability Incident Management Capacity Planning Cloud Infrastructure DevOps Machine Learning Software Architecture Automation System Reliability Disaster Recovery Data-driven Problem Solving Leadership Communication
Onsite (Sunnyvale, CA) $262k - $364k/yr Full-time Senior Level 3 benefits + 2 perks
Posted 2 days ago

About the role

Google's Site Reliability Engineering team builds and runs large-scale, fault-tolerant systems to ensure reliability and performance for its collaboration platforms. This role leads a multi-site engineering organization to drive operational excellence and architectural resilience.

Skills

Software Engineering Site Reliability Engineering Distributed Systems Automation Project Management People Management System Design Capacity Planning Performance Optimization Algorithms Complexity Analysis Mentorship Production Strategy Architectural Resilience Operational Excellence
Onsite (Milpitas, California) $214k - $309k/yr Full-time Senior Level 8 benefits + 2 perks
Posted 2 days ago

About the role

Cisco's Collaboration Business Unit seeks a Site Reliability Engineering Technical Leader to design and operate resilient large-scale distributed data platforms. You will lead cloud engineering initiatives, automate infrastructure, and drive capacity planning to support Webex's global hybrid workforce.

Skills

Cassandra Kafka OpenSearch Terraform Ansible Kubernetes AWS Python CI/CD Jenkins GitHub Actions GitLab CI PostgreSQL Cloud architecture Capacity planning Incident management
Hybrid (GM Automation - Sunnyvale - GM Automation - Sunnyvale, Texas) $172k - $300k/yr Full-time Senior Level 12 benefits + 4 perks
Posted 2 days ago

About the role

GM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property of the systems used to build, validate, release, and operate autonomous-vehicle software. As a founding SRE, you will establish the core governance layer, define pragmatic standards, and write production automation.

Skills

Site Reliability Engineering Distributed systems Go Python Java C++ Linux Networking Kubernetes Cloud infrastructure CI/CD Observability SLI/SLO definition Error budgets Incident management Production architecture
Hybrid (San Francisco, California) $203k - $275k/yr Full-time Senior Level 5 benefits + 7 perks
Posted 2 days ago

About the role

Altruist is building an AI platform for wealth management professionals. The Staff SRE role focuses on the performance and resilience of backend systems, collaborating with cross-functional teams to solve technical challenges and implement observability standards.

Skills

Site Reliability Engineering Java Spring Boot RESTful APIs Microservices architecture Distributed systems SQL Postgres MySQL Docker Kubernetes Helm Observability Open Telemetry Incident Response Automated testing
Onsite (Berkeley, Missouri) $198k - $267k/yr Full-time Senior Level 7 benefits + 2 perks
Posted 2 days ago

About the role

Boeing is seeking a Lead Site Reliability Engineer to define the reliability strategy and architecture for mission-critical developer platforms used by Air Dominance engineering teams. This role involves leading technical investigations, mentoring engineers, and driving operational excellence through automation and infrastructure as code.

Skills

Site Reliability Engineering GitLab CI/CD PostgreSQL Jira Confluence Infrastructure as Code Ansible Cloud architecture Root cause analysis Technical leadership Observability Security controls Disaster recovery Kubernetes Docker
Hybrid (Los Angeles, California) $181k - $265k/yr Full-time Senior Level 5 benefits + 7 perks
Posted 2 days ago

About the role

Altruist is building an AI platform for wealth management professionals. The Staff SRE will focus on the performance and resilience of backend systems, mentoring teams, and owning incident response processes.

Skills

Java Spring Boot RESTful APIs Microservices architecture Distributed systems SQL Postgres MySQL Docker Kubernetes Helm Observability Open Telemetry Incident response Automated testing Cloud infrastructure
Onsite (Pittsburgh, Pennsylvania) $182k - $247k/yr Full-time Senior Level 1 perks
Posted 2 days ago

About the role

Duolingo is the world's most popular language learning app with a mission to make education universally available. As a Senior Site Reliability Engineer, you will ensure the reliability and scalability of distributed systems supporting hundreds of millions of users.

Skills

Site Reliability Engineering DevOps Distributed Systems Java Kotlin Python Go Docker Kubernetes Automation Incident Response System Design Database Troubleshooting MySQL PostgreSQL DynamoDB
Hybrid (San Francisco, California) $200k - $240k/yr Full-time Mid Level 5 benefits + 7 perks
Posted 2 days ago

About the role

Altruist is building an AI platform for wealth professionals to transform the wealth management industry. The Senior Site Reliability Engineer will focus on the performance and resilience of backend systems, collaborating with cross-functional teams to deliver robust infrastructure solutions.

Skills

Java Spring Boot RESTful APIs Microservices architecture Distributed systems SQL Postgres MySQL Docker Kubernetes Helm Observability Open Telemetry Incident response Automated testing Secure coding
Onsite (Frisco, Texas) $123k - $229k/yr Full-time Senior Level 9 benefits + 2 perks
Posted 2 days ago

About the role

McAfee is a leader in personal security protecting people in an always-online world. This role leads the North American SRE team, owning reliability strategy across Cloud and Kubernetes platforms while managing incident processes and driving automation.

Skills

Site Reliability Engineering Team Leadership Incident Management Problem Management AWS GCP Kubernetes EKS GKE Python Terraform Observability Grafana SQL CloudWatch Infrastructure-as-code
A

Site Reliability Engineer

Anduril Industries

Onsite (Waltham, Massachusetts) $166k - $220k/yr Full-time Mid Level 2 benefits + 1 perks
Posted 2 days ago

About the role

Anduril Industries is a defense technology company building advanced imaging and sensor systems powered by Lattice OS. The Site Reliability Engineer joins the Imaging team to own the reliability and uptime of fielded systems, triaging issues across the full stack and creating self-service tooling to reduce recurring incidents.

Skills

Site Reliability Engineering Linux Networking Troubleshooting Python Bash Systemd Observability PagerDuty Hardware Support Sensor Calibration Nix NixOS Field Engineering DevOps
Remote (Remote - USA) $186k - $218k/yr Full-time Senior Level 5 benefits + 2 perks
Posted 2 days ago

About the role

Coinbase is seeking a Senior Infrastructure Engineer to build and operate low-latency, event-sourced trading systems for its institutional exchange. This role involves owning core infrastructure that processes high-throughput financial data, ensuring reliability and correctness for billions in daily volume.

Skills

Java C++ Go Distributed systems Low-latency Trading systems Financial infrastructure Linux performance tuning JVM Event-sourced systems Market data Order management Risk management Cloud environments Colocated environments System reliability
Hybrid (Charlotte, North Carolina) $136k - $218k/yr Full-time Senior Level 3 benefits + 4 perks
Posted 2 days ago

About the role

Capital Group seeks a Senior Platform Engineer to design and operate observability capabilities using open telemetry standards. The role focuses on building unified, vendor-flexible platforms that provide actionable insights across metrics, logs, and traces to improve system reliability and developer experience.

Skills

Observability OpenTelemetry Python Go Terraform Kubernetes AWS Prometheus Grafana Datadog SRE Cloud-native architecture CI/CD AIOps Distributed tracing Infrastructure-as-Code

Finding more jobs