Skip to content
DevOps Jobs
Hybrid (Santa Clara, California) $168k - $333k/yr Full-time Senior Level 3 benefits + 2 perks
Posted 2 days ago

About the role

NVIDIA is seeking a Senior Storage Platform Engineer to design, deploy, and operate multi-vendor storage platforms powering its EDA FARM and AI/ML teams. The role focuses on building automation, CI/CD pipelines, and self-service infrastructure to support large-scale engineering workflows.

Skills

NetApp ONTAP Pure Storage Cloudian DDN Ansible Terraform Python Go CI/CD Infrastructure as Code NFS S3 Lustre Kubernetes GitOps Observability
Hybrid (San Francisco-HQ, California) $160k - $180k/yr Full-time Mid Level 3 benefits + 3 perks
Posted 2 days ago

About the role

Forage is a mission-driven payments company empowering merchants to serve underserved communities by enabling government benefit acceptance. The Infrastructure Engineer role focuses on owning critical compute and data services, evolving CI/CD pipelines, and enhancing developer experience to support rapid, reliable growth.

Skills

Infrastructure engineering DevOps SRE AWS Terraform CDK Pulumi CloudFormation CI/CD Monitoring Autoscaling Security Containerized deployments Database management Incident response Load testing
Crunchyroll, LLC
Hybrid (Los Angeles, California) $210k - $263k/yr Full-time Senior Level 12 benefits + 3 perks
Posted 3 days ago

About the role

Crunchyroll is hiring a Staff Site Reliability Engineer to lead reliability, scalability, and security initiatives for its consumer-facing data platforms. This senior role involves driving operational excellence through automation, observability, and cross-functional collaboration within the Center for Data & Insights.

Skills

Site Reliability Engineering Kubernetes Google Cloud Platform Terraform Infrastructure as Code Linux Go Python Java Prometheus Grafana OpenTelemetry Datadog Incident Management Capacity Planning Security Operations
Hybrid (San Mateo, California) $196k - $243k/yr Full-time Senior Level 2 benefits + 2 perks
Posted 3 days ago

About the role

Roblox is building the tools and platform that empower a global community of developers to create 3D immersive digital experiences. The Infrastructure Compute SRE team owns the underlying cell infrastructure, private cloud, and Kubernetes-based systems to ensure production readiness and platform stability.

Skills

Site Reliability Engineering Go Java C# Kubernetes Nomad Vault Consul Cloud infrastructure System observability Automation Load testing Capacity planning Fault-tolerance Software development
Hybrid (Irvine, California) $119k - $200k/yr Full-time Senior Level 8 benefits + 3 perks
Posted 3 days ago

About the role

Panasonic Avionics Corporation is a global leader in inflight entertainment and communications systems, serving over 300 airline customers. This role leads the design and lifecycle management of virtualized infrastructure, automation frameworks, and shared network services at scale.

Skills

Virtualized infrastructure Automation frameworks Network services VMware Hybrid cloud AWS DNS Load balancing PKI Observability platform Infrastructure-as-code Data center design Network architecture Linux administration Storage architecture Disaster recovery
Hybrid (Santa Clara, California) $248k - $396k/yr Full-time Senior Level 1 benefits + 2 perks
Posted 4 days ago

About the role

NVIDIA is seeking a Principal Site Reliability Engineer to define the technical vision and architecture for reliability across its AI Platform Runtime. This role involves leading intelligent automation initiatives and establishing standards for scalable, resilient distributed systems.

Skills

Site Reliability Engineering Distributed Systems Kubernetes Cloud Architecture Python Go Terraform OpenTelemetry Infrastructure-as-code Capacity Management Incident Management AI/ML Platforms Automation Observability Linux Networking
Hybrid (Santa Clara, California) $168k - $333k/yr Full-time Senior Level 2 benefits + 2 perks
Posted 4 days ago

About the role

NVIDIA is seeking a Staff Site Reliability Engineer to lead technical strategy for large-scale SRE initiatives across its AI Platform Runtime. The role involves designing resilient distributed systems, driving automation, and enhancing observability to support next-generation AI-driven enterprise products.

Skills

Site Reliability Engineering Kubernetes Python Typescript JavaScript Go AWS Azure GCP Infrastructure-as-code Terraform OpenTelemetry System Architecture Automation Observability Cloud Computing
Ford Motor Company
Hybrid (Long Beach, CA) $150k - $283k/yr Full-time Senior Level 7 benefits + 5 perks
Posted 4 days ago

About the role

Ford Motor Company is seeking a Staff Embedded Platform Engineer to architect and develop foundational software platforms, including real-time drivers and middleware, for next-generation automotive electronic control units. This role ensures functional safety and reliability across vehicle programs.

Skills

Embedded C MISRA C ISO 26262 RTOS CAN LIN Automotive Ethernet SPI I2C Embedded systems Functional safety JTAG GDB Hardware-software interface Technical leadership Mentoring
Hybrid (Sunnyvale, California) $143k - $243k/yr Full-time Senior Level 3 benefits + 2 perks
Posted 4 days ago

About the role

Intuitive is a global leader in robotic-assisted surgery, developing systems like the da Vinci surgical platform to transform patient care. This role supports the embedded software build environment for the DaVinci Single Port robot, ensuring robust infrastructure for surgical innovation.

Skills

Embedded build infrastructure Bazel CMake Python Ansible Jenkins Git Linux Windows CI/CD Infrastructure as code C++ Bash ElasticSearch TCP/IP Debugging
Hybrid (San Jose, California) $60 - $90/hr Contract Mid Level 2 perks
Posted 4 days ago

About the role

C-Serv is seeking a Contract Site Reliability Engineer II to manage real-time monitoring, triage, and root cause investigation within a FedRAMP-authorized AWS and EKS environment. This role offers significant exposure to GovCloud operations and serves as a stepping stone toward senior SRE positions.

Skills

Site Reliability Engineering AWS GovCloud Kubernetes EKS Linux Bash Python Kibana Elasticsearch GitLab CI/CD ArgoCD Prometheus Grafana ChatOps FedRAMP
Hybrid (US - CA - Foster City, California) $20 - $55/hr Full-time Entry Level 2 benefits + 2 perks
Posted 4 days ago

About the role

Gilead Sciences is a global biotechnology company dedicated to creating a healthier world by developing therapies for HIV, hepatitis, cancer, and other major health challenges. This internship offers a unique opportunity to support business strategy projects within the Development organization, driving performance improvements and change management initiatives.

Skills

Business strategy DevOps Project management Change management Analytical skills Strategic thinking Problem-solving Communication Microsoft Office Smartsheet Data analysis Collaboration Detail-oriented Process improvement
Hybrid (San Francisco, California) $45 - $50/hr Intern Entry Level 3 perks
Posted 4 days ago

About the role

Gallup is seeking a Site Reliability Engineer Intern to strengthen the performance and reliability of its global technology platform. The role involves building observability systems, creating dashboards, and supporting incident response alongside experienced engineers.

Skills

AWS Python Bash PowerShell Terraform Docker Amazon ECS GitHub Dynatrace PagerDuty Observability Automation Incident response Networking Linux Windows
Hybrid (San Francisco, California) $190k - $235k/yr Full-time Senior Level 9 benefits + 3 perks
Posted 5 days ago

About the role

IREN is a vertically integrated AI Cloud provider delivering large-scale data centers and GPU clusters for AI training and inference using 100% renewable energy. The Azure DevOps Lead Engineer will own the technical direction, architecture, and governance of the company's Azure platform and DevOps practices.

Skills

Azure DevOps Infrastructure as code Terraform Bicep CI/CD Entra ID PowerShell Python Azure networking Security governance SRE practices Observability SOX compliance ISO 27001 Cloud architecture
Hybrid (Mountain View, California) $298k - $368k/yr Full-time Senior Level 3 benefits + 3 perks
Posted 1 week ago

About the role

Waymo is building the world's most trusted autonomous driver to improve mobility and save lives. As a Pipeline SRE Lead, you will drive reliability, observability, and incident management for critical release pipelines supporting a large-scale autonomous vehicle fleet.

Skills

Site Reliability Engineering C++ Distributed Systems Observability Incident Management Capacity Planning Cloud Infrastructure DevOps Machine Learning Software Architecture Automation System Reliability Disaster Recovery Data-driven Problem Solving Leadership Communication
Hybrid (USA-CA-Office-Mountain View, California) $205k - $342k/yr Full-time Senior Level 4 benefits + 2 perks
Posted 1 week ago

About the role

Omnissa is a leader in digital work platforms, empowering organizations with secure, seamless access from anywhere. As a Staff DevOps Engineer, you will advance the reliability, scalability, and automation of the Horizon Cloud platform through global infrastructure management and CI/CD pipeline optimization.

Skills

Python CI/CD Jenkins GitHub Actions ArgoCD Kubernetes Helm Azure Terraform Elastic Stack Prometheus Grafana SRE Infrastructure as Code GitOps
Hybrid (Long Beach, California) $205k - $310k/yr Full-time Senior Level 8 benefits + 2 perks
Posted 1 week ago

About the role

True Anomaly is a defense technology company building autonomous spacecraft and mission software for space superiority. This role involves building security tooling and controls to protect multi-cloud infrastructure across Azure and AWS, enabling safe engineering operations in a mission-critical environment.

Skills

Cloud Security Azure AWS Python Go Terraform Kubernetes PKI HashiCorp Vault IAM Network Security Infrastructure-as-code DevSecOps Threat Detection Incident Response CSPM
Hybrid (Foster City, California) $230k - $277k/yr Full-time Senior Level 5 benefits + 3 perks
Posted 1 week ago

About the role

Zoox is developing the first ground-up, fully autonomous vehicle fleet to reinvent personal transportation. This role involves architecting and scaling a high-throughput data fabric to unify enterprise and AI data foundations, driving performance tuning and governance for AI-ready infrastructure.

Skills

Data Platform Engineering Software Engineering Data Governance CI/CD Pipelines Cloud Optimization Databricks AWS EMR RBAC Data Architecture Infrastructure-as-Code Terraform Kubernetes EKS Performance Tuning Metadata Management Schema Evolution
Hybrid (San Francisco, California) $203k - $275k/yr Full-time Senior Level 5 benefits + 7 perks
Posted 1 week ago

About the role

Altruist is building an AI platform for wealth management professionals. The Staff SRE role focuses on the performance and resilience of backend systems, collaborating with cross-functional teams to solve technical challenges and implement observability standards.

Skills

Site Reliability Engineering Java Spring Boot RESTful APIs Microservices architecture Distributed systems SQL Postgres MySQL Docker Kubernetes Helm Observability Open Telemetry Incident Response Automated testing
Hybrid (Los Angeles, California) $181k - $265k/yr Full-time Senior Level 5 benefits + 7 perks
Posted 1 week ago

About the role

Altruist is building an AI platform for wealth management professionals. The Staff SRE will focus on the performance and resilience of backend systems, mentoring teams, and owning incident response processes.

Skills

Java Spring Boot RESTful APIs Microservices architecture Distributed systems SQL Postgres MySQL Docker Kubernetes Helm Observability Open Telemetry Incident response Automated testing Cloud infrastructure
Hybrid (San Francisco, California) $180k - $260k/yr Full-time Senior Level 7 benefits + 7 perks
Posted 1 week ago

About the role

Sprig is an AI-powered customer survey platform trusted by leading tech companies. As a Senior Platform Engineer, you will own the build-and-test loop for a large monorepo, optimizing CI/CD processes and developer productivity to support both human and AI agent workflows.

Skills

Bazel TypeScript Node.js Go CI/CD Monorepo Build Systems Test Automation React Kubernetes Cloud Infrastructure Telemetry System Architecture Software Engineering Performance Optimization

Finding more jobs