Skip to content
DevOps Jobs
Thinking Machines Lab

AI Infrastructure Engineer

Thinking Machines Lab

Onsite (San Francisco, California) $350k - $475k/yr Full-time Mid Level 5 benefits + 1 perks
Posted 1 day ago

About the role

Thinking Machines Lab is building frontier AI models and human-AI communication interfaces. This role focuses on ensuring the reliability and performance of large-scale post-training and reinforcement learning systems, partnering directly with research teams to debug failures and optimize cluster utilization.

Skills

Python Go C++ Linux Distributed systems GPU TPU PyTorch Ray Kubernetes Slurm InfiniBand RDMA NCCL Site reliability engineering Infrastructure engineering
Onsite (Sunnyvale, CA) $207k - $300k/yr Full-time Senior Level 3 benefits + 2 perks
Posted 1 day ago

About the role

Google's Site Reliability Engineering team builds and runs large-scale, fault-tolerant distributed systems to ensure reliability and performance for its global services. This role leads engineers in optimizing infrastructure, automating operations, and managing on-call rotations to maintain high availability.

Skills

Software Development Distributed Systems Site Reliability Engineering Automation Artificial Intelligence Machine Learning Large Language Models System Design Complexity Analysis Algorithms Scalability Latency Optimization Performance Tuning Mentorship Team Management Infrastructure Engineering
Hybrid (Washington, District of Columbia) $174k - $267k/yr Full-time Senior Level 7 benefits + 1 perks
Posted 1 day ago

About the role

Okta is the world's identity company securing access for AI and human users. The SRE role involves architecting and managing scalable Kubernetes platforms on AWS to ensure high availability and performance for cloud-native applications.

Skills

Kubernetes AWS Terraform Helm Istio Karpenter Python Bash Go CI/CD Prometheus Grafana CloudWatch ELK Stack Docker Infrastructure as Code
Onsite (Santa Clara, California) $144k - $230k/yr Full-time Senior Level 1 perks
Posted 1 day ago

About the role

NVIDIA is seeking an AI Tools Engineer to join the SRE Data Team, building AI-powered tools and LLM-based systems to optimize GeForce NOW operations. The role focuses on automating root cause analysis and improving operational efficiency through advanced data pipelines and machine learning.

Skills

Python Kubernetes AWS LLM Machine Learning Site Reliability Engineering Data Pipelines Grafana Automation Go Root Cause Analysis AI Frameworks Cloud Computing Data Management Operational Efficiency
Onsite (United States) $85 - $105/hr Contract Senior Level
Posted 1 day ago

About the role

C-Serv is seeking a Senior Site Reliability Engineer to own reliability and design decisions for a multi-cluster AWS infrastructure within a FedRAMP Moderate-authorized environment. The role focuses on maintaining an audit-ready platform through advanced incident response, observability architecture, and compliance controls.

Skills

Site Reliability Engineering AWS Kubernetes EKS Terraform GitLab CI/CD GitOps FedRAMP Incident Response Observability Prometheus Grafana Vulnerability Management DISA STIG NIST SP 800-53 Infrastructure Architecture
Ford Motor Company

Site Reliability Engineer

Ford Motor Company

Onsite (United States) $85k - $192k/yr Full-time Mid Level 5 benefits + 4 perks
Posted 1 day ago

About the role

Ford Motor Company is seeking a Site Reliability Engineer to develop and expand its global monitoring and observability platform. This role blends software and systems engineering to ensure the uptime, scalability, and maintainability of critical cloud services.

Skills

Site Reliability Engineering Software Engineering DevOps Golang Scripting Monitoring Observability OpenTelemetry Dynatrace Kubernetes Google Cloud Platform Relational databases Document databases Debugging Code optimization Automation
Onsite (USA - Remote, Colorado) $136k - $181k/yr Full-time Senior Level 4 benefits + 2 perks
Posted 1 day ago

About the role

Ping Identity is a global enterprise security company specializing in identity and access management. This role involves designing and maintaining cloud-based infrastructure and optimizing CI/CD pipelines to ensure the reliability and scalability of their mission-critical platform.

Skills

Site Reliability Engineering Go Docker Kubernetes CI/CD Cloud Platforms Distributed Systems Infrastructure as Code Automation Observability Networking Identity and Access Management Git Security Deployment Automation
Onsite (United States) $70 - $80/hr Contract Mid Level
Posted 1 day ago

About the role

C-Serv is partnering with a client to staff a Site Reliability Engineer for a 24/7 FedRAMP-authorized cloud operations team. The role focuses on real-time monitoring, log-level root cause investigation, and deployment execution within a high-stakes AWS GovCloud environment.

Skills

Site Reliability Engineering DevOps Linux Administration Bash Python Kubernetes EKS AWS Prometheus Grafana Kibana Elasticsearch GitLab CI/CD ArgoCD STIG Hardening FedRAMP
Hybrid (Tampa, FL) $75k - $150k/yr Full-time Senior Level 6 benefits + 2 perks
Posted 1 day ago

About the role

DTCC is the premier post-trade market infrastructure for the global financial services industry. This role ensures the stability and reliability of mission-critical applications by applying Site Reliability Engineering principles to drive operational excellence.

Skills

Site Reliability Engineering Linux Windows Python Bash Shell Scripting Dynatrace Splunk Grafana SQL ServiceNow Jira Kafka Autosys AWS OpenShift
Trimble Inc.
Hybrid (Lake Oswego, Oregon) $105k - $145k/yr Full-time Senior Level 8 benefits + 3 perks
Posted 1 day ago

About the role

Trimble is seeking a Site Reliability Engineer to support and optimize public cloud infrastructure within AWS GovCloud and Azure environments. This role focuses on enhancing security posture, operational excellence, and architectural scalability for critical FedRAMP PaaS solutions.

Skills

AWS Azure Python PowerShell Bash Perl System Administration FedRAMP Infrastructure-as-code Ansible Terraform Kubernetes CI/CD Git Jira Jenkins
Trimble Inc.
Hybrid (Lake Oswego, Oregon) $91k - $125k/yr Full-time Mid Level 8 benefits + 1 perks
Posted 1 day ago

About the role

Trimble is a global technology company connecting physical and digital worlds to transform industries like construction and transportation. This role involves shaping and securing their FedRAMP PaaS portfolio while driving system resilience and scalability.

Skills

AWS Azure Python PowerShell Bash Perl System Administration Networking Storage Automation Infrastructure-as-code Ansible Terraform Kubernetes FedRAMP Cloud Security
Onsite (Charlotte NC - 2320 Cascade Pointe Boulevard, North Carolina) Full-time Senior Level 11 benefits
Posted 1 day ago

About the role

Truist is seeking an Observability Lead to define and scale enterprise observability capabilities, transitioning the organization toward proactive, intelligence-driven monitoring. This role combines technical leadership with hands-on engineering to improve system reliability and reduce resolution times across complex infrastructure.

Skills

Observability Infrastructure engineering Software development Kafka Event streaming OpenTelemetry Technical leadership SRE Cloud platforms Kubernetes Python Go Java CI/CD Automation Performance engineering
Onsite (SSC Irving TX, Texas) Full-time Mid Level
Posted 1 day ago

About the role

7-Eleven is the world's largest convenience retailer, revolutionizing the industry through innovation and customer-centric services. The SRE RunOps Engineer role ensures the reliability and performance of the 7NOW delivery platform, combining software engineering with operations expertise to maintain high-availability systems.

Skills

Python Go Bash PowerShell AWS Azure Terraform Ansible Prometheus Grafana Datadog CI/CD Linux Kubernetes SQL NoSQL
Remote (St. Louis, MO) Full-time Mid Level 4 benefits + 2 perks
Posted 1 day ago

About the role

Enterprise Mobility is a leading global mobility provider managing major car rental brands. This role ensures the availability and performance of mission-critical rental systems through Site Reliability Engineering, bridging development and operations to support high-availability infrastructure.

Skills

Site Reliability Engineering Java Web Services C/C++ PL/SQL Tuxedo WebLogic TomCat AIX Linux Shell Bash Python Splunk Dynatrace Capacity planning
Onsite (USA-GA-Alpharetta-1500BluegrassLakesPkwy, Georgia) Full-time Senior Level
Posted 1 day ago

About the role

Scientific Games is a global leader in lottery games, sports betting, and technology. This role enhances the stability, performance, and reliability of production systems through monitoring, observability, and automation.

Skills

Site Reliability Engineering AWS Kubernetes Terraform Python Bash New Relic Graylog HashiCorp Vault CI/CD GitHub Actions GitLab CI/CD Helm ArgoCD Incident Management Observability
Onsite (Barrington, Rhode Island) Contract Senior Level
Posted 1 day ago

About the role

Workiy is seeking a Site Reliability Engineer to design, deploy, and maintain highly available, scalable, and secure production systems. The role focuses on improving system reliability, automating operational processes, and implementing robust observability solutions in collaboration with development teams.

Skills

Site Reliability Engineering Go Python Java Rust AWS Azure GCP Kubernetes Docker OpenTelemetry Prometheus Grafana Terraform Ansible CI/CD
Thinking Machines Lab

Site Reliability Engineer, RL Infra

Thinking Machines Lab

Onsite (San Francisco, California) $350k - $475k/yr Full-time Senior Level 6 benefits
Posted 2 days ago

About the role

Thinking Machines Lab is building frontier AI models and human-AI interfaces. This role drives end-to-end reliability for Reinforcement Learning infrastructure, ensuring robust performance for concurrent multi-tenant workloads across distributed training systems.

Skills

Site reliability engineering Distributed systems Cloud infrastructure Kubernetes GPU workload management Incident response Observability CI/CD Reinforcement learning Python System architecture Automation Performance tuning Multi-tenant isolation Checkpointing Data synchronization
Onsite (San Francisco, California) $350k - $475k/yr Full-time Senior Level 6 benefits
Posted 2 days ago

About the role

Thinking Machines Lab is building frontier AI models and tools to extend human will and judgment. This role drives end-to-end reliability for Model Post Training workloads, including supervised fine-tuning and reinforcement learning, ensuring platform stability and rapid incident response.

Skills

Site reliability engineering Distributed systems Cloud infrastructure Machine learning systems Python Go Rust Kubernetes PyTorch FSDP DDP Megatron DeepSpeed Incident response Observability Distributed training
Onsite (Santa Clara, CALIFORNIA) $190k - $334k/yr Full-time Senior Level 4 benefits + 4 perks
Posted 2 days ago

About the role

ServiceNow is seeking a Senior Staff Reliability Engineer to drive infrastructure automation and operational resilience across its hybrid cloud platform. This role focuses on designing auto-remediation systems and Kubernetes architectures to eliminate toil and maintain high availability for global engineering teams.

Skills

Kubernetes SRE AIOps Infrastructure-as-Code GitOps AWS Azure GCP Python Go Bash Linux Observability Incident management Distributed systems Automation
Onsite (San Jose, CA) $207k - $300k/yr Full-time Senior Level 3 benefits + 2 perks
Posted 2 days ago

About the role

Join Google's Site Reliability Engineering team to build and run large-scale, fault-tolerant systems. You will manage the full service lifecycle, ensuring reliability, uptime, and performance while optimizing infrastructure through automation and system design.

Skills

Software Development Distributed Systems Project Leadership Capacity Planning Automation System Design Troubleshooting Algorithms Complexity Analysis Incident Response Monitoring Latency Analysis Availability Management Infrastructure Engineering Performance Optimization

Finding more jobs