Site Reliability Engineering (SRE) Manager, Apple Maps
Apple
- Location
- Onsite (Cupertino, California)
- Employment
- Full-time
- Level
- Senior Level
Posted 1 week ago
About the Role
Join Apple Maps as an SRE Manager to lead the strategic direction and operational excellence of a massive distributed serving infrastructure used by hundreds of millions globally. Drive the evolution of SRE practices by integrating AI/ML solutions for enhanced reliability and scalability.
Skills
SRE Leadership
Distributed Systems
Linux Fundamentals
AI/ML Tooling
LLM Integration
Infrastructure Scaling
Networking
Capacity Planning
Performance Engineering
Kubernetes
Cloud Infrastructure
AIOps
Executive Communication
Operational Excellence
Strategic Planning
Production Engineering
Full job details
Apple Maps and location services are used by hundreds of millions of people every day to navigate the world. Behind every search, route, and point of interest is a massive distributed serving infrastructure that must be fast, reliable, and always available.
This SRE org is responsible for the availability and automation of some of the most visible and widely used services that power Apple Maps.
If you are passionate about reliability at planet scale and are excited to help us build, grow, manage and deliver infrastructure that scales with Apple Maps, this is the opportunity for you!
We are looking for a senior SRE leader to set the strategic direction for our Maps serving infrastructure. This is not just a "keep the lights on" role — this is a leadership position that defines where our infrastructure goes next, how our SRE practice evolves, and how we build the teams and partnerships to get there. You will work closely with engineering, product, and operations partners across Apple to shape the roadmap for our serving platform. You will lead an organization of SREs and hold responsibility for the reliability, scalability, and operational excellence of some of Apple's most visible services. We believe AI will fundamentally reshape how SRE is practiced — from incident detection and resolution to capacity planning and toil elimination — and we're looking for a leader who shares that conviction and can drive that transformation across the organization.
10-15+ years of experience in SRE or adjacent disciplines (systems engineering, infrastructure engineering, production engineering), with at least 10 years in senior management roles Demonstrated experience leading SRE organizations supporting large-scale, user-facing distributed services Strong technical proficiency in Linux fundamentals, distributed systems concepts, networking, and infrastructure at scale Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operations Ability to read and understand code produced by LLMs and evaluate its suitability for production use Has defined or is actively executing an AI strategy for a large SRE organization Proven ability to communicate at the executive level and negotiate across organizational boundaries
Experience with cloud infrastructure (AWS, GCP) and Kubernetes at scale Background in capacity planning, performance engineering, or infrastructure architecture Track record of driving cultural and process transformation within SRE organizations Experience building or deploying AI-powered operational tooling (AIOps, intelligent alerting, automated diagnostics) Hands-on experience with LLM-based developer/SRE productivity tools Track record of driving AI adoption within engineering teams Experience operating services at Apple-scale user volumes
Description
We are looking for a senior SRE leader to set the strategic direction for our Maps serving infrastructure. This is not just a "keep the lights on" role — this is a leadership position that defines where our infrastructure goes next, how our SRE practice evolves, and how we build the teams and partnerships to get there. You will work closely with engineering, product, and operations partners across Apple to shape the roadmap for our serving platform. You will lead an organization of SREs and hold responsibility for the reliability, scalability, and operational excellence of some of Apple's most visible services. We believe AI will fundamentally reshape how SRE is practiced — from incident detection and resolution to capacity planning and toil elimination — and we're looking for a leader who shares that conviction and can drive that transformation across the organization.
Minimum Qualifications
10-15+ years of experience in SRE or adjacent disciplines (systems engineering, infrastructure engineering, production engineering), with at least 10 years in senior management roles Demonstrated experience leading SRE organizations supporting large-scale, user-facing distributed services Strong technical proficiency in Linux fundamentals, distributed systems concepts, networking, and infrastructure at scale Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operations Ability to read and understand code produced by LLMs and evaluate its suitability for production use Has defined or is actively executing an AI strategy for a large SRE organization Proven ability to communicate at the executive level and negotiate across organizational boundaries
Preferred Qualifications
Experience with cloud infrastructure (AWS, GCP) and Kubernetes at scale Background in capacity planning, performance engineering, or infrastructure architecture Track record of driving cultural and process transformation within SRE organizations Experience building or deploying AI-powered operational tooling (AIOps, intelligent alerting, automated diagnostics) Hands-on experience with LLM-based developer/SRE productivity tools Track record of driving AI adoption within engineering teams Experience operating services at Apple-scale user volumes