Staff Site Reliability Engineer, AQI Data SRE
- Location
- Onsite (Pittsburgh, PA)
- Compensation
- $207k - $300k/yr
- Employment
- Full-time
- Level
- Senior Level
About the Role
Join the Ads Quality Infrastructure Data SRE team to ensure reliability and performance for Google Ads data processing across Search, Maps, and SAGE/SPARK platforms. You will define architectural roadmaps, manage complex incidents, and drive automation to support massive-scale distributed systems.
Skills
Benefits
- Health insurance
Perks
- Bonus target
- Equity
Full job details
Minimum qualifications:
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 8 years of experience with software development in one or more programming languages.
- 3 years of experience leading projects.
- 3 years of experience designing, analyzing, and troubleshooting distributed systems.
- Experience with large-scale data processing.
Preferred qualifications:
- Master's degree in Computer Science or Engineering.
- Experience in distributed, real-time systems and high-throughput data pipelines.
- Knowledge or experience with agentic flows/development and with AI-assisted production/reliability tooling.
- Proven track record of driving cross-functional architectural alignment and influencing technical decisions across large engineering organizations.
- Excellent collaboration skills, with a track record of building trusted, blameless partnerships with development teams.
- Excellent non-abstract large systems design skills, with experience building scalable, self-service reliability and observability.
About the job:
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
Ads Quality Infrastructure (AQI) Data Site Reliability Engineering (SRE) keeps Search Ads and Ads on Google Experiences (SAGE/SPARK) data processing reliable, so those teams can deliver useful ads that create advertiser value. We own the core infrastructure behind the Google Ads business across Google-owned surfaces, including Google Search, Google Maps, and Feed Ads, and have a portion of Google business.
AQI SRE owns: serving, data extraction and copy, and the platforms that keep those systems fast and healthy. We partner with development teams to co-build and support this infrastructure.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Define the technical goal, architectural roadmap, and reliability strategy for AQI Data SRE across the SAGE/SPARK production stack.
- Partner with dev leadership to lead system design reviews, drive backend simplification and resource isolation, and ensure production readiness for flagship Ads launches.
- Architect and lead the implementation of critical infrastructure projects, including data-pipeline resilience, automated rollback/restart platforms, and consolidated observability.
- Mentor and grow engineers on the team, and push for software engineering excellence, AI-first tooling, and SRE best practices; drive the operational maturity handover that helps partner dev teams mature.
- Lead the response to complex production incidents, correlate production signals with customer and business impact, and drive high-impact post-incident architectural improvements and blameless postmortems and participate in a healthy tier 2 oncall rotation for core shared SAGE/SPARK infrastructure.