Data Platform SRE, AI & Data Platforms (AiDP)
Apple
- Location
- Onsite (Austin, Texas)
- Employment
- Full-time
- Level
- Mid Level
Posted 2 weeks ago
About the Role
Join Apple’s AI & Data Platforms team to build and operate large-scale big data infrastructure supporting analytics, reporting, and AI/ML applications. This role focuses on optimizing performance, automating operations, and ensuring reliability for critical enterprise systems.
Skills
Java
Scala
Python
Go
Apache Spark
Apache Iceberg
AWS
GCP
Unix
Linux
Kubernetes
Airflow
DBT
Data Modeling
Data Warehousing
Generative AI
Full job details
Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless. Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.
AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning — including generative AI — along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.
As a Data Platform SRE, you will be responsible for developing and operating our big data platform using open source or other solutions to aid critical applications, such as analytics, reporting, and AI/ML apps. This includes working to optimize performance and cost, automate operations, and identifying and resolving production errors and issues to ensure the best data platform experience
3+ years of professional software engineering experience with large-scale big data platforms, including strong programming skills in Java, Scala, Python, or Go. Proven expertise in operating large-scale distributed data processing systems with a strong focus on Apache Spark. Hands-on experience with table formats and data lake technologies such as Apache Iceberg, ensuring scalability, reliability, and optimized query performance. Strong background in incident management, including troubleshooting, root cause analysis, and performance optimization in complex production environments. Proficient with cloud technologies such as AWS and GCP Experience with Unix/Linux systems and command-line tools for debugging and operational support.
Expertise in designing, building, and operating critical, large-scale distributed systems with a focus on low latency, fault-tolerance, and high availability. Experience with contribution to Open Source projects is a plus. Experience with multiple public cloud infrastructure, managing multi-tenant Kubernetes clusters at scale and debugging Kubernetes/Spark issues. Experience with workflow and data pipeline orchestration tools (e.g., Airflow, DBT). Understanding of data modeling and data warehousing concepts. Familiarity with the AI/ML stack, including GPUs, MLFlow, or Large Language Models (LLMs). A learning attitude to continuously improve the self, team, and the organization. Solid understanding of software engineering best practices, including the full development lifecycle, secure coding, and experience building reusable frameworks or libraries.
Description
As a Data Platform SRE, you will be responsible for developing and operating our big data platform using open source or other solutions to aid critical applications, such as analytics, reporting, and AI/ML apps. This includes working to optimize performance and cost, automate operations, and identifying and resolving production errors and issues to ensure the best data platform experience
Minimum Qualifications
3+ years of professional software engineering experience with large-scale big data platforms, including strong programming skills in Java, Scala, Python, or Go. Proven expertise in operating large-scale distributed data processing systems with a strong focus on Apache Spark. Hands-on experience with table formats and data lake technologies such as Apache Iceberg, ensuring scalability, reliability, and optimized query performance. Strong background in incident management, including troubleshooting, root cause analysis, and performance optimization in complex production environments. Proficient with cloud technologies such as AWS and GCP Experience with Unix/Linux systems and command-line tools for debugging and operational support.
Preferred Qualifications
Expertise in designing, building, and operating critical, large-scale distributed systems with a focus on low latency, fault-tolerance, and high availability. Experience with contribution to Open Source projects is a plus. Experience with multiple public cloud infrastructure, managing multi-tenant Kubernetes clusters at scale and debugging Kubernetes/Spark issues. Experience with workflow and data pipeline orchestration tools (e.g., Airflow, DBT). Understanding of data modeling and data warehousing concepts. Familiarity with the AI/ML stack, including GPUs, MLFlow, or Large Language Models (LLMs). A learning attitude to continuously improve the self, team, and the organization. Solid understanding of software engineering best practices, including the full development lifecycle, secure coding, and experience building reusable frameworks or libraries.