Origami Risk provides integrated SaaS solutions for risk, compliance, and insurance ecosystems. The Site Reliability Engineer improves system scalability, stability, and incident resolution through post-incident investigations and observability tool management.
Skills
Site Reliability EngineeringIncident ManagementObservability ToolsNew RelicData DogSumoLogicJavaScript.NETSQLAWSAzureC#SaaS OperationsCI/CDInfrastructure as CodeTroubleshooting
Arctiq is a global technology services firm seeking a Site Reliability Engineer to support a client's vulnerability remediation program. The role focuses on triaging and remediating infrastructure and configuration vulnerabilities across cloud and on-premise environments using automated solutions.
Wayve is a leading developer of Embodied AI technology for autonomous driving. This role leads the Japan Fleet Reliability SRE team, ensuring the safety, availability, and reliability of autonomous vehicle operations while collaborating with global engineering teams.
Wayve is building a global driving intelligence for autonomous vehicles. This role leads the Japan Fleet Reliability SRE team to ensure the safety, reliability, and availability of the autonomous vehicle fleet through engineering management and operational excellence.
Indeed is the world's leading job site, helping people get jobs through AI and real-time data. This role partners with product teams to design, code, and maintain scalable, high-performance systems while driving modern reliability practices and operational excellence.
Skills
Site Reliability EngineeringCloud InfrastructureKubernetesArgoCDTerraformGitLabDatadogJavaKotlinService Level ObjectivesError BudgetsIncident ResponseObservabilityAutomationGenerative AISystem Architecture
Onsite (New York, NY)
$250k - $500k/yr
Full-time
Senior Level
5 benefits
+ 2 perks
Posted 2 days ago
About the role
Polymarket is the world's largest prediction market platform, enabling users to trade on real-world event outcomes. This role supports the growth engineering team by building and maintaining the infrastructure layer for user acquisition, onboarding, and referral systems to ensure scalability and reliability.
Anthropic is building reliable, interpretable, and steerable AI systems. This role on the Safeguards ML Infra team designs and operates production infrastructure to ensure safety classifiers are correctly deployed and verified across all platforms for every model launch.
Waymo is building the world's most trusted autonomous driver to improve mobility and save lives. As a Pipeline SRE Lead, you will drive reliability, observability, and incident management for critical release pipelines supporting a large-scale autonomous vehicle fleet.
Skills
Site Reliability EngineeringC++Distributed SystemsObservabilityIncident ManagementCapacity PlanningCloud InfrastructureDevOpsMachine LearningSoftware ArchitectureAutomationSystem ReliabilityDisaster RecoveryData-driven Problem SolvingLeadershipCommunication
Google's Site Reliability Engineering team builds and runs large-scale, fault-tolerant systems to ensure reliability and performance for its collaboration platforms. This role leads a multi-site engineering organization to drive operational excellence and architectural resilience.
Cisco's Collaboration Business Unit seeks a Site Reliability Engineering Technical Leader to design and operate resilient large-scale distributed data platforms. You will lead cloud engineering initiatives, automate infrastructure, and drive capacity planning to support Webex's global hybrid workforce.
GM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property of the systems used to build, validate, release, and operate autonomous-vehicle software. As a founding SRE, you will establish the core governance layer, define pragmatic standards, and write production automation.
Skills
Site Reliability EngineeringDistributed systemsGoPythonJavaC++LinuxNetworkingKubernetesCloud infrastructureCI/CDObservabilitySLI/SLO definitionError budgetsIncident managementProduction architecture
Altruist is building an AI platform for wealth management professionals. The Staff SRE role focuses on the performance and resilience of backend systems, collaborating with cross-functional teams to solve technical challenges and implement observability standards.
Skills
Site Reliability EngineeringJavaSpring BootRESTful APIsMicroservices architectureDistributed systemsSQLPostgresMySQLDockerKubernetesHelmObservabilityOpen TelemetryIncident ResponseAutomated testing
Boeing is seeking a Lead Site Reliability Engineer to define the reliability strategy and architecture for mission-critical developer platforms used by Air Dominance engineering teams. This role involves leading technical investigations, mentoring engineers, and driving operational excellence through automation and infrastructure as code.
Skills
Site Reliability EngineeringGitLabCI/CDPostgreSQLJiraConfluenceInfrastructure as CodeAnsibleCloud architectureRoot cause analysisTechnical leadershipObservabilitySecurity controlsDisaster recoveryKubernetesDocker
Altruist is building an AI platform for wealth management professionals. The Staff SRE will focus on the performance and resilience of backend systems, mentoring teams, and owning incident response processes.
Duolingo is the world's most popular language learning app with a mission to make education universally available. As a Senior Site Reliability Engineer, you will ensure the reliability and scalability of distributed systems supporting hundreds of millions of users.
Skills
Site Reliability EngineeringDevOpsDistributed SystemsJavaKotlinPythonGoDockerKubernetesAutomationIncident ResponseSystem DesignDatabase TroubleshootingMySQLPostgreSQLDynamoDB
Altruist is building an AI platform for wealth professionals to transform the wealth management industry. The Senior Site Reliability Engineer will focus on the performance and resilience of backend systems, collaborating with cross-functional teams to deliver robust infrastructure solutions.
McAfee is a leader in personal security protecting people in an always-online world. This role leads the North American SRE team, owning reliability strategy across Cloud and Kubernetes platforms while managing incident processes and driving automation.
Skills
Site Reliability EngineeringTeam LeadershipIncident ManagementProblem ManagementAWSGCPKubernetesEKSGKEPythonTerraformObservabilityGrafanaSQLCloudWatchInfrastructure-as-code
Anduril Industries is a defense technology company building advanced imaging and sensor systems powered by Lattice OS. The Site Reliability Engineer joins the Imaging team to own the reliability and uptime of fielded systems, triaging issues across the full stack and creating self-service tooling to reduce recurring incidents.
Skills
Site Reliability EngineeringLinuxNetworkingTroubleshootingPythonBashSystemdObservabilityPagerDutyHardware SupportSensor CalibrationNixNixOSField EngineeringDevOps
Coinbase is seeking a Senior Infrastructure Engineer to build and operate low-latency, event-sourced trading systems for its institutional exchange. This role involves owning core infrastructure that processes high-throughput financial data, ensuring reliability and correctness for billions in daily volume.
Capital Group seeks a Senior Platform Engineer to design and operate observability capabilities using open telemetry standards. The role focuses on building unified, vendor-flexible platforms that provide actionable insights across metrics, logs, and traces to improve system reliability and developer experience.
Origami Risk provides integrated SaaS solutions for risk, compliance, and insurance ecosystems. The Site Reliability Engineer improves system scalability, stability, and incident resolution through post-incident investigations and observability tool management.
Skills
Site Reliability EngineeringIncident ManagementObservability ToolsNew RelicData DogSumoLogicJavaScript.NETSQLAWSAzureC#SaaS OperationsCI/CDInfrastructure as CodeTroubleshooting
Arctiq is a global technology services firm seeking a Site Reliability Engineer to support a client's vulnerability remediation program. The role focuses on triaging and remediating infrastructure and configuration vulnerabilities across cloud and on-premise environments using automated solutions.
Wayve is a leading developer of Embodied AI technology for autonomous driving. This role leads the Japan Fleet Reliability SRE team, ensuring the safety, availability, and reliability of autonomous vehicle operations while collaborating with global engineering teams.
Wayve is building a global driving intelligence for autonomous vehicles. This role leads the Japan Fleet Reliability SRE team to ensure the safety, reliability, and availability of the autonomous vehicle fleet through engineering management and operational excellence.
Indeed is the world's leading job site, helping people get jobs through AI and real-time data. This role partners with product teams to design, code, and maintain scalable, high-performance systems while driving modern reliability practices and operational excellence.
Skills
Site Reliability EngineeringCloud InfrastructureKubernetesArgoCDTerraformGitLabDatadogJavaKotlinService Level ObjectivesError BudgetsIncident ResponseObservabilityAutomationGenerative AISystem Architecture
Onsite (New York, NY)
$250k - $500k/yr
Full-time
Senior Level
5 benefits
+ 2 perks
Posted 2 days ago
About the role
Polymarket is the world's largest prediction market platform, enabling users to trade on real-world event outcomes. This role supports the growth engineering team by building and maintaining the infrastructure layer for user acquisition, onboarding, and referral systems to ensure scalability and reliability.
Anthropic is building reliable, interpretable, and steerable AI systems. This role on the Safeguards ML Infra team designs and operates production infrastructure to ensure safety classifiers are correctly deployed and verified across all platforms for every model launch.
Waymo is building the world's most trusted autonomous driver to improve mobility and save lives. As a Pipeline SRE Lead, you will drive reliability, observability, and incident management for critical release pipelines supporting a large-scale autonomous vehicle fleet.
Skills
Site Reliability EngineeringC++Distributed SystemsObservabilityIncident ManagementCapacity PlanningCloud InfrastructureDevOpsMachine LearningSoftware ArchitectureAutomationSystem ReliabilityDisaster RecoveryData-driven Problem SolvingLeadershipCommunication
Google's Site Reliability Engineering team builds and runs large-scale, fault-tolerant systems to ensure reliability and performance for its collaboration platforms. This role leads a multi-site engineering organization to drive operational excellence and architectural resilience.
Cisco's Collaboration Business Unit seeks a Site Reliability Engineering Technical Leader to design and operate resilient large-scale distributed data platforms. You will lead cloud engineering initiatives, automate infrastructure, and drive capacity planning to support Webex's global hybrid workforce.
GM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property of the systems used to build, validate, release, and operate autonomous-vehicle software. As a founding SRE, you will establish the core governance layer, define pragmatic standards, and write production automation.
Skills
Site Reliability EngineeringDistributed systemsGoPythonJavaC++LinuxNetworkingKubernetesCloud infrastructureCI/CDObservabilitySLI/SLO definitionError budgetsIncident managementProduction architecture
Altruist is building an AI platform for wealth management professionals. The Staff SRE role focuses on the performance and resilience of backend systems, collaborating with cross-functional teams to solve technical challenges and implement observability standards.
Skills
Site Reliability EngineeringJavaSpring BootRESTful APIsMicroservices architectureDistributed systemsSQLPostgresMySQLDockerKubernetesHelmObservabilityOpen TelemetryIncident ResponseAutomated testing
Boeing is seeking a Lead Site Reliability Engineer to define the reliability strategy and architecture for mission-critical developer platforms used by Air Dominance engineering teams. This role involves leading technical investigations, mentoring engineers, and driving operational excellence through automation and infrastructure as code.
Skills
Site Reliability EngineeringGitLabCI/CDPostgreSQLJiraConfluenceInfrastructure as CodeAnsibleCloud architectureRoot cause analysisTechnical leadershipObservabilitySecurity controlsDisaster recoveryKubernetesDocker
Altruist is building an AI platform for wealth management professionals. The Staff SRE will focus on the performance and resilience of backend systems, mentoring teams, and owning incident response processes.
Duolingo is the world's most popular language learning app with a mission to make education universally available. As a Senior Site Reliability Engineer, you will ensure the reliability and scalability of distributed systems supporting hundreds of millions of users.
Skills
Site Reliability EngineeringDevOpsDistributed SystemsJavaKotlinPythonGoDockerKubernetesAutomationIncident ResponseSystem DesignDatabase TroubleshootingMySQLPostgreSQLDynamoDB
Altruist is building an AI platform for wealth professionals to transform the wealth management industry. The Senior Site Reliability Engineer will focus on the performance and resilience of backend systems, collaborating with cross-functional teams to deliver robust infrastructure solutions.
McAfee is a leader in personal security protecting people in an always-online world. This role leads the North American SRE team, owning reliability strategy across Cloud and Kubernetes platforms while managing incident processes and driving automation.
Skills
Site Reliability EngineeringTeam LeadershipIncident ManagementProblem ManagementAWSGCPKubernetesEKSGKEPythonTerraformObservabilityGrafanaSQLCloudWatchInfrastructure-as-code
Anduril Industries is a defense technology company building advanced imaging and sensor systems powered by Lattice OS. The Site Reliability Engineer joins the Imaging team to own the reliability and uptime of fielded systems, triaging issues across the full stack and creating self-service tooling to reduce recurring incidents.
Skills
Site Reliability EngineeringLinuxNetworkingTroubleshootingPythonBashSystemdObservabilityPagerDutyHardware SupportSensor CalibrationNixNixOSField EngineeringDevOps
Coinbase is seeking a Senior Infrastructure Engineer to build and operate low-latency, event-sourced trading systems for its institutional exchange. This role involves owning core infrastructure that processes high-throughput financial data, ensuring reliability and correctness for billions in daily volume.
Capital Group seeks a Senior Platform Engineer to design and operate observability capabilities using open telemetry standards. The role focuses on building unified, vendor-flexible platforms that provide actionable insights across metrics, logs, and traces to improve system reliability and developer experience.