Skip to content
Skip to content
DevOps Jobs
Longbridge Group

Site Reliability Engineer

Longbridge Group

Location
Hybrid (New York, New York)
Employment
Full-time
Level
Senior Level
Posted 3 days ago

About the Role

Longbridge is an AI-driven online brokerage redefining the investment journey for global retail investors. As a Site Reliability Engineer, you will design, scale, and safeguard the reliability of cloud-native financial platforms while driving incident response and infrastructure automation.

Skills

Site Reliability Engineering DevOps AWS Kubernetes Docker Python Go Linux CI/CD Terraform Ansible Helm Prometheus Incident Management Infrastructure-as-code Disaster Recovery

Full job details

About Us
Longbridge is a new-generation, AI-driven online brokerage on a mission to make investing smarter, simpler, and more accessible for everyone. Headquartered in Singapore, we are redefining the investment journey by connecting the stages of "Discovery → Learning → Trading." With our proprietary AI assistant, Longbridge AI, and a cloud-native infrastructure, we provide retail investors with institutional-grade insights and a seamless global trading network. At Longbridge, you won’t just be working for a brokerage; you’ll be building the future of financial infrastructure. 


As part of our global expansion, we’re looking for a hands-on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our next-generation financial platforms. This is a high-impact role where you’ll partner closely with product and engineering teams across the globe.
  • Own system reliability: Design, implement, and operate highly available, secure distributed systems to meet strict uptime and performance targets.
  • Build automation at scale: Develop and enforce best practices in monitoring, alerting, and infrastructure-as-code (e.g., Terraform, Ansible, Helm).
  • Partner globally: Work with development teams from design through deployment, ensuring reliability and resiliency are built in from day one.
  • Lead incident response: Drive on-call processes, conduct root-cause analysis, and continuously reduce MTTR and failure recurrence.
  • Future-proof our stack: Evaluate and adopt modern cloud-native technologies (e.g., Kubernetes, Prometheus, AWS/GCP) to keep systems secure and scalable.
  • Stress-test and safeguard: Lead disaster recovery, chaos testing, and capacity planning for critical wealth management services.

What We’re Looking For
  • 5+ years of experience in SRE, DevOps, or production engineering roles.
  • Strong background in AWS (or GCP/Azure) and container orchestration (Docker, Kubernetes).
  • Proficiency in at least one programming language (Python, Go, or similar) for automation and tooling.
  • Solid Linux administration skills and experience with CI/CD pipelines.
  • Proven ability in incident management and troubleshooting distributed systems.
  • Strong collaboration and communication skills across remote/global teams.
  • Comfortable working in a fast-moving fintech/tech startup environment.
  • Bonus: Experience supporting regulated financial systems; ability to communicate in Mandarin to collaborate with Asia-based colleagues.

Why Join Us
  • Shape the reliability foundation of our U.S. product launch in wealth tech.
  • Opportunity to build systems from the ground up and influence technical direction.
  • Competitive compensation package and growth opportunities.