Skip to content
Skip to content
DevOps Jobs
Anza Mortgage Insurance Corporation

Lead DevOps (Site Reliability Engineer)

Anza Mortgage Insurance Corporation

Location
Onsite (McLean, Virginia · McLean, Virginia)
Employment
Full-time
Level
Senior Level
Posted 1 week ago

About the Role

Anza Mortgage Insurance Corporation is a fintech startup revolutionizing the US mortgage market through technology and analytics. The Lead SRE role involves designing scalable AWS infrastructure, leading an engineering team, and ensuring system reliability and security.

Skills

AWS Terraform Terragrunt Kubernetes Argo CD Argo Workflows GitHub Actions CI/CD pipelines Docker Infrastructure as Code Monitoring Incident management Disaster recovery Cloud infrastructure Root cause analysis Scripting

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
  • 401(k)
  • PTO

Perks

  • Performance bonuses
  • Team events
  • Wellness programs

Full job details

About the role

We are seeking a highly skilled and motivated Lead Site Reliability Engineer (SRE) to join our team. The Lead SRE will play a critical role in designing, implementing, and maintaining the reliability, scalability, and performance of our cloud-based systems hosted in AWS. You will collaborate closely with software engineers, operations teams, and other stakeholders to enhance system reliability and developer productivity through automation, monitoring, and incident response


What you'll do

  • Lead a team of SREs consisting of FTEs and contractors
  • Define and assign the tasks to SREs, review the PRs and provide the feedback
  • Design and implement scalable, reliable, and secure cloud infrastructure in AWS.
  • Develop and maintain monitoring, alerting, and dashboarding solutions to ensure system health and uptime.
  • Automate infrastructure provisioning and configuration management using tools like Terraform and Terragrunt
  • Implement CI/CD pipelines to streamline deployments and improve development workflows.
  • Respond to incidents, perform root cause analysis, and implement permanent fixes to prevent recurring issues.
  • Optimize system performance, reliability, and cost-effectiveness in collaboration with engineering teams.
  • Drive infrastructure improvements and advocate for best practices in system design and operations.
  • Establish and manage disaster recovery plans, ensuring system availability during unexpected events.


Qualifications

Minimum Education

  • BS/BA in Computer Science or equivalent experience

Preferred Education

  • MA in Computer Science

Minimum Skills

  • Argo CD and Argo Workflows
  • IaC: Terraform and Terragrunt
  • Kubernetes and and related plugins
  • GitHub Actions
  • AWS (EKS, Fargate, Aurora)
  • Security and Compliance
  • Containerization (Docker)
  • Logging and Monitoring Tools
  • Programming Scripting Language
  • DB Management
  • Version Control (Git)
  • Incident Management

Preferred Skills

  • Experience with Datadog
  • Experience with Cloudflare
  • Mortgage Domain Knowledge
  • Advanced Security Practices (GuardDuty, Security Hub)
  • Disaster Recovery Planning
  • Experience with workflow automation tools such as Camunda

Preferred Certification(s)

  • AWS Associate, AWS DevOps Professional

What we offer
We’re committed to creating an environment where our team members can thrive both professionally and personally. We currently offer:

  • Competitive Compensation – Including salary and performance bonuses.
  • Comprehensive Benefits – Health, dental, vision, and mental wellness support.
  • Retirement Savings – 401(k) with company matching.
  • Career advancement opportunities with business growth. 
  • Inclusive Culture – A diverse, collaborative, and supportive workplace where every voice is valued.
  • Perks & Extras – Generous PTO, team events, wellness programs, and more.