Skip to content
Skip to content
DevOps Jobs
Volta

Platform Engineer

Volta

Location
Onsite (Palo Alto, California)
Compensation
$200k - $215k/yr
Employment
Full-time
Level
Mid Level
Posted 2 days ago

About the Role

Volta builds large-scale GPU compute infrastructure for AI workloads using a Kubernetes-native platform. This role involves designing and implementing platform capabilities, including Kubernetes operators, confidential computing, and observability, to support product roadmaps and operational needs.

Skills

Kubernetes Python Go Rust Infrastructure as code API design Linux Networking Distributed systems Confidential computing Observability CI/CD Cloud native GPU infrastructure Storage systems Agile

Benefits

  • Comprehensive benefits

Perks

  • Discretionary bonus
  • Equity awards

Full job details

About The Role

Volta builds and operates large-scale GPU compute infrastructure for AI workloads. Our platform is Kubernetes-native, spans multiple regions, and delivers virtual machines, storage, and networking through a fully automated infrastructure stack built on custom Kubernetes operators.

Platform Engineers work at the intersection of infrastructure and software development. You will translate three key inputs into durable platform capabilities: product roadmap requirements from the product team, operational learnings from the bring-up team, and security guidance from the security engineering team.

What You Will Be Doing

  • Design and implement Kubernetes operators and controllers that manage the lifecycle of compute, storage, and networking resources

  • Work closely with the product team to understand roadmap requirements and implement the platform capabilities that support them

  • Collaborate with the bring-up team to identify operational pain points and turn them into scalable platform features

  • Improve and extend the northbound API layer — the interface between user-facing services and the underlying infrastructure platform

  • Build and extend confidential computing capabilities across the platform stack — from secure bare metal and confidential VMs to Confidential Containers (CoCo)

  • Integrate security guidance from the security engineering team into platform-level controls and remediate security findings at the platform layer

  • Build platform capabilities around networking: reliability, performance, and observability of the overlay and underlay network stack

  • Contribute to storage platform improvements: provisioning workflows, attachment reliability, performance tuning, and failure handling

  • Own observability as a platform concern — instrument services, define meaningful metrics, and build tooling that gives the team visibility into platform health

  • Participate in code review, technical design discussions, and cross-team collaboration in an Agile (Kanban or Scrum) environment

What You Bring

  • 3–5 years of software engineering experience, with a meaningful portion spent on infrastructure or platform systems

  • Working proficiency in at least one relevant language — Python, Go, or Rust — with experience writing production-grade backend services or automation, and a willingness to work across languages as the codebase evolves

  • Solid understanding of Kubernetes internals: the control loop model, CRDs, controllers/operators, and reliable reconciliation logic

  • Comfortable working close to the infrastructure layer — Linux, networking fundamentals, and distributed systems behaviour

  • Experience designing and building APIs or service interfaces that other teams depend on

  • Strong engineering fundamentals: clean code, testing, version control, code review, and CI/CD practices

Nice to Have

  • Fluency with AI-assisted development, and interest in scaling agent-assisted workflows across the team (agentic CLI tools, MCP, skills, APIs) to amplify delivery.

  • Familiarity with confidential computing technologies: TEEs, AMD SEV, Intel TDX, or Confidential Containers (CoCo)

  • Experience integrating security requirements into platform or infrastructure systems

  • Familiarity with high-performance networking: overlay protocols, BGP, RDMA, or packet-processing frameworks

  • Hands-on experience with distributed storage systems (Ceph or similar) at an engineering level

  • Background building Kubernetes operators using frameworks such as Kopf, controller-runtime, or similar

  • Experience with observability tooling: Prometheus, Grafana, OpenTelemetry, or structured logging in distributed systems

  • Exposure to GPU infrastructure or HPC environments