Site Reliability Engineer
LTD Global
- Location
- Hybrid (Berkeley, CA)
- Compensation
- $75 - $80/hr
- Employment
- Contract
- Level
- Mid Level
Posted 1 week ago
About the Role
Join a national HPC facility supporting over 11,000 scientists in energy, physics, and materials science. This role enables world-class research by maintaining high-performance computing infrastructure and facility systems.
Skills
Linux
Python
C
C++
Perl
Java
Kubernetes
Prometheus
VictoriaMetrics
Alertmanager
Network security
SSH
ServiceNow
ITSM
Automation
Monitoring
Full job details
📍 Hybrid — Berkeley, CA
📅 1 Year Contract Assignment with possibility of extension based on performance and organizational needs.
đź’° $80/hr
Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption.
If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.
What You'll Own
📅 1 Year Contract Assignment with possibility of extension based on performance and organizational needs.
đź’° $80/hr
Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption.
If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.
What You'll Own
- Monitor and triage alerts across compute, storage, network, and facility systems in real time
- Build automation that prevents issues before they become outages
- Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
- Walk the data center floor to keep power, cooling, and environmental systems humming
- Coordinate maintenance activities across teams and keep incidents accurately tracked
- Dig into complex, ambiguous problems and drive them to resolution
- Comfort working Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA
- Solid Linux/command-line (SSH) chops
- Programming/scripting experience:Â Python, C, C++, Perl, or Java
- A self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
- Network security fundamentals (ACLs, firewalls)
- Strong cross-team communication and collaboration skills
- Experience building or deploying Agentic AI / autonomous automation for technical workflows
- ServiceNow implementation experience
- ITSM best-practice know-how