Phizenix seeks a Director of Site Reliability Engineering to build and lead our AI accelerator infrastructure from the ground up, spanning colocation, on‑prem lab clusters, and multi‑cloud environments. You will define reliability architecture, own SLOs, and act as the senior escalation point for critical incidents.
You will hire and grow a 3–5 person SRE team, partner with DevOps, and ensure a unified platform across hardware, software, and executive stakeholders while driving continuous
#J-18808-Ljbffr

Director, SRE — AI Accelerator Infrastructure
Phizenix · Santa Clara, CA, USA ·
- Pay:
- 195.000 - 285.000
- Job type:
- Full Time