What you'll do
Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
Debug across the full stack, hardware to fabric to scheduler, when a trainin
Run customer-facing incident communication with technical depth and no spin
Turn recurring customer pain into engineering fixes with the production team
You've supported large-scale compute customers (HPC, cloud, or AI labs) at a