Current opportunity
Site Reliability Engineer
STN Incorporated
San Francisco, United StatesRemoteNot disclosedExternal listing
About the opportunity
Site Reliability EngineerPlatform and software · shared across customersReports to: Director, Site ReliabilityLocation: Remote (US)Department: Cloud Platform Engineering / SRE/ReliabilityPosition summaryThe Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.
Key responsibilities
- Define and operate Service Level Objectives (SLOs) aligned with customer SLAsBuild and maintain the observability stack including metrics, logs, traces, and alertingLead incident response and chair post-incident reviewsDrive automation to reduce toil and improve mean-time-to-recover (MTTR)Author and maintain operational runbooks alongside the NOCManage on-call rotation, escalation paths, and incident-management toolingCoordinate cross-functionally with NOC, Platform Engineering, and Network EngineeringDrive chaos engineering, game days, and reliability testing programsProduce SLA performance reports in coordination with the SLA ManagerMentor junior engineers and contribute to engineering cultureRequired qualifications5+ years in SRE, DevOps, or production engineering rolesStrong programming skills in Go, Python, or bothHands-on experience operating Kubernetes-based platforms at scaleDeep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)Strong incident management experience including major-incident commandPreferred qualificationsGPU or HPC platform operational experienceFamiliarity with SLA-driven customer environments and credit calculationsExperience with chaos engineering tools (Gremlin, Litmus, or similar)Published SRE content or contributions
