SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
San Francisco Compute operates GPU clusters and supercomputers, leasing them to enterprise customers with flexible contracts that allow subleasing. The company was founded by former leaders from Lambda, Crusoe, Digital Ocean, AWS, and Hut8, with a CTO who previously founded and led Voltage Park.
As a Customer Support Engineer, you will own inbound support tickets and live escalations for enterprise customers running AI and GPU workloads on the company's infrastructure. You will triage and resolve technical issues spanning compute, networking, and platform layers—including scheduling, orchestration, and performance problems. You'll communicate clearly with technical customers under pressure, set realistic expectations, provide status updates, and ensure clean handoffs across timezones for 24/7 follow-the-sun coverage.
You will use monitoring and alerting tools (Prometheus, Grafana, Datadog) to diagnose issues proactively. You'll escalate hardware, data-center, and facility-level issues to the appropriate internal teams or external partners (engineering, colo partners, OEMs) with clear, well-documented handoffs. As a first responder on incidents, you'll work alongside engineering through resolution.
You will help build and improve runbooks, SOPs, and the knowledge base, flagging gaps rather than working around them. You'll use AI-assisted tooling to work faster without sacrificing quality or judgment. You'll track your own CSAT, first-response time, and time-to-resolve metrics as your personal feedback loop.
REQUIREMENTS:
- 3–5 years in technical support, customer support engineering, or similar customer-facing technical role
- Comfortable troubleshooting infrastructure-level issues: Linux administration, basic shell or Python scripting, hands-on use of monitoring/observability tools (Prometheus, Grafana, Datadog); GPU/AI/HPC experience is a strong plus but not required
- Ability to explain technical problems clearly to both technical customers and internal engineering teams
- Calm under pressure; able to stay composed when customers are frustrated or systems are down
- Comfortable with shift-based hours as part of 24/7 global coverage, including occasional after-hours, weekend, or holiday coverage during incidents
- Genuine commitment to getting customers to a good outcome, not just closing tickets
- Excited to help build process and coverage from scratch, not just operate within existing structures
- Excellent written communication skills for both customers and internal teams
NICE TO HAVES:
- Experience with GPU/AI infrastructure, specialized hardware, or managed services
- Familiarity with data-center or colocation operations
- Experience with modern support tooling (ticketing, monitoring/alerting platforms) and AI-assisted workflows
- Scripting or basic programming ability for diagnostics and automation
- Experience with incident management