J
JobQuip
DE
채용공고 목록으로

Customer Reliability Engineer

ashby:andromedaGlobal Remote / San Francisco, CA급여 협의정규직

직무 설명

Customer Reliability EngineerLocation: Remote/SF-Hybrid · Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers. We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible. Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth. Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering. The Role Our customers run large AI training and inference workloads on GPU clusters we source from providers worldwide. When a node goes dark or a job dies eight hours into a run, the Customer Reliability Engineer is who they hear from, and who gets it sorted. The job has three parts. You triage incoming issues and debug them at the Linux and Kubernetes layer. You work provider-side to figure out whose fault something actually is and push external providers to fix it. And you build the monitoring and scripts that catch problems before a customer has to tell us. You need to be comfortable in a Linux shell and know how Kubernetes works. You don't need GPU or HPC experience. Most people pick that up here. What You’ll Do Triage and fix customer issues Own issues start to finish: reproduce, diagnose, fix or escalate, close the loop Debug at the Linux layer: processes, networking, storage, kernel logs, resource contention, systemd, journald Dig into Kubernetes problems like pods stuck pending or crash-looping, node conditions, scheduling failures, resource limits Work GPU failures: driver and device-plugin issues, XID errors, thermal throttling, nodes that need cordoning or draining, jobs failing across multiple nodes Escalate when you're past your depth, with the evidence already gathered Handle incidents Take part in a 24/7 on-call rotation First response on alerts and customer-reported outages: assess impact, set severity, pull in the right people Keep customers updated during incidents.

바로 지원

게시일 2025. 11. 7.
지원 페이지가 잠겨 있습니다

가입 후 전체 공고를 보고 지원하세요

무료 가입으로 지원 페이지를 열고 공고 저장과 진행 상황 확인을 이용하세요.