Build infrastructure to monitor and manage OpenAI's global GPU compute fleet health.
Build infrastructure to monitor and manage GPU compute fleet health across global operations. You'll work with Python, SQL, PromQL, and Linux to develop monitoring and management systems for GPU and InfiniBand hardware. This is a senior-level, full-time onsite position in San Francisco requiring fluency in English.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 27 July 2026 · InsideJobs links you to the original posting.