Staff engineer building large-scale GPU training infrastructure and distributed systems.
You'll design and build large-scale GPU training infrastructure and distributed systems, working with PyTorch, FSDP, DeepSpeed, and Megatron across GPU clusters connected via NVLink, InfiniBand, and RoCE. This is a senior-level full-time position based onsite in Mountain View or San Francisco, requiring expertise in distributed systems at scale.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 14 August 2026 · InsideJobs links you to the original posting.