Build petabyte-scale distributed storage systems for AI model training.
Design and operate distributed storage infrastructure that serves massive training workloads across GPU clusters. You'll solve networking, I/O, and consistency challenges when moving datasets and checkpoints at scale, working closely with research and training teams. The role requires strong storage fundamentals, proficiency in Python or Go, and hands-on Kubernetes experience with stateful systems.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 11 September 2026 · InsideJobs links you to the original posting.