Build and optimize inference layers for AI models on edge cloud infrastructure.
Join a global infrastructure company to develop production-grade inference systems supporting large language and multimodal models. You will work with frameworks like vLLM and TensorRT-LLM, optimize GPU performance, and collaborate across platform and infrastructure teams. The role requires 5+ years of production Python experience, hands-on PyTorch and Kubernetes skills, and deep expertise in distributed systems or GPU computing.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 10 September 2026 · InsideJobs links you to the original posting.