Build LLM inference infrastructure powering enterprise-scale generative AI workloads.
You'll build and optimize LLM inference infrastructure that handles enterprise-scale generative AI workloads, working with PyTorch, MLflow, Ray, vLLM, and SGLang across GPU systems. This is a permanent, full-time senior role based in San Francisco requiring onsite work and English proficiency.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 14 August 2026 · InsideJobs links you to the original posting.