Lead GenAI inference engine architecture and optimization at scale.
You will lead the architecture and optimization of generative AI inference engines at scale. This senior-level position is based in San Francisco and requires on-site work with a focus on performance optimization using CUDA, cuBLAS, cuDNN, NCCL, and deep learning frameworks like PyTorch and TensorFlow. The role is a permanent, full-time position requiring fluent English and experience with low-level GPU acceleration technologies.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 14 August 2026 · InsideJobs links you to the original posting.