← All roles

Staff Software Engineer - GenAI inference

Permanentonsite5+ yrs📍 San Francisco, California, 🇺🇸 United States🗣 English

Lead GenAI inference engine architecture and optimization at scale.

You will lead the architecture and optimization of generative AI inference engines at scale. This senior-level position is based in San Francisco and requires on-site work with a focus on performance optimization using CUDA, cuBLAS, cuDNN, NCCL, and deep learning frameworks like PyTorch and TensorFlow. The role is a permanent, full-time position requiring fluent English and experience with low-level GPU acceleration technologies.

Tech stack
CUDAcuBLAScuDNNNCCLPyTorchTensorFlow
Rate
Not stated by the agency
via

Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.

Found at a specialist agency · listed 14 August 2026 · InsideJobs links you to the original posting.