Build high-performance LLM inference systems and optimization techniques at scale.
Join a research team developing next-generation inference systems that optimize both speed and efficiency for large language models. You'll work across the full stack—from distributed serving and GPU kernels to model-system co-design—using techniques like speculative decoding, adaptive parallelism, and quantization. This role suits experienced systems engineers comfortable with performance profiling, GPU optimization, and independent problem-solving in fast-moving, research-driven environments.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 10 September 2026 · InsideJobs links you to the original posting.