← All roles

Senior/Staff Software Engineer – LLM Inference & Reinforcement Learning Platform

Permanentremote5+ yrs📍 US-WA-Bellevue, 🇺🇸 United States🗣 English

Build high-performance LLM inference systems and optimization techniques at scale.

Join a research team developing next-generation inference systems that optimize both speed and efficiency for large language models. You'll work across the full stack—from distributed serving and GPU kernels to model-system co-design—using techniques like speculative decoding, adaptive parallelism, and quantization. This role suits experienced systems engineers comfortable with performance profiling, GPU optimization, and independent problem-solving in fast-moving, research-driven environments.

Tech stack
CUDATritonvLLMTensorRT-LLMcuBLAScuDNNCUTLASSGPU
Rate
Not stated by the agency
via

Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.

Found at a specialist agency · listed 10 September 2026 · InsideJobs links you to the original posting.