Build LLM inference runtime for custom AI accelerators at scale.
Design and implement a production-grade inference engine that runs frontier language models on OpenAI's custom silicon, optimizing for throughput, latency, and hardware efficiency. You'll work across model architecture, distributed systems, and kernel-level optimization, collaborating with hardware and compiler teams to translate complex workloads into performant execution. The role requires deep systems programming experience and understanding of modern LLM inference patterns, batching strategies, and memory management.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 25 August 2026 · InsideJobs links you to the original posting.