Staff SRE managing AI infrastructure reliability, monitoring, and incident response.
You'll manage AI infrastructure reliability, monitoring, and incident response as a Staff SRE, working with AWS, Kubernetes, Docker, Terraform, Python, Go, Git, and Ansible. This is a permanent, full-time remote position in the USA with hybrid flexibility for a senior-level role. You'll work in English with a team focused on core AI infrastructure stability and operations.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 22 July 2026 · InsideJobs links you to the original posting.