Design evaluation frameworks for agent-based AI systems at enterprise scale.
Build automated quality assessment and testing infrastructure for LLM-powered agents, combining deterministic grading with LLM-as-judge approaches. You'll architect multi-layer evaluation strategies, implement deployment gates, and create feedback loops from production to testing. Requires 4+ years building evaluation platforms for ML/LLM systems, hands-on LangGraph experience, and strong Python skills.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 1 September 2026 · InsideJobs links you to the original posting.