← All roles

Senior ML Evaluation Engineer with AWS Agent Evaluation

Permanentremote5+ yrs📍 Armenia, Bulgaria, Cyprus, Georgia, 🇬🇧 United Kingdom🗣 English

Senior ML engineer designing evaluation frameworks for enterprise AI agents and systems.

Build and operationalize quality assessment pipelines for AI agents and LLM-based systems at enterprise scale. You'll define evaluation standards, implement custom Python-based validators, design CI/CD quality gates, and collaborate with platform teams on governance and monitoring. This role suits experienced ML engineers comfortable with LLM evaluation methodologies, AWS services, and translating technical quality requirements into automated processes.

Tech stack
AWS Agent Evaluation ServicesPythonAWS LambdaAWS BedrockOpenTelemetryCloudWatchCI/CD PipelinesLarge Language Models
Rate
Not stated by the agency
via

Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.

Found at a specialist agency · listed 1 September 2026 · InsideJobs links you to the original posting.