SEE ALL VACANCIES

Python Engineer — Evaluator Library

Vacancy details
Software Engineering
Python Engineer
Strong Middle
Bulgaria, Canada, Croatia, Cyprus, Egypt, Germany, India, Japan, Malta, Poland, Portugal, Saudi Arabia, Spain, Ukraine, United Arab Emirates
Remote
APPLY NOW
REFER A FRIEND

We are looking for a Python Engineer – Evaluator Library to design and implement reusable evaluation components that ensure the quality, safety, and compliance of enterprise AI agents and LLM-powered workflows. You will build custom evaluation capabilities used across AI platforms, focusing on automated quality validation, workflow compliance, PII protection, and structured output verification. Working closely with AI Platform Engineers, ML Engineers, and DevOps teams, you will help establish reliable evaluation standards and scalable quality assurance mechanisms for agent-based systems.

What project we have for you

Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
Intellia’s mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.

What you will do

  • Design, develop, and maintain reusable Python-based evaluator libraries for AI agents and LLM-powered workflows.
  • Implement AWS Lambda-based custom evaluators to perform deterministic quality, compliance, and validation checks.
  • Develop LLM-as-a-judge evaluation logic to assess subjective dimensions such as relevance, helpfulness, consistency, and response quality.
  • Build automated PII detection evaluators using regex-based techniques and AWS Bedrock Guardrails integrations.
  • Implement TOOL_CALL-level validation mechanisms to verify structured outputs, JSON schema compliance, and tool response correctness.
  • Develop SESSION-level evaluators to validate workflow contract compliance, execution integrity, and cross-step behavioral expectations.
  • Create TRACE-level evaluators for numerical accuracy verification, calculation consistency, and deterministic result validation.
  • Integrate evaluator components with AWS AgentCore Evaluation workflows and enterprise AI quality pipelines.
  • Collaborate with AI platform teams to define evaluation standards, scoring methodologies, and quality acceptance criteria.
  • Design and maintain unit, integration, and validation tests for evaluator libraries and quality frameworks.
  • Support observability and troubleshooting by integrating evaluation outputs with CloudWatch logging and monitoring capabilities.
  • Contribute to enterprise AI governance initiatives by improving evaluation coverage, auditability, and compliance controls.

What you need for this

Skills:

• Python (Lambda functions as AWS AgentCore custom code-based evaluators)
• LLM-as-judge prompt engineering for subjective evaluation dimensions
• PII detection (regex-based + AWS Bedrock Guardrails)
• Tool response schema validation (JSON Schema — TOOL_CALL level evaluator)
• Workflow contract compliance checking (SESSION level evaluator)
• Numerical accuracy validation logic (TRACE level evaluator) 

Experience:

• 4+ years Python engineering
• LLM evaluation or quality assurance for AI/ML systems
• AWS Lambda function development and deployment 

Nice-to-have

• AWS AgentCore Evaluation custom evaluator Lambda registration
• AWS Bedrock Guardrails for PII detection integration
• CloudWatch Logs as evaluator output sink

What it’s like to work at Intellias

At Intellias, where technology takes center stage, people always come before processes. By creating a comfortable atmosphere in our team, we empower individuals to unlock their true potential and achieve extraordinary results. That’s why we offer a range of benefits that support your well-being and charge your professional growth.
We are committed to fostering equity, diversity, and inclusion as an equal opportunity employer. All applicants will be considered for employment without discrimination based on race, color, religion, age, gender, nationality, disability, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law.
We welcome and celebrate the uniqueness of every individual. Join Intellias for a career where your perspectives and contributions are vital to our shared success.

Have not found the most
suitable position yet?
Leave your resume and we will select a cool option for you.
Find me a job
Good news!
Link copied
Good news!
You did it.
Bad news!
Something went wrong. Please try again.