SEE ALL VACANCIES

QA / ML Tester – Evaluation Framework

Vacancy details
Test Engineering
Automation Test Engineer (Python)
Strong Middle
Bulgaria, Canada, Croatia, Cyprus, Egypt, Germany, India, Japan, Malta, Poland, Portugal, Saudi Arabia, Spain, Ukraine, United Arab Emirates
Remote
APPLY NOW
REFER A FRIEND

We are looking for a QA / ML Tester – Evaluation Framework to ensure the quality, reliability, and correctness of evaluation systems used across enterprise AI and agent-based platforms. In this role, you will design and execute validation strategies for evaluation frameworks, test evaluator behavior across multiple scenarios, and verify the accuracy of automated quality assessment pipelines. You will work closely with AI engineers, platform teams, and quality specialists to establish confidence in evaluation results and support enterprise-grade AI governance.

What project we have for you

Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
Intellia’s mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.

What you will do

  • Design, implement, and maintain automated test suites for AI evaluation frameworks and evaluation pipelines.
  • Develop Python-based test automation using pytest to validate evaluator behavior, quality scoring, and framework reliability.
  • Create and maintain known-good and known-bad test datasets, sessions, and workflows for evaluator correctness validation.
  • Validate the accuracy and consistency of evaluation results across different agent workflows, prompts, tools, and execution scenarios.
  • Design and execute integration tests for on-demand evaluation workflows integrated into CI/CD pipelines.
  • Verify online evaluation behavior, sampling accuracy, and evaluation result consistency in production-like environments.
  • Conduct functional testing of evaluation components, including evaluator execution flows, scoring logic, and result aggregation.
  • Collaborate with AI and platform engineering teams to identify edge cases, failure scenarios, and evaluation blind spots.
  • Validate workflow compliance, tool execution assessment, and end-to-end quality evaluation processes.
  • Support feasibility assessments for applying evaluation frameworks to non-AgentCore runtimes and alternative AI execution environments.
  • Analyze defects, inconsistencies, and quality regressions within evaluation systems and provide actionable recommendations.
  • Contribute to quality assurance standards, testing methodologies, and best practices for AI evaluation platforms.

What you need for this

Skills:

• Python test automation (pytest)
• Evaluator correctness testing (known-good / known-bad session pairs)
• On-demand mode integration testing with CI/CD
• Online mode sampling accuracy validation
• Non-AgentCore runtime feasibility assessment methodology 

Experience:

• 4+ years QA or ML testing engineering
• AI/LLM system quality testing
• Integration test design for evaluation pipelines 

Nice-to-have

• AWS AgentCore Evaluation API testing
• OpenTelemetry trace-based evaluation input testing
• Multi-evaluator execution correctness testing

What it’s like to work at Intellias

At Intellias, where technology takes center stage, people always come before processes. By creating a comfortable atmosphere in our team, we empower individuals to unlock their true potential and achieve extraordinary results. That’s why we offer a range of benefits that support your well-being and charge your professional growth.
We are committed to fostering equity, diversity, and inclusion as an equal opportunity employer. All applicants will be considered for employment without discrimination based on race, color, religion, age, gender, nationality, disability, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law.
We welcome and celebrate the uniqueness of every individual. Join Intellias for a career where your perspectives and contributions are vital to our shared success.

Have not found the most
suitable position yet?
Leave your resume and we will select a cool option for you.
Find me a job
Good news!
Link copied
Good news!
You did it.
Bad news!
Something went wrong. Please try again.