SEE ALL VACANCIES

Senior AI/RAG Engineer (Document Intelligence)

Vacancy details
Data Engineering
Data Engineer
Senior
Poland, Spain
Remote
APPLY NOW
REFER A FRIEND

Our client is a leading global investment management company headquartered in London. It manages over $228 billion in assets and serves institutional investors, pension funds, wealth managers, and other sophisticated clients worldwide. The firm specializes in quantitative investing, alternative investments, systematic trading strategies, and technology-driven asset management. Data science, machine learning, and AI are core components of its investment and research processes.

As part of our collaboration we will focus on two foundational capabilities required to enable safe and scalable AI adoption across the enterprise: Agentic Security and AI-Ready Data Foundations.

What project we have for you

We build the data foundations that make AI useful and safe inside regulated financial firms. The value of AI is capped by the data its agents can reach: if an agent cannot find, interpret, trace or be correctly permissioned against data, the capability is useless, or worse, unsafe. Your job is to close that gap.

This is a hands-on senior role for an excellent Python engineer with strong data-engineering skills who is genuinely comfortable building with AI agents. You will design and build the catalogue, semantic, entitlement and analytical layers that turn large on-premise data estates into something agents can use.

What you will do

  • Build the evaluation foundation at the start: convert the client’s existing question bank and diagnostic results into a versioned automated test suite with expert-confirmed outcomes, establish the baseline, and mine historical support records as a second ground-truth set.
  • Improve the client’s existing RAG service in place: hybrid retrieval with reranking, structure-aware chunks, context notes, early metadata filters, reliability monitoring, with every change measured against the suite before release.
  • Build the metadata and entity extraction pipeline under a governed three-tier schema (universal envelope, versioned per-collection specifications, open discovery tier), with a stratified pilot, confidence-routed human review and coverage dashboards.
  • Build the synchronisation plane over the client’s content platform (change notifications, delta polling, scheduled full reviews) shared by retrieval, extraction and downstream views.
  • Build document lineage and temporal views: supersession and amendment chains extracted only where stated in the text, a bitemporal effective-terms view, derived document status, and a queryable obligations register, with human verification for high-stakes chains.
  • Contribute to the controlled knowledge graph and query orchestration: typed nodes and edges carrying provenance and confidence, routing between structured lookup, filtered retrieval and lineage views, explicit completeness statements, and typed gaps reported as answers.
  • Operate the delivered capabilities: releases gated on evaluation results, freshness bounds enforced by withholding stale data, coverage and quality dashboards, corpus health checks reported to document owners.

What you need for this

  • 5+ years of production Python development, including 2+ years of LLM and RAG engineering in production: retrieval pipelines, vector stores, structured extraction, and the surrounding operational tooling.
  • Strong experience building evaluation harnesses: versioned test suites derived from real question banks, separate scoring of retrieval and answers, scoring where a correct “not found” counts as a pass, and suites wired into delivery as release gates (e.g. Langfuse, RAGAS, DeepEval or similar, plus custom metrics).
  • Hybrid retrieval engineering: keyword and semantic search combined, result merging, cross-encoder reranking, tuning against measured baselines, working within a platform-fixed embedding model and index.
  • Structure-aware document processing: layout-aware parsing and chunking that keeps tables intact (e.g. Docling, Tika or similar), including OCR handling for scanned documents and multilingual content.
  • LLM extraction at scale: schema-driven extraction of attributes, entities, clauses and relationships with per-field confidence, calibrated thresholds and a human review loop (e.g. Label Studio or similar), piloted and measured before scale-out.
  • Strong PostgreSQL: typed relational modelling plus JSONB, schema-as-code with migration tooling, derived views managed as tested transformations.
  • Provenance and citation discipline: every extracted fact traceable to its source document and passage; answers that state explicitly when something could not be confirmed.
  • Fluent English for written and spoken communication with client teams.

Will be a plus

  • Integration with SharePoint and Microsoft Graph APIs or an equivalent enterprise content platform: change notifications, delta polling, permission metadata.
  • Bitemporal modelling (execution versus effective dates, as-of queries) and document lineage or supersession modelling.
  • Graph engines (e.g. Apache AGE, Neo4j or similar); controlled knowledge graphs with provenance on every element.
  • Legal, contract or fund-documentation domain knowledge, or demonstrated ability to learn a document domain in depth.
  • Permission-aware retrieval: access lists stored in the index, group resolution at query time, denial on uncertainty (the entitlement design is owned by a parallel workstream; this role implements against it).
  • Token-efficiency engineering: corpus preparation, section slicing, cost measurement per query.
  • Exposing capabilities to an assistant platform as MCP tools or skills; day-to-day use of AI coding agents.
  • Experience in financial services or other regulated on-premise environments; client-facing experience.

What it’s like to work at Intellias

At Intellias, where technology takes center stage, people always come before processes. By creating a comfortable atmosphere in our team, we empower individuals to unlock their true potential and achieve extraordinary results. That’s why we offer a range of benefits that support your well-being and charge your professional growth.
We are committed to fostering equity, diversity, and inclusion as an equal opportunity employer. All applicants will be considered for employment without discrimination based on race, color, religion, age, gender, nationality, disability, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law.
We welcome and celebrate the uniqueness of every individual. Join Intellias for a career where your perspectives and contributions are vital to our shared success.

Skills

LLM
Python
Rag
SQL
Have not found the most
suitable position yet?
Leave your resume and we will select a cool option for you.
Find me a job
Good news!
Link copied
Good news!
You did it.
Bad news!
Something went wrong. Please try again.