SEE ALL VACANCIES

Senior Data Engineer (healthcare solution)

Vacancy details
Data Engineering
Data Engineer
Senior
Bulgaria, Colombia, Croatia, Egypt, India, Poland, Portugal, Spain, Ukraine, United States
Remote
APPLY NOW
REFER A FRIEND

Intellias is looking for a Senior Healthcare Data Engineer to join a large-scale digital healthcare initiative focused on transforming fragmented healthcare data into clean, standardized, and trustworthy health information.

This is a hands-on, high-ownership engineering role combining distributed data engineering, healthcare interoperability, and AI-native software development. You will own healthcare data end-to-end — from onboarding new data sources and building transformation pipelines to delivering production-ready, standards-conformant FHIR data.

We are looking for an engineer who combines strong Python, Spark/PySpark, and data engineering expertise with practical FHIR/HL7 knowledge and an AI-native engineering mindset.

What project we have for you

The project is building a FHIR-based health data platform at scale, integrating information from EHRs, healthcare organizations, national and reference datasets, and other clinical data sources.

The platform ingests heterogeneous healthcare data, transforms and normalizes it, performs patient and entity linking, terminology normalization and deduplication, and produces clean, standards-conformant FHIR resources for downstream healthcare applications and services.

The engineering environment combines Python, Apache Spark/PySpark, Databricks, Delta Lake, SQL, ETL/ELT pipelines, cloud-native infrastructure, FHIR R4, HL7, and healthcare terminology standards.

This is also an AI-native engineering environment. Coding agents and reusable AI capabilities are expected to be part of the regular engineering workflow — accelerating implementation, testing, mappings, data analysis, and repetitive engineering tasks while maintaining strict requirements around correctness, reliability, and PHI safety.

What you will do

  • Design, build, and operate healthcare data pipelines that transform heterogeneous source data into standards-conformant FHIR resources.
  • Own data sources from ingestion through production delivery, including file parsing, decryption, mapping, validation, transformation, enrichment, and downstream integration.
  • Develop scalable distributed processing solutions using Python and Spark/PySpark.
  • Build and maintain ETL/ELT workflows supporting batch, incremental, CDC, and streaming processing patterns.
  • Implement patient/record linking, deduplication, terminology normalization, enrichment, and data-quality controls.
  • Transform healthcare data across standards and formats, including HL7, C-CDA, FHIR, and proprietary source formats.
  • Build reliable processing around healthcare terminology systems including SNOMED CT, LOINC, and RxNorm.
  • Improve pipeline resilience through retry strategies, idempotency, automated validation, monitoring, alerting, and failure recovery.
  • Optimize data-processing workloads for performance, scalability, throughput, and cloud cost efficiency.
  • Investigate production data issues across multi-stage pipelines and drive them through root-cause analysis to permanent resolution.
  • Write and maintain unit, integration, and data-quality tests.
  • Handle PHI and encrypted healthcare data according to applicable security and privacy requirements.
  • Use AI coding agents and reusable AI capabilities as a regular part of the engineering workflow to accelerate development, testing, mappings, analysis, and repetitive engineering activities.
  • Apply engineering judgment when selecting between deterministic/rule-based and AI/LLM-based approaches, balancing correctness, reliability, cost, latency, and PHI safety.
  • Develop and contribute reusable AI-enabled engineering capabilities that can be adopted by other engineers and teams.
  • Collaborate with engineering, product, data, and healthcare-domain stakeholders to turn complex source data into reliable and usable health information.

What you need for this

 

  • 6+ years of professional experience in data engineering, backend/microservices development, HealthTech, or healthcare interoperability.
  • Strong hands-on expertise in Python, including production-grade data processing and pipeline development.
  • Proven experience owning data pipelines end-to-end in production, covering ingestion, transformation, validation, delivery, monitoring, troubleshooting, and operational support.
  • Strong experience with Apache Spark/PySpark at scale, including partitioning, performance optimization, memory tuning, and distributed processing.
  • Experience with modern lakehouse technologies such as Databricks and Delta Lake, or comparable platforms.
  • Strong experience designing ETL/ELT pipelines, including CDC, incremental/delta processing, idempotent processing, schema evolution, and batch/streaming workloads.
  • Strong SQL skills and solid understanding of data modeling, transformation, validation, and data-quality principles.
  • Hands-on experience with FHIR R4, including resources, profiles, Bundles, validation, and healthcare data transformation.
  • Working knowledge of HL7 v2.x and healthcare interoperability concepts.
  • Understanding of US Core and USCDI standards and their application to healthcare data.
  • Experience with healthcare data transformation patterns such as HL7/C-CDA to FHIR.
  • Understanding of healthcare terminology standards such as SNOMED CT, LOINC, and RxNorm.
  • Experience working with healthcare data domains such as clinical, claims, coverage, eligibility, or related datasets.
  • Strong understanding of production engineering practices including Git, CI/CD, automated testing, observability, reliability, and incident troubleshooting.
  • Proven ability to independently take a new data source or ambiguous engineering problem from initial analysis through production deployment and operation.
  • Practical experience using AI coding agents such as Claude Code or comparable tools as an integral part of the engineering workflow.
  • Ability to apply AI-assisted development with appropriate engineering judgment, ensuring generated or assisted solutions remain tested, deterministic where required, secure, and PHI-safe.
  • Strong ownership, analytical thinking, and communication skills.
  • Professional English sufficient for direct collaboration with U.S.-based engineering, product, and domain stakeholders.

Nice to Have

  • Deep experience with FHIR-based healthcare interoperability platforms operating at scale.
  • Experience with patient matching, record linkage, entity resolution, deduplication, and identity management.
  • Experience with healthcare terminology normalization and terminology services.
  • Experience with C-CDA/CDA and complex clinical document transformation.
  • Experience integrating data from multiple EHR/EMR platforms.
  • Experience handling PHI and regulated healthcare data, including HIPAA-related security and privacy considerations.
  • Experience with provider/reference datasets such as national provider registries.
  • Experience optimizing cloud data-processing workloads for cost, throughput, and scalability.
  • Experience building reusable AI skills, agents, or AI-assisted engineering workflows for data engineering.
  • Familiarity with LLM evaluation, guardrails, data privacy, and determining when deterministic/rule-based processing is preferable to an LLM-based solution.

What it’s like to work at Intellias

At Intellias, where technology takes center stage, people always come before processes. By creating a comfortable atmosphere in our team, we empower individuals to unlock their true potential and achieve extraordinary results. That’s why we offer a range of benefits that support your well-being and charge your professional growth.
We are committed to fostering equity, diversity, and inclusion as an equal opportunity employer. All applicants will be considered for employment without discrimination based on race, color, religion, age, gender, nationality, disability, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law.
We welcome and celebrate the uniqueness of every individual. Join Intellias for a career where your perspectives and contributions are vital to our shared success.

Skills

Python
Spark
Have not found the most
suitable position yet?
Leave your resume and we will select a cool option for you.
Find me a job
Good news!
Link copied
Good news!
You did it.
Bad news!
Something went wrong. Please try again.