SEE ALL VACANCIES

Senior Data Engineer

Meet your recruiter
Sara Abdelaziz Sherif
Vacancy details
Data Engineering
Data Engineer
Senior
Bulgaria, Colombia, Croatia, Egypt, India, Poland, Portugal, Spain, Ukraine
Remote
APPLY NOW
REFER A FRIEND

Intellias is looking for a Senior Data Engineer to join a digital healthcare initiative focused on building a large-scale health data platform that helps millions of people discover healthcare providers, connect with their health information, and make better-informed care decisions.

This is a hands-on, high-ownership engineering role for someone experienced in building and operating data-intensive systems at scale. You will own critical data ingestion and processing capabilities, improve data quality and reliability, and develop AI-powered solutions for complex real-world healthcare data challenges.

We are looking for an engineer who combines strong Python and data engineering expertise with cloud-native development, production ownership, and an AI-first engineering mindset.

What project we have for you

The project is building a FHIR-based healthcare data platform that aggregates provider information from national registries, EHR systems, healthcare networks, and partner data feeds.

The platform ingests fragmented and often inconsistent healthcare data, transforms it into standardized FHIR resources, resolves duplicate identities and relationships, and makes trusted provider information available through high-scale search and data services.

The engineering challenge goes well beyond traditional ETL. Provider data comes from hundreds of heterogeneous sources and contains duplicates, outdated records, conflicting identities, incomplete relationships, and inconsistent schemas. The platform must determine how these sources relate and continuously improve the accuracy and confidence of the resulting data.

The technology landscape includes Python, Apache Spark, Prefect, AWS, Kubernetes, MongoDB, OpenSearch, and FHIR, with AI increasingly used both as part of the engineering lifecycle and within data-processing solutions.

What you will do

  • Design, build, and operate scalable data ingestion pipelines that onboard healthcare provider data from EHR systems, national registries, healthcare networks, and partner feeds.
  • Transform heterogeneous source data into standardized, trusted data models, including FHIR-native resources.
  • Own data quality end-to-end, including validation rules, quality thresholds, confidence scoring, automated monitoring, and alerting.
  • Design and implement solutions for entity resolution, deduplication, relationship inference, classification, enrichment, and data quality scoring.
  • Build and evolve distributed data processing pipelines using Python, Spark, and modern orchestration technologies.
  • Improve pipeline reliability, scalability, observability, and performance so that new data sources can be onboarded efficiently and safely.
  • Establish and evolve data governance standards, including schemas, staging models, validation processes, and data refinement workflows.
  • Design and operate data storage and search solutions supporting large-scale healthcare datasets.
  • Build AI-powered data tooling that improves the accuracy, automation, and intelligence of data processing workflows.
  • Use modern AI development tools as an integral part of the engineering workflow to accelerate implementation, testing, debugging, documentation, and data analysis.
  • Own production operations for the solutions you build, including monitoring, troubleshooting, incident response, and root-cause analysis.
  • Partner with Analytics and Data Science teams to develop data quality dashboards, metrics, and reporting.
  • Collaborate with Product and Business stakeholders to prioritize data sources and improvements based on customer and business impact.
  • Contribute to technical design reviews and help establish scalable engineering patterns across the broader data platform.

What you need for this

 

  • 5+ years of professional experience building and operating data-intensive backend or data engineering systems at scale.
  • Strong hands-on expertise in Python, including development of production-grade data processing and backend services.
  • Strong experience designing and building scalable data ingestion and ETL/ELT pipelines for large and heterogeneous datasets.
  • Hands-on experience with Apache Spark or Databricks for distributed data processing.
  • Experience with workflow orchestration technologies such as Prefect, Apache Airflow, or equivalent.
  • Strong understanding of data modeling, schema design, validation strategies, data transformation, and data quality management.
  • Experience designing solutions for data deduplication, entity resolution, data enrichment, classification, or record linkage.
  • Experience with both relational and document-oriented databases; practical knowledge of technologies such as MongoDB or comparable NoSQL platforms.
  • Experience with search technologies such as OpenSearch, Elasticsearch, or comparable solutions.
  • Strong AWS experience, including services such as S3, ECS/EKS, Lambda, or equivalent cloud-native technologies.
  • Hands-on experience with Docker, Kubernetes, CI/CD, and operating cloud-native workloads in production.
  • Strong understanding of production data pipeline concerns including reliability, fault tolerance, observability, monitoring, alerting, and incident response.
  • Experience establishing and maintaining data governance standards, schemas, validation rules, and repeatable data refinement processes.
  • Practical experience using AI-assisted software engineering tools as part of the regular development lifecycle.
  • Strong ownership mindset and ability to independently identify data or engineering problems, propose solutions, implement them, and operate them in production.
  • Strong communication skills and ability to collaborate directly with engineering, analytics, product, data science, and business stakeholders.
  • Professional English sufficient for direct collaboration with U.S.-based teams.

Nice to Have

  • Experience working with healthcare data, HealthTech platforms, or healthcare interoperability.
  • Knowledge of FHIR, particularly resources such as Practitioner, Organization, PractitionerRole, Endpoint, and Location.
  • Experience integrating data from EHR/EMR systems, healthcare registries, provider networks, or similar complex data ecosystems.
  • Experience designing data platforms handling PHI/PII or other regulated and sensitive information.
  • Familiarity with HIPAA-related data protection and security considerations.
  • Experience building AI-powered data processing solutions, particularly entity resolution, relationship inference, classification, enrichment, anomaly detection, or data quality scoring.
  • Practical experience integrating LLMs into production data workflows, including evaluation, reliability, latency, cost, and data-protection considerations.
  • Experience with AI engineering tools such as Claude Code, GitHub Copilot, Cursor, or similar platforms.

What it’s like to work at Intellias

At Intellias, where technology takes center stage, people always come before processes. By creating a comfortable atmosphere in our team, we empower individuals to unlock their true potential and achieve extraordinary results. That’s why we offer a range of benefits that support your well-being and charge your professional growth.
We are committed to fostering equity, diversity, and inclusion as an equal opportunity employer. All applicants will be considered for employment without discrimination based on race, color, religion, age, gender, nationality, disability, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law.
We welcome and celebrate the uniqueness of every individual. Join Intellias for a career where your perspectives and contributions are vital to our shared success.

Have not found the most
suitable position yet?
Leave your resume and we will select a cool option for you.
Find me a job
Good news!
Link copied
Good news!
You did it.
Bad news!
Something went wrong. Please try again.