- Must-Have Requirements
- Strong hands-on experience with AWS
- Experience working with data sets, data sources, and AWS data services
- Strong AI & LLM experience
- Excellent communication skills
- Active and complete LinkedIn profile
- Healthcare experience is a plus, but not required
- Role Overview
- We are looking for a Generalist Data Engineer to support a healthcare-focused AI benchmark and evaluation platform.
- The engineer will be responsible for the data lifecycle before a model receives it and after the model generates a response. This includes building ingestion, normalization, packaging, storage, and analysis layers to transform raw healthcare data into standardized benchmark inputs and actionable evaluation results.
- Key Responsibilities
- Build ingestion and normalization pipelines for:
- DICOM radiology studies
- Whole-slide pathology images
- Tabular EHR data, including labs, vitals, encounters, and medication records
- Design a canonical benchmark record format that can support different task configurations and model adapters
- Develop strategies for packaging large medical imaging datasets within third-party API limitations
- Build tiling, region selection, downsampling, and compression workflows while maintaining diagnostic information
- Maintain provenance metadata to ensure results are reproducible and defensible
- Develop cohort and label pipelines for clinical prediction tasks such as sepsis onset, survival horizons, and longitudinal lab trends
- Build results storage and analysis layers for per-run, per-model, and per-task outputs
- Enforce PHI handling requirements, including encryption, least-privilege access, audit logging, and de-identification
- Ensure clear controls around data leaving the VPC when interacting with third-party APIs
- Required Skills:
- Data Engineering
- Production-grade data pipeline development
- Strong testing discipline
- Experience building deterministic, idempotent, and re-runnable jobs
- AWS Data Stack
- Deep hands-on experience with S3
- AWS Glue and/or Spark on EMR
- Athena
- Step Functions
- Lambda
- AWS Batch
- Understanding of storage layout, lifecycle policies, and storage economics at scale
- SQL & Data Modeling
- Complex temporal joins
- Point-in-time correctness
- Strong understanding of preventing label leakage in time-series data
- Data Quality & Lineage
- Data validation frameworks
- Schema enforcement
- Versioned datasets
- Data lineage and reproducibility
- AWS Security & Governance
- IAM policy design
- KMS
- VPC endpoints and PrivateLink
- Experience working within HIPAA-eligible AWS architectures
- Desirable Skills
- Healthcare data standards including DICOM, FHIR, HL7v2, OMOP CDM
- Familiarity with clinical coding systems such as ICD, LOINC, RxNorm, and SNOMED
- Medical imaging experience with pydicom and OpenSlide
- Understanding of WSI pyramid structures and tiling
- Terraform / Infrastructure as Code
- Docker, ECR, and CI/CD
- Familiarity with LLM APIs and multimodal payload construction
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
LLM Data Engineer
Xaxis Solutions
Similar Engineering Jobs
View all Engineering jobs→Collins Aerospace
Data Engineer (Hybrid)
New
Cedar Rapids, Iowa (USA)
Northland Power Inc.
Data Engineer/Architect
New
Toronto, Ontario (Canada)
CLEAR
Senior Software Engineer, Data
New
New York, New York (USA)
Jobright.ai
Machine Learning Engineer, Model Development — Entry Level
New
USA
Jobright.ai
Data Engineer (Early Career) (Canada)
New
Canada
Bevertec
Senior Machine Learning Engineer
New
RemoteCanada
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.