Data First Jobs

SARACA

Data Network Engineer

Contract · In Office · Denver, Colorado (USA)

Posted Sep 30, 2026

  • Company Description SARACA is a global engineering R&D services company supporting 25+ Fortune 500 customers across industries such as MedTech, aerospace, rail, automotive, semiconductor, defense, and farm equipment. The company is ISO 13485 certified and has extensive experience in medical device design and development, including embedded software, UI/UX, mechanical systems, and product testing. SARACA offers specialized expertise in IEC 62304, EU MDR, and a wide range of scientific and regulatory documentation for MedTech. A leadership team with deep product engineering experience and a technically strong team of 400+ engineers and consultants deliver solutions to complex onsite and offsite projects worldwide. SARACA is an equal-opportunity employer that promotes innovation, learning, and customer success through faster business growth and industry leadership.
  • ob Title: Data Developer / Data Engineer
  • Location: Denver, CO – Onsite
  • Start Date: Immediate
  • Background Check: Mandatory
  • Work Arrangement: Candidate must work from the customer's Denver, CO office from Day 1.
  • Job Summary
  • We are seeking a skilled Data Developer / Data Engineer to join the ETL and Operations team responsible for designing, developing, optimizing, and supporting large-scale data pipelines on AWS.
  • The engineer will work with data from diverse sources including network telemetry, Wi-Fi platforms, billing, and provisioning systems, processing billions of records per day to support analytics and data-driven solutions.
  • The ideal candidate should have strong hands-on experience with AWS, PySpark, Spark SQL, Athena/Trino/Presto, Python, ETL pipelines, and Airflow/MWAA, along with the ability to troubleshoot and optimize production data workflows.
  • Key Responsibilities
  • Design, develop, and maintain scalable data ingestion pipelines from network, Wi-Fi, billing, and provisioning data sources.
  • Build and enhance production ETL jobs running on AWS.
  • Develop and optimize Spark SQL and PySpark transformations using AWS EMR at large data volumes.
  • Write and optimize Athena/Trino/Presto SQL for data analysis, validation, troubleshooting, and stakeholder requirements.
  • Develop scalable solutions for processing billions of records per day across partitioned data lakes.
  • Perform end-to-end pipeline optimization, including partition pruning, join strategies, shuffle optimization, and Adaptive Query Execution (AQE).
  • Contribute to migration from legacy AWS Step Functions orchestration to Apache Airflow / AWS MWAA DAGs and a centralized Python job framework.
  • Monitor production data pipelines and troubleshoot job failures, data quality issues, and performance problems.
  • Implement fixes, enhancements, and performance improvements across existing ETL workflows.
  • Participate in GitLab merge request reviews, branching strategies, peer reviews, and CI/CD deployment processes.
  • Collaborate closely with data analysts, data scientists, architects, and other engineering teams.
  • Develop Bash/Linux automation scripts and support data processing environments running on AWS EMR.
  • Learn and adopt new technologies and engineering practices as the data platform evolves.
  • Required Technical Skills
  • Strong hands-on SQL experience, particularly Spark SQL and Athena (Trino/Presto).
  • Strong experience with AWS data engineering services, including:
  • AWS EMR
  • Amazon S3
  • AWS Glue Data Catalog
  • Amazon Athena
  • AWS Lambda
  • AWS Step Functions
  • Apache Airflow / AWS MWAA
  • Strong Python development and scripting skills.
  • Hands-on experience with PySpark and distributed data processing.
  • Experience building and tuning large-scale ETL/data pipelines.
  • Experience working with billions of records per day and partitioned data lakes.
  • Strong understanding of Parquet and S3-based data lake architectures.
  • Experience with query and Spark performance optimization, including:
  • Partition pruning
  • Join strategies
  • Shuffle optimization
  • AQE tuning
  • Experience with Git/GitLab, branching, merge requests, code reviews, and CI/CD workflows.
  • Strong Linux/Bash scripting experience.
  • Excellent troubleshooting, communication, and collaboration skills.
  • Preferred Qualifications
  • Experience with Apache Iceberg, including table formats, MERGE operations, and Spark-managed DDL.
  • Knowledge of dimensional data modeling and Slowly Changing Dimensions (SCD).
  • Experience with Medallion Architecture such as Bronze, Silver, and Gold layers.
  • Familiarity with Hadoop/Hive/HiveQL and SQL-on-Hadoop environments.
  • Exposure to AI/ML data pipelines.
  • Experience using approved AI-assisted development tools such as Amazon Q, Kiro, or GitLab Duo.
  • Experience working in highly technical, cross-functional data engineering environments.
  • Education & Experience
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience.
  • 2+ years of hands-on experience in Data Engineering, ETL development, or a related data engineering role.
  • Experience supporting production-grade data pipelines in a cloud environment is highly preferred.
  • Ideal Candidate Profile
  • The ideal candidate is a hands-on AWS Data Engineer / Data Developer with strong experience in PySpark, Spark SQL, Athena/Trino/Presto, Python, and AWS EMR/S3.
  • Candidates should have experience working with large-scale data lakes and high-volume ETL pipelines, along with strong troubleshooting and performance-tuning capabilities. Experience with Airflow/MWAA, GitLab CI/CD, Linux/Bash, and AWS data services will be important for success in this role.

Mention you found this on Data First Jobs — it helps us bring you more roles like this.

Data Network Engineer

SARACA

Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.

Free, no spam. Unsubscribe anytime.