- Data Engineer
- Experience Level: 5–7 Years
- Department: Data Engineering / Data & Analytics
- About the Role
- We are looking for an experienced Data Engineer with 5–7 years of hands-on experience building and optimizing large-scale data pipelines and platforms. The ideal candidate has deep expertise in Databricks on AWS (Azure Databricks experience also considered), along with strong skills in Apache Airflow, Python, PySpark, and SQL. You will play a key role in designing, building, and maintaining scalable data infrastructure that powers analytics, reporting, and data science initiatives across the organization.
- Key Responsibilities
- Design, develop, and maintain robust, scalable, and efficient ETL/ELT data pipelines using Databricks, PySpark, and SQL.
- Build and orchestrate complex data workflows using Apache Airflow, including DAG design, scheduling, monitoring, and error handling.
- Work extensively within the Databricks on AWS ecosystem (Delta Lake, Unity Catalog, Databricks Workflows, cluster/job optimization); Azure Databricks experience is a plus.
- Write clean, efficient, and reusable Python code for data transformation, automation, and pipeline development.
- Optimize SQL queries and data models for performance, scalability, and cost efficiency.
- Design and implement data lakehouse architectures leveraging Delta Lake best practices (schema evolution, partitioning, Z-ordering, vacuuming, etc.).
- Integrate data from multiple sources (APIs, databases, streaming platforms, third-party systems) into centralized data platforms.
- Ensure data quality, integrity, and governance through validation frameworks, monitoring, and alerting.
- Collaborate closely with Data Analysts, Data Scientists, and Business stakeholders to understand data requirements and deliver reliable datasets.
- Implement and maintain CI/CD pipelines for data engineering workflows (e.g., using Git, Jenkins, GitHub Actions, or similar).
- Monitor and troubleshoot production data pipelines, ensuring high availability and minimal downtime.
- Contribute to architectural decisions around cloud infrastructure, cost optimization, and data platform scalability.
- Document technical designs, data flows, and operational runbooks.
- Mentor junior data engineers and contribute to best practices, coding standards, and design patterns within the team.
- Required Skills & Experience
- 5–7 years of overall experience in Data Engineering roles.
- Strong hands-on experience with Databricks (AWS preferred; Azure Databricks acceptable) — including Delta Lake, cluster management, job scheduling, and notebook-based development.
- Proficiency in Apache Airflow for workflow orchestration — DAG authoring, sensors, operators, and custom plugins.
- Strong programming skills in Python, with experience writing production-grade, modular, and testable code.
- Deep expertise in PySpark for distributed data processing, including performance tuning and optimization techniques.
- Advanced SQL skills — complex joins, window functions, query optimization, and data modeling (dimensional modeling, star/snowflake schemas).
- Solid understanding of AWS cloud services relevant to data engineering (S3, IAM, EC2, Glue, Lambda, EMR, Redshift, etc.); Azure equivalents (ADLS, ADF, Synapse) a plus.
- Experience with Delta Lake concepts — ACID transactions, time travel, schema enforcement/evolution.
- Familiarity with version control (Git) and CI/CD practices for data pipelines.
- Understanding of data warehousing concepts, data lake architectures, and modern lakehouse paradigms.
- Experience working with structured, semi-structured, and unstructured data (JSON, Parquet, Avro, CSV, etc.).
- Strong debugging, performance tuning, and problem-solving skills in distributed data processing environments.
- Good understanding of data governance, security, and compliance practices
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
Senior Data Engineer
Acumenz Consulting
Similar Engineering Jobs
View all Engineering jobs→name
Senior Machine Learning Engineer, ML Efficiency
New
RemoteUSA$216,700 - $303,400
name
Data Platform Engineer
New
RemoteUSA$120,000 - $160,000
Meta
Construction Manager - Data Center Design, Engineering, & Construction
New
Eagle Mountain, Utah (USA)$123,000 - $176,000
Vista Bank Tchad
Data Developer
New
RemoteUSA
hackajob
ML Engineer (Coding Agent Experience) (Train AI Models Part Time!)
New
USA
iPivot
Data Engineer with Data Mining & Python || W2 Only || Remote
New
California (USA)
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.