Data First Jobs

Central Business Solutions Inc.

Principal Data Engineer – (Databricks, PySpark, AWS, EMR, Airflow)

Contract · In Office · Irvine, California (USA)

Posted Aug 14, 2026

Work Options
Cloud Stack
Industry
Positions
Job Type
Position Group
  • Position: Principal Data Engineer – (Databricks, PySpark, AWS, EMR, Airflow)
  • Location: Irvine, CA - Hybrid
  • Position Type: Contract to Hire

Responsibilities

  • You will be responsible for all aspects of data acquisition, data transformation, analytics scheduling and operationalization to drive high-visibility, cross-division outcomes. Investigate, evaluate, test and recommend technical solutions for future systems.
  • They will support software developers, database architects, data scientists on data initiatives and will contribute optimal data delivery architecture.
  • What you will be doing:
  • Data Operations
  • Lead the creation of data environments and/or data sets to serve a wide range of data users, including but not limited to Data Scientists, Data Analysts, Business Analysts etc.
  • Perform offline analysis of large data sets using components of a big data software ecosystem.
  • Validate the solution of root cause analysis escalated by various technical staff in multiple organizations and with differing levels of expertise.
  • Investigate, evaluate, test and recommend technical solutions for future systems.
  • Data Management
  • Own product data sets from the definition phase through to production deployment (end-to-end).
  • Provide solutions for the design and implementation of Hadoop EMR Cluster/ Big Data Infrastructure.
  • Deploy Hadoop/Big Data/Spark and database storage Infrastructures in AWS cloud.
  • Monitor HDFS/Hadoop/Spark and related software releases, third-party utilities with emphasis on overall system performance.
  • Lead and develop tools and procedures to monitor and automate system tasks on servers and clusters
  • Lead and collaborate with other teams to design, develop, and deploy data tools that support both operations and product use cases.
  • Data Design
  • Lead and design distributed, scalable, and reliable data pipelines that ingest and process data at scale and in real-time.
  • Lead and design big data technologies and prototype solutions to improve data processing architecture.

Qualifications

  • What we are looking for:
  • Bachelor’s degree in computer science, computer engineering, or a related technical field.
  • 12+ years of professional experience as a data software engineer; or 15+ years of related experience as a data software engineer in lieu of 4-year degree.
  • 2+ years of experience with AWS cloud or other cloud Big Data computing design, provisioning, and tuning.
  • Related AWS certification, preferred.
  • Previous experience as a Data Engineer / Database Administrator and/or Business Intelligence Analyst.
  • Expertise in database concepts, object and data modeling techniques and design principles.
  • Expertise in database architectures, software, and facilities
  • Expertise with programming languages - Python (required), Scala, Ruby, R Database technologies - SQL, performance tuning concepts, AWS RDS, RedShift, MySQL
  • Expertise with big data batch processing tools: Hadoop MapReduce, ElasticSearch, PIG, Hive, Cascading/Scalding, Apache Spark, AWS EMR
  • Expertise with stream-processing systems: Kinesis, Kafka, MQTT
  • Expertise with relational NoSQL databases including DyanamoDB
  • Expertise in writing JSON, XML, YAML and other data definition schemas
  • Excellent verbal and written communication skills necessary to effectively collaborate in a team environment and present and explain technical information and provide advice to management.
  • Ability to work on advanced complex technical projects or business issues requiring state of the art technical knowledge or industry.
  • Ability to work on significant and unique issues where analysis of situations or data requires an evaluation of intangibles.
  • Exercises independent judgment in methods, techniques and evaluation criteria for obtaining results.
  • Ability to lead and mentor junior engineers and colleagues.

Mention you found this on Data First Jobs — it helps us bring you more roles like this.

Principal Data Engineer – (Databricks, PySpark, AWS, EMR, Airflow)

Central Business Solutions Inc.

Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.

Free, no spam. Unsubscribe anytime.