Data First Jobs

Bitus Labs

Data Engineer – AWS Lakehouse (Mandarin Required)

Full Time · In Office · Irvine, California (USA)

Posted Aug 19, 2026

Job description:

  • About the Role
  • We are looking for a mid-level Data Engineer to join our Data Platform team and take ownership of building and scaling our AWS-based data lakehouse. You will architect and deliver robust, production-grade data pipelines, work closely with data scientists, analytics engineers, and product teams, and set the technical direction for how data flows across the organization. This is a hands-on engineering role — you will write production code in Java and Python every day, while also contributing to platform design decisions, mentoring junior engineers, and driving best practices around data quality, reliability, and governance.
  • Key Responsibilities
  • Data Lakehouse Development
  • Build and extend medallion-architecture data lakehouse layers (Bronze /
  • Silver / Gold) on AWS S3 using the Apache Iceberg table format.
  • Develop and maintain high-throughput ETL/ELT pipelines using AWS
  • Glue, EMR (Spark), and Lambda.
  • Implement schema evolution, partitioning strategies, and compaction
  • processes for Iceberg tables to optimize storage and query performance.
  • Write production-quality pipeline code in Java and Python, following team
  • conventions for structure, testing, and maintainability.
  • Real-Time & Batch Streaming
  • Build and operate event-driven data pipelines using Amazon Kinesis Data
  • Streams, Kinesis Firehose, or Apache Kafka (MSK).
  • Implement exactly-once or at-least-once processing semantics for
  • streaming workloads using Apache Flink or Spark Structured Streaming on EMR.
  • AWS Platform Engineering
  • Contribute infrastructure-as-code definitions using AWS CDK or
  • Terraform for repeatable, auditable data platform deployments.
  • Monitor and help tune cost and performance across AWS services
  • including S3, Glue, Athena, Redshift Spectrum, EMR, Lambda, Step Functions, and
  • EventBridge.
  • Follow platform security standards in day-to-day work: IAM least-
  • privilege policies, KMS encryption, and VPC networking.
  • Maintain and extend CI/CD pipelines for data workloads using AWS
  • CodePipeline, GitHub Actions, or equivalent.
  • Data Quality & Governance
  • Add data quality checks using frameworks such as Great Expectations or
  • Deequ, and integrate validation steps into pipeline orchestration.
  • Help maintain and uphold data contracts between producing and
  • consuming systems.
  • Contribute to data cataloguing and lineage tracking using AWS Glue Data
  • Catalog or Apache Atlas.
  • Collaboration & Ways of Working
  • Partner with data scientists, ML engineers, and analysts to understand data
  • requirements and deliver performant, well-documented datasets.
  • Take an active part in code reviews, design discussions, and pair
  • programming, and support junior engineers where you can.
  • Document your pipeline designs and decisions, and contribute to the
  • internal engineering knowledge base.
  • Required Qualifications
  • Experience
  • 3+ years of professional data engineering experience, with at least 1–2
  • years on AWS cloud platforms.
  • Experience building and supporting production data pipelines, ideally on
  • large datasets with defined freshness or reliability SLAs.
  • Working knowledge of data lakehouse concepts — medallion pattern and
  • open table formats (Iceberg preferred; Delta Lake or Hudi acceptable).
  • Programming Languages
  • Java: Working proficiency in Java (8+) for Spark jobs and pipeline
  • components, with familiarity with Maven or Gradle build systems.
  • Python: Solid Python 3 skills for AWS Glue scripts, orchestration logic,
  • data quality checks, and automation tooling. Experience with pandas, PySpark, and
  • boto3. Strong skills in one language and a willingness to ramp up on the other is
  • acceptable.
  • AWS Core Services
  • Storage & Compute: S3, Glue (jobs, crawlers, Data Catalog), EMR
  • (Spark/Flink), Lambda, EC2.
  • Streaming: Kinesis Data Streams, Kinesis Firehose, or MSK (Managed
  • Kafka).
  • Orchestration: Step Functions, MWAA (Managed Airflow), or
  • EventBridge Scheduler.
  • Querying: Athena, Redshift, or Redshift Spectrum.
  • Security & Governance: Comfortable working within IAM, KMS, Secrets
  • Manager, and VPC setups.
  • DevOps: Exposure to AWS CDK or CloudFormation, and to CodePipeline
  • or equivalent CI/CD tools.
  • Data Processing Frameworks
  • Apache Spark (PySpark and/or Spark Java API) — distributed
  • transformations and a working grasp of performance tuning.
  • Apache Iceberg — reading and writing tables, time travel, and basic table
  • maintenance.
  • SQL — strong SQL for data transformation, including window functions,
  • CTEs, and query tuning.
  • Must be Chinese Mandarin fluent.
  • Preferred Qualifications
  • AWS Certified Data Engineer – Associate or AWS Certified Solutions
  • Architect certification.
  • Experience with dbt for SQL-based transformation layers on top of the
  • lakehouse.
  • Familiarity with ML platform integration: feature stores (SageMaker
  • Feature Store), model serving data needs, or MLflow experiment tracking.
  • Experience with real-time OLAP engines such as Apache Druid or
  • ClickHouse.
  • Experience with Lake Formation fine-grained access control, or with data
  • cataloguing and lineage tooling.
  • Exposure to data mesh or data product thinking — domain ownership and
  • data contracts.
  • Tech Stack at a Glance
  • Languages
  • Java/ Python 3
  • Cloud Platform
  • AWS (S3, Glue, EMR, Kinesis, Athena, Lambda, Step Functions, Lake Formation, CDK)
  • Processing
  • Apache Spark, Apache Flink, Spark Structured Streaming
  • Table Format
  • Apache Iceberg (primary), Delta Lake / Hudi (familiarity)
  • Streaming
  • Amazon Kinesis, MSK (Kafka), Kinesis Firehose
  • Orchestration
  • Apache Airflow (MWAA), AWS Step Functions
  • IaC & CI/CD
  • AWS CDK / Terraform, GitHub Actions / CodePipeline
  • Job Type: Full-time
  • Benefits:
  • 401(k)
  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Life insurance
  • Paid time off
  • Parental leave
  • Retirement plan
  • Vision insurance

Mention you found this on Data First Jobs — it helps us bring you more roles like this.

Data Engineer – AWS Lakehouse (Mandarin Required)

Bitus Labs

Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.

Free, no spam. Unsubscribe anytime.