- Big Data & Spark Engineer
- Duration – 1 year
- Location – Toronto
- Work Model – Hybrid (4 days work from office)
- Key Responsibilities
- Big Data & Spark
- Design and develop Spark-Scala applications for large-scale data processing on Hadoop/CDP clusters
- Build and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL
- Tune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization)
- Migrate Spark 2 applications to Spark 3 on Cloudera CDP platforms
- Work with Parquet, ORC, Avro file formats on HDFS
- SQL & Data Engineering
- Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations
- Design and maintain Hive external/managed tables and partitioned datasets
- Optimize slow-running queries and resolve correlated subquery issues
- Work with HDFS encryption zones and data governance requirements
- Unix / Shell Scripting
- Develop and maintain bash shell scripts for job orchestration and automation
- Handle error management, return codes, logging and alerting in shell scripts
- Manage HDFS operations (hdfs dfs commands), file transfers, and data validation
- Manage Kerberos authentication (kinit, keytab handling)
- API Extraction & Integration
- Build scripts and pipelines to extract data from REST APIs using curl and Python
- Handle OAuth2 token generation, bearer token refresh and API health checks
- Parse and process JSON API responses and load into HDFS/Hive
- Manage pagination, error handling and retry logic for API calls
- Work with enterprise API gateways and URL parameter construction
- AI & Copilot Capabilities
- Leverage GitHub Copilot / AI coding assistants to accelerate development
- Use AI tools for code review, SQL generation, script debugging and documentation
- Contribute to AI-assisted data quality and anomaly detection pipelines
- Explore and implement LLM-based automation for repetitive data engineering tasks
- Scheduling & Orchestration
- Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cron
- Build and maintain Ansible playbooks for automated deployments
- Manage deployment pipelines including artifact versioning, Vault secret injection and environment-specific configuration
- Monitor job health, handle failures and implement alerting
- Nice to Have
- Experience with Cloudera CDP (7.x) and migration from HDP
- Knowledge of Kerberos, Vault, HDFS encryption zones
- Familiarity with CI/CD pipelines (Helios, GitHub Actions)
- Experience with MSSQL / JDBC connectivity from Spark
- Understanding of AML / Financial regulatory data domains
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
Big Data & Spark Engineer
TheCorporate
Similar Engineering Jobs
View all Engineering jobs→Americon Corporation for Public Well Being
Senior BI Developer
New
Milwaukee, Wisconsin (USA)
Enhance IT
Data Engineer
New
Mississippi (USA)
Rise up movement Congo
Data Developer
New
RemoteUSA
Cardinal Health
Sr Engineering Analyst - Project Management - Strategic Commercialization
New
RemoteUSA$68,500 - $97,800
Bright Vision Technologies
Data Engineering Specialist – AI
New
North Carolina (USA)
Job Discovery Platform - CyOpsPath
Data Engineer
New
USA
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.