- Title: Data Engineer – Macro Financial Data & Analytics
- Location: Parsippany, NJ (Hybrid) Need Local!!
- Onsite Interview!!
Requires strong expertise in Databricks and PySpark
- About the Role
- We are seeking an experienced Lead Data Engineer with a blend of approximately 70% data engineering and 30% analytical responsibilities. This individual will have strong proficiency in modern cloud-native data platforms, large-scale data processing, and the ability to contribute meaningfully to analytical and data science workflows.
- The ideal candidate combines deep Databricks and PySpark expertise with finance or payroll domain experience and the analytical fluency to partner closely with data scientists on macroeconomic and financial data research.
- This role will work across production pipeline development, platform optimization, exploratory data analysis, and close collaboration with data scientists, economists, and business stakeholders to build and maintain scalable data platforms for financial analytics and research.
- The right candidate will be comfortable moving fluidly between engineering and analytical work — building production ETL pipelines, exploring and profiling new datasets, and helping translate complex analytical requirements into scalable solutions.
- What You'll Do
- Data Engineering & Platform Development
- Design, develop, and maintain scalable data pipelines for the ingestion, transformation, and distribution of payroll, macroeconomic, and financial datasets.
- Build and support Databricks-based platforms that enable financial research and analytical workloads.
- Implement ETL/ELT frameworks using PySpark and Delta Lake for structured and unstructured data from internal and external sources.
- Develop data models and data marts optimized for analytical, reporting, and machine learning use cases.
- Ensure data quality, consistency, lineage, governance, and observability across data assets.
- Optimize performance for large-scale datasets, including billions of records, multi-terabyte environments, and time-series data.
- Translate business requirements into scalable and maintainable data solutions.
- Analytical & Data Science Support
- Perform exploratory data analysis, including profiling datasets and identifying distributions, outliers, missing patterns, and data drift.
- Translate data scientist logic into efficient, scalable PySpark implementations, including cross-sectional metrics and time-windowed aggregations.
- Build validation dashboards and exploratory notebooks to verify pipeline outputs and data quality.
- Support feature engineering by implementing complex aggregation and transformation logic at scale.
- Independently verify and sanity-check analytical outputs and identify results that warrant further investigation.
- Conduct ad hoc analytical work using pandas and NumPy alongside PySpark to support research initiatives.
- Modern Data Architecture
- Work within established lakehouse architecture using Databricks and Delta Lake.
- Contribute to architecture design discussions and evaluate tradeoffs related to catalog design, medallion architecture, and data mesh concepts.
- Implement CI/CD pipelines using Databricks Asset Bundles, Bitbucket Pipelines, and Jenkins.
- Manage Unity Catalog governance, access patterns, and schema design.
- Ensure security, compliance, and data governance standards are maintained.
- AI-Assisted Development
- Leverage AI coding tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent platforms to accelerate development.
- Critically review AI-generated code for correctness, performance, scalability, and maintainability.
- Integrate AI-assisted workflows into day-to-day engineering and analytical activities.
- Required Qualifications
- Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, Statistics, Economics, Finance, or a related field.
- 5+ years of experience in Data Engineering or Data Platform development.
- Finance or payroll domain experience, including familiarity with payroll data structures, pay-period logic, compensation and deduction relationships, or financial-services data environments.
- Experience handling large-scale datasets, including billions of records, multi-terabyte environments, and time-series data.
- Strong Databricks proficiency with a deep understanding of:
- Unity Catalog, including governance, access patterns, and catalog/schema design.
- Delta Lake internals, including optimization, clustering, change data feed, and versioning.
- Databricks Workflows, including orchestration, dependencies, and monitoring.
- Databricks Asset Bundles or equivalent deployment frameworks.
- Experience participating in architecture-level design decisions and evaluating technical tradeoffs.
- Strong proficiency in:
- Python
- SQL
- PySpark
- Data Modeling
- ETL/ELT Development
- Analytical fluency, including exploratory data analysis, basic statistical concepts, distributions, correlations, time-series patterns, and feature engineering.
- Proficiency with pandas and NumPy for ad hoc analytical work alongside production PySpark.
- CI/CD implementation experience using tools such as Bitbucket Pipelines, Jenkins, and automated deployment frameworks.
- Experience with AI-assisted development tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent.
- Experience implementing data quality and validation frameworks.
- Preferred Qualifications
- Experience with macroeconomic, capital markets, or financial-services data.
- Exposure to data architecture design, including lakehouse patterns, data mesh concepts, and medallion architecture.
- Experience with streaming and event-driven pipelines using Kafka or Structured Streaming.
- Census or geographic data processing experience, including TIGER and FIPS codes.
- Infrastructure-as-code experience with Terraform, CDK, or similar technologies.
- Databricks certifications at the Associate or Professional level.
- Experience migrating legacy data platforms to modern technology stacks, including Glue, EMR, or HDInsight to Databricks.
- Experience supporting machine learning and AI-driven analytics solutions.
- Data visualization experience using Power BI, Tableau, Databricks Dashboards, or similar platforms.
- Scala experience is a plus.
- Technical Skills
- Programming: Python, SQL, PySpark, Scala preferred
- Data Engineering & Analytics: Databricks, Apache Spark, Delta Lake, Unity Catalog, Databricks Asset Bundles, Kafka
- Databases: SQL Server, PostgreSQL, Delta Tables, NoSQL databases
- Visualization & Reporting: Power BI, Databricks Dashboards, Python visualization libraries including Matplotlib and Plotly
- DevOps & Automation: Git, Bitbucket Pipelines, Jenkins, Terraform, CI/CD pipelines, Databricks Asset Bundles
- Key Competencies
- Comfortable moving between engineering and analytical work, from building production ETL pipelines to profiling and investigating datasets.
- Strong analytical and problem-solving skills with statistical literacy.
- Builds reliable, scalable, and maintainable pipelines with consideration for long-term ownership and support.
- Uses AI tools effectively to accelerate engineering and analytical workflows.
- Has enough domain knowledge to identify questionable data, investigate anomalies, and propose data-driven improvements.
- Can independently take an exploratory analysis from an initial dataset through findings and recommendations.
- Excellent communication and documentation skills.
- Strong understanding of data governance, data quality, and metadata management.
- Ability to work effectively in a fast-paced, data-driven environment.
- Nice-to-Have Domain Experience
- Payroll data structures and processing
- Macroeconomic analysis and forecasting
- Financial markets and alternative data
- Large-scale time-series data engineering
- Census and geographic data processing
- What We're Looking For
- We're looking for someone who brings strong technical depth while remaining curious, collaborative, and comfortable solving ambiguous problems. This person should be able to take ownership of complex data initiatives, work closely with technical and business stakeholders, and develop scalable solutions that support sophisticated financial and economic analysis.
- The successful candidate will combine hands-on engineering expertise with analytical thinking and an ability to understand the broader business and research questions behind the data.
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
Data Engineer
Marlabs
Similar Engineering Jobs
View all Engineering jobs→TriWest Healthcare Alliance
Data Engineer Sr.
New
RemoteUSA
Value Innovation Labs
Data Engineering & Reporting Developer (GCP)
New
Texas (USA)
Accenture Federal Services
Data Operations Engineer
New
Arlington, Virginia (USA)
Accenture Federal Services
Senior Data Operations Engineer
New
Washington, District of Columbia (USA)
VeriiPro
Data Engineer
New
Chicago, Illinois (USA)
RemoteHunter
Sr. Data Engineer
New
USA
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.