About The Role
You will engineer the critical path of CYBERIA's Data Platform — ingestion, transformation, feature pipelines, and vector retrieval — with SLOs for latency, freshness, and quality.
Responsibilities
- Build and operate batch and streaming pipelines with offline + online parity
- Implement feature stores with sub-50 ms p99 reads and point-in-time correctness
- Develop pipeline designers that turn documents (PDF, DOCX) into structured, AI-ready data
- Instrument everything: lineage, data quality checks, and cost telemetry by default
- Harden prototypes from applied ML teams into production-grade platform features
Requirements
- 4+ years engineering production data systems
- Strong Python plus SQL; experience with Spark, dbt, Airflow/Dagster, or similar
- Familiarity with vector databases, embeddings, or retrieval systems a strong plus
- Comfort with distributed systems, observability, and on-call
Benefits
- Fully remote within the US
- Competitive base + equity
- Learning and conference budget
- Premium health benefits and 401(k) match
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
Data Engineer
CYBERIA Software
Similar Engineering Jobs
View all Engineering jobs→Meta
Capacity Engineer, Infrastructure Data Governance
New
Menlo Park, California (USA)$184,000 - $257,000
Meta
Software Engineer, Systems ML - Compilers / Backend
New
Austin, Texas (USA)$154,003 - $217,000
Meta
IP Validation Engineer - Machine Learning Accelerators
New
Austin, Texas (USA)$178,000 - $250,000
Meta
Data Engineer, Teen & Family Experience (TFE)
New
New York, New York (USA)$147,000 - $208,000
CACI International Inc
Reliability and Maintainability Engineering Analyst
New
Crane, Indiana (USA)
SPECTRAFORCE
Machine Learning Engineer
New
San Francisco, California (USA)$180,000 - $230,000
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.