- Sigmatic builds an operating system for surgery centres. We pull operational data out of the fragmented systems a facility runs on — clinical, financial, scheduling, supply chain — and turn it into something AI agents can answer questions from.
- That last part shapes the whole job. Our primary consumer is a set of agents, not a BI tool. When a model is subtly wrong the error does not surface as a broken chart; it gets stated confidently to a customer who has no way to check it. So the standard for correctness here sits higher than typical reporting work, and the proof has to be built in rather than inspected afterwards.
- You would own new sources end to end: assess the source, map it into our canonical model, build the extraction and loading, write the dbt models it lands in, produce the semantic metadata the agents read, and prove the whole path with tests.
- What makes this interesting
- We would rather be straight about the difficulty than sell you a tidy version of it.
- • Conformance across rival sources. Two clinical systems describe the same surgery in incompatible ways. Both map into one canonical model, and neither maps to the other. Get this wrong and you produce marts that look perfectly healthy and are wrong.
- • Multi-tenancy with teeth. Schema-per-tenant Postgres, where tenants differ in source systems, enabled capabilities and fiscal calendars. A surrogate key that quietly drops the tenant dimension will merge two customers' data. We have shipped that bug three times, which is why the tests guarding against it are the ones we care most about.
- • Sources that fight back. A QuickBooks Desktop connector speaking SOAP to software on somebody's office PC, which once returned a 111 MB response and took the sync down for five days. Vendor APIs with 30× latency variance between identical calls. Reports whose row counts legitimately change depending on where you place the window boundary.
- • Documentation as an interface. Marts feed a generated catalogue that agents read to decide what to query and how to interpret it. Your column descriptions, declared grain and documented units are what the agent reasons over. A vague description is a model the agent will use wrongly, and no test will catch it.
- How we build
- Most of our code is now written with AI assistance, including substantial autonomous work. The scarce skill on this team is no longer producing code. It is specifying work precisely enough to delegate, and verifying output rigorously enough to trust.
- In practice that means we ablate tests rather than admiring them: delete the guard, confirm the test goes red. A test that still passes with the code removed is worse than no test. It means a bug fix starts by proving the test fails against the unfixed code. And it means that when we do not know how a vendor behaves, we probe it and write down the number instead of reasoning about it.
- You do not need experience with any particular tool. You do need to be comfortable in a codebase where much of the diff was not typed by a human, and to have real opinions about how to verify it. If that is already how you work you will move fast here. If it is not, this will be a frustrating role.
- Our stack
- • Transformation — dbt on PostgreSQL: staging, intermediate and marts, with contracts and tests enforced in CI
- • Warehouse — multi-tenant PostgreSQL on AWS RDS, schema per tenant
- • Extraction and loading — Python on AWS Lambda, Glue PySpark, Step Functions, EventBridge, with SAM and Terraform
- • Sources — EHR and clinical systems, accounting (QuickBooks, NetSuite), scheduling, supply chain, IoT and scanner event streams
- • AI layer — AWS Bedrock, a generated semantic catalogue, Qdrant retrieval, Python agent services
- • Engineering — GitHub with gitflow, mandatory review, CI gates, Linear, SOC 2 Type II
- What we are looking forRequired
- • Five or more years in data engineering or analytics engineering, having owned pipelines a business depended on
- • Strong SQL and Python: production transformations, testing, debugging, performance work
- • Serious dbt experience. Multi-tenant or multi-source projects are what we most want to hear about
- • ETL and ELT against messy real sources: APIs, databases, files, cloud storage, with full and incremental loads, CDC and idempotent reruns
- • A way of working that already assumes AI assistance, and a clear account of how you verify what you did not write
- • Precision in writing. Much of your output is read by a model before a person sees it, so unambiguous descriptions of what a column means are a deliverable, not an afterthought
- • A record of finishing: edge cases, validation, documentation, deployment, production support
- Strong plus
- • AWS data stack with operational ownership: Glue, Lambda, Step Functions, EventBridge
- • Multi-tenant SaaS platforms where customer-specific schemas map into a shared model
- • Layered architecture with write-audit-publish or equivalent quality gating
- • Data products consumed by LLMs or agents: semantic layers, catalogues, retrieval
- Useful
- • Healthcare operational, revenue-cycle, claims, scheduling or EHR-adjacent data; HIPAA-aware architecture, de-identification, least privilege
- • Accounting and financial system data, operational KPI development, IoT or time-series
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
Data Engineer
Sigmatic
Similar Engineering Jobs
View all Engineering jobs→Curate Partners
Data Engineer
New
Boston, Massachusetts (USA)
Jobright.ai
Machine Learning Engineer, Model Development — Entry Level
New
USA
Concept Reply US
Data Engineer - GenAI
New
Chicago, Illinois (USA)
Drift AI
Senior Software Engineer, Data Governance
New
San Francisco, California (USA)
Drift AI
Engineering Manager, Data
New
San Francisco, California (USA)
HNI Corporation
Manager, Data Engineering - Platform
New
Chicago, Illinois (USA)
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.