The New York Consulting Group
AI Data Architect
Full Time · In Office · New York, New York (USA)
Posted Jul 28, 2026
We are seeking a Senior AI Data Architect to define the data architecture powering AI/ML products from raw ingestion through training, inference, retrieval, and model-serving. In this full-time, hybrid role, you will own architectural direction rather than simply operating pipelines: you will establish the contracts, platform primitives, governance controls, and operating models that make AI data reliable, explainable, secure, and reusable.
You will design production architectures spanning batch and streaming ingestion, raw/bronze, standardized/silver, curated/gold, feature, embedding, retrieval, and model-serving layers. Your work will enable enterprise RAG over knowledge bases, policies, technical documentation, contracts, tickets, and internal operational content; customer-support copilots; semantic search; recommendation and personalization; fraud, risk, anomaly, and abuse detection; document intelligence; forecasting; classification; ranking; propensity models; and agentic AI systems with authorization-aware context and traceable execution history.
The role requires deep ownership of vector retrieval using Pinecone, Weaviate, and pgvector, including index and collection design, metadata filtering, tenancy, lifecycle management, migration strategy, embedding-model versioning, hybrid lexical-vector retrieval, reranking, citation provenance, and stale-index controls. You will also architect feature stores with offline historical storage and online low-latency serving, point-in-time correct joins, materialization, feature freshness, training-serving skew detection, and reproducible backfills.
You will partner with ML engineers, data engineers, platform engineering, product, security, and leadership to translate freshness, relevance, explainability, privacy, availability, and cost requirements into measurable architecture decisions. We help place you into paid roles via our referral network, giving you the opportunity to apply senior-level architecture judgment to consequential AI data platforms.
Key Responsibilities
Architect end-to-end AI/ML data flows across batch and streaming ingestion, object storage, relational databases, warehouses/lakehouses, raw/bronze, standardized/silver, curated/gold, feature, embedding, retrieval, and model-serving layers, explicitly separating offline training paths from online inference paths.
Define ingestion contracts, landing zones, schema registries, partitioning, watermarks, deduplication, idempotency, replay, backfills, dead-letter handling, and late-arriving-data behavior for CDC, change streams, JDBC/ODBC, bulk exports, Kafka, and event-driven integrations.
Design governed data contracts for source, curated, feature, embedding, retrieval, and model-serving datasets, including schema evolution, ownership, SLAs, quality rules, backward compatibility, stable identifiers, and metadata for documents, chunks, tenants, permissions, sources, timestamps, ACLs, embeddings, models, feature values, and retrieval results.
Build RAG data flows covering document acquisition, parsing, normalization, semantic chunking, overlap, parent-child relationships, metadata and ACL propagation, embedding generation, batch/stream indexing, hybrid lexical-vector retrieval with Elasticsearch/OpenSearch, reranking, context assembly, citation/provenance, and evaluation.
Architect Pinecone, Weaviate, and pgvector implementations, making explicit decisions about collections, indexes, schemas, namespaces or tenants, distance metrics, HNSW/IVF-style indexes, dimensionality compatibility, metadata filtering, upserts, deletes, compaction, re-indexing, embedding-model migration, dual writes, shadow reads, reconciliation, cutover, and rollback.
Define retrieval correctness and operational behavior, including top-k selection, filtering before or after ANN search, score thresholds, diversity or maximal marginal relevance, freshness, stale-index handling, fallback behavior, tenant isolation, and document-level authorization enforcement at retrieval time.
Design feature-store architecture using tooling such as Feast, with offline historical storage, online serving through Redis, DynamoDB, or Cassandra, entity keys, point-in-time correct joins, materialization jobs, feature freshness, versioning, backfills, offline/online parity, training-serving skew detection, and drift monitoring.
Establish training-data and evaluation pipelines for supervised fine-tuning, preference data, pretraining or domain adaptation, golden sets, red-team cases, regression suites, classification, ranking, recommendation, fraud, risk, anomaly, and abuse detection, including immutable snapshots, dataset manifests, label/version tracking, train-validation-test isolation, temporal splitting, leakage prevention, deduplication, contamination controls, PII removal, balancing, and population monitoring.
Implement privacy, security, and regulated-model governance controls, including PII/PHI classification, masking or tokenization, encryption, IAM/RBAC or ABAC, row/document-level authorization, tenant isolation, secrets management, network boundaries, audit trails, retention and deletion propagation into feature stores, embeddings, vector indexes, caches, prompts, training corpora, and model artifacts, with evidence aligned where relevant to GDPR, CCPA, HIPAA, PCI DSS, SOC 2, ISO 27001, NIST AI RMF, and ISO/IEC 42001.
Produce architecture decision records, logical and physical data models, C4 or equivalent diagrams, data-flow and trust-boundary diagrams, interface contracts, capacity and cost models, failure-mode analyses, migration plans, runbooks, observability designs, and staged architecture roadmaps; defend tradeoffs involving consistency, availability, latency, throughput, durability, exactly-once versus at-least-once processing, synchronous versus asynchronous indexing, and replayability.
Required Skills and Qualifications
8–12 years of experience designing and architecting production data systems at scale, including distributed storage, compute, batch and streaming pipelines, and ML/LLM data platforms; architecture ownership must extend beyond pipeline operations.
Deep practical experience with Pinecone, Weaviate, and pgvector, including collection/index/schema design, HNSW/IVF-style indexing, distance metrics, dimensionality compatibility, metadata filtering, namespaces or tenants, upserts, deletes, compaction, re-indexing, lifecycle management, operational limits, and migration tradeoffs.
Proven ability to architect RAG systems with document parsing, semantic chunking, parent-child retrieval, embedding generation, hybrid Elasticsearch/OpenSearch retrieval, cross-encoder or hosted reranking, ACL-aware retrieval, citation provenance, retrieval evaluation, stale-index handling, and embedding-model version migration.
Strong distributed-systems design competency covering partitioning, replication, consistency and availability choices, ordering, idempotency, retries, backpressure, fault isolation, disaster recovery, checkpointing, watermarking, replay, backfills, dead-letter queues, and cost/performance analysis.
Hands-on experience designing feature stores and ML data layers with offline and online storage, point-in-time correctness, materialization, entity keys, feature freshness, feature versioning, training-serving skew detection, online feature lookup, and batch, asynchronous, and request-time inference paths.
Experience integrating data pipelines with ML/LLM training and inference through REST and gRPC services, Kafka consumer groups and replay, workflow/orchestration systems such as Airflow or Dagster, distributed processing with Spark, Beam, or Flink, hosted LLM APIs or self-hosted inference services, and model-serving interfaces with versioned schemas, authentication, timeouts, retries, idempotency keys, and error contracts.
Demonstrated implementation of data governance covering cataloging, lineage from source systems through transformations, feature computation, embedding generation, vector indexing, retrieval, prompt assembly, and model outputs; quality gates must address completeness, freshness, validity, duplication, referential integrity, semantic correctness, label quality, chunk quality, embedding coverage, retrieval relevance, and feature freshness.
Ability to translate AI product requirements into concrete architecture decisions and interface contracts while implementing security controls for PII/PHI, consent, purpose limitation, data minimization, subject access/deletion, encryption, retention, tenant isolation, prompt/data exfiltration prevention, auditability, and reproducible dataset, code, model, and embedding-version provenance.
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
AI Data Architect
The New York Consulting Group
Similar Other Jobs
View all Other jobs→name
Senior Manager, Embodied AI
name
AI Enablement Lead
name
Data Lakehouse/BI Specialist - FULLY REMOTE - Contractor in USD
LVI Associates
Traveling Sr Data Center Scheduler
CTC
Senior Data Architect
Project Evident
Associate Director, Data Hub
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.