Data First Jobs

XChange Software Inc

Remote Role :: Conversational AI Data Engineer

Contract · In Office · USA

Posted Sep 9, 2026

Work Options
Cloud Stack
Positions
Job Type
Position Group
  • Conversational Data Engineer
  • 100% Remote Role
  • Client:UHG
  • Remote
  • Job Requirements
  • Experience: 7–10 years
  • Minimum Relevant Experience: 4+ years
  • Primary Focus: Conversational AI, Synthetic Data, RAG/Knowledge Pipelines, NLP, Data Engineering
  • Cloud Platforms: GCP and AWS
  • Role Summary
  • We are looking for an experienced Conversational Data Engineer to create high-quality, customer-specific synthetic data and own RAG/knowledge pipelines supporting CCAI, voice, and chat deployments.
  • The engineer will design scalable data generation and ingestion pipelines and load data into the appropriate GCP and AWS services. The role requires a strong understanding of conversational AI data, knowledge management, retrieval systems, synthetic data generation, and data privacy.
  • The goal is to enable each customer deployment to be configured, grounded, demonstrated, and validated without using real customer PII.
  • What Success Looks Like
  • Each customer engagement has a documented synthetic dataset covering all channels within scope.
  • Each in-scope customer has a functional RAG/knowledge pipeline, including:
  • Corpus preparation
  • Indexing
  • Retrieval
  • Evaluation
  • Data and retrieval quality are sufficient for:
  • Configuration
  • Evaluation and testing
  • Stakeholder demonstrations
  • Synthetic data meets privacy, isolation, and compliance expectations.
  • Data generation and indexing processes are parameterized, repeatable, and scalable, rather than manual, one-off processes.
  • Key Responsibilities
  • Synthetic Data Generation
  • Analyze each customer's domain, including:
  • Intents
  • Entities
  • Knowledge topics
  • Document types
  • Languages
  • Tone
  • Edge cases
  • Generate synthetic conversation transcripts for voice and chat.
  • Create synthetic CCAI training and evaluation dialogues.
  • Generate supporting content, including:
  • Customer and agent profiles
  • Knowledge-base articles
  • FAQs
  • Structured entity values
  • Select appropriate synthetic-data generation techniques and document parameters, assumptions, and limitations.
  • RAG & Knowledge Pipelines
  • Design and maintain RAG and knowledge ingestion pipelines.
  • Prepare, transform, index, retrieve, and evaluate customer-specific knowledge corpora.
  • Load data into the appropriate GCP and AWS services.
  • Schedule and document index-refresh processes when customer knowledge changes.
  • Ensure data and indexes are correctly integrated into deployed conversational experiences.
  • Data Quality & Privacy
  • Validate synthetic data for:
  • Realism
  • Coverage
  • Diversity
  • Consistency
  • Retrieval quality
  • Verify the absence of residual real-world PII in synthetic datasets and source corpora.
  • Ensure data meets appropriate privacy, isolation, and compliance requirements.
  • Develop and maintain reusable data-quality checklists.
  • Automation & Reusability
  • Build and maintain reusable:
  • Synthetic-data generators
  • Data ingestion jobs
  • Indexing pipelines
  • Quality-validation processes
  • Parameterize pipelines to support different customers and use cases.
  • Minimize manual, customer-specific copy-paste processes through automation and repeatable workflows.
  • Collaboration
  • Partner with the Conversational Platform Specialist to ensure loaded data and indexes effectively support the deployed conversational experience.
  • Partner with DevOps to automate:
  • Pipeline jobs
  • Data stores
  • Secrets management
  • Ensure pipeline jobs, data stores, and other resources are properly isolated for each customer.
  • Required Qualifications
  • 4+ years of experience in one or more of the following:
  • Data Engineering
  • Conversation Design Operations
  • Applied NLP Data
  • Knowledge-Pipeline Engineering
  • Working knowledge of how conversational platforms consume:
  • Training data
  • FAQs
  • Conversation transcripts
  • Retrieval-grounded knowledge
  • Strong understanding of synthetic-data quality and generation.
  • Strong understanding of RAG and retrieval quality.
  • Strong judgment regarding data privacy, PII protection, and security.
  • Experience designing or maintaining data generation and ingestion pipelines.
  • Experience working with cloud-based data platforms, particularly GCP and/or AWS.
  • Preferred Qualifications
  • Experience with LLM-assisted synthetic data generation in a production or customer implementation environment.
  • Familiarity with:
  • Google BigQuery
  • Amazon S3
  • Document stores used as knowledge sources
  • Experience with RAG, vector search, semantic search, knowledge retrieval, or indexing.
  • Experience with multilingual data generation or evaluation.
  • Experience supporting conversational AI, voice, chat, or CCAI platforms.
  • Experience working on enterprise or customer-specific implementations.
  • Key Skills
  • Conversational AI | Synthetic Data | RAG | Knowledge Pipelines | NLP | LLMs | CCAI | Data Engineering | BigQuery | AWS S3 | GCP | AWS | Data Ingestion | Indexing | Knowledge Base | Data Quality | PII Protection | Multilingual Data

Mention you found this on Data First Jobs — it helps us bring you more roles like this.

Remote Role :: Conversational AI Data Engineer

XChange Software Inc

Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.

Free, no spam. Unsubscribe anytime.