- Conversational Data Engineer
- 100% Remote Role
- Client:UHG
- Remote
- Job Requirements
- Experience: 7–10 years
- Minimum Relevant Experience: 4+ years
- Primary Focus: Conversational AI, Synthetic Data, RAG/Knowledge Pipelines, NLP, Data Engineering
- Cloud Platforms: GCP and AWS
- Role Summary
- We are looking for an experienced Conversational Data Engineer to create high-quality, customer-specific synthetic data and own RAG/knowledge pipelines supporting CCAI, voice, and chat deployments.
- The engineer will design scalable data generation and ingestion pipelines and load data into the appropriate GCP and AWS services. The role requires a strong understanding of conversational AI data, knowledge management, retrieval systems, synthetic data generation, and data privacy.
- The goal is to enable each customer deployment to be configured, grounded, demonstrated, and validated without using real customer PII.
- What Success Looks Like
- Each customer engagement has a documented synthetic dataset covering all channels within scope.
- Each in-scope customer has a functional RAG/knowledge pipeline, including:
- Corpus preparation
- Indexing
- Retrieval
- Evaluation
- Data and retrieval quality are sufficient for:
- Configuration
- Evaluation and testing
- Stakeholder demonstrations
- Synthetic data meets privacy, isolation, and compliance expectations.
- Data generation and indexing processes are parameterized, repeatable, and scalable, rather than manual, one-off processes.
- Key Responsibilities
- Synthetic Data Generation
- Analyze each customer's domain, including:
- Intents
- Entities
- Knowledge topics
- Document types
- Languages
- Tone
- Edge cases
- Generate synthetic conversation transcripts for voice and chat.
- Create synthetic CCAI training and evaluation dialogues.
- Generate supporting content, including:
- Customer and agent profiles
- Knowledge-base articles
- FAQs
- Structured entity values
- Select appropriate synthetic-data generation techniques and document parameters, assumptions, and limitations.
- RAG & Knowledge Pipelines
- Design and maintain RAG and knowledge ingestion pipelines.
- Prepare, transform, index, retrieve, and evaluate customer-specific knowledge corpora.
- Load data into the appropriate GCP and AWS services.
- Schedule and document index-refresh processes when customer knowledge changes.
- Ensure data and indexes are correctly integrated into deployed conversational experiences.
- Data Quality & Privacy
- Validate synthetic data for:
- Realism
- Coverage
- Diversity
- Consistency
- Retrieval quality
- Verify the absence of residual real-world PII in synthetic datasets and source corpora.
- Ensure data meets appropriate privacy, isolation, and compliance requirements.
- Develop and maintain reusable data-quality checklists.
- Automation & Reusability
- Build and maintain reusable:
- Synthetic-data generators
- Data ingestion jobs
- Indexing pipelines
- Quality-validation processes
- Parameterize pipelines to support different customers and use cases.
- Minimize manual, customer-specific copy-paste processes through automation and repeatable workflows.
- Collaboration
- Partner with the Conversational Platform Specialist to ensure loaded data and indexes effectively support the deployed conversational experience.
- Partner with DevOps to automate:
- Pipeline jobs
- Data stores
- Secrets management
- Ensure pipeline jobs, data stores, and other resources are properly isolated for each customer.
- Required Qualifications
- 4+ years of experience in one or more of the following:
- Data Engineering
- Conversation Design Operations
- Applied NLP Data
- Knowledge-Pipeline Engineering
- Working knowledge of how conversational platforms consume:
- Training data
- FAQs
- Conversation transcripts
- Retrieval-grounded knowledge
- Strong understanding of synthetic-data quality and generation.
- Strong understanding of RAG and retrieval quality.
- Strong judgment regarding data privacy, PII protection, and security.
- Experience designing or maintaining data generation and ingestion pipelines.
- Experience working with cloud-based data platforms, particularly GCP and/or AWS.
- Preferred Qualifications
- Experience with LLM-assisted synthetic data generation in a production or customer implementation environment.
- Familiarity with:
- Google BigQuery
- Amazon S3
- Document stores used as knowledge sources
- Experience with RAG, vector search, semantic search, knowledge retrieval, or indexing.
- Experience with multilingual data generation or evaluation.
- Experience supporting conversational AI, voice, chat, or CCAI platforms.
- Experience working on enterprise or customer-specific implementations.
- Key Skills
- Conversational AI | Synthetic Data | RAG | Knowledge Pipelines | NLP | LLMs | CCAI | Data Engineering | BigQuery | AWS S3 | GCP | AWS | Data Ingestion | Indexing | Knowledge Base | Data Quality | PII Protection | Multilingual Data
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
Remote Role :: Conversational AI Data Engineer
XChange Software Inc
Similar Engineering Jobs
View all Engineering jobs→AgileGrid Solutions
Machine Learning Developer
New
USA
Nexwave
Senior Data Engineering & Platform Engineer (DBT)
New
RemoteUSA
Arkhya Tech. Inc.
Hadoop Data Engineer - Scottsdale AZ (100% Onsite)
New
Scottsdale, Arizona (USA)
Miracle Software Systems, Inc
Sr. Full Stack Data AI Engineer – GCP & Agentic AI
New
Novi, Michigan (USA)
Skill
Full-Stack Engineer – Data Experience, Backstage Portal [SK-17196]
New
RemoteUSA$64,000 - $75,000
Jain Global
Solutions Analyst & Engineer
New
New York, New York (USA)
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.