Data First Jobs

Ampstek

Data Analyst

Contract · In Office · USA

Posted Aug 14, 2026

Work Options
Cloud Stack
Positions
Job Type
Position Group
  • Job Role : Data Analyst
  • Location : USA (Remote)
  • Long Term Contract

Job Description :

  • 1. Relational Data Modeling
  • Design realistic, normalized relational data models (dimension and fact tables) for four
  • industry verticals: Financial, Healthcare, Manufacturing, and Retail.
  • Ensure schema design supports representative industry use cases and natural-language
  • querying patterns without exposing PHI/PII or sensitive data.
  • Map data requirements to align with the Oracle AI Database Agent metadata indexing
  • and query execution engine.
  • 2. Synthetic Data Generation & Loading
  • Develop and execute reusable Python/SQL scripts to programmatically synthesize data
  • at scale.
  • Generate target volume thresholds (~50–500 rows per dimension table;
  • ~10,000–100,000 rows per fact table per industry).
  • Perform one-time seed data loads into designated schemas within the shared Oracle
  • Autonomous Database.
  • 3. Query Validation & Golden Dataset Development
  • Validate data quality, primary/foreign key integrity, and representative query coverage
  • across all four schemas.
  • Collaborate with the AI/Gemini engineering team to build prompt libraries, golden
  • datasets, and sample Critical User Journeys (CUJs).
  • Perform end-to-end testing (Question → SQL → Result Execution) to confirm the
  • accuracy of generated SQL queries.
  • Participate in cross-industry isolation testing to ensure dataset security across schema
  • boundaries.
  • 4. Post-Deployment & UAT Support
  • Assist with User Acceptance Testing (UAT) bug fixing and schema refinements based on
  • demo feedback.
  • Track query performance and remediate gap accounts/schemas during the adoption
  • push phase.
  • Required Qualifications & Experience
  • Data Analysis & Modeling: 3+ years of experience in data analysis, relational data
  • modeling (star schema, snowflake, transactional), and data warehousing.
  • Domain Knowledge: Understanding of schema structures across at least two target
  • industries (Financial, Healthcare, Manufacturing, or Retail).
  • Programming & SQL: Strong proficiency in SQL (Oracle dialect preferred) and Python
  • (Pandas, Faker, or synthetic data generation frameworks).
  • Data Synthesis & Quality: Experience generating synthetic/mock data for testing, demo
  • environments, or AI model training while avoiding PHI/PII regulatory exposure.
  • AI & Analytical Tools: Familiarity with AI/LLM concepts, prompt engineering baseline
  • evaluation, or Natural-Language-to-SQL workflows is highly desirable.

Mention you found this on Data First Jobs — it helps us bring you more roles like this.

Data Analyst

Ampstek

Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.

Free, no spam. Unsubscribe anytime.