- Job Role : Data Analyst
- Location : USA (Remote)
- Long Term Contract
Job Description :
- 1. Relational Data Modeling
- Design realistic, normalized relational data models (dimension and fact tables) for four
- industry verticals: Financial, Healthcare, Manufacturing, and Retail.
- Ensure schema design supports representative industry use cases and natural-language
- querying patterns without exposing PHI/PII or sensitive data.
- Map data requirements to align with the Oracle AI Database Agent metadata indexing
- and query execution engine.
- 2. Synthetic Data Generation & Loading
- Develop and execute reusable Python/SQL scripts to programmatically synthesize data
- at scale.
- Generate target volume thresholds (~50–500 rows per dimension table;
- ~10,000–100,000 rows per fact table per industry).
- Perform one-time seed data loads into designated schemas within the shared Oracle
- Autonomous Database.
- 3. Query Validation & Golden Dataset Development
- Validate data quality, primary/foreign key integrity, and representative query coverage
- across all four schemas.
- Collaborate with the AI/Gemini engineering team to build prompt libraries, golden
- datasets, and sample Critical User Journeys (CUJs).
- Perform end-to-end testing (Question → SQL → Result Execution) to confirm the
- accuracy of generated SQL queries.
- Participate in cross-industry isolation testing to ensure dataset security across schema
- boundaries.
- 4. Post-Deployment & UAT Support
- Assist with User Acceptance Testing (UAT) bug fixing and schema refinements based on
- demo feedback.
- Track query performance and remediate gap accounts/schemas during the adoption
- push phase.
- Required Qualifications & Experience
- Data Analysis & Modeling: 3+ years of experience in data analysis, relational data
- modeling (star schema, snowflake, transactional), and data warehousing.
- Domain Knowledge: Understanding of schema structures across at least two target
- industries (Financial, Healthcare, Manufacturing, or Retail).
- Programming & SQL: Strong proficiency in SQL (Oracle dialect preferred) and Python
- (Pandas, Faker, or synthetic data generation frameworks).
- Data Synthesis & Quality: Experience generating synthetic/mock data for testing, demo
- environments, or AI model training while avoiding PHI/PII regulatory exposure.
- AI & Analytical Tools: Familiarity with AI/LLM concepts, prompt engineering baseline
- evaluation, or Natural-Language-to-SQL workflows is highly desirable.
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
Data Analyst
Ampstek
Similar Analytics Jobs
View all Analytics jobs→Fengate Asset Management
Financial Analyst, Infrastructure
New
Oakville, Ontario (Canada)
Galapagos Federal Systems
Program Security OPSEC Analyst
New
Colorado Springs, Colorado (USA)
Galapagos Federal Systems
Operations Research Analyst
New
Los Angeles, California (USA)
Galapagos Federal Systems
Data Analyst
New
Washington, District of Columbia (USA)
Jobright.ai
Cloud Security Analyst, Entry Level
New
USA
Jobright.ai
SOC Analyst, Entry Level
New
USA
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.