Data First Jobs

Cobalt

Principal Machine Learning Scientist, AI Model Evaluation (PhD)

Full Time · Remote · USA

Posted Sep 20, 2026

  • About the role
  • Cobalt builds expert data and evaluation infrastructure for AI developers. We are recruiting senior machine learning scientists for a contract project that evaluates how well frontier AI models perform real machine learning work in a command line environment. Your job is to design original, realistic machine learning tasks that a leading AI model fails to solve.
  • What you will do
  • Design self-contained command line tasks drawn from real machine learning research and engineering, such as diagnosing a training run that fails to converge, reproducing a published result from a paper, finding data leakage in an evaluation pipeline, fixing a numerically unstable implementation, optimizing inference under memory or latency limits, and debugging distributed training code.
  • Build the task environment, including code, data, and dependencies, write a reference solution, and write automated tests that verify whether a solution is correct. Because machine learning results can vary between runs, tests need to be deterministic or use well-justified tolerances.
  • Test your task against a frontier AI model and refine it until the model fails for substantive reasons rather than because of ambiguity, trick wording, or excessive compute requirements.
  • Respond to reviewer feedback until the task is accepted.
  • Who we are looking for
  • A PhD in machine learning, computer science, statistics, or a closely related field.
  • Industry or academic hands-on experience in machine learning research or applied machine learning.
  • At least one publication, either academic (for example a peer-reviewed paper at a recognized machine learning venue) or professional (for example a technical report, a conference talk, or a widely used open-source library).
  • Deep working knowledge of Python and at least one major framework, such as PyTorch or JAX.
  • Fluency in the Linux command line, shell scripting, Docker, and Git.
  • The ability to write precise task specifications that another scientist could follow without asking questions.

Why Cobalt AI:

  • Advance frontier AI where it counts. Apply your research expertise to the data that frontier labs cannot obtain any other way, where your reasoning directly shapes how the next generation of models works through technical problems.
  • Grow professionally. Expand your influence through evaluation projects, advisory roles, and research collaborations, while deepening your understanding of how frontier models are trained and assessed.
  • Work with a top-tier network. Collaborate with researchers from leading institutions and labs on high-impact, flexible work.
  • Set your own schedule. Flexible 10 to 40 hour weeks that fit around your research position and your life.
  • Competitive pay. Rates vary by project and are determined by a number of factors, including scope, skillset, and experience.

Mention you found this on Data First Jobs — it helps us bring you more roles like this.

Principal Machine Learning Scientist, AI Model Evaluation (PhD)

Cobalt

Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.

Free, no spam. Unsubscribe anytime.