Provena Research Partners

What would you learn from
100× more data experiments?

Provena’s AutoCurator explored 161 data interventions on one model, finding a 5.6% subset that improved ARC-Easy by 5.2pp.

Now, we’re looking for teams training ML models to explore which data interventions could improve theirs.

Selected partners receive up to $10K in compute + research support from Provena.

You bring

  • 01A model
  • 02Training data
  • 03A benchmark that matters to you

We bring

  • 01Compute
  • 02Data experimentation
  • 03Research + engineering support

01

How it works

  1. 01

    Bring us a research question

    What capability do you want to improve?

  2. 02

    We explore what works

    Systematically test data interventions

  3. 03

    We measure what moves performance

    Come away with concrete results and learnings

02

Why we’re doing this

Data can make or break model performance. But figuring out what works is still largely manual and empirical.

We believe an autonomous training data researcher can help teams test more ideas, get more from their data, and retain what they learn.

03

Who should apply

We’re looking for model builders and research teams with a real performance question where data could be part of the answer.

For now, we’re focused on language models trained on text data.

Priority deadline

October 1, 2026

Applications considered on a rolling basis thereafter.