September 2026

What if AutoResearch experimented with data?

Provena's AutoCurator experimented with how to curate training data — and found a surprisingly small subset that performed better.

161

Experiments

5.6%

Data selected

+5.2pp

ARC-Easy benchmark

Gain held

At 1-3x training budgets

AutoCurator

Improving Karpathy's AutoResearch with data curation

DiscardedSuccessful
AutoCurator experiment progressA stepped line chart showing two retained improvements across 161 experiments, ending with 5.6 percent of training data selected and a 5.2 percentage point benchmark gain.−0.10−0.11−0.12−0.13−0.14−0.15−0.16BENCHMARK MARGIN (HIGHER IS BETTER)0326496128161EXPERIMENTS #10% data selected5.6% data selected

01

Making data the variable

Karpathy's AutoResearch showed that AI agents can accelerate the path to better models by experimenting with model architecture and training.

Using the same NVIDIA ClimbMix training dataset, AutoCurator explored ways to select and curate the data, while keeping the model and training compute fixed.

Could an autonomous researcher discover better training data for the same model?

02

What AutoCurator found

Across 161 experiments, most interventions failed. Two improved on the running best.

The strongest intervention combined three data-quality signals in our platform. AutoCurator discovered the conditions and thresholds that worked best, retaining just 5.6% of the original training data.

At equal training compute, ARC-Easy improved from 32.5% to 37.7% (+5.2pp).

BenchmarkBaselineAutoCuratorChange
ARC-Easy32.5%37.7%+5.2pp
HellaSwag31.5%31.3%−0.2pp
PIQA50.8%49.1%−1.7pp
3-task mean38.3%39.4%+1.1pp

The ARC-Easy advantage also held at 1×, 2× and 3× training budgets.

03

What we learned

  1. 01

    More data isn't always better data. AutoCurator found a much smaller, targeted subset that performed better on the capability it was optimizing.

  2. 02

    The right data depends on what you're trying to improve. ARC-Easy improved materially while other benchmarks were flat or slightly down. There may be no universally “best” dataset — only data better suited to particular capabilities.

  3. 03

    Most experiments won't work — and that's the point. AutoCurator tried 161 interventions; only two improved on the running best. Autonomous research makes it practical to explore a much larger hypothesis space — and retain what works.

04

What we don't know yet

This is an early result on an 11.5M-parameter proxy model, and the strongest gain is on one benchmark: ARC-Easy. We don't yet know how well the selection strategy will generalize across larger models, datasets or capabilities.

For now, we see this as evidence that AutoCurator can discover promising data curation strategies — not yet a production-scale recipe.

Next, we're testing larger models, more compute and more research questions.

05 · Help us test it

Push the next experiment further with us

We’re looking for model builders and research teams to push this further with us.

Through Provena Research Partners, we'll contribute up to $10K in compute + research support to experiment on a model, training data and benchmark you care about.

Become a Research Partner