September 2026
What if AutoResearch experimented with data?
Provena's AutoCurator experimented with how to curate training data — and found a surprisingly small subset that performed better.
161
Experiments
5.6%
Data selected
+5.2pp
ARC-Easy benchmark
Gain held
At 1-3x training budgets
AutoCurator
Improving Karpathy's AutoResearch with data curation
01
Making data the variable
Karpathy's AutoResearch showed that AI agents can accelerate the path to better models by experimenting with model architecture and training.
Using the same NVIDIA ClimbMix training dataset, AutoCurator explored ways to select and curate the data, while keeping the model and training compute fixed.
Could an autonomous researcher discover better training data for the same model?
02
What AutoCurator found
Across 161 experiments, most interventions failed. Two improved on the running best.
The strongest intervention combined three data-quality signals in our platform. AutoCurator discovered the conditions and thresholds that worked best, retaining just 5.6% of the original training data.
At equal training compute, ARC-Easy improved from 32.5% to 37.7% (+5.2pp).
| Benchmark | Baseline | AutoCurator | Change |
|---|---|---|---|
| ARC-Easy | 32.5% | 37.7% | +5.2pp |
| HellaSwag | 31.5% | 31.3% | −0.2pp |
| PIQA | 50.8% | 49.1% | −1.7pp |
| 3-task mean | 38.3% | 39.4% | +1.1pp |
The ARC-Easy advantage also held at 1×, 2× and 3× training budgets.
03
What we learned
- 01
More data isn't always better data. AutoCurator found a much smaller, targeted subset that performed better on the capability it was optimizing.
- 02
The right data depends on what you're trying to improve. ARC-Easy improved materially while other benchmarks were flat or slightly down. There may be no universally “best” dataset — only data better suited to particular capabilities.
- 03
Most experiments won't work — and that's the point. AutoCurator tried 161 interventions; only two improved on the running best. Autonomous research makes it practical to explore a much larger hypothesis space — and retain what works.
04
What we don't know yet
This is an early result on an 11.5M-parameter proxy model, and the strongest gain is on one benchmark: ARC-Easy. We don't yet know how well the selection strategy will generalize across larger models, datasets or capabilities.
For now, we see this as evidence that AutoCurator can discover promising data curation strategies — not yet a production-scale recipe.
Next, we're testing larger models, more compute and more research questions.
05 · Help us test it
Push the next experiment further with us
We’re looking for model builders and research teams to push this further with us.
Through Provena Research Partners, we'll contribute up to $10K in compute + research support to experiment on a model, training data and benchmark you care about.