Find the hardest tasks. Train, then find harder tasks.
Recursive self-improvement at your fingertips, repeated. Campaign does the finding using our proprietary self-play technology to build verifiable tasks from your data, measure them, mark the ones worth training on, over and over again.
Campaign tells you how well your model did. The frontier tasks it finds tell you what training your model needs next. We move beyond classic reinforcement learning with verifiable rewards. The old way is slow, because you have to attempt arbitrary tasks before their solve rate is known.
Today, Campaign runs the search as self-play. A setter model writes tasks in your domain, each with its own environment and verifier, and a solver model attempts each one. Each frontier task it finds is one more place training can move your model past its current limit. And we can do this over and over again. It’s quite powerful.
Campaign trials
Pack authoring and taxonomy trials prepare the campaign. Each trial has its own entry in the table of contents, with the commands and output below.
Click READ ALL to have me read you all the messages we wrote while running the campaign trials.
Built from your work
Start from one written description of the tasks you care about, grounded in your own code, data, or evals. A setter model writes each task with its own environment and verifier.
Measured against the frontier
A solver model attempts every task several times. Campaign marks the tasks it sometimes solves and sometimes fails: that contrast is what reinforcement learning learns from.
Scores you can audit
Every verifier is certified against known answers before it scores anything, and every rollout keeps its files, logs, and full trajectory.
- MethodTask pack sourceSandbox backendRuntime imageSetter network accessSolver network accessTasksIterationsRollouts per taskConcurrent trialsCompute timeout (s)s
Add private environment variables and choose which roles can use each one.
No secrets were supplied.
ModelReasoning effortHarnessCPUsCPUMemory (MB)MBStorage (MB)MBAgent timeout (s)sAction timeout (s)sStep limitstepMax output tokenstokTemperatureModel defaultVerifier timeout (s)sContext token budgettokModelReasoning effortHarnessCPUsCPUMemory (MB)MBStorage (MB)MBAgent timeout (s)sAction timeout (s)sStep limitstepMax output tokenstokTemperatureModel defaultVerifier timeout (s)sContext token budgettokSpend you control
Set the number of tasks, attempts, parallel runs, and a time limit before launch. Nothing is spent until you press Run, and nothing needs you after.
How recursive self-learning works
What the frontier is worth
An asset that compounds
Recursive self-improvement compounds only as fast as new verifiable tasks arrive at a model’s edge, and a task is only known to sit there once it has been measured. Campaign writes and measures those tasks from your own code, data, and evals, so every round of the loop adds to a dataset that belongs to you and no one else has.
Training data with signal
Run reinforcement learning on tasks your model solves some of the time, not the ones it always passes or never can. Download every task with its full instruction and measured solve rate as JSONL, ready for your training pipeline. When your model outgrows them, raise the bar and run the next campaign from its recorded settings.
Built for long runs
A long-horizon campaign runs self-play across many iterations, unattended for up to a day, inside the time and attempt budget you set. With the Mulberry method, each new batch of tasks sees the history of the work before it and leans harder on purpose. Each run leaves a certified task pack that starts the next, so the loop can keep turning for weeks without rebuilding anything.
Watch a run find the frontier.
See every step while it runs.
Campaign certifies its task setting before any task exists, then records every trial as it happens. The graph grows with the run, from the instruction to each task, trial, and reward, so it is never a black box.
Know exactly where your agent breaks.
Every step of a run is recorded, including the ones that fail. Here an independent check catches a task that could not tell a working fix from no fix, the repair fails too, and the task is discarded before a single attempt is spent on it.
Trace any result to the evidence.
Open any attempt and follow it step by step, from the task it was given to the files it changed and how it was graded. When a number is questioned, the answer is one click away.
Get a clear answer when it finishes.
One chart shows the whole run: tasks started and concluded, your agent’s solve rate, and the share of measured tasks that landed in the frontier band, over the campaign’s elapsed time. You see what the budget bought.
- Iteration 1, measuredfrontier-task-000117m 0s3 of 8 passedFrontier
- Iteration 1, measuredfrontier-task-000224m 0s2 of 8 passedFrontier
- Iteration 1, measuredfrontier-task-000322m 0s4 of 8 passedFrontier
- Iteration 1, measuredfrontier-task-000420m 0s1 of 8 passedFrontier
- Iteration 2, measuredfrontier-task-000518m 0s3 of 8 passedFrontier
- Iteration 2, measuredfrontier-task-000625m 0s2 of 8 passedFrontier
- Iteration 2, measuredfrontier-task-000723m 0s5 of 8 passed
- Iteration 2, measuredfrontier-task-000821m 0s3 of 8 passedFrontier
- Iteration 3, measuredfrontier-task-000919m 0s2 of 8 passedFrontier
- Iteration 3, measuredfrontier-task-001017m 0s4 of 8 passedFrontier
- Iteration 3, measuredfrontier-task-001124m 0s1 of 8 passedFrontier
- Iteration 3, measuredfrontier-task-001222m 0s3 of 8 passedFrontier
- Iteration 4, measuredfrontier-task-001320m 0s0 of 8 passed
- Iteration 4, measuredfrontier-task-001418m 0s2 of 8 passedFrontier
- Iteration 4, measuredfrontier-task-001525m 0s3 of 8 passedFrontier
- Iteration 4, measuredfrontier-task-001623m 0s4 of 8 passedFrontier
- Iteration 5, measuredfrontier-task-001721m 0s2 of 8 passedFrontier
- Iteration 5, measuredfrontier-task-001819m 0s6 of 8 passed
- Iteration 5, measuredfrontier-task-001917m 0s3 of 8 passedFrontier
- Iteration 5, measuredfrontier-task-002024m 0s1 of 8 passedFrontier
- Iteration 6, measuredfrontier-task-002122m 0s2 of 8 passedFrontier
- Iteration 6, measuredfrontier-task-002220m 0s4 of 8 passedFrontier
- Iteration 6, measuredfrontier-task-002318m 0s3 of 8 passedFrontier
- Iteration 6, measuredfrontier-task-002425m 0s2 of 8 passedFrontier
Own everything each round produces.
Export every task, score, and trial record as files you can train on, start the next round from the certified task pack, or share a finished campaign with a link you can revoke.
Campaign data
Play a whole campaign through.
The Vmax Campaigns Overview replays every stage of a campaign, from the first instruction to the frontier. Press play, pick a scenario, or step through it with the arrow keys once you click inside.
Vmax Campaigns V3.9.4 Overview
Overview of the native self-play system built by the Vmax team.
- Spine
- Fan
- Tree
- Converge
- Data
Start the loop with one page of instructions.
Find the hard tasks, train on them, then find harder ones. Campaign is open to approved Vmax accounts: join the waitlist for early access, or sign in if you already have one.