PROPEL: Breaking the Solver Bottleneck in Task-Generator RL
PROPEL replaces solver trials with an activation probe on a frozen generator's internals, doubling frontier-task generation across math, code, and SWE.
Vmax is developing AI capable of open-ended learning. We build systems that exceed human performance across every domain, optimizing beyond the local maxima of learning from human expertise. Our approach lets agents define and optimize their own goals. We do not seek to replace human labor with machine labor, but to discover radically new ways of working.
We are hiring researchers with deep expertise in reinforcement learning and a keen interest in bringing its more esoteric aspects to real-world applications.
PROPEL replaces solver trials with an activation probe on a frozen generator's internals, doubling frontier-task generation across math, code, and SWE.
We introduce unix-ctf, a procedural generator of capture-the-flag environments for training and evaluating Unix competence in language-model shell agents.
We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs.