Vmax is developing AI capable of open-ended learning. We build systems that exceed human performance across every domain, optimizing beyond the local maxima of learning from human expertise. Our approach lets agents define and optimize their own goals. We do not seek to replace human labor with machine labor, but to discover radically new ways of working.

We are hiring researchers with deep expertise in reinforcement learning and a keen interest in bringing its more esoteric aspects to real-world applications.

June 10, 2026

PROPEL: Breaking the Solver Bottleneck in Task-Generator RL

  • Augustine N. Mavor-Parker
  • Connor Watts
  • Geoffrey Bradway
  • Lorenz Wolf
  • Matthew Daborn-Sargent
  • Maxwill Lin
  • Roger Creus Castanyer

PROPEL replaces solver trials with an activation probe on a frozen generator's internals, doubling frontier-task generation across math, code, and SWE.

May 20, 2026

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-⁠Play

  • Augustine N. Mavor-Parker
  • Geoffrey Bradway
  • Lorenz Wolf
  • Matthew James Sargent
  • Maxwill Lin
  • Roger Creus Castanyer

We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs.