Get started

We build RL environments for professional software to investigate where models fail and turn those failures into tasks for evaluation and RL training.

Use traces to understand why models fail and design tasks that reduce reward hacking while helping your team hillclimb model performance. Read more.

Ask Claude or Codex to read usedesktop.com/setup.txt and set up your environments, runtime, and SDK.

Preparing