RL environments

Reinforcement Learning environments for AI agents:

  • Real tasks
  • Real tools
  • Rewards you can verify

For agents that reason, code and use computers.

TALK TO US

01 / THE LOOP

Environments turn practice into skill.

An environment is the practice field: a place where a model can try, fail, get feedback and try again.

01

TASK

A goal drawn from real work. Fix the bug. Finish the flow. Solve the problem.

02

ENVIRONMENT

A sandbox to act in, with the files, apps and tools the task needs.

03

VERIFIER

An automatic check of the outcome. It scores what actually happened, not how it sounds.

04

REWARD

A clean signal for the trainer. Thousands of attempts later, the model is better at the job.

02 / ENVIRONMENTS

Three environments to start.

Each has outcomes a program can check (RLVR).

LVL 1

CODE

Real repositories, failing tests and working toolchains. The model ships a fix and the test suite decides.

  • repos
  • tests
  • terminals
LVL 2

COMPUTER USE

Sandboxed browsers and desktop apps. Multi-step goals with the final state checked at the end.

  • browsers
  • forms
  • apps
LVL 3

EXPERT REASONING

Hard problems across expert domains with answers that can be verified.

  • proofs
  • calculations
  • analysis

03 / PRINCIPLES

What makes an environment worth training on.

[+]

VERIFIABLE

Every reward traces back to a check you can read.

[+]

HARD TO GAME

runenv environments prevent both shortcuts and reward hacks before a model ever finds them.

[+]

RIGHT DIFFICULTY

Tasks sit at the edge of what the model can do. That is where the learning signal is.

[+]

PLUG-IN READY

Containerized and scalable interfaces that drop into open-source RL trainers.

04 / CONTACT

Using reinforcement learning?

Let's build the environments your agents learn in. We're working with a small number of early teams:

  • Startups training agents with RL
  • Research groups exploring RL for reasoning and agents
  • Companies turning their own workflows into RL tasks
GUILLERMO@RUNENV.SI