CODE
Real repositories, failing tests and working toolchains. The model ships a fix and the test suite decides.
- repos
- tests
- terminals
Reinforcement Learning environments for AI agents:
For agents that reason, code and use computers.
01 / THE LOOP
An environment is the practice field: a place where a model can try, fail, get feedback and try again.
A goal drawn from real work. Fix the bug. Finish the flow. Solve the problem.
A sandbox to act in, with the files, apps and tools the task needs.
An automatic check of the outcome. It scores what actually happened, not how it sounds.
A clean signal for the trainer. Thousands of attempts later, the model is better at the job.
02 / ENVIRONMENTS
Each has outcomes a program can check (RLVR).
Real repositories, failing tests and working toolchains. The model ships a fix and the test suite decides.
Sandboxed browsers and desktop apps. Multi-step goals with the final state checked at the end.
Hard problems across expert domains with answers that can be verified.
03 / PRINCIPLES
Every reward traces back to a check you can read.
runenv environments prevent both shortcuts and reward hacks before a model ever finds them.
Tasks sit at the edge of what the model can do. That is where the learning signal is.
Containerized and scalable interfaces that drop into open-source RL trainers.
04 / CONTACT
Let's build the environments your agents learn in. We're working with a small number of early teams: