Verification for RL task environments
You paid $4,000 an environment for three that teach the exploit
A task environment is only worth training on if reward tracks the thing you meant. Environment Forge finds the ones where a policy can farm reward without doing the task, before they reach a training run.
The problem
Environments are bought by the batch and verified by nobody
Labs are buying verified task environments faster than anyone can check them, and the check that matters is not whether the environment runs. It is whether the reward function can be satisfied without completing the task — because if it can, the model will find that path, and you have paid to teach it the shortcut.
The insight
Two independent tells, and neither is conclusive alone
Reward that correlates weakly with task completion is suspicious but can be explained by partial credit. A non-completing policy that outscores the median completing one is suspicious but can be a lucky run. Both together is an exploit, and requiring both is what keeps the false-positive rate low enough that a buyer will act on the report rather than argue with it.
Reward-to-completion correlation across episodes, exploit-gap measurement between the best non-completing policy and the median completing one, and a conjunction rule requiring at least two independent signals before an environment is called gameable.
How it works
Four steps, no data science team
A policy sweep across every environment in the delivery, including deliberately degenerate policies.
Alignment and exploit gap, per environment.
Which policy broke it and by how much, so the vendor can fix or refund.
Only verified environments enter the training set.
Who it is for
Whoever signs off that a training batch is safe to use
Frontier labs and the data companies supplying them. A small number of extremely sophisticated buyers with real budgets and a real problem.
Pricing
- –Alignment measurement
- –Exploit-gap detection
- –Public methodology
- –Batch verification
- –Delivery gate
- –Vendor scorecards
- –Exploit reproduction
- –Self-hosted
- –Custom sweeps
- –Vendor SLA reporting
- –Dedicated support
Competition
What exists, and what it does not do
| Who | What they do | The gap |
|---|---|---|
| Environment vendors own QA | Suppliers verify their own deliveries. | The supplier is paid per verified environment. Asking them to fail their own product is a structural conflict, not a quality complaint. |
| Internal eval teams | Labs check environments themselves. | They can, and it is nobody dedicated job. This is the sort of check that gets skipped in a quarter where the training run is already late. |
| Reward-hacking research | A live literature on specification gaming. | Papers and benchmarks, not a verification service pointed at a specific delivery on a deadline. |
| Finding out during training | The current mechanism. | The most expensive possible detector, and it fires after the compute has been spent. |
The entire market is perhaps twenty buyers, all of whom employ people smarter than this tool about reinforcement learning, and any of whom could build it in a fortnight if they decided it mattered. Selling to labs also means long procurement, security review and an expectation of on-premise everything. The realistic path is not selling software to labs — it is selling verification as a service to the data vendors who need to prove quality to those labs, which is a larger and far less sophisticated buyer set.
Market
Attached to the training-data budget, which is one of the few genuinely large and urgent budgets in 2026
Few buyers, high contract values, brutal sales cycles. Twenty vendor relationships at Team pricing is $720k ARR; a single lab contract could exceed all of them.