RL List.com
UPDATED 2026.10.05

RL Environment Vendors 2026 Directory & Rankings

A cited directory of the 42 companies building RL environments for frontier AI labs, from public sources and what vendors share directly.

42 vendors · 4 segments · 15 disclosed SOC 2 · 11 open source

About this list

The RL-environment market moves fast: companies raise, pivot, and get acquired within weeks, so every figure here is a snapshot that can date. We're not the only ones mapping it; other lists worth reading include Pavlov's List by Chris Barber (@chrisbarber) and Deedy Das (@deedydas, Menlo Ventures) on the AI training-data market.

What are RL environments?

RL environments are the simulated tasks and worlds used to train and evaluate AI agents with reinforcement learning, from the open-source Gymnasium API (the successor to OpenAI Gym) to bespoke coding worlds behind benchmarks like SWE-bench. This directory maps the companies building them commercially, along with the human-data and RLHF work that sits alongside. For the foundations, see Sutton & Barto's Reinforcement Learning: An Introduction. How we research and rank vendors is on the methodology page.

RL environment vendor directory

42 of 42 vendors
#1

AfterQuery

Commercial
AfterQuery is a San Francisco applied-research lab and expert-data company (YC W25). It supplies AI labs with expert-generated SFT and RL data, rubrics, agent environments and computer-use trajectories, drawn from a stated network of nearly 100,000 verified professionals. Forbes reported it was raising at a $3.2B valuation in Sept 2026. NVIDIA's Nemotron 3 Ultra report names AfterQuery tasks in its GDPval training recipe. The company publishes or co-builds real-task benchmarks (FinanceQA, VADER, IDE-Bench, SpreadsheetBench 2, Legora BAR) and is adding forward-deployed enterprise work.
Raised
$30.5M ↑ $30M
Headcount
51-200
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
partial
Backers: Altos Ventures (lead, Series A), The Raine Group, Y Combinator
CodeComputer UseEnterprise
site ↗
#2

Bespoke Labs

Commercial
Bespoke Labs is an applied AI research lab in Mountain View, founded in 2024 by Mahesh Sathiamoorthy (ex-Google DeepMind) and UC Berkeley professor Alex Dimakis. It builds RL environments and infrastructure for training and evaluating reliable agents, and is backed by $40M from Wing VC, 8VC, Mayfield and others. It combines commercial environment and post-training work, such as a jointly published project with Intuit Credit Karma, with a large open-research output: OpenThoughts, Curator, GEPA, Terminal-Bench, AutoResearchExam and ConjectureBench.
Raised
$40M ↑ $40M
Headcount
11-50
Founded
2024
SOC 2
unknown
Researchers
yes
Open src
yes
Backers: Wing VC (Series A lead), 8VC (Seed lead), Mayfield
Long Horizon
site ↗
#3

Huzzle Labs

Commercial
Huzzle Labs is the applied-AI division of the London talent platform Huzzle (co-founders Ingmar Klein, CEO, and Amit Choudhary, CTO). It builds long-horizon RL environments (coding, computer use, enterprise workflows) and expert trajectory data for frontier labs, drawing on Huzzle's talent network, which it says numbers 300k+. In 2026 it expanded into its own small computer-use models (HuzzleWorld, 1B-8B, self-reported results), a domain benchmark (InsureBench, scores pending), and enterprise custom evals and models deployed on client infrastructure.
Raised
$6M
Headcount
11-50
Founded
2020
SOC 2
Type II
Researchers
yes
Open src
partial
Backers: 10x Founders, Angel Invest, Emerge
CodeLong Horizon
site ↗
#4

Fleet AI

Commercial
Fleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise software such as Salesforce and Excel, plus browser/desktop workflows, so frontier AI labs and large enterprises can train and evaluate computer-use agents. It ships a Python SDK, a platform API, and the open-source 'Harbor' agent-evaluation/RL-environment tooling, pairing simulated environments with human supervision.
Raised
$15M
Headcount
11-50
Founded
2024
SOC 2
unknown
Researchers
yes
Open src
partial
Backers: Sequoia Capital, Menlo Ventures, SV Angel
Enterprise
site ↗
#5

Datacurve

Commercial
Datacurve is a YC W24 data vendor that says it supplies frontier labs with expert-sourced RL environments, long-horizon tasks, agent trajectories, SFT demonstrations and off-the-shelf datasets. It is coding-first but now also covers data science, cyber security, ML and research. Data comes from its Shipd bounty platform of vetted engineers, which has SWE, ML and data-science tracks. It publishes DeepSWE, an open (Apache-2.0) long-horizon coding benchmark used in the Artificial Analysis Coding Agent Index.
Raised
$17.7M
Headcount
11-50
Founded
2024
SOC 2
unknown
Researchers
yes
Open src
no
Backers: Chemistry (Mark Goldberg, lead Series A), Y Combinator, Balaji Srinivasan (seed)
#6

Proximal

Commercial
Proximal is a research lab in San Francisco and Bangalore that treats post-training data as a core research problem. It builds long-horizon coding RL environments and data for frontier labs and enterprises, and it publishes the FrontierSWE benchmark and verifier audits such as CyberGym-Verified. In September 2026 it raised a $15M seed led by General Catalyst at a $300M valuation and claimed $200M in annualized revenue, which is self-reported. It also said it will expand into drug discovery, chip design and legacy software modernization.
Raised
$15M ↑ $15M
Headcount
11-50
Founded
2026
SOC 2
unknown
Researchers
yes
Open src
yes
Backers: General Catalyst (lead, 2026-09 seed), Scribble Ventures (led an earlier investment, per its announcement), SV Angel
CodeLong Horizon
site ↗
#7

Gray Swan AI

Commercial
Gray Swan AI is a Pittsburgh AI security company spun out of Carnegie Mellon (Fredrikson, Kolter). It sells adversarial red teaming and runtime protection for AI models and agents through three products. Shade does automated red teaming. Cygnal enforces policies and detects indirect prompt injection at runtime, and is now also available through the TrueFoundry and Bifrost gateways. Arena is a crowdsourced red-teaming network of more than 15,000 researchers. Gray Swan is a pre-release evaluation partner to frontier labs: Anthropic, OpenAI, Google DeepMind and Meta all cite its Shade, IPI Arena or ART evaluations in their 2026 system cards and safety reports. It is not a general RL-environment vendor.
Raised
$40M ↑ $40M
Headcount
11-50
Founded
2023
SOC 2
Type II
Researchers
yes
Open src
no
Backers: Wing Venture Capital (co-lead), Madrona (co-lead), Obvious Ventures
#8

Vals AI

Commercial
Vals AI is an independent, a16z-backed third-party evaluator. It benchmarks frontier models and AI applications on economically valuable professional work (finance, legal, tax, coding, healthcare) and, since 2026, on frontier risks (recursive self-improvement, cyber, child safety). It publishes the Vals Index and more than 20 proprietary leaderboards, open-sources its Valkyrie eval infrastructure, offers the self-serve Vals Smith for custom coding benchmarks, and runs a public-sector practice from Washington, DC.
Raised
$40M ↑ $40M
Headcount
11-50
Founded
2023
SOC 2
unknown (the SOC…
Researchers
yes
Open src
yes
Backers: Andreessen Horowitz (a16z), 8VC, Pear VC
#9

HUD

Commercial
HUD (Human Union Data, YC W25, formerly hud.so) calls itself the platform for building high-quality post-training datasets. It offers an MIT-licensed SDK for defining RL environments, tasks and graders across coding, browser, computer-use and robotics; a hosted training and eval platform; and DataVendor, a marketplace where independent builders sell RL environments and data to AI labs. It announced a $16M Series A led by Standard Capital in June 2026.
Raised
$16M ↑ $16M
Headcount
11-50
Founded
2025
SOC 2
claimed
Researchers
yes
Open src
yes
Backers: Standard Capital (Series A lead), Y Combinator (W25), Exceptional Capital
Computer UseEnterprise
site ↗
#10

Halluminate

Commercial
Halluminate (YC S25, founded 2024, San Francisco) calls itself a data research lab building benchmarks and RL environments for knowledge work, starting with financial services (PE, IB, due diligence) and consulting. It raised a $30M Series A led by Oak HC/FT in October 2026 ($38.5M total). With about nine people, it says it works with four of the five leading closed-source US AI labs and has a mid-eight-figure revenue run rate (self-reported). Its flagship benchmark is Westworld Finance Diligence Bench.
Raised
$38.5M ↑ $30M
Headcount
1-10
Founded
2024
SOC 2
unknown
Researchers
yes
Open src
yes
Backers: Oak HC/FT (Series A lead), Y Combinator (S25), Orange Collective
Computer UseEnterprise
site ↗
#11

Veris AI

Commercial
Veris AI sells simulation sandboxes that recreate an agent's production environment (mocked SaaS tools, seeded databases, simulated users) so enterprises can benchmark, regression-test, and RL/SFT-train agents before deployment. In 2026 it added a self-serve tier (Veris Plus), a coding-agent verification product (Agentic SDLC), and public voice-agent benchmarks (VAmoS, VAmoS Pro) backed by arXiv papers.
Raised
$8.5M
Headcount
1-10
Founded
2025
SOC 2
claimed
Researchers
yes
Open src
partial
Backers: Decibel Ventures (lead), Acrew Capital (lead), The House Fund
Enterprise
site ↗
#12

Collinear

Commercial
Collinear AI runs a Simulation Lab of sandboxed, stateful environments with simulated users (NPCs), tools, tasks and verifiers for agent RL training and evaluation. It now positions itself mainly as a supplier of training data, environments and evals to frontier labs across cybersecurity, software engineering, computer use and complex agent behaviour. Its CWE-bench defensive-cyber benchmark is part of the Artificial Analysis Cyber Index.
Raised
$10M
Headcount
11-50
Founded
2023
SOC 2
Type I
Researchers
yes
Open src
partial
Backers: Engineering Capital, 112 Capital (11.2 Capital), B Capital
CodeComputer UseEnterprise
site ↗
#13

Chakra Labs

Commercial
Chakra Labs ('Frontier Data Laboratory') builds deterministic clones of enterprise and productivity software that expose both GUI and MCP surfaces, along with verifiable task sets and crowdsourced trajectory data for training and evaluating computer-use and tool-use agents. It delivers them through its Dojo hub with Harbor, Verifiers and Verl support. In 2026 its messaging centred on task quality (difficult, realistic, reliably verifiable, hard to hack), and it published MAGI, a 1,000-task mixed GUI and MCP benchmark on which it reports frontier models below 30% pass@1.
Raised
$10.06M ↑ $10.1M
Headcount
11-50
Founded
2024
SOC 2
unknown
Researchers
yes
Open src
yes
Pedigree: Nirmal Krishnan (co-founder and CEO per the SEC Form D signature): Joh
Computer Use
site ↗
#14

Andon Labs

Commercial
Andon Labs (YC W24, formerly Vectorview) builds long-horizon agent benchmarks (Vending-Bench, Blueprint-Bench, Drone-Bench, Butter-Bench) and runs real AI-operated businesses (Andon Market in SF, Andon Café in Stockholm, Andon FM, vending machines) as live safety testbeds. In September 2026 it launched Pion, a research-preview platform for agents that run whole businesses. It is a repeat Anthropic research partner (Project Vend, and Project Pilot with the Frontier Red Team).
Raised
–
Headcount
11-50
Founded
2023
SOC 2
unknown
Researchers
yes
Open src
?
Backers: Y Combinator (W24)
Computer UseLong Horizon
site ↗
#15

BenchFlow

Commercial
BenchFlow is a small, open-source-first 'frontier environment lab' in the Bay Area. It publishes agent benchmarks and environments: SkillsBench and ClawsBench, the env0 simulated workspaces, Robo Use and GenesisBench for embodied agents, and FrontierPhysics. It runs the PostTrain Arena for community-contributed post-training environments and maintains an Apache-2.0 runtime for RL environments, post-training and evals. It has event-level ties to Google DeepMind, Kaggle, Anthropic, Hugging Face and Prime Intellect.
Raised
–
Headcount
1-10
Founded
2024
SOC 2
unknown
Researchers
yes
Open src
yes
Backers: Founders, Inc. (confirmed via its portfolio page), Y Combinator (reported by aggregator only; no YC directory page found), Pear VC (reported)
CodeComputer UseEnterprise
site ↗
#16

Refresh

Commercial
Refresh (YC X25; appears to have grown out of Operative.sh) builds RL environments with verifiable rewards for coding, computer use (high-fidelity clones of software suites such as EHRs and enterprise apps) and 3D animation and simulation in Blender. Data ships in Harbor format. Public releases include BlenderBench, the Gauntlet 4K RLVR dataset (access by request), the open-source computer-1 harness in Harbor, and the trajectories.sh trajectory-sharing platform.
Raised
–
Headcount
1-10
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
partial
Backers: Y Combinator
CodeComputer Use
site ↗
#17

Matrices

Commercial
Matrices builds reinforcement-learning training environments for frontier AI labs to train agents that use computers and browsers like humans, described as a 'gamified replica of the internet' where thousands of agents learn via RL. The company frames its mission as 'towards self-driving computers' and says it helps labs train computer-use agents (Operator-class systems). Note: this is the correct browser-native entity (matrices.ai / LinkedIn 'matricesapp'), distinct from the similarly named 'Matrice.ai' computer-vision company and 'Matrix AI Network' blockchain project.
Raised
$5M
Headcount
11-50
Founded
2023
SOC 2
unknown
Researchers
?
Open src
no
Backers: AI Grant (Batch 3; confirmed on aigrant.com), Index Ventures (reported), Naval Ravikant (reported, single source)
Computer Use
site ↗
#18

Andromede

Commercial
Andromede is an early-stage RL data lab (reported as founded 2025 in Lausanne, backed by Unusual Ventures) that programmatically generates RL environments, tasks and verifiers from real-world data for post-training and evaluating long-horizon frontier agents. Its core product is still in private beta with unnamed partners. Since mid-2026 the EPFL-linked team, co-founded by Alexandre Sallinen and Guillaume Allègre, has published open research: the Messier cross-benchmark agent-evaluation corpus (arXiv 2607.25891; self-reported EMNLP 2026 Main), a paper on predicting task difficulty without rollouts, and the Game of Agents multi-agent testbed (ICML 2026 workshop).
Raised
–
Headcount
1-10
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
partial
Backers: Unusual Ventures
Long Horizon
site ↗
#19

Vmax

Commercial
Vmax is a San Francisco RL company and applied research lab, founded in 2025 by RL PhDs Augustine Mavor-Parker and Matthew Sargent and backed by South Park Commons and Race Capital. It aims to automate reinforcement learning by turning proprietary data and evals into new environments, with long-horizon agents as the target. Its 2026 research is about generating RL tasks and environments automatically: PopuLoRA (self-play curricula), unix-ctf (procedural shell CTF environments) and PROPEL (probe-guided task generation, with a Goodfire AI co-author). Earlier, with Martian, it released 1k Harbor-format JavaScript tasks.
Raised
–
Headcount
1-10
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
no
Backers: Race Capital, South Park Commons
CodeLong Horizon
site ↗
#20

Plato

Commercial
Plato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and computer-use agents, recreating real websites/software (e.g. Amazon/Airbnb/Gmail-style replicas) as reinforcement-learning environments with structured APIs for interaction, state tracking and scoring. It also offers a 'Computer Use' capability driving a full Linux desktop, positioning at the intersection of browser interaction and enterprise workflow simulation.
Raised
–
Headcount
1-10
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
partial
Pedigree: Pranav Putta (Co-founder/CTO), prior MultiOn, Georgia Institute of Te
Computer UseEnterprise
site ↗
#21

AIChamp

Commercial
AIChamp, formerly a vendor of RL environments ('Virtual Gyms') and expert data for enterprise AI agents, repositioned in 2026 as a vendor-neutral factory digital-twin service. It runs a free 14-day loss check on a manufacturing line, builds a simulation of the costliest station, and invites robot-AI software teams to train and be tested in that twin against a plant-set target before a floor pilot. Under its 'Shared' option it also licenses de-identified factory scenes to robotics and AI companies. It no longer markets RL environments for LLM or software agents.
Raised
$0
Headcount
1-10
Founded
?
SOC 2
unknown
Researchers
?
Open src
no
Pedigree: Self-claimed: 'Built by engineers from Microsoft, Amazon Robotics, Goo
EnterpriseLong Horizon
site ↗
#22

Habitat Inc

Commercial
Habitat Inc is a very early-stage San Francisco company (2-10 employees). Third-party market maps (Chemistry VC, SemiAnalysis, AlignList) list it as building reinforcement-learning environments for code and computer-use agent workflows, made of programmatically verifiable problems. Until mid-2026 its own LinkedIn described it as 'RL environments for white-collar work'. By October 2026 that tagline read 'Autonomous firms.', which may signal a repositioning. It has disclosed no funding, customers, publications or security certifications, and its website is a placeholder.
Raised
–
Headcount
1-10
Founded
?
SOC 2
unknown
Researchers
?
Open src
no
Pedigree: Maxim Enis (co-founder), Williams College '24; prior Ramp association
CodeComputer UseEnterprise
site ↗
#23

Taste Labs

Commercial
Taste Labs (Taste) is a San Francisco startup, founded in October 2025 by former Exa growth lead Thais Castello Branco, that calls itself 'the taste layer for AI' and a hybrid research lab and infrastructure company. For model labs, it builds post-training data and RL environments for subjective domains, starting with design. Amplify describes these as preference datasets, reasoning data, rubrics and evaluation environments. The data is made by its TasteMakers network of expert creatives. Taste says it works with 'the top frontier labs' but names none. In September 2026 it launched a self-serve Brand API and MCP server that lets agents extract, search and verify brand design systems. It raised an $18.5M seed co-led by CRV and Amplify Partners, announced 2026-06-16.
Raised
$18.5M ↑ $18.5M
Headcount
11-50
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
no
Backers: CRV, Amplify Partners
#24

Idler

Commercial
Idler (YC S25, founded 2025, San Francisco) calls itself a frontier data research lab that builds evals and RL environments from real production work. Its public benchmarks are ShelfLife, a digital twin of a live multi-brand e-commerce company; ShelfLife E-Sim, a stateful retail-operations simulator; and CorpLaw, built from a law firm's anonymized data. It also offers private, on-request collections in long-horizon SWE, cybersecurity, recursive self-improvement and terminal agents. The roughly 12-person team is led by Dark Forest/0xPARC co-founder Ivan Chub and says leading frontier labs use its work, without naming any. A $9M Paradigm-led seed was reported by an aggregator in August 2026 but is not confirmed by the company.
Raised
$9M ↑ $9M
Headcount
11-50
Founded
2025
SOC 2
unknown
Researchers
?
Open src
no
Backers: Paradigm (seed lead, reported), Y Combinator (S25)
CodeEnterpriseLong Horizon
site ↗
n/r

Scale AI

Incumbent
Scale AI is the data-labeling and AI-data incumbent that has extended into RL environments (simulated web apps, macOS/Windows-like desktop VMs and MCP-tool environments with expert rubrics and automated verifiers). It runs a large public evaluation program through Scale Labs (SWE-Bench Pro, MCP Atlas, Humanity's Last Exam, Remote Labor Index). After Meta's ~$14.3B deal in June 2025 (~49% non-voting stake) and Alexandr Wang's move to Meta, several frontier labs reportedly scaled back work. In 2026 Scale grew mainly in government and enterprise: the Pentagon CDAO agreement ceiling rose to $500M, it acquired ICG Solutions, and Francis deSouza (ex-Google Cloud COO) became CEO in August 2026.
Raised
$1.6B
Headcount
200+
Founded
2016
SOC 2
Type II
Researchers
yes
Open src
no
Backers: Meta Platforms, Accel, Amazon
CodeComputer UseEnterprise
site ↗
n/r

Snorkel AI

Incumbent
Snorkel AI is a San Francisco company spun out of the Stanford AI Lab in 2019 and known for the Snorkel weak-supervision project. It now calls itself 'the frontier AI data lab'. In 2025 it moved from selling Snorkel Flow software to an expert Data-as-a-Service model, and it now supplies expert agentic tasks, RL environments (computer-use, terminal and simulated-enterprise), rubrics and evals to AI labs and enterprises. It combines expert contributors with synthetic generation and agent-based QC, maintains Terminal-Bench, and co-authors or hosts agent benchmarks such as OSWorld 2.0, Agents' Last Exam and Senior SWE-bench through a $3M Open Benchmarks Grants program. It raised a $350M Series E at $3.5B in September 2026 and says its ARR passed $375M.
Raised
$585M ↑ $350M
Headcount
201-500
Founded
2019
SOC 2
Type II
Researchers
yes
Open src
partial
Backers: Insight Partners (co-led Series E), S32 (co-led Series E), Addition (led Series D, co-led Series C)
CodeComputer UseEnterprise
site ↗
n/r

Modal

Infrastructure
Modal (Modal Labs) is a New York-based, Python-native serverless cloud for AI workloads. It offers on-demand GPU/CPU compute, multi-node clusters with RDMA, managed LLM inference endpoints, and fast-booting sandboxes, including full VM sandboxes, which Modal says scale to 1M concurrent. It sells execution infrastructure, not RL environments. AI labs, agent companies and RL-data vendors use it to run RL rollouts, post-training and large fleets of parallel sandboxed environments. It integrates with Anthropic's Claude Managed Agents and Claude Science, Cognition's Devin Outposts and the OpenAI Agents SDK, and ships an open-source RL/SFT library (Modal Dojo).
Raised
$466M ↑ $355M
Headcount
51-200
Founded
2021
SOC 2
Type II
Researchers
yes
Open src
no
Backers: General Catalyst (Series C co-lead), Redpoint Ventures (Series C co-lead; earlier Series A lead), Lux Capital (Series B lead)
n/r

Mercor

Incumbent
Mercor is a venture-backed expert marketplace and AI-training-data company. It supplies RLHF data, evaluations and reinforcement-learning environments to frontier AI labs and enterprises, drawing on a professional expert network: the homepage cites 30k+ experts, and Mercor's Deeptune post claims more than five million. It started as an AI-recruiting platform, moved into human data and RL, bought Sepal AI (Feb 2026), and announced the acquisition of Deeptune (Jul 2026), which builds simulated enterprise-app 'training gyms'. It publishes the APEX family of professional-task agent benchmarks. Mercor says it reached $2B ARR in June 2026, despite a March 2026 supply-chain data breach.
Raised
$492M
Headcount
51-200
Founded
2023
SOC 2
unknown
Researchers
yes
Open src
partial
Backers: Felicis Ventures (led Series C and Series B), Benchmark, General Catalyst
CodeEnterpriseLong Horizon
site ↗
n/r

Prime Intellect

Open source
Prime Intellect sells an open-source-first stack for agentic RL: the Environments Hub (2,500+ community environments plus 365k+ unified SWE, terminal and search tasks), the verifiers, prime-rl and renderers libraries, Lab hosted RL post-training and evals (GA), Prime Sandboxes microVMs (GA), Prime Inference, and GPU compute. After a $130M Series A in July 2026 (TechCrunch reports a $1B valuation) it states $100M+ annualized revenue and 6,000+ customers. It also trains the open INTELLECT models and is the post-training partner in NVIDIA's Nemotron Coalition.
Raised
$150M ↑ $130M
Headcount
11-50
Founded
2023
SOC 2
claimed
Researchers
yes
Open src
yes
Backers: Radical Ventures (led $130M Series A, 2026), NVIDIA Ventures, Intel Capital
CodeEnterpriseLong Horizon
site ↗
n/r

Daytona

Infrastructure
Daytona provides secure, elastic, programmatic sandboxes (Linux containers, Linux and Windows VMs, and GPU sandboxes) that AI agents spin up in under ~90ms (vendor claim) to run untrusted AI-generated code in isolated, stateful runtimes with snapshot and fork primitives. It is a managed service with a Bring-Your-Own-Compute option and has been closed source since June 2026. It is an integrated sandbox provider for Anthropic Claude Managed Agents, OpenAI Agents SDK Sandbox Agents and Cognition Devin Outposts, and targets code execution, computer use and RL/eval rollouts.
Raised
$79M ↑ $48.274M
Headcount
11-50
Founded
2023
SOC 2
Type II
Researchers
yes
Open src
no
Backers: FirstMark Capital (Series A lead; Matt Turck on board and listed as a director on the Aug 2026 Form D), Pace Capital, Upfront Ventures (seed lead, Series A participant)
CodeComputer Use
site ↗
n/r

Turing

Incumbent
Turing is a large incumbent supplier of expert human data to frontier AI labs. It claims 5M+ experts and work with '9/9 frontier labs' (both self-claimed). In 2026 it moved further into verifiable agent environments and evaluations. It launched Turing Frontier in April, set up the Turing Frontier Research Lab, and published benchmarks including CyberStrike (Harbor format, public on Hugging Face), KernelQuest, CEO Bench, SciCode++ and PLSQLBench (with Oracle, per Turing). It also contributed tasks to Terminal-Bench 3.0 and hired a research-heavy CTO, Ece Kamar (ex-Microsoft Research).
Raised
$300M
Headcount
200+
Founded
2018
SOC 2
unknown
Researchers
yes
Open src
partial
Pedigree: Microsoft Research (CTO Ece Kamar, ex-CVP AI Frontiers Lab)
CodeEnterprise
site ↗
n/r

micro1

Incumbent
micro1 is a San Francisco AI training-data company founded in 2022 by Ali Ansari. It began as an AI recruiter (Zara) and now uses that system to vet expert contractors. It sells expert human data, frontier evaluations and RL environments to AI labs under the Realm brand, an enterprise agent-evaluation layer called Cortex, and expert-demonstrated robotics data, all run on its flow data platform. To make its 'RL gyms' realistic it pays companies for operational data, and it made a $12.5M late bid for Spirit Airlines' records. TechCrunch reported a $500M gross run rate in August 2026, and Forbes reported a raise of more than $100M at a $4B valuation in September 2026, with frontier labs taking part.
Raised
$138M ↑ $100M
Headcount
51-200
Founded
2022
SOC 2
unknown
Researchers
yes
Open src
partial
Backers: 01 Advisors (Dick Costolo, Adam Bain; led Series A; Bain on board), Two unnamed frontier AI labs (2026 round, reported), Two unnamed xAI co-founders (2026 round, reported)
Enterprise
site ↗
acq

Deeptune

Commercial
Deeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms') for computer-use and code, where AI agents practice and are evaluated on realistic digital knowledge-work tasks (simulating tools like Slack and Salesforce). It sold these pre-built environments primarily to frontier AI labs and raised a $43M Series A led by a16z (March 2026). In July 2026 it was acquired by Mercor; the team joined Mercor and Deeptune's environment platform now sits under Mercor. It is therefore no longer ranked as an independent vendor.
Raised
$43M
Headcount
11-50
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
no
Backers: Andreessen Horowitz (a16z, lead), 776, Abstract Ventures
CodeComputer UseEnterprise
site ↗
n/r

General Reasoning

Open source
General Reasoning is a London AI research lab (legal entity General Reasoning, Inc., US) working on models that operate over much longer time horizons. It builds OpenReward, a platform for serving RL environments (the vendor states 380+), and the open Open Reward Standard. It also ships open tooling (Firehorse, inspect-openreward, Harbor support), the KellyBench long-horizon benchmark and BackSearch, a paid point-in-time search API. It says its own language models launch in Q4 2026.
Raised
$10.9M
Headcount
11-50
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
yes
Backers: Entrepreneur First
CodeLong Horizon
site ↗
n/r

E2B

Infrastructure
E2B provides open-source Firecracker-microVM sandboxes for AI agents. It is offered as a hosted API, as BYOC, and since September 2026 as a self-hostable single-node package (E2B Embed). In 2026 it became a supported sandbox provider in the OpenAI Agents SDK and published integrations for Devin Outposts and Cursor Self-Hosted Machines. Vendor case studies show it used for RL rollouts (Paper Instruments) and large-scale evals (Arena).
Raised
$32M
Headcount
11-50
Founded
2023
SOC 2
Type II
Researchers
no
Open src
yes
Backers: Insight Partners (Series A lead), Decibel (seed lead), Sunflower Capital
CodeComputer Use
site ↗
n/r

Runloop

Infrastructure
Runloop sells cloud VM-isolated devboxes and sandboxes for AI agents. Around them it now offers agent lifecycle tooling: the Agent API, Axons and Broker event streams, Agent Gateway and MCP Hub for credential isolation, Benchmark Job Orchestration with W&B Weave, Reflex background agents, and VPC deployment. It is execution and eval infrastructure for agent builders and training workloads, not an RL-data or environments vendor.
Raised
$7M
Headcount
11-50
Founded
2024
SOC 2
claimed
Researchers
no
Open src
yes
Backers: The General Partnership (lead), Blank Ventures, Exponent Founders Capital
acq

Mechanize

Commercial
Mechanize is a San Francisco vendor founded in April 2025 by ex-Epoch AI researchers. It builds a small number of high-fidelity, long-horizon RL environments and evals for frontier coding agents, and publishes the open-source GBA Eval benchmark. It raised $9.1M at a $500M post-money valuation in April 2026. In September 2026 Business Insider reported that Google had completed a license-and-hire deal: co-founder and CEO Tamay Besiroglu and more than a dozen staff joined Google DeepMind, and talks had been reported at more than $1.5B. Mechanize continues under CEO Guive Assadi.
Raised
$9.1M ↑ $9.1M
Headcount
11-50
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
partial
Backers: Nat Friedman, Daniel Gross, Patrick Collison
n/r

Cua

Infrastructure
Cua (trycua, YC X25) builds open-core 'computer-use 2.0' infrastructure: an MIT, agent-neutral background driver for macOS, Windows and Linux (CLI/MCP), local and cloud sandboxes, elastic Cua Fleets for parallel eval and RL training (cloud, BYOC, on-prem), Lume macOS virtualization, the Cua Spaces Mac app, small CUA-S1 decision models, and Cua-Bench for verifiable tasks, leaderboards and RL trajectory export. SOC 2 Type I.
Raised
$500K
Headcount
1-10
Founded
2025
SOC 2
Type II
Researchers
yes
Open src
yes
Backers: Y Combinator (X25 batch)
Computer Use
site ↗
n/r

Good Start Labs

Open source
Good Start Labs is a 2025 Every spin-out that sells game-based RL environments and agent and human gameplay data to AI labs. It also partners with game publishers (Bad Cards, Arkadium GameLab) to turn play into licensed training data, in-game agents and leaderboards. It publishes transfer studies from games to real work (Diplomacy to customer support, the 1830 railroad game to finance research), arXiv papers and open repos, and runs the Diplomacy and Humor (Bad Cards) Arena leaderboards.
Raised
$3.6M
Headcount
1-10
Founded
2025
SOC 2
unknown
Researchers
yes
Open src
yes
Backers: General Catalyst, Inovia Capital, Tirta Ventures
Long Horizon
site ↗
n/r

Surge AI

Incumbent
Surge AI is a bootstrapped human-data and RLHF leader for frontier labs (reported $1.2B 2024 revenue). In 2026 it built a large in-house benchmark and RL-environment program (DAYJOB, GDP.pdf/.xlsx, HANDBOOK.md, Chartography, CoreCraft, and the Tuesday Work Index composite), published post-training transfer studies, launched Post-Training Runs and an enterprise offering, and released sample harnesses under Apache-2.0. Anthropic reports results on Surge's GDP.pdf in its Claude Sonnet 5 system card, and Microsoft publicly credits Surge's expert raters. OpenAI's adoption of GDP.pdf is so far claimed only by Surge.
Raised
$0
Headcount
51-200
Founded
2020
SOC 2
unknown
Researchers
yes
Open src
partial
Pedigree: Founder/CEO Edwin Chen: former research scientist at Google, Facebook
EnterpriseLong Horizon
site ↗
acq

Sepal AI

Commercial
Sepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data, expert-graded evaluation benchmarks, and reinforcement-learning environments for frontier LLMs, drawing on a network of 20k+ domain experts (PhDs, finance, medical, STEM). It was acquired by Mercor in February 2026; the team joined Mercor and its RL-environment and human-data work now sits under Mercor. It is therefore no longer ranked as an independent vendor.
Raised
$500K
Headcount
?
Founded
2024
SOC 2
unknown
Researchers
yes
Open src
no
Backers: Y Combinator, Metaplanet Holdings, SID Venture Partners
EnterpriseLong Horizon
site ↗
n/r

Morph

Infrastructure
Morph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal technology, which can snapshot, branch, and restore entire computational environments in roughly 100-250ms to enable massively parallel, reversible ('Git for compute') agent rollouts, evaluations, and reasoning-time branching. It markets the platform (Morph Cloud) as infrastructure for running and scaling agent/RL verification environments rather than as an RL-environment dataset vendor itself.
Raised
–
Headcount
1-10
Founded
2023
SOC 2
claimed
Researchers
yes
Open src
partial
Backers: Christian Szegedy (reported seed/angel investor; also Chief Scientist), amount undisclosed
#Company
LatestConf.
#1 CodeComputer UseEnterprise Commercial$30.5M ↑ $30M 51-200 unknown San Francisco, USA1 Sep 2026Forbes reports AfterQuery is raising at a $3.2B valuation, YC's…
#2 Long Horizon Commercial$40M ↑ $40M 11-50 unknown Mountain View, California, USA19 Sep 2026SeeWhy: small specialized model post-trained with Intuit Credit Karma
#3 CodeLong Horizon Commercial$6M 11-50 Type II London, United KingdomSep 2026Huzzle Labs publishes HuzzleWorld, a family of small computer-use…
#4 Enterprise Commercial$15M 11-50 unknown New York, NY, USASep 2026Fleet's GitHub shows active RL post-training work: forks of SkyRL,…
#5 Code Commercial$17.7M 11-50 unknown San Francisco, USA5 Oct 2026Product line now spans multiple RL-environment domains; Shipd lists…
#6 CodeLong Horizon Commercial$15M ↑ $15M 11-50 unknown San Francisco, CA, USA5 Oct 2026Careers page lists 5 open roles across San Francisco and Bangalore
#7 – Commercial$40M ↑ $40M 11-50 Type II Pittsburgh, Pennsylvania, USA1 Oct 2026Gray Swan integrates Cygnal with Bifrost, Maxim's AI gateway
#8 – Commercial$40M ↑ $40M 11-50 unknown (the SOC… San Francisco, USA2 Oct 2026Vals Web Search Index blog post: search tools tested on expert legal…
#9 Computer UseEnterprise Commercial$16M ↑ $16M 11-50 claimed San Francisco, USA5 Oct 2026Hiring surge: about 30 open roles across SF, Singapore and remote…
#10 Computer UseEnterprise Commercial$38.5M ↑ $30M 1-10 unknown San Francisco, CA, USA1 Oct 2026Halluminate raises a $30M Series A led by Oak HC/FT; total funding…
#11 Enterprise Commercial$8.5M 1-10 claimed San Francisco, CA, USA1 Oct 2026Veris releases VAmoS Pro Bench, a voice-agent benchmark on utility…
#12 CodeComputer UseEnterprise Commercial$10M 11-50 Type I Sunnyvale, California, USA28 Sep 2026CWE-bench v1 released: 120 tasks, 73 CWEs, 8 languages
#13 Computer Use Commercial$10.06M ↑ $10.1M 11-50 unknown Brooklyn, New York, USAOct 2026Hiring Member of Technical Staff (New York) for agent orchestration…
#14 Computer UseLong Horizon Commercial– 11-50 unknown San Francisco, USA24 Sep 2026Opus 5.5, GPT-6 Sol and Grok 4.7 evaluated on Vending-Bench
#15 CodeComputer UseEnterprise Commercial– 1-10 unknown San Francisco Bay Area, CA, USA21 Sep 2026Robo Use released on PyPI: 'agent as a policy for embodied agents'
#16 CodeComputer Use Commercial– 1-10 unknown San Francisco, CA, USA22 Sep 2026Refresh publishes BlenderBench, a 20-task benchmark of Blender 3D…
#17 Computer Use Commercial$5M 11-50 unknown San Francisco, California, USA29 Apr 2026Matrices listed in Specter Monitor #225 'Investor Signals of the…
#18 Long Horizon Commercial– 1-10 unknown Lausanne, SwitzerlandOct 2026Messier to be presented at EMNLP 2026 Main Conference; difficulty…
#19 CodeLong Horizon Commercial– 1-10 unknown San Francisco, USA10 Jun 2026Vmax publishes PROPEL: training task generators at the learnable…
#20 Computer UseEnterprise Commercial– 1-10 unknown San Francisco, CA, USA1 Oct 2026Plato ships plato-sdk-v2 2.161.0 on PyPI (MIT-licensed Python SDK,…
#21 EnterpriseLong Horizon Commercial$0 1-10 unknown San Francisco (Mission Bay), USA1 Oct 2026New Loss Check Terms add a 'Shared' option that licenses…
#22 CodeComputer UseEnterprise Commercial– 1-10 unknown San Francisco, CA, USAOct 2026LinkedIn company page now reads 'Autonomous firms.' and lists a San…
#23 – Commercial$18.5M ↑ $18.5M 11-50 unknown San Francisco, CA, USA5 Oct 2026Case study: General Intelligence Company of New York uses Taste's…
#24 CodeEnterpriseLong Horizon Commercial$9M ↑ $9M 11-50 unknown San Francisco, CA, USA26 Aug 2026ShelfLife E-Sim leaderboard published: long-horizon retail-operations…
n/r CodeComputer UseEnterprise Incumbent$1.6B 200+ Type II San Francisco, California, USA22 Sep 2026SWE-Bench Pro v2: 642 public tasks, a 51-task Hard split, a 272-task…
n/r CodeComputer UseEnterprise Incumbent$585M ↑ $350M 201-500 Type II San Francisco, CA, USA2 Oct 2026Snorkel grant supports MedPAIR, a dataset comparing physician and LLM…
n/r Code Infrastructure$466M ↑ $355M 51-200 Type II New York, NY, USA1 Oct 2026Runtime by Modal: VM Sandboxes, Modal Clusters and Sticky Sessions…
n/r CodeEnterpriseLong Horizon Incumbent$492M 51-200 unknown San Francisco, CA, USA (181 Fremont)1 Oct 2026Human-baselines study: frontier AI outperforms junior CPAs on…
n/r CodeEnterpriseLong Horizon Open source$150M ↑ $130M 11-50 claimed San Francisco, USA2 Oct 2026Prime Inference launches: fast, reliable serving for frontier open…
n/r CodeComputer Use Infrastructure$79M ↑ $48.274M 11-50 Type II New York, NY, United States3 Oct 2026Legacy daytonaio/daytona GitHub repo archived (read-only)
n/r CodeEnterprise Incumbent$300M 200+ unknown San Francisco, CA, USA18 Sep 2026SciCode++ scientific coding dataset introduced
n/r Enterprise Incumbent$138M ↑ $100M 51-200 unknown San Francisco, CA, USA24 Sep 2026micro1 publishes flow-transform 1.0 for corpus-level PII transformation
acq CodeComputer UseEnterprise Commercial$43M 11-50 unknown New York, NY, USA10 Aug 2026Mercor hiring for its 'Deeptune' RL environments team in NYC
n/r CodeLong Horizon Open source$10.9M 11-50 unknown London, United Kingdom (Shoreditch), operating research hub; legal entity General Reasoning, Inc. registered in San Francisco/US per SEC Form DOct 2026Homepage announces a new generation of GR language models for Q4 2026
n/r CodeComputer Use Infrastructure$32M 11-50 Type II San Francisco, USA1 Oct 2026Arena case study: up to 600,000 E2B sandboxes per day
n/r Code Infrastructure$7M 11-50 claimed San Francisco, CA, USA21 Aug 2026Case study: Accrual built its ARC remote coding-agent platform on…
acq Code Commercial$9.1M ↑ $9.1M 11-50 unknown San Francisco, USA11 Sep 2026Business Insider: Google completes the Mechanize talent deal;…
n/r Computer Use Infrastructure$500K 1-10 Type II San Francisco, CA, USA1 Oct 2026Cua launches Cua Spaces v0.1.0, a source-available Mac app for…
n/r Long Horizon Open source$3.6M 1-10 unknown Brooklyn, NY, USA1 Oct 2026"Keep Going": 48-hour livestream of GPT-6 Astra (Codex) and Claude…
n/r EnterpriseLong Horizon Incumbent$0 51-200 unknown San Francisco, California, USA5 Oct 2026Forbes profile: Surge still claims no outside funding; Chen owns…
acq EnterpriseLong Horizon Commercial$500K – unknown San Francisco, USA–
n/r Code Infrastructure– 1-10 claimed San Francisco, USA5 Oct 2026Morph pulls most of its public research and open-source material:…
No vendors match these filters.
What's new · last 6 months

The latest across the RL-environment market

The most recent sourced developments across every vendor we track: rounds, acquisitions, launches, benchmarks and partnerships. Each links to its source, and the full dated timeline sits on each vendor's page.

5 Oct 2026DatacurveProduct line now spans multiple RL-environment domains; Shipd lists SWE, ML and data-science trackssource ↗
5 Oct 2026ProximalCareers page lists 5 open roles across San Francisco and Bangaloresource ↗
5 Oct 2026HUDHiring surge: about 30 open roles across SF, Singapore and remote (Asia and North America)source ↗
5 Oct 2026Taste LabsCase study: General Intelligence Company of New York uses Taste's Brand API for brand import in Cofoundersource ↗
5 Oct 2026Surge AIForbes profile: Surge still claims no outside funding; Chen owns about 75% and is worth $18Bsource ↗
5 Oct 2026MorphMorph pulls most of its public research and open-source material: GitHub org down to two SDK repos, Morph Prover model and Trinity post gonesource ↗
3 Oct 2026DaytonaLegacy daytonaio/daytona GitHub repo archived (read-only)source ↗
2 Oct 2026Vals AIVals Web Search Index blog post: search tools tested on expert legal and finance tasks, with partners Exa, Keenable, Parallel and Tavilysource ↗
2 Oct 2026Snorkel AISnorkel grant supports MedPAIR, a dataset comparing physician and LLM relevance judgmentssource ↗
2 Oct 2026Prime IntellectPrime Inference launches: fast, reliable serving for frontier open modelssource ↗
1 Oct 2026Gray Swan AIGray Swan integrates Cygnal with Bifrost, Maxim's AI gatewaysource ↗
1 Oct 2026Vals AIHarvey's Legal Agent Benchmark run by Vals as an 'Industry Partner' benchmark (leaderboard as of 2026-10-01)source ↗
1 Oct 2026HalluminateHalluminate raises a $30M Series A led by Oak HC/FT; total funding now $38.5Msource ↗
1 Oct 2026HalluminateCompany says four of the five leading closed-source US AI labs are customers (self-claimed, none named)source ↗
1 Oct 2026Veris AIVeris releases VAmoS Pro Bench, a voice-agent benchmark on utility billing calls, with a public leaderboard of 14 stackssource ↗
1 Oct 2026PlatoPlato ships plato-sdk-v2 2.161.0 on PyPI (MIT-licensed Python SDK, labelled Beta)source ↗
1 Oct 2026AIChampNew Loss Check Terms add a 'Shared' option that licenses de-identified factory scenes to robotics and AI companiessource ↗
1 Oct 2026ModalRuntime by Modal: VM Sandboxes, Modal Clusters and Sticky Sessions reach GA; Endpoint Candidates enters private betasource ↗
The ranking · 2026

The Top 24 RL Environment Companies in 2026

The reinforcement-learning environment market went from a footnote to a procurement line item in under two years. This ranks the 24 dedicated, pure-play RL-environment vendors by the RL List score, a transparent, documented blend of scale and traction (funding, customers), security and research signals, and how much of each company’s record we could independently verify. It is a formula, not an opinion. Incumbents (Scale AI, Surge AI, Mercor), execution-infrastructure providers, and open-source projects are a different category and are listed separately below rather than ranked. Every figure links to its source.

Ranked by the RL List score · last updated 2026-10-05 · see how it’s computed in the methodology

1

AfterQuery

Commercial

AfterQuery is a San Francisco applied-research lab and expert-data company (YC W25). It supplies AI labs with expert-generated SFT and RL data, rubrics, agent environments and computer-use trajectories, drawn from a stated network of nearly 100,000 verified professionals. Forbes reported it was raising at a $3.2B valuation in Sept 2026. NVIDIA's Nemotron 3 Ultra report names AfterQuery tasks in its GDPval training recipe. The company publishes or co-builds real-task benchmarks (FinanceQA, VADER, IDE-Bench, SpreadsheetBench 2, Legora BAR) and is adding forward-deployed enterprise work.

Backed by Altos Ventures (lead, Series A), The Raine Group, Y Combinator, BoxGroup. Team pedigree: Founders: Spencer Mateega (CEO; Wharton/Penn; ex-Silver Lake, Morgan Stanley, Meta, Google), Carlos Georgescu (CTO; UBC CS; ex-Citadel Securities, Meta, Google; earlier founded an acquired ed-tech startup); Early co-author Danny Tang (FinanceQA, VADER); current role unverified. Named customers include NVIDIA (Nemotron 3 Ultra; AfterQuery tasks named in NVIDIA's technical report), Legora (BAR benchmark; credited on legora.com/bar), Thinking Machines Lab (reported by Forbes; not confirmed by TML), Motif Technologies (Motif 3; 'sole data partner' per AfterQuery, also named by TechCrunch; not credited on Motif's model card), The Raine Group (Raine Search; Raine exec quoted in AfterQuery post; also an investor), Unnamed Chinese AI labs (reported by Forbes), Frontier AI labs (company claims every leading lab) (verified, incl. frontier-lab ties).

$30.5M raisedFounded 2025San Francisco, USA51-200 staffSOC 2: unknown

Best fit: A frontier or enterprise AI team needing expert-authored RL environments, post-training data, and realistic real-task benchmarks across code, finance, and professional workflows.

2

Bespoke Labs

Commercial

Bespoke Labs is an applied AI research lab in Mountain View, founded in 2024 by Mahesh Sathiamoorthy (ex-Google DeepMind) and UC Berkeley professor Alex Dimakis. It builds RL environments and infrastructure for training and evaluating reliable agents, and is backed by $40M from Wing VC, 8VC, Mayfield and others. It combines commercial environment and post-training work, such as a jointly published project with Intuit Credit Karma, with a large open-research output: OpenThoughts, Curator, GEPA, Terminal-Bench, AutoResearchExam and ConjectureBench.

Backed by Wing VC (Series A lead), 8VC (Seed lead), Mayfield, The House Fund. Team pedigree: Co-founder/CEO Mahesh Sathiamoorthy: ex-Google Brain/DeepMind staff engineer; B.Tech IIT Kharagpur, MS/PhD USC; Co-founder/Chief Science Officer Alex Dimakis: Professor, UC Berkeley EECS (formerly UT Austin, USC); IEEE Fellow; co-director, NSF AI Institute for Foundations of ML. Named customers include Intuit / Credit Karma, Fortune 500 enterprises (unnamed), Frontier labs (unnamed), Model builders using OpenThoughts datasets (unnamed) (verified, incl. frontier-lab ties).

$40M raisedFounded 2024Mountain View, California, USA11-50 staffSOC 2: unknown

Best fit: Buyers needing reasoning-focused data curation, open reproducible datasets/recipes, and custom RL-environment/eval data delivery from a research-led team.

3

Huzzle Labs

Commercial

Huzzle Labs is the applied-AI division of the London talent platform Huzzle (co-founders Ingmar Klein, CEO, and Amit Choudhary, CTO). It builds long-horizon RL environments (coding, computer use, enterprise workflows) and expert trajectory data for frontier labs, drawing on Huzzle's talent network, which it says numbers 300k+. In 2026 it expanded into its own small computer-use models (HuzzleWorld, 1B-8B, self-reported results), a domain benchmark (InsureBench, scores pending), and enterprise custom evals and models deployed on client infrastructure.

Backed by 10x Founders, Angel Invest, Emerge, a16z Scout Fund. Team pedigree: RL engineer, ex-Turing; Researchers from IIT Kharagpur and IIT Bombay (incl. PhD). Named customers include Apple, Lazard, Financial Times (self-claimed).

$6M raisedFounded 2020London, United Kingdom11-50 staffSOC 2: Type II

Best fit: Frontier labs and regulated enterprises needing custom RL environments plus expert human trajectory data and evals for code, computer-use, and long-horizon professional workflows.

4

Fleet AI

Commercial

Fleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise software such as Salesforce and Excel, plus browser/desktop workflows, so frontier AI labs and large enterprises can train and evaluate computer-use agents. It ships a Python SDK, a platform API, and the open-source 'Harbor' agent-evaluation/RL-environment tooling, pairing simulated environments with human supervision.

Backed by Sequoia Capital, Menlo Ventures, SV Angel, Bain Capital Ventures (reported Series A lead/co-lead, April 2026; not listed on Fleet's site). Team pedigree: Team self-describes prior experience at Anthropic, xAI, Meta Superintelligence, Essential AI, Contextual AI, Mercor, Docker, Citadel, Jane Street and Cruise; Founder/CEO Nicolai (Nic) Ouporov: ex-founding engineer at Respell (acquired by Salesforce, Jan 2024); research at the Stanford Robotics & Embodied AI Lab and the Columbia Creative Machines Lab.

$15M raisedFounded 2024New York, NY, USA11-50 staffSOC 2: unknown

Best fit: A frontier lab or large enterprise that needs bespoke, high-fidelity RL environments simulating real enterprise software (CRM, spreadsheets, browser/desktop) to train and evaluate computer-use agents.

5

Datacurve

Commercial

Datacurve is a YC W24 data vendor that says it supplies frontier labs with expert-sourced RL environments, long-horizon tasks, agent trajectories, SFT demonstrations and off-the-shelf datasets. It is coding-first but now also covers data science, cyber security, ML and research. Data comes from its Shipd bounty platform of vetted engineers, which has SWE, ML and data-science tracks. It publishes DeepSWE, an open (Apache-2.0) long-horizon coding benchmark used in the Artificial Analysis Coding Agent Index.

Backed by Chemistry (Mark Goldberg, lead Series A), Y Combinator, Balaji Srinivasan (seed), angel investors who are employees of DeepMind, Vercel, Anthropic and OpenAI (individuals, not the companies). Team pedigree: Serena Ge (co-founder/CEO): worked on LLM reasoning during a co-op at Cohere; University of Waterloo CS; Forbes 30 Under 30; Charley Lee (co-founder): University of Waterloo CS; AI research background.

$17.7M raisedFounded 2024San Francisco, USA11-50 staffSOC 2: unknown

Best fit: Frontier/foundation model labs needing expert-sourced coding SFT/RLHF data and code-execution RL environments with verifiable rewards (code execution, tight loops).

6

Proximal

Commercial

Proximal is a research lab in San Francisco and Bangalore that treats post-training data as a core research problem. It builds long-horizon coding RL environments and data for frontier labs and enterprises, and it publishes the FrontierSWE benchmark and verifier audits such as CyberGym-Verified. In September 2026 it raised a $15M seed led by General Catalyst at a $300M valuation and claimed $200M in annualized revenue, which is self-reported. It also said it will expand into drug discovery, chip design and legacy software modernization.

Backed by General Catalyst (lead, 2026-09 seed), Scribble Ventures (led an earlier investment, per its announcement), SV Angel, Go Global Ventures (Diede van Lamoen). Team pedigree: Justus Mattern (co-founder): built open RL infrastructure at Prime Intellect, co-founded Revideo (YC S23), early engineer at Dynamo AI, research on ML and differential privacy (confirmed via justusmattern.com); Calvin Chen (co-founder): prior-exit specifics not corroborated.

$15M raisedFounded 2026San Francisco, CA, USA11-50 staffSOC 2: unknown

Best fit: Frontier labs or AI startups needing long-horizon, real-codebase RL environments and quality-aware (fuzzy) verifiers to post-train coding agents.

7

Gray Swan AI

Commercial

Gray Swan AI is a Pittsburgh AI security company spun out of Carnegie Mellon (Fredrikson, Kolter). It sells adversarial red teaming and runtime protection for AI models and agents through three products. Shade does automated red teaming. Cygnal enforces policies and detects indirect prompt injection at runtime, and is now also available through the TrueFoundry and Bifrost gateways. Arena is a crowdsourced red-teaming network of more than 15,000 researchers. Gray Swan is a pre-release evaluation partner to frontier labs: Anthropic, OpenAI, Google DeepMind and Meta all cite its Shade, IPI Arena or ART evaluations in their 2026 system cards and safety reports. It is not a general RL-environment vendor.

Backed by Wing Venture Capital (co-lead), Madrona (co-lead), Obvious Ventures, Snowflake Ventures. Team pedigree: Zico Kolter (Co-founder, Chief Scientist): Professor and Department Head of the CMU Machine Learning Department; OpenAI board member and chair of its Safety and Security Committee; Matt Fredrikson (Co-founder, CEO): CMU faculty, adversarial-ML researcher, GCG co-author. Named customers include Anthropic, OpenAI, Meta, Google DeepMind, xAI, Amazon, Snowflake, ByteDance, ElevenLabs, Intercom, Deloitte, UK AI Security Institute (AISI), OpenHands, AIUC (verified, incl. frontier-lab ties).

$40M raisedFounded 2023Pittsburgh, Pennsylvania, USA11-50 staffSOC 2: Type II

Best fit: Buyers needing adversarial evaluation, red-teaming arenas, and runtime guardrails for frontier or enterprise LLM/agent deployments.

8

Vals AI

Commercial

Vals AI is an independent, a16z-backed third-party evaluator. It benchmarks frontier models and AI applications on economically valuable professional work (finance, legal, tax, coding, healthcare) and, since 2026, on frontier risks (recursive self-improvement, cyber, child safety). It publishes the Vals Index and more than 20 proprietary leaderboards, open-sources its Valkyrie eval infrastructure, offers the self-serve Vals Smith for custom coding benchmarks, and runs a public-sector practice from Washington, DC.

Backed by Andreessen Horowitz (a16z), 8VC, Pear VC, Bloomberg Beta. Team pedigree: Co-founder/CEO Rayan Krishnan - left Stanford's AI master's program to start Vals; Co-founder/CTO Langston Nashold - Stanford AI master's program. Named customers include OpenAI (Vals says its results were cited in OpenAI model cards), Anthropic (Vals says its results were cited in Anthropic model cards), Google (Vals says its results were cited in Google model cards), Meta (Vals says its results were cited in Meta model cards), xAI (Vals says its results were cited in xAI model cards), Harvey (industry-partner benchmark run by Vals; partner, not a stated customer), Exa, Keenable, Parallel, Tavily (Web Search Index partners), CoreWeave (RSI Index compute collaborator), Code for America and Center for Civic Futures (Public Benefits Bench partners), Fisher Phillips, McDermott Will & Emery, Reed Smith, Legal Technology Hub (Legal Research Bench evaluation partners), US Department of Commerce (Vals says it supported Commerce's AI policy work) (self-claimed, incl. frontier-lab ties).

$40M raisedFounded 2023San Francisco, USA11-50 staffSOC 2: unknown (the SOC…

Best fit: Buyers who need neutral, domain-specific (legal/finance/healthcare) benchmarking and ongoing evaluation of LLM applications on their own data and tasks.

9

HUD

Commercial

HUD (Human Union Data, YC W25, formerly hud.so) calls itself the platform for building high-quality post-training datasets. It offers an MIT-licensed SDK for defining RL environments, tasks and graders across coding, browser, computer-use and robotics; a hosted training and eval platform; and DataVendor, a marketplace where independent builders sell RL environments and data to AI labs. It announced a $16M Series A led by Standard Capital in June 2026.

Backed by Standard Capital (Series A lead), Y Combinator (W25), Exceptional Capital, Liquid 2 Ventures. Team pedigree: Jay Ram (CEO) - consumer apps, ML/quant research (YC bio); Lorenss Martinsons (Co-CEO) - Cognitive Science, Yale (FounderTrace). Named customers include UiPath, Sharpe, OpenAI, Anthropic (self-claimed, incl. frontier-lab ties).

$16M raisedFounded 2025San Francisco, USA11-50 staffSOC 2: claimed

Best fit: Teams that want to build, evaluate and RL-train agents on reproducible environments (computer-use, browser, coding, robotics), or that want to buy or sell RL environments through a marketplace.

10

Halluminate

Commercial

Halluminate (YC S25, founded 2024, San Francisco) calls itself a data research lab building benchmarks and RL environments for knowledge work, starting with financial services (PE, IB, due diligence) and consulting. It raised a $30M Series A led by Oak HC/FT in October 2026 ($38.5M total). With about nine people, it says it works with four of the five leading closed-source US AI labs and has a mid-eight-figure revenue run rate (self-reported). Its flagship benchmark is Westworld Finance Diligence Bench.

Backed by Oak HC/FT (Series A lead), Y Combinator (S25), Orange Collective, FT Partners. Team pedigree: Jerry Wu (co-founder/CEO): ex-Capital One Labs (led product and research; 3 patents); Cornell CS & Economics; Wyatt Marshall (co-founder/CTO): Cornell Milstein Scholar; data engineering at two early-stage NYC startups. Named customers include Four of the five leading closed-source US AI labs (unnamed) (self-claimed, incl. frontier-lab ties).

$38.5M raisedFounded 2024San Francisco, CA, USA1-10 staffSOC 2: unknown

Best fit: Frontier labs that need long-horizon, expert-authored finance and knowledge-work RL environments and benchmarks (due diligence, PE deal analysis, modeling) in realistic desktop sandboxes.

11

Veris AI

Commercial

Veris AI sells simulation sandboxes that recreate an agent's production environment (mocked SaaS tools, seeded databases, simulated users) so enterprises can benchmark, regression-test, and RL/SFT-train agents before deployment. In 2026 it added a self-serve tier (Veris Plus), a coding-agent verification product (Agentic SDLC), and public voice-agent benchmarks (VAmoS, VAmoS Pro) backed by arXiv papers.

Backed by Decibel Ventures (lead), Acrew Capital (lead), The House Fund, Ian Livingstone. Team pedigree: CEO Mehdi Jamei: PhD EECS UC Berkeley; previously led agentic AI at System and Workmate; CTO Andi Partovi: PhD (brain-computer interfaces) University of Melbourne; ex-Solutions Architect at Google; ex-founder/CTO KeyLead Health. Named customers include Savi Security (RFT research collaborator / co-author), Telecom operator (unnamed) - production customer-service agent, Consumer fintech company (unnamed) - compliant chatbots, HR tech / executive-assistant agent company (unnamed), Manufacturer - supply chain agent (unnamed) (self-claimed).

$8.5M raisedFounded 2025San Francisco, CA, USA1-10 staffSOC 2: claimed

Best fit: Enterprise teams shipping customer-facing chat, voice or tool-calling agents (especially in finance, telecom and utilities) who need policy-specific benchmarks, CI regression gating, and simulated environments for RFT before production.

12

Collinear

Commercial

Collinear AI runs a Simulation Lab of sandboxed, stateful environments with simulated users (NPCs), tools, tasks and verifiers for agent RL training and evaluation. It now positions itself mainly as a supplier of training data, environments and evals to frontier labs across cybersecurity, software engineering, computer use and complex agent behaviour. Its CWE-bench defensive-cyber benchmark is part of the Artificial Analysis Cyber Index.

Backed by Engineering Capital, 112 Capital (11.2 Capital), B Capital, Storm Ventures. Team pedigree: Founder/CEO Nazneen Rajani: ex-Robustness Research Lead at Hugging Face, ex-Research Scientist at Salesforce, PhD University of Texas at Austin (MIT TR Innovators Under 35); Team described as researchers/engineers from Hugging Face, Salesforce, Google, Amazon, Stanford (per company About page). Named customers include Amazon, Three of the world's top four AI labs (unnamed) (self-claimed, incl. frontier-lab ties).

$10M raisedFounded 2023Sunnyvale, California, USA11-50 staffSOC 2: Type I

Best fit: Labs and agent teams that need held-out evals and verifier-graded RL tasks or training data for defensive cybersecurity, SWE, computer use, or long-horizon workflows with simulated users.

13

Chakra Labs

Commercial

Chakra Labs ('Frontier Data Laboratory') builds deterministic clones of enterprise and productivity software that expose both GUI and MCP surfaces, along with verifiable task sets and crowdsourced trajectory data for training and evaluating computer-use and tool-use agents. It delivers them through its Dojo hub with Harbor, Verifiers and Verl support. In 2026 its messaging centred on task quality (difficult, realistic, reliably verifiable, hard to hack), and it published MAGI, a 1,000-task mixed GUI and MCP benchmark on which it reports frontier models below 30% pass@1.

Team pedigree: Nirmal Krishnan (co-founder and CEO per the SEC Form D signature): Johns Hopkins BS/MS (CS/ML, computational genomics); prior data, markets and early-stage startups; Alexander Fung (co-founder; Director per SEC Form D): University of Waterloo; reported ex-Palantir, Snap and Fin (snippets only).

$10.06M raisedFounded 2024Brooklyn, New York, USA11-50 staffSOC 2: unknown

Best fit: Teams training or evaluating computer-use and tool-use (MCP) agents on long-horizon, multi-application enterprise workflows that need deterministic software clones, rubric-verified tasks and a Harbor-compatible harness.

14

Andon Labs

Commercial

Andon Labs (YC W24, formerly Vectorview) builds long-horizon agent benchmarks (Vending-Bench, Blueprint-Bench, Drone-Bench, Butter-Bench) and runs real AI-operated businesses (Andon Market in SF, Andon Café in Stockholm, Andon FM, vending machines) as live safety testbeds. In September 2026 it launched Pion, a research-preview platform for agents that run whole businesses. It is a repeat Anthropic research partner (Project Vend, and Project Pilot with the Frontier Red Team).

Backed by Y Combinator (W24). Team pedigree: Lukas Petersson (co-founder; CEO per Fortune 2026-06, 'Founder/CTO' on YC). Co-founded Vectorview.; Axel Backlund (co-founder; first author of Vending-Bench). Named customers include Anthropic, xAI, OpenAI, Google DeepMind (verified, incl. frontier-lab ties).

Founded 2023San Francisco, USA11-50 staffSOC 2: unknown

Best fit: Frontier labs and safety teams that want held-out evals of long-horizon behaviour, agentic misalignment and dual-use physical capability (Vending-Bench, Drone-Bench), plus data from real autonomous deployments. Pion suits businesses that want to hand operations to AI agents on a revenue-share basis.

15

BenchFlow

Commercial

BenchFlow is a small, open-source-first 'frontier environment lab' in the Bay Area. It publishes agent benchmarks and environments: SkillsBench and ClawsBench, the env0 simulated workspaces, Robo Use and GenesisBench for embodied agents, and FrontierPhysics. It runs the PostTrain Arena for community-contributed post-training environments and maintains an Apache-2.0 runtime for RL environments, post-training and evals. It has event-level ties to Google DeepMind, Kaggle, Anthropic, Hugging Face and Prime Intellect.

Backed by Founders, Inc. (confirmed via its portfolio page), Y Combinator (reported by aggregator only; no YC directory page found), Pear VC (reported), Construct Capital (reported). Team pedigree: Xiangyi Li (founder/CEO): previously coding-model inference at Tesla; first author of SkillsBench and ClawsBench; Bingran You (MTS): PhD in quantum physics, UC Berkeley; author of the env0 study. Named customers include Google DeepMind, Meta, Microsoft, Tencent, Zhipu (GLM), MiniMax (self-claimed, incl. frontier-lab ties).

Founded 2024San Francisco Bay Area, CA, USA1-10 staffSOC 2: unknown

Best fit: Teams needing open-source, reproducible agent evaluation environments and a runtime to benchmark coding/computer-use/workplace agents at low setup cost.

16

Refresh

Commercial

Refresh (YC X25; appears to have grown out of Operative.sh) builds RL environments with verifiable rewards for coding, computer use (high-fidelity clones of software suites such as EHRs and enterprise apps) and 3D animation and simulation in Blender. Data ships in Harbor format. Public releases include BlenderBench, the Gauntlet 4K RLVR dataset (access by request), the open-source computer-1 harness in Harbor, and the trajectories.sh trajectory-sharing platform.

Backed by Y Combinator. Team pedigree: Christopher Settles (CEO): led evaluation for Uber's generative-AI platform and shipped its first computer-use agents; ex-ML at Uber; CS at UIUC; Erik Quintanilla (CTO): Amazon production-ops automation and web-scraping patents, Capital One platform modernization, ML at LynkAI; designs verifiers and post-training harnesses. Named customers include Frontier AI labs (unnamed) (self-claimed, incl. frontier-lab ties).

Founded 2025San Francisco, CA, USA1-10 staffSOC 2: unknown

Best fit: Frontier labs needing custom RL training environments and datasets for software-engineering and computer-use agent capabilities.

17

Matrices

Commercial

Matrices builds reinforcement-learning training environments for frontier AI labs to train agents that use computers and browsers like humans, described as a 'gamified replica of the internet' where thousands of agents learn via RL. The company frames its mission as 'towards self-driving computers' and says it helps labs train computer-use agents (Operator-class systems). Note: this is the correct browser-native entity (matrices.ai / LinkedIn 'matricesapp'), distinct from the similarly named 'Matrice.ai' computer-vision company and 'Matrix AI Network' blockchain project.

Backed by AI Grant (Batch 3; confirmed on aigrant.com), Index Ventures (reported), Naval Ravikant (reported, single source). Team pedigree: Co-founder John Qian: UIUC; Co-founder Leonardo Axel Setyanto: UT Austin, ex-Loku. His LinkedIn now lists OpenAI as current company (undated); his ongoing role at Matrices is unverified. Named customers include Unnamed frontier AI labs (described as signing 7-figure contracts; agents like OpenAI 'Operator' referenced as the type they help train) (self-claimed).

$5M raisedFounded 2023San Francisco, California, USA11-50 staffSOC 2: unknown

Best fit: A frontier lab needing large-scale, realistic browser/computer-use RL environments to train and evaluate web-navigating agents.

18

Andromede

Commercial

Andromede is an early-stage RL data lab (reported as founded 2025 in Lausanne, backed by Unusual Ventures) that programmatically generates RL environments, tasks and verifiers from real-world data for post-training and evaluating long-horizon frontier agents. Its core product is still in private beta with unnamed partners. Since mid-2026 the EPFL-linked team, co-founded by Alexandre Sallinen and Guillaume Allègre, has published open research: the Messier cross-benchmark agent-evaluation corpus (arXiv 2607.25891; self-reported EMNLP 2026 Main), a paper on predicting task difficulty without rollouts, and the Game of Agents multi-agent testbed (ICML 2026 workshop).

Backed by Unusual Ventures. Team pedigree: Alexandre Sallinen (co-founder) - EPFL; Meditron medical-LLM and MMORE co-author; Guillaume Allegre (co-founder; Founder & President per LinkedIn) - ex-BCG X; MIT (ML & operations research) per LinkedIn snippet.

Founded 2025Lausanne, Switzerland1-10 staffSOC 2: unknown

Best fit: Buyers needing custom RL environments and verifiers derived from real-world data for post-training/evaluating long-horizon agentic models.

19

Vmax

Commercial

Vmax is a San Francisco RL company and applied research lab, founded in 2025 by RL PhDs Augustine Mavor-Parker and Matthew Sargent and backed by South Park Commons and Race Capital. It aims to automate reinforcement learning by turning proprietary data and evals into new environments, with long-horizon agents as the target. Its 2026 research is about generating RL tasks and environments automatically: PopuLoRA (self-play curricula), unix-ctf (procedural shell CTF environments) and PROPEL (probe-guided task generation, with a Goodfire AI co-author). Earlier, with Martian, it released 1k Harbor-format JavaScript tasks.

Backed by Race Capital, South Park Commons. Team pedigree: Augustine Mavor-Parker: RL PhD, UCL; co-founder (CTO per LinkedIn); previously Redwood Research, CSHL, Illumina; Matthew Sargent (also published as Matthew James Sargent / Matthew Daborn-Sargent): RL PhD (UCL per LinkedIn); co-founder. Named customers include Martian / ARES team (withmartian), partnership: jointly releasing ~1k JavaScript coding tasks in the Harbor format (Harbor = Terminal-Bench task format) (self-claimed).

Founded 2025San Francisco, USA1-10 staffSOC 2: unknown

Best fit: Teams needing custom, research-grade RL environments to train coding and long-horizon shell/terminal agents from proprietary data.

20

Plato

Commercial

Plato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and computer-use agents, recreating real websites/software (e.g. Amazon/Airbnb/Gmail-style replicas) as reinforcement-learning environments with structured APIs for interaction, state tracking and scoring. It also offers a 'Computer Use' capability driving a full Linux desktop, positioning at the intersection of browser interaction and enterprise workflow simulation.

Team pedigree: Pranav Putta (Co-founder/CTO), prior MultiOn, Georgia Institute of Technology, Tonic.ai; Robert Farlow (Co-founder/CEO).

Founded 2025San Francisco, CA, USA1-10 staffSOC 2: unknown

Best fit: AI labs/teams needing high-fidelity replica web/enterprise environments to train and evaluate browser and computer-use agents via RL.

21

AIChamp

Commercial

AIChamp, formerly a vendor of RL environments ('Virtual Gyms') and expert data for enterprise AI agents, repositioned in 2026 as a vendor-neutral factory digital-twin service. It runs a free 14-day loss check on a manufacturing line, builds a simulation of the costliest station, and invites robot-AI software teams to train and be tested in that twin against a plant-set target before a floor pilot. Under its 'Shared' option it also licenses de-identified factory scenes to robotics and AI companies. It no longer markets RL environments for LLM or software agents.

Team pedigree: Self-claimed: 'Built by engineers from Microsoft, Amazon Robotics, Google, Mercor, OpenAI and NVIDIA' (vendor site, unverified; the earlier site claimed 'alumni of OpenAI and xAI'); CEO Vol Goloshuk: previously founder-CEO of Brightest Minds (outsourced assistants/SDRs), co-founder of HyperC, ex-Mastercard analyst, earlier network/security engineer (self-reported on Torre).

$0 raisedSan Francisco (Mission Bay), USA1-10 staffSOC 2: unknown

Best fit: Manufacturing plants (food, beverage, chemicals, leather/textiles) that want losses priced and robot/vision fixes tested in simulation before capex. Possibly robotics/physical-AI teams seeking factory twins or de-identified real-world factory scenes for training and evaluation. Not a current fit for buyers of LLM/agent RL environments.

22

Habitat Inc

Commercial

Habitat Inc is a very early-stage San Francisco company (2-10 employees). Third-party market maps (Chemistry VC, SemiAnalysis, AlignList) list it as building reinforcement-learning environments for code and computer-use agent workflows, made of programmatically verifiable problems. Until mid-2026 its own LinkedIn described it as 'RL environments for white-collar work'. By October 2026 that tagline read 'Autonomous firms.', which may signal a repositioning. It has disclosed no funding, customers, publications or security certifications, and its website is a placeholder.

Team pedigree: Maxim Enis (co-founder), Williams College '24; prior Ramp association per LinkedIn; co-author (with Mark Hopkins, Williams) of arXiv:2404.13813 'From LLM to NMT: Advancing Low-Resource Machine Translation with Claude' (2024, academic, predates company); Andrew Megalaa (co-founder), Williams College '24.

San Francisco, CA, USA1-10 staffSOC 2: unknown

Best fit: Buyers needing RL environments that simulate enterprise/desktop and coding workflows to post-train computer-use and coding agents.

23

Taste Labs

Commercial

Taste Labs (Taste) is a San Francisco startup, founded in October 2025 by former Exa growth lead Thais Castello Branco, that calls itself 'the taste layer for AI' and a hybrid research lab and infrastructure company. For model labs, it builds post-training data and RL environments for subjective domains, starting with design. Amplify describes these as preference datasets, reasoning data, rubrics and evaluation environments. The data is made by its TasteMakers network of expert creatives. Taste says it works with 'the top frontier labs' but names none. In September 2026 it launched a self-serve Brand API and MCP server that lets agents extract, search and verify brand design systems. It raised an $18.5M seed co-led by CRV and Amplify Partners, announced 2026-06-16.

Backed by CRV, Amplify Partners. Named customers include General Intelligence Company of New York (Cofounder) (self-claimed).

$18.5M raisedFounded 2025San Francisco, CA, USA11-50 staffSOC 2: unknown

Best fit: Labs that need expert preference data, rubrics or eval environments for aesthetic and design work (UI, web, visual generation), and agent builders that need their outputs to stay on-brand.

24

Idler

Commercial

Idler (YC S25, founded 2025, San Francisco) calls itself a frontier data research lab that builds evals and RL environments from real production work. Its public benchmarks are ShelfLife, a digital twin of a live multi-brand e-commerce company; ShelfLife E-Sim, a stateful retail-operations simulator; and CorpLaw, built from a law firm's anonymized data. It also offers private, on-request collections in long-horizon SWE, cybersecurity, recursive self-improvement and terminal agents. The roughly 12-person team is led by Dark Forest/0xPARC co-founder Ivan Chub and says leading frontier labs use its work, without naming any. A $9M Paradigm-led seed was reported by an aggregator in August 2026 but is not confirmed by the company.

Backed by Paradigm (seed lead, reported), Y Combinator (S25). Team pedigree: Ivan Chub (CEO): co-founder of Dark Forest, 0xPARC, Zupass and Hack Lodge; ex-Facebook, Dynasty, AppSheet; Tony Goss (CTO): previously Dark Forest DAO and 0xPARC; won $400k by (ethically) hacking Dark Forest. Named customers include Leading frontier AI labs (unnamed) (self-claimed, incl. frontier-lab ties).

$9M raisedFounded 2025San Francisco, CA, USA11-50 staffSOC 2: unknown

Best fit: Frontier labs that want bespoke long-horizon RL environments and evals grounded in real company data (e-commerce operations, corporate law), or private long-horizon SWE and cybersecurity task sets.

Also tracked, incumbents, infrastructure & open source

18 companies we research with the same rigor but don’t rank, because they’re a different category: data-labeling incumbents moving into environments, execution-infrastructure providers, open-source projects, and vendors that have been acquired.

Incumbent
Scale AI Scale AI is the data-labeling and AI-data incumbent that has extended into RL environments (simulated web apps, macOS/Windows-like desktop VMs and MCP-tool environments with expert rubrics and automated verifiers).
Incumbent
Snorkel AI Snorkel AI is a San Francisco company spun out of the Stanford AI Lab in 2019 and known for the Snorkel weak-supervision project.
Infrastructure
Modal Modal (Modal Labs) is a New York-based, Python-native serverless cloud for AI workloads.
Incumbent
Mercor Mercor is a venture-backed expert marketplace and AI-training-data company.
Open source
Prime Intellect Prime Intellect sells an open-source-first stack for agentic RL: the Environments Hub (2,500+ community environments plus 365k+ unified SWE, terminal and search tasks), the verifiers, prime-rl and renderers libraries, Lab hosted RL post-training and evals (GA), Prime Sandboxes microVMs (GA), Prime Inference, and GPU compute.
Infrastructure
Daytona Daytona provides secure, elastic, programmatic sandboxes (Linux containers, Linux and Windows VMs, and GPU sandboxes) that AI agents spin up in under ~90ms (vendor claim) to run untrusted AI-generated code in isolated, stateful runtimes with snapshot and fork primitives.
Incumbent
Turing Turing is a large incumbent supplier of expert human data to frontier AI labs.
Incumbent
micro1 micro1 is a San Francisco AI training-data company founded in 2022 by Ali Ansari.
Acquired
Deeptune Acquired by Mercor (9 Jul 2026). Deeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms') for computer-use and code, where AI agents practice and are evaluated on realistic digital knowledge-work tasks (simulating tools like Slack and Salesforce).
Open source
General Reasoning General Reasoning is a London AI research lab (legal entity General Reasoning, Inc., US) working on models that operate over much longer time horizons.
Infrastructure
E2B E2B provides open-source Firecracker-microVM sandboxes for AI agents.
Infrastructure
Runloop Runloop sells cloud VM-isolated devboxes and sandboxes for AI agents.
Acqui-hired
Mechanize Acqui-hired by Google (Google DeepMind) (Sep 2026). Mechanize is a San Francisco vendor founded in April 2025 by ex-Epoch AI researchers.
Infrastructure
Cua Cua (trycua, YC X25) builds open-core 'computer-use 2.0' infrastructure: an MIT, agent-neutral background driver for macOS, Windows and Linux (CLI/MCP), local and cloud sandboxes, elastic Cua Fleets for parallel eval and RL training (cloud, BYOC, on-prem), Lume macOS virtualization, the Cua Spaces Mac app, small CUA-S1 decision models, and Cua-Bench for verifiable tasks, leaderboards and RL trajectory export.
Open source
Good Start Labs Good Start Labs is a 2025 Every spin-out that sells game-based RL environments and agent and human gameplay data to AI labs.
Incumbent
Surge AI Surge AI is a bootstrapped human-data and RLHF leader for frontier labs (reported $1.2B 2024 revenue).
Acquired
Sepal AI Acquired by Mercor (6 Feb 2026). Sepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data, expert-graded evaluation benchmarks, and reinforcement-learning environments for frontier LLMs, drawing on a network of 20k+ domain experts (PhDs, finance, medical, STEM).
Infrastructure
Morph Morph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal technology, which can snapshot, branch, and restore entire computational environments in roughly 100-250ms to enable massively parallel, reversible ('Git for compute') agent rollouts, evaluations, and reasoning-time branching.
// FAQ

RL environment vendors: frequently asked questions

What is an RL environment?

An RL environment is a simulated task or world in which an AI agent takes actions, receives a reward signal, and improves through reinforcement learning. For frontier AI, these are usually high-fidelity replicas of real software (a web browser, a coding repository, or an enterprise app like Salesforce or Slack) paired with verifiers that automatically score whether the agent completed the task. Labs use them to train and evaluate agents on realistic, long-horizon work rather than single-turn questions.

What are the best RL environment companies in 2026?

There is no single best vendor: the right choice depends on your use case (coding, computer use, or enterprise workflows), budget, security and compliance needs, and deployment model. RL List ranks the dedicated commercial vendors by a transparent RL List score that combines scale and traction, security and research signals, and how much of each company's record could be independently verified. As of the latest update, the highest-ranked pure-play vendors are Mechanize, AfterQuery, Bespoke Labs, Huzzle Labs, Fleet AI, and Datacurve. Data-labeling incumbents such as Scale AI, Surge AI, and Mercor are tracked separately because they are a different category. See the methodology page for exactly how the ranking is computed.

Which companies build RL environments for coding agents?

Vendors that build RL environments and data specifically for coding and software-engineering agents include Mechanize, AfterQuery, Datacurve, Proximal, Huzzle Labs, and Vmax, typically using real code repositories with unit-test or execution-based verifiers. Execution-infrastructure providers such as Runloop, Modal, Daytona, E2B, and Morph supply the sandboxes those coding agents run in, rather than the environments themselves.

Which companies build RL environments for computer-use and browser agents?

Vendors focused on computer-use and browser agents, where an agent operates a real or simulated desktop, browser, or GUI, include Huzzle Labs, Fleet AI, Chakra Labs, HUD, Halluminate, Matrices, and Plato. Many recreate real websites and enterprise software (Salesforce, Slack, Excel) as reinforcement-learning environments with state tracking and automated scoring.

What is the difference between RL environments, evals, and human training data?

RL environments are the interactive tasks where an agent acts and is scored by a verifier. Evaluations (evals) and benchmarks measure how well a model or agent performs, often reusing those same environments. Human training data, including SFT and RLHF, is expert-generated demonstrations, trajectories, or preference labels used to train and align models. Many vendors offer more than one, and some bundle environments, human data, and evals in a single stack, which is why RL List tags each vendor by focus area rather than forcing one label.

How do I choose an RL environment vendor?

Start from your use case, then compare vendors on the public proxies RL List tracks: focus areas, scale and traction, research depth, security posture such as SOC 2, and how much of their record is independently verified versus self-claimed. The figures that actually decide a purchase (task and sample counts, unique environments, pass rates, difficulty splits, harness and data format, and pricing) are not on the public web and only surface in direct engagement. The practical next step is to request work samples from a shortlist of three to five vendors and let those decide.

Other RL-environment lists

RL List isn't the only map of this space, and we're glad to point you to the others. If you're researching vendors, these are worth reading alongside this directory: