#1Commercial
AfterQuery is a San Francisco applied-research lab and expert-data company (YC W25). It supplies AI labs with expert-generated SFT and RL data, rubrics, agent environments and computer-use trajectories, drawn from a stated network of nearly 100,000 verified professionals. Forbes reported it was raising at a $3.2B valuation in Sept 2026. NVIDIA's Nemotron 3 Ultra report names AfterQuery tasks in its GDPval training recipe. The company publishes or co-builds real-task benchmarks (FinanceQA, VADER, IDE-Bench, SpreadsheetBench 2, Legora BAR) and is adding forward-deployed enterprise work.
Backers: Altos Ventures (lead, Series A), The Raine Group, Y Combinator
CodeComputer UseEnterprise
site ↗
#2Commercial
Bespoke Labs is an applied AI research lab in Mountain View, founded in 2024 by Mahesh Sathiamoorthy (ex-Google DeepMind) and UC Berkeley professor Alex Dimakis. It builds RL environments and infrastructure for training and evaluating reliable agents, and is backed by $40M from Wing VC, 8VC, Mayfield and others. It combines commercial environment and post-training work, such as a jointly published project with Intuit Credit Karma, with a large open-research output: OpenThoughts, Curator, GEPA, Terminal-Bench, AutoResearchExam and ConjectureBench.
Backers: Wing VC (Series A lead), 8VC (Seed lead), Mayfield
#3Commercial
Huzzle Labs is the applied-AI division of the London talent platform Huzzle (co-founders Ingmar Klein, CEO, and Amit Choudhary, CTO). It builds long-horizon RL environments (coding, computer use, enterprise workflows) and expert trajectory data for frontier labs, drawing on Huzzle's talent network, which it says numbers 300k+. In 2026 it expanded into its own small computer-use models (HuzzleWorld, 1B-8B, self-reported results), a domain benchmark (InsureBench, scores pending), and enterprise custom evals and models deployed on client infrastructure.
Backers: 10x Founders, Angel Invest, Emerge
#4Commercial
Fleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise software such as Salesforce and Excel, plus browser/desktop workflows, so frontier AI labs and large enterprises can train and evaluate computer-use agents. It ships a Python SDK, a platform API, and the open-source 'Harbor' agent-evaluation/RL-environment tooling, pairing simulated environments with human supervision.
Backers: Sequoia Capital, Menlo Ventures, SV Angel
#5Commercial
Datacurve is a YC W24 data vendor that says it supplies frontier labs with expert-sourced RL environments, long-horizon tasks, agent trajectories, SFT demonstrations and off-the-shelf datasets. It is coding-first but now also covers data science, cyber security, ML and research. Data comes from its Shipd bounty platform of vetted engineers, which has SWE, ML and data-science tracks. It publishes DeepSWE, an open (Apache-2.0) long-horizon coding benchmark used in the Artificial Analysis Coding Agent Index.
Backers: Chemistry (Mark Goldberg, lead Series A), Y Combinator, Balaji Srinivasan (seed)
#6Commercial
Proximal is a research lab in San Francisco and Bangalore that treats post-training data as a core research problem. It builds long-horizon coding RL environments and data for frontier labs and enterprises, and it publishes the FrontierSWE benchmark and verifier audits such as CyberGym-Verified. In September 2026 it raised a $15M seed led by General Catalyst at a $300M valuation and claimed $200M in annualized revenue, which is self-reported. It also said it will expand into drug discovery, chip design and legacy software modernization.
Backers: General Catalyst (lead, 2026-09 seed), Scribble Ventures (led an earlier investment, per its announcement), SV Angel
Gray Swan AI is a Pittsburgh AI security company spun out of Carnegie Mellon (Fredrikson, Kolter). It sells adversarial red teaming and runtime protection for AI models and agents through three products. Shade does automated red teaming. Cygnal enforces policies and detects indirect prompt injection at runtime, and is now also available through the TrueFoundry and Bifrost gateways. Arena is a crowdsourced red-teaming network of more than 15,000 researchers. Gray Swan is a pre-release evaluation partner to frontier labs: Anthropic, OpenAI, Google DeepMind and Meta all cite its Shade, IPI Arena or ART evaluations in their 2026 system cards and safety reports. It is not a general RL-environment vendor.
Backers: Wing Venture Capital (co-lead), Madrona (co-lead), Obvious Ventures
#8Commercial
Vals AI is an independent, a16z-backed third-party evaluator. It benchmarks frontier models and AI applications on economically valuable professional work (finance, legal, tax, coding, healthcare) and, since 2026, on frontier risks (recursive self-improvement, cyber, child safety). It publishes the Vals Index and more than 20 proprietary leaderboards, open-sources its Valkyrie eval infrastructure, offers the self-serve Vals Smith for custom coding benchmarks, and runs a public-sector practice from Washington, DC.
Backers: Andreessen Horowitz (a16z), 8VC, Pear VC
#9Commercial
HUD (Human Union Data, YC W25, formerly hud.so) calls itself the platform for building high-quality post-training datasets. It offers an MIT-licensed SDK for defining RL environments, tasks and graders across coding, browser, computer-use and robotics; a hosted training and eval platform; and DataVendor, a marketplace where independent builders sell RL environments and data to AI labs. It announced a $16M Series A led by Standard Capital in June 2026.
Backers: Standard Capital (Series A lead), Y Combinator (W25), Exceptional Capital
#10Commercial
Halluminate (YC S25, founded 2024, San Francisco) calls itself a data research lab building benchmarks and RL environments for knowledge work, starting with financial services (PE, IB, due diligence) and consulting. It raised a $30M Series A led by Oak HC/FT in October 2026 ($38.5M total). With about nine people, it says it works with four of the five leading closed-source US AI labs and has a mid-eight-figure revenue run rate (self-reported). Its flagship benchmark is Westworld Finance Diligence Bench.
Backers: Oak HC/FT (Series A lead), Y Combinator (S25), Orange Collective
Veris AI sells simulation sandboxes that recreate an agent's production environment (mocked SaaS tools, seeded databases, simulated users) so enterprises can benchmark, regression-test, and RL/SFT-train agents before deployment. In 2026 it added a self-serve tier (Veris Plus), a coding-agent verification product (Agentic SDLC), and public voice-agent benchmarks (VAmoS, VAmoS Pro) backed by arXiv papers.
Backers: Decibel Ventures (lead), Acrew Capital (lead), The House Fund
#12Commercial
Collinear AI runs a Simulation Lab of sandboxed, stateful environments with simulated users (NPCs), tools, tasks and verifiers for agent RL training and evaluation. It now positions itself mainly as a supplier of training data, environments and evals to frontier labs across cybersecurity, software engineering, computer use and complex agent behaviour. Its CWE-bench defensive-cyber benchmark is part of the Artificial Analysis Cyber Index.
Backers: Engineering Capital, 112 Capital (11.2 Capital), B Capital
CodeComputer UseEnterprise
site ↗
#13Commercial
Chakra Labs ('Frontier Data Laboratory') builds deterministic clones of enterprise and productivity software that expose both GUI and MCP surfaces, along with verifiable task sets and crowdsourced trajectory data for training and evaluating computer-use and tool-use agents. It delivers them through its Dojo hub with Harbor, Verifiers and Verl support. In 2026 its messaging centred on task quality (difficult, realistic, reliably verifiable, hard to hack), and it published MAGI, a 1,000-task mixed GUI and MCP benchmark on which it reports frontier models below 30% pass@1.
Pedigree: Nirmal Krishnan (co-founder and CEO per the SEC Form D signature): Joh
#14Commercial
Andon Labs (YC W24, formerly Vectorview) builds long-horizon agent benchmarks (Vending-Bench, Blueprint-Bench, Drone-Bench, Butter-Bench) and runs real AI-operated businesses (Andon Market in SF, Andon Café in Stockholm, Andon FM, vending machines) as live safety testbeds. In September 2026 it launched Pion, a research-preview platform for agents that run whole businesses. It is a repeat Anthropic research partner (Project Vend, and Project Pilot with the Frontier Red Team).
Backers: Y Combinator (W24)
Computer UseLong Horizon
site ↗
#15Commercial
BenchFlow is a small, open-source-first 'frontier environment lab' in the Bay Area. It publishes agent benchmarks and environments: SkillsBench and ClawsBench, the env0 simulated workspaces, Robo Use and GenesisBench for embodied agents, and FrontierPhysics. It runs the PostTrain Arena for community-contributed post-training environments and maintains an Apache-2.0 runtime for RL environments, post-training and evals. It has event-level ties to Google DeepMind, Kaggle, Anthropic, Hugging Face and Prime Intellect.
Backers: Founders, Inc. (confirmed via its portfolio page), Y Combinator (reported by aggregator only; no YC directory page found), Pear VC (reported)
CodeComputer UseEnterprise
site ↗
#16Commercial
Refresh (YC X25; appears to have grown out of Operative.sh) builds RL environments with verifiable rewards for coding, computer use (high-fidelity clones of software suites such as EHRs and enterprise apps) and 3D animation and simulation in Blender. Data ships in Harbor format. Public releases include BlenderBench, the Gauntlet 4K RLVR dataset (access by request), the open-source computer-1 harness in Harbor, and the trajectories.sh trajectory-sharing platform.
Backers: Y Combinator
#17Commercial
Matrices builds reinforcement-learning training environments for frontier AI labs to train agents that use computers and browsers like humans, described as a 'gamified replica of the internet' where thousands of agents learn via RL. The company frames its mission as 'towards self-driving computers' and says it helps labs train computer-use agents (Operator-class systems). Note: this is the correct browser-native entity (matrices.ai / LinkedIn 'matricesapp'), distinct from the similarly named 'Matrice.ai' computer-vision company and 'Matrix AI Network' blockchain project.
Backers: AI Grant (Batch 3; confirmed on aigrant.com), Index Ventures (reported), Naval Ravikant (reported, single source)
Andromede is an early-stage RL data lab (reported as founded 2025 in Lausanne, backed by Unusual Ventures) that programmatically generates RL environments, tasks and verifiers from real-world data for post-training and evaluating long-horizon frontier agents. Its core product is still in private beta with unnamed partners. Since mid-2026 the EPFL-linked team, co-founded by Alexandre Sallinen and Guillaume Allègre, has published open research: the Messier cross-benchmark agent-evaluation corpus (arXiv 2607.25891; self-reported EMNLP 2026 Main), a paper on predicting task difficulty without rollouts, and the Game of Agents multi-agent testbed (ICML 2026 workshop).
Backers: Unusual Ventures
#19Commercial
Vmax is a San Francisco RL company and applied research lab, founded in 2025 by RL PhDs Augustine Mavor-Parker and Matthew Sargent and backed by South Park Commons and Race Capital. It aims to automate reinforcement learning by turning proprietary data and evals into new environments, with long-horizon agents as the target. Its 2026 research is about generating RL tasks and environments automatically: PopuLoRA (self-play curricula), unix-ctf (procedural shell CTF environments) and PROPEL (probe-guided task generation, with a Goodfire AI co-author). Earlier, with Martian, it released 1k Harbor-format JavaScript tasks.
Backers: Race Capital, South Park Commons
#20Commercial
Plato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and computer-use agents, recreating real websites/software (e.g. Amazon/Airbnb/Gmail-style replicas) as reinforcement-learning environments with structured APIs for interaction, state tracking and scoring. It also offers a 'Computer Use' capability driving a full Linux desktop, positioning at the intersection of browser interaction and enterprise workflow simulation.
Pedigree: Pranav Putta (Co-founder/CTO), prior MultiOn, Georgia Institute of Te
#21Commercial
AIChamp, formerly a vendor of RL environments ('Virtual Gyms') and expert data for enterprise AI agents, repositioned in 2026 as a vendor-neutral factory digital-twin service. It runs a free 14-day loss check on a manufacturing line, builds a simulation of the costliest station, and invites robot-AI software teams to train and be tested in that twin against a plant-set target before a floor pilot. Under its 'Shared' option it also licenses de-identified factory scenes to robotics and AI companies. It no longer markets RL environments for LLM or software agents.
Pedigree: Self-claimed: 'Built by engineers from Microsoft, Amazon Robotics, Goo
#22Commercial
Habitat Inc is a very early-stage San Francisco company (2-10 employees). Third-party market maps (Chemistry VC, SemiAnalysis, AlignList) list it as building reinforcement-learning environments for code and computer-use agent workflows, made of programmatically verifiable problems. Until mid-2026 its own LinkedIn described it as 'RL environments for white-collar work'. By October 2026 that tagline read 'Autonomous firms.', which may signal a repositioning. It has disclosed no funding, customers, publications or security certifications, and its website is a placeholder.
Pedigree: Maxim Enis (co-founder), Williams College '24; prior Ramp association
CodeComputer UseEnterprise
site ↗
Taste Labs (Taste) is a San Francisco startup, founded in October 2025 by former Exa growth lead Thais Castello Branco, that calls itself 'the taste layer for AI' and a hybrid research lab and infrastructure company. For model labs, it builds post-training data and RL environments for subjective domains, starting with design. Amplify describes these as preference datasets, reasoning data, rubrics and evaluation environments. The data is made by its TasteMakers network of expert creatives. Taste says it works with 'the top frontier labs' but names none. In September 2026 it launched a self-serve Brand API and MCP server that lets agents extract, search and verify brand design systems. It raised an $18.5M seed co-led by CRV and Amplify Partners, announced 2026-06-16.
Backers: CRV, Amplify Partners
#24Commercial
Idler (YC S25, founded 2025, San Francisco) calls itself a frontier data research lab that builds evals and RL environments from real production work. Its public benchmarks are ShelfLife, a digital twin of a live multi-brand e-commerce company; ShelfLife E-Sim, a stateful retail-operations simulator; and CorpLaw, built from a law firm's anonymized data. It also offers private, on-request collections in long-horizon SWE, cybersecurity, recursive self-improvement and terminal agents. The roughly 12-person team is led by Dark Forest/0xPARC co-founder Ivan Chub and says leading frontier labs use its work, without naming any. A $9M Paradigm-led seed was reported by an aggregator in August 2026 but is not confirmed by the company.
Backers: Paradigm (seed lead, reported), Y Combinator (S25)
CodeEnterpriseLong Horizon
site ↗
n/rIncumbent
Scale AI is the data-labeling and AI-data incumbent that has extended into RL environments (simulated web apps, macOS/Windows-like desktop VMs and MCP-tool environments with expert rubrics and automated verifiers). It runs a large public evaluation program through Scale Labs (SWE-Bench Pro, MCP Atlas, Humanity's Last Exam, Remote Labor Index). After Meta's ~$14.3B deal in June 2025 (~49% non-voting stake) and Alexandr Wang's move to Meta, several frontier labs reportedly scaled back work. In 2026 Scale grew mainly in government and enterprise: the Pentagon CDAO agreement ceiling rose to $500M, it acquired ICG Solutions, and Francis deSouza (ex-Google Cloud COO) became CEO in August 2026.
Backers: Meta Platforms, Accel, Amazon
CodeComputer UseEnterprise
site ↗
Snorkel AI is a San Francisco company spun out of the Stanford AI Lab in 2019 and known for the Snorkel weak-supervision project. It now calls itself 'the frontier AI data lab'. In 2025 it moved from selling Snorkel Flow software to an expert Data-as-a-Service model, and it now supplies expert agentic tasks, RL environments (computer-use, terminal and simulated-enterprise), rubrics and evals to AI labs and enterprises. It combines expert contributors with synthetic generation and agent-based QC, maintains Terminal-Bench, and co-authors or hosts agent benchmarks such as OSWorld 2.0, Agents' Last Exam and Senior SWE-bench through a $3M Open Benchmarks Grants program. It raised a $350M Series E at $3.5B in September 2026 and says its ARR passed $375M.
Backers: Insight Partners (co-led Series E), S32 (co-led Series E), Addition (led Series D, co-led Series C)
CodeComputer UseEnterprise
site ↗
Modal (Modal Labs) is a New York-based, Python-native serverless cloud for AI workloads. It offers on-demand GPU/CPU compute, multi-node clusters with RDMA, managed LLM inference endpoints, and fast-booting sandboxes, including full VM sandboxes, which Modal says scale to 1M concurrent. It sells execution infrastructure, not RL environments. AI labs, agent companies and RL-data vendors use it to run RL rollouts, post-training and large fleets of parallel sandboxed environments. It integrates with Anthropic's Claude Managed Agents and Claude Science, Cognition's Devin Outposts and the OpenAI Agents SDK, and ships an open-source RL/SFT library (Modal Dojo).
Backers: General Catalyst (Series C co-lead), Redpoint Ventures (Series C co-lead; earlier Series A lead), Lux Capital (Series B lead)
n/rIncumbent
Mercor is a venture-backed expert marketplace and AI-training-data company. It supplies RLHF data, evaluations and reinforcement-learning environments to frontier AI labs and enterprises, drawing on a professional expert network: the homepage cites 30k+ experts, and Mercor's Deeptune post claims more than five million. It started as an AI-recruiting platform, moved into human data and RL, bought Sepal AI (Feb 2026), and announced the acquisition of Deeptune (Jul 2026), which builds simulated enterprise-app 'training gyms'. It publishes the APEX family of professional-task agent benchmarks. Mercor says it reached $2B ARR in June 2026, despite a March 2026 supply-chain data breach.
Backers: Felicis Ventures (led Series C and Series B), Benchmark, General Catalyst
CodeEnterpriseLong Horizon
site ↗
n/rOpen source
Prime Intellect sells an open-source-first stack for agentic RL: the Environments Hub (2,500+ community environments plus 365k+ unified SWE, terminal and search tasks), the verifiers, prime-rl and renderers libraries, Lab hosted RL post-training and evals (GA), Prime Sandboxes microVMs (GA), Prime Inference, and GPU compute. After a $130M Series A in July 2026 (TechCrunch reports a $1B valuation) it states $100M+ annualized revenue and 6,000+ customers. It also trains the open INTELLECT models and is the post-training partner in NVIDIA's Nemotron Coalition.
Backers: Radical Ventures (led $130M Series A, 2026), NVIDIA Ventures, Intel Capital
CodeEnterpriseLong Horizon
site ↗
n/rInfrastructure
Daytona provides secure, elastic, programmatic sandboxes (Linux containers, Linux and Windows VMs, and GPU sandboxes) that AI agents spin up in under ~90ms (vendor claim) to run untrusted AI-generated code in isolated, stateful runtimes with snapshot and fork primitives. It is a managed service with a Bring-Your-Own-Compute option and has been closed source since June 2026. It is an integrated sandbox provider for Anthropic Claude Managed Agents, OpenAI Agents SDK Sandbox Agents and Cognition Devin Outposts, and targets code execution, computer use and RL/eval rollouts.
Backers: FirstMark Capital (Series A lead; Matt Turck on board and listed as a director on the Aug 2026 Form D), Pace Capital, Upfront Ventures (seed lead, Series A participant)
n/rIncumbent
Turing is a large incumbent supplier of expert human data to frontier AI labs. It claims 5M+ experts and work with '9/9 frontier labs' (both self-claimed). In 2026 it moved further into verifiable agent environments and evaluations. It launched Turing Frontier in April, set up the Turing Frontier Research Lab, and published benchmarks including CyberStrike (Harbor format, public on Hugging Face), KernelQuest, CEO Bench, SciCode++ and PLSQLBench (with Oracle, per Turing). It also contributed tasks to Terminal-Bench 3.0 and hired a research-heavy CTO, Ece Kamar (ex-Microsoft Research).
Pedigree: Microsoft Research (CTO Ece Kamar, ex-CVP AI Frontiers Lab)
micro1 is a San Francisco AI training-data company founded in 2022 by Ali Ansari. It began as an AI recruiter (Zara) and now uses that system to vet expert contractors. It sells expert human data, frontier evaluations and RL environments to AI labs under the Realm brand, an enterprise agent-evaluation layer called Cortex, and expert-demonstrated robotics data, all run on its flow data platform. To make its 'RL gyms' realistic it pays companies for operational data, and it made a $12.5M late bid for Spirit Airlines' records. TechCrunch reported a $500M gross run rate in August 2026, and Forbes reported a raise of more than $100M at a $4B valuation in September 2026, with frontier labs taking part.
Backers: 01 Advisors (Dick Costolo, Adam Bain; led Series A; Bain on board), Two unnamed frontier AI labs (2026 round, reported), Two unnamed xAI co-founders (2026 round, reported)
acqCommercial
Deeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms') for computer-use and code, where AI agents practice and are evaluated on realistic digital knowledge-work tasks (simulating tools like Slack and Salesforce). It sold these pre-built environments primarily to frontier AI labs and raised a $43M Series A led by a16z (March 2026). In July 2026 it was acquired by Mercor; the team joined Mercor and Deeptune's environment platform now sits under Mercor. It is therefore no longer ranked as an independent vendor.
Backers: Andreessen Horowitz (a16z, lead), 776, Abstract Ventures
CodeComputer UseEnterprise
site ↗
n/rOpen source
General Reasoning is a London AI research lab (legal entity General Reasoning, Inc., US) working on models that operate over much longer time horizons. It builds OpenReward, a platform for serving RL environments (the vendor states 380+), and the open Open Reward Standard. It also ships open tooling (Firehorse, inspect-openreward, Harbor support), the KellyBench long-horizon benchmark and BackSearch, a paid point-in-time search API. It says its own language models launch in Q4 2026.
Backers: Entrepreneur First
E2B provides open-source Firecracker-microVM sandboxes for AI agents. It is offered as a hosted API, as BYOC, and since September 2026 as a self-hostable single-node package (E2B Embed). In 2026 it became a supported sandbox provider in the OpenAI Agents SDK and published integrations for Devin Outposts and Cursor Self-Hosted Machines. Vendor case studies show it used for RL rollouts (Paper Instruments) and large-scale evals (Arena).
Backers: Insight Partners (Series A lead), Decibel (seed lead), Sunflower Capital
n/rInfrastructure
Runloop sells cloud VM-isolated devboxes and sandboxes for AI agents. Around them it now offers agent lifecycle tooling: the Agent API, Axons and Broker event streams, Agent Gateway and MCP Hub for credential isolation, Benchmark Job Orchestration with W&B Weave, Reflex background agents, and VPC deployment. It is execution and eval infrastructure for agent builders and training workloads, not an RL-data or environments vendor.
Backers: The General Partnership (lead), Blank Ventures, Exponent Founders Capital
acqCommercial
Mechanize is a San Francisco vendor founded in April 2025 by ex-Epoch AI researchers. It builds a small number of high-fidelity, long-horizon RL environments and evals for frontier coding agents, and publishes the open-source GBA Eval benchmark. It raised $9.1M at a $500M post-money valuation in April 2026. In September 2026 Business Insider reported that Google had completed a license-and-hire deal: co-founder and CEO Tamay Besiroglu and more than a dozen staff joined Google DeepMind, and talks had been reported at more than $1.5B. Mechanize continues under CEO Guive Assadi.
Backers: Nat Friedman, Daniel Gross, Patrick Collison
n/rInfrastructure
Cua (trycua, YC X25) builds open-core 'computer-use 2.0' infrastructure: an MIT, agent-neutral background driver for macOS, Windows and Linux (CLI/MCP), local and cloud sandboxes, elastic Cua Fleets for parallel eval and RL training (cloud, BYOC, on-prem), Lume macOS virtualization, the Cua Spaces Mac app, small CUA-S1 decision models, and Cua-Bench for verifiable tasks, leaderboards and RL trajectory export. SOC 2 Type I.
Backers: Y Combinator (X25 batch)
n/rOpen source
Good Start Labs is a 2025 Every spin-out that sells game-based RL environments and agent and human gameplay data to AI labs. It also partners with game publishers (Bad Cards, Arkadium GameLab) to turn play into licensed training data, in-game agents and leaderboards. It publishes transfer studies from games to real work (Diplomacy to customer support, the 1830 railroad game to finance research), arXiv papers and open repos, and runs the Diplomacy and Humor (Bad Cards) Arena leaderboards.
Backers: General Catalyst, Inovia Capital, Tirta Ventures
n/rIncumbent
Surge AI is a bootstrapped human-data and RLHF leader for frontier labs (reported $1.2B 2024 revenue). In 2026 it built a large in-house benchmark and RL-environment program (DAYJOB, GDP.pdf/.xlsx, HANDBOOK.md, Chartography, CoreCraft, and the Tuesday Work Index composite), published post-training transfer studies, launched Post-Training Runs and an enterprise offering, and released sample harnesses under Apache-2.0. Anthropic reports results on Surge's GDP.pdf in its Claude Sonnet 5 system card, and Microsoft publicly credits Surge's expert raters. OpenAI's adoption of GDP.pdf is so far claimed only by Surge.
Pedigree: Founder/CEO Edwin Chen: former research scientist at Google, Facebook
Sepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data, expert-graded evaluation benchmarks, and reinforcement-learning environments for frontier LLMs, drawing on a network of 20k+ domain experts (PhDs, finance, medical, STEM). It was acquired by Mercor in February 2026; the team joined Mercor and its RL-environment and human-data work now sits under Mercor. It is therefore no longer ranked as an independent vendor.
Backers: Y Combinator, Metaplanet Holdings, SID Venture Partners
n/rInfrastructure
Morph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal technology, which can snapshot, branch, and restore entire computational environments in roughly 100-250ms to enable massively parallel, reversible ('Git for compute') agent rollouts, evaluations, and reasoning-time branching. It markets the platform (Morph Cloud) as infrastructure for running and scaling agent/RL verification environments rather than as an RL-environment dataset vendor itself.
Backers: Christian Szegedy (reported seed/angel investor; also Chief Scientist), amount undisclosed
No vendors match these filters.