There is a next wave of AI development happening underneath the headlines about chatbots and image generators. Researchers are trying to build AI that understands how the physical world actually works: how a box falls off a shelf, how a robot arm grabs an object, how a vehicle should respond when a person steps into the road. The technical name for these systems is world models, and the companies chasing this goal have raised extraordinary sums. Yann LeCun's AMI Labs raised over a billion dollars after leaving Meta. Fei-Fei Li's World Labs also raised over a billion. Goldman Sachs published a report this year calling this the missing link in AI development.
The core problem is data. Chatbots were trained on text scraped from the internet, and there was plenty of it. But there is no equivalent pile of ready-made data that teaches a machine how physical objects behave. Robotics companies have to physically collect that data in the real world, which is slow and expensive. Researchers working in labs have pointed out that this data scarcity is one of the main walls standing between current AI and anything that works reliably in the physical world.
Video games are a surprisingly logical answer. Game studios have spent decades building detailed, physics-accurate virtual environments. The objects in a modern video game behave according to rules that closely mirror reality. They cast shadows correctly, they collide with each other in believable ways, and they exist in three-dimensional space with depth and scale. From an AI perspective, that is enormously valuable raw material.
Origin Lab, which just raised $8 million in seed funding led by Lightspeed Ventures, wants to act as a marketplace connecting game companies to AI labs. The game companies get a new revenue stream from assets they have already built and own. The AI labs get clean, legally licensed data. Origin Lab sits in the middle, converting the raw game files into a format that AI training pipelines can actually use.
The legal angle matters a lot here. OpenAI's video generation tool Sora launched in December 2024 and immediately drew scrutiny when users discovered it could closely replicate footage from Super Mario, Call of Duty, Counter-Strike, and even the visual style of popular Twitch streamers. OpenAI never disclosed exactly where it got its training data. Legal experts at the time said that if game footage was indeed included without licensing, it created serious copyright exposure, because video game content has multiple layers of rights holders: the game developer, sometimes the publisher, and even the player who recorded the footage.
That controversy is essentially the demand case for what Origin Lab is selling. AI labs know they need this data. They also now know that simply taking it from publicly available video is legally risky. A licensed marketplace gives them both the data and the legal cover.
The broader context is the data supply business becoming a serious industry. Scale AI, which processes and labels data for AI companies, went from roughly $290 million in annual revenue in 2022 to an estimated $2 billion in 2025, and was valued at $29 billion after a major investment deal. Investors who backed Scale early watched it become critical infrastructure for the AI industry. Origin Lab's backers are making an early bet on the same pattern repeating for physical world data, specifically for this next generation of AI systems.
For industries that involve physical processes, including manufacturing, logistics, construction, and retail, this matters more than it might initially seem. The AI systems that will eventually manage warehouse robots, quality control cameras, and autonomous delivery vehicles all need to understand physical reality. The quality and variety of data used to train them today will shape how capable they are when deployed in real environments in the next few years. A small, well-positioned data marketplace today could become essential infrastructure tomorrow.