As founders figure out how AI translates into the physical world, we’re getting some of the funniest business models I’ve seen in a while.
In New York, a company called Shift is now cleaning private apartments for free. Just one catch: The cleaners wear a headset that records first-person footage of the work, and Shift licenses that footage to train the next generation of household robots. You get a clean apartment, they get training data.
In case you’re surprised, this isn’t entirely new. Shift already runs in 15+ countries with 10k+ households and businesses, and NYC is just the latest market. And the underlying model - giving away (or underpricing) a service to capture training data - already exists elsewhere. comma.ai for example has 20k+ drivers within its openpilot system, generating them miles of real-world driving data. DoorDash launched an app paying their couriers to strap on a body cam and film themselves folding laundry and making beds… all to generate training data for humanoids.
So Shift is not a novelty, but the consequent, consumer-facing version of a growing data generation boom.
I could not resist but ask: When does this hit healthcare? And in which form and shape? Let’s walk through some ideas, and what could come next.
Note: I’ve been speaking to many robotics founders and data people in the past year. But I’m not an engineer or a data guy - so I still feel like I’m at the very beginning in understanding this stuff. So take this piece as a first take on a complex space. You’ve been warned.
The robotic data gap
The foundation for the huge breakthroughs we’ve all witnessed in LLMs was the sheer amount of data available. OpenAI and Co fed the whole internet into their models - Wikipedia, Reddit, and so on. Data that already existed in masses.
But what if you train something that isn’t text? That’s where robotics companies hit a wall in the physical world. There simply isn’t enough useful footage or sensor data to train robots at scale. Which is why we get extreme business models like Shift. People will do almost anything to manufacture the right training data!
From what I’m seeing, video is the most popular data type right now, but haptic feedback, location data and others can be valuable too. And the gap applies to most industries - it’s just brutal in healthcare, where privacy and risk make it much harder to go out and capture data.
But beware, not all data is created equal. First-person video (as Shift uses) is a decent start. But the most solid data type is probably teleoperated robot data: It gives you video footage plus the exact movements that produced them. That includes motor commands, force used, and corrections when something slips. The model can learn the action instead of just watching. That difference counts - and we’ll likely only find out in hindsight how useful which type of data really is.
Note: This leads us to another buzz word you’ll keep hearing: “world models”. A world model is a system’s learned internal model of how the world behaves, so it can predict what happens next if it takes a given action. The data sets we’re describing here could be an essential part of such a world model in healthcare, which is currently missing.
So next, let’s look at the strategies people are using to attack that data gap.
Two bridges to autonomy
Judging from our dealflow, the most common approach to building a fully autonomous robot is “training-on-the-go”. You put a less-than-autonomous robot to work today, and make money while a human is still in the loop. You use the data that usage generates to automate the robot over time. That’s often pitched as the “Tesla approach” - building cool cars, getting them on the roads, and only later automating the driving part. Although I’m not sure Tesla is still a positive role model?!
There are two sub-flavours of this approach:
In the first, you sell a less automated version of your product and the customer operates it. ForSight Robotics is one example. A surgeon drives the robot through cataract surgery (they did the world’s first robot-assisted cataract surgery this April), and every procedure is data on the path toward more autonomy. A lesser-known one is Enabling Robotics, where they attach robotic arms to wheelchairs for daily assistance. The users operate that arm constantly in real life (eating, opening doors, picking things up) and generate incredibly rich robotics data simply by living their day
In the second flavour, you operate the solution yourself and sell it as a service. Starlife does this with teleoperated humanoids: a remote pilot puts on a headset and works through the robot at the customer’s site, and over time that builds the dataset to automate the work. They deployed first MVPs in clinics. Devanthro is a European example in the space, working on bringing humanoids to people’s homes for companionship and everyday support. Shoutout to Rafael for his input on this topic!
Both strategies result in a stream of data to train your robot while already being live with a real customer. Both control the long-term robot solution and their own data pipeline. Not all of them build their own hardware, but they do own the intelligence to run it. Important to mention: Most founders I’ve met here took this route out of necessity - there was no external dataset to train on, so they had to create their own.
The next breed: data sales companies
Now that building autonomous robots looks less crazy, we’re starting to see purely data-focused companies. Back to the LLM analogy, people are asking: Is there a potential VC outlier in doing for robots what Scale AI did for the big labs? Produce exactly the data they need in masses? This is where Shift came from. They look at well-funded humanoid players like Figure, 1X and Neura Robotics raising enough money to buy external data, and they go build it.
In healthcare, I can see that data creation game going three ways:
For specialized, complex environments - think specific surgeries - companies will keep building their own pipelines. I can’t see a data vendor touching this any time soon
For lower-risk, general movements, I’m genuinely curious whether a Shift-like model could work in healthcare. Just imagine nurses or technical staff arranging gadgets, prepping machines, sorting supplies, running logistics. If you equip them with cameras and sensors, you unlock potentially valuable data - and extra revenue for the clinic. I would start with non-patient-facing and non-clinical workflows, for obvious reasons
Another interesting version is where the data is a side product. Bricca in Sweden could be one example: They build an AI wearable for doctors, embedded in the service badge, that listens to and documents the clinical day. The main modality today is audio, but it shouldn’t be hard to include cameras and other sensors in the future. Technically, it could create very interesting datasets on top of its actual job. The catch with this side-product model is that your data pipeline is only as good as the core product’s adoption. No adoption of the underlying business, no data. But if the core business works, the data multiplies your upside
Of course there are some risks to these businesses:
First of all, how certain is the demand for such data from large robotics companies? Scale AI worked because the big LLM labs couldn’t label data fast enough on their own. It’s not a given the humanoid players behave the same way. Figure, Neura and Tesla are building their own pipelines today. Neura is literally constructing “robot gyms” to generate training data in-house, and they treat that as core IP. Despite that, there may well come a point where staying ahead means buying external data on top of your own, the way the labs eventually did. This is worth monitoring.
The other challenge is (obviously!) data privacy. In their household settings, Shift blurs faces and ID cards and that’s enough. In a clinic, getting the privacy right is 10x harder and 100x more risky with all the confidential information flying around. Plus, trust is an immense barrier. You can see that as a challenge, or as a potential moat...
Should we invest?
I’ll be honest: I am still forming my opinion on where and when to invest here. There’s a real timing question around when this data market actually opens. Right now you’ll barely find a buyer for healthcare-specific robotics data.
In other words, it remains a secondary market to robots actually being adopted in hospitals. It’s fuel for cars that are not truly on the streets yet. so the thing to watch is adoption. Wherever robots go into hospitals (Elvio Robotics being a great example here in Europe), a data need shows up somewhere in the value chain.
My guess is we’ll likely see a dedicated healthcare data-pipeline company getting funded within the next 3 years.
If you’re working on healthcare robots and the datasets underneath them, please reach out! If your pitch deck is titled “Shift for Healthcare,” you get fast-tracked straight to our IC.
Speak soon,
Lucas
P.S. jk, I obviously can’t promise the IC part. But I will eagerly talk to you




