back to insights page
back to fundamentals page
Thesis
AI
blog
10
September
,
2026
4 mins

The Last 35 Points

Somewhere in a warehouse, a humanoid robot is sorting poly bags and flat envelopes so a scanner can read the barcodes. Four seconds per pack. Barcode read rate sits around 95%, up from 70% six months ago. It runs around the clock. By any fair measure, it works.

What it cannot do is walk to the next station and tend to a different machine. The robot is general-purpose hardware. The deployment is a single fixed task - object types known, lighting tuned, workflow rebuilt to fit the robot's reach envelope. The only real variability is the shape and weight of the bag. The robot is general-purpose. The deployment is not.

That gap - between a robot that could in principle do anything and a deployment where it does exactly one thing - is where the largest opportunities in physical AI sit for the next few years.

Capital is not the bottleneck. Robotics venture funding was around $15B in 2025 and has passed $18B in 2026 so far. A humanoid hardware company is trading north of $60B in the public market; the leading private ones are marked between $5B and $39B. Against all that, the deployed base is negligible. The most advanced humanoid program in the west has a few hundred units in production settings. Most industrial robots working today still run pre-programmed routines - no learned behavior at all.

The model is not the bottleneck either. The architecture is converging: a small action model on the robot converting instructions to motor commands, a large reasoning model in the cloud handling planning, video frames flowing upward to both. Everyone serious is arriving at this shape from different directions.

What doesn't exist is the data to make that recipe work at any particular site, and collecting that data is physical. It happens one warehouse, one factory floor, one loading dock at a time.

Which turns the adoption bottleneck into something unglamorous and VERY investable.

What "35 points" means

Take a leading action model. Fine-tune it on a specific commercial task — say, machine tending, bin picking, or kitting. In controlled conditions, you land somewhere around 55% task success. Production reliability - the number a customer will pay for - is around 90%. Closing those 35 points is what every robotics team is really spending its money on right now, and almost none of that spend is research. It is site engineering: prompt tuning, custom fixtures, camera repositioning, task decomposition, failure triage, targeted retraining on the failures that showed up in production.

Last mile deployments are the way to bridge this gap between a demo and a production robot, and we think some of the most successful companies in physical AI will start this way.

Why no one owns this layer

The frontier labs are racing toward a general foundation model. Site-specific deployment engineering is a distraction from that race, at least until the model plateaus. Their incentive is to ship a model, not to stand up a field team at every customer site.

The customers don't want to own it either. We talked to engineers at contract manufacturers and 3PL operators; the consistent message was that they will buy access to a model, they will pay for a deployed outcome, but they will not hire a robotics team to babysit the integration. They want this to be outsourced.

That leaves a layer with real demand, real willingness to pay, and no incumbent positioned to capture it well. We think it will get built as a combination of a product and a service.

What a deployment company actually does

A customer shows up with a task and a setting: pack these orders, tend this CNC machine, sort these returns, palletize these cases. The deployment company selects the action model, tunes it against that task, builds the fixtures and the camera setup, instruments the failures, runs the correction loop, and measures a success rate. What the customer gets is uptime on a task at a site.

This is what a typical loop looks like in a physical AI deployment:

  1. Detect. Today there is no automated failure detection in deployments. Oftentimes, a human watches live feeds of thousands of robot attempts a day and manually flags failures. Failures are caught after they happen, not predicted, and if the robot's training data is full of try-fail-retry sequences, it learns that pattern as normal. An ideal detection system will watch and solve for anomalies in real time.
  2. Correct. When something breaks, a human intervenes - teleoperation, in-person demonstration, remote guidance, structured retries with variation. These interventions are the most valuable data in the entire system. The deployment company's job is to make these interventions systematic, low-latency, and well-labeled.
  3. Compound. This is where the business either scales or doesn't. After every deployment, the company has to answer: which corrections generalize to other sites, which are site-specific, and what can the next customer skip?

A deployment company has the harness already, because it needed one for the last engagement. Over time, this becomes a benchmarking and recommendation layer - "for deformable-object manipulation in the 50-500g range, Model X outperforms at 20% lower compute" - that the model providers themselves want access to.

This is the entire difference between a services business that scales linearly and a product business with compounding margins. If you are rebuilding the failure taxonomy and the evaluation pipeline from scratch at every new customer, you are a body shop. If site five can skip 40% of the work that site one required, you have a software business with a services motion.

Keeping a narrow focus early is important. One task family across many sites, not many tasks at one site. Depth in tasks early beats a general capability in the short-medium term.

Over a period of time you compound. Failure and intervention data across customers, task families and hardware types, collected in production rather than in a lab. Evaluation infrastructure that says what a model will actually do on a real floor, which is a thing the labs cannot generate for themselves and increasingly need.

The obvious risk is that this layer is temporary. If reasoning models get good enough at physics and safety to automate their own site setup, the deployment loop compresses into the model providers and the opportunity closes. We think that takes longer than people expect, but we'd price it as a 5 year window rather than a decade-long moat.

If you have run a deployment or data-operations team inside a robotics company, or built the internal tooling that turns robot failures into training data and watched it get rebuilt from scratch at every new site, we would like to talk.

insights

BLOG
Why AI Native Services is the Next Big Bet
BLOG
The Agentic Analyst: Why We Invested in Pinegap
BLOG
10x'ing one of the hardest problems in clinical AI. Why we doubled down on CARPL.
BLOG
Why AI Native Services is the Next Big Bet
BLOG
The Agentic Analyst: Why We Invested in Pinegap
BLOG
10x'ing one of the hardest problems in clinical AI. Why we doubled down on CARPL.

insights delivered straight to your inbox

*Click the checkbox below to enable it
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.