Artificial intelligence

The AI Bottleneck Keeps Moving Up the Stack

AI research taught

Five years ago, the hardest part of working in AI was proving it was worth paying for. Today, the hardest part is something almost nobody outside the field talks about making sure a capable model does not fall apart the moment it is asked to act on its own.

I’ve spent the last five years chasing this moving bottleneck to commercializing AI, first at Apple and Samsung Research, building products on top of AI research; then at Niantic, teaching models to understand the physical world; and now at Turing, building the post-training pipelines, benchmarks, and reinforcement learning (RL) environments that frontier labs use to develop models reliable enough to ship.

In those three jobs, I have worked on three layers of the same stack – and in the process, I have encountered three different versions of “the hard part.” Usually, there are the same two constraints across altitudes: what it takes to build a good model (data, researcher talent, computing) and where that model actually solves a real-world problem. However, that bottleneck never truly disappeared, it just moved and understood why it is more useful than any prediction about where AI goes next.

Applied AI: the bottleneck was product-market fit

As a product manager working on visual intelligence and search at Apple, and later on computer vision at Samsung Research, the model itself was rarely the constraint. The real question was which of a thousand technically feasible ideas were also commercially viable — identifying user problems ambiguous enough that no one had named them yet, then proving there was enough demand to justify building a solution.

In this sense, this AI research resembled venture investing: brainstorm pain points across a dozen verticals, back the shortlist with quantitative proxy data, prototype with a small team of researchers and engineers, and kill anything that does not survive contact with real usage. That is the application-identification constraint in its most visible form.

The model-building constraint was there too, just harder to see. The technical challenges were present — data scarcity, models trained on synthetic data misbehaving in the real world, infrastructure costs that made a promising demo unshippable — but they were downstream of a harder question: was this worth building at all?

Foundational AI: the bottleneck moved to generalization

By the time I was working on spatial AI at Niantic, the constraint had shifted. The market had stopped asking “should we build this” and started asking “can the model generalize.” Spatial AI — the subset of spatial computing concerned with getting machines to perceive, reason about, and act in 3D space — was maturing from narrow, task-specific systems (SLAM, LiDAR-based mapping, point-in-time computer vision) toward foundational models that could generalize across both real and virtual environments.

The data constraint became an access problem: spatial foundation models need large-scale, real-world 3D data that no one can simply scrape off the internet. Meanwhile, the application-identification constraint got harder in the opposite direction — it was no longer about finding a use case for an existing model, but deciding which spatial problems were worth training a foundation model to solve. Foundation models are expensive to build and even more expensive to get wrong; the teams that succeeded understood early which capabilities were about to become table stakes because a foundation model was about to swallow their category.

Infrastructure: where both bottlenecks become the entire business

Which brings me to where I sit today at Turing, working with frontier AI labs on the layer underneath both of those eras: post-training pipelines, evaluation benchmarks, and RL environments. This is the part of AI that almost never makes it into a keynote; it is also currently the part that determines whether anything built on top of it actually works.

Here, the data-and-talent constraint is not a supporting cost center, but the product itself. Instead of lacking computing or model architecture ideas, frontier labs now lack the specific, high-signal data needed to make post-training work — expert-authored reasoning tasks, human preference data for RLHF, PhD-level rubrics that let a model’s output be graded correctly. “Talent” here means domain experts who can write a task hard enough that a frontier model actually fails at it — a scarcer skill than it sounds.

The application-identification constraint shows up in a form I did not expect it is no longer about finding a market for the model; it is about finding the specific failure mode worth building a benchmark or RL environment around. A model can clear every existing eval and still fall apart the moment it is deployed as an agent — generating a malformed document or losing coherence on step forty of a multi-step task — because nobody had identified that scenario as worth testing. Designing a test problem hard enough that a capable model actually fails at it, then feeding that failure back into training, is now one of the highest-leverage things a frontier lab can do.

The throughline

None of these eras replaced what came before it. Product-market fit still matters, and foundation models still need someone to decide what’s worth building. What changed is where the constraint sits, and every time it moved, it moved up a layer of abstraction: from the application to the model, to the infrastructure that trains and assesses the model.

If there is a single piece of advice, I would give researchers or developers to decide where to spend the next few years, it is this: do not assume the bottleneck you learned to solve is still the one worth solving. The AI industry has relocated “the hard part” at least twice in five years – and betting on where it lands next is a lot more useful than getting good at where it used to be.

Comments

TechBullion

FinTech News and Information

Copyright © 2026 TechBullion. All Rights Reserved.

To Top

Pin It on Pinterest

Share This