Blog
Category

Heading for the Edge: Where AI Spending Goes Next

By
Jeff Pehl
Director, Investments
August 2026
7
min read
Share this post
“Now, you don’t necessarily want to run everything in the cloud — because if you can run it locally, it’s free.”
— Jensen Huang, CEO of NVIDIA
Download the accompanying presentation here.

For the past few years, the AI sector has been about one thing: training. The building of bigger and bigger models took enormous amounts of computing power, and the companies that sold the picks and shovels, such as Nvidia, for that were the runaway winners.  

Now, that's changed. The market has woken up to the fact that the next phase of AI is not about just building frontier models, it's about using them. That's called inference, and it's the part where AI actually does useful work — answering a question, writing an email, flagging a fraudulent transaction, helping a doctor read a scan. Training a model is a one-time cost. Inference is a forever cost. Every time someone uses an AI product, the meter runs.  

Our view is that inference moving to the edge, custom chips are moving into specific applications, and agentic AI driving a new monetization layer. Less of the value will sit with one or two vendors. More value will be created by the companies that own the device, the workload, and the chip designed for it.

Lets dig into the implications and opportunity.

The shift is already showing up in the numbers: inference’s share of all AI compute is on track to double.  

While GPUs currently dominate, ASICs (Application Specific Integrated Circuit), custom built for a specific task with speed and power benefits, are expected to grow from representing 10–30% of internal hyperscaler workloads today to 30–50% within five years.

The reason hyperscalers are willing to spend hundreds of millions of dollars per chip design is simple economics: custom silicon offers up to a 65% total cost of ownership advantage over conventional GPUs for inference at production scale.  

When you are going to run the same workload billions of times, a chip built specifically for that workload wins.  

It is worth being clear about who really benefits in public markets today: Broadcom and Marvell together control roughly 95% of the ASIC co-design market, and Taiwan Semiconductor manufactures it. The "custom silicon" story is, in practice, a story about a very small number of companies.  

In memory, three companies produce HBM, and SK Hynix holds roughly 58% of the market by revenue, roughly triple either Samsung or Micron. In compute infrastructure, CoreWeave has become a capacity provider for many frontier training runs, while Nebius plays the same role among the neoclouds serving the next tier of demand.  

And at the model layer itself, the concentration is starkest of all: there is no pure public vehicle for frontier models with Alphabet is the only listed company running a frontier lab outright, with Microsoft and Amazon offering exposure only obliquely, through their stakes in OpenAI and Anthropic.  

In private markets, capital is flowing toward inference-optimized chip startups, edge AI hardware, and the application-layer software that turns models into actual products. Funding rounds for inference infrastructure, such as serving, routing, optimization, have moved up the priority list, and valuations for pure model labs have started to compress relative to the companies that put models to work.  

At BFA, we target exposure to companies at each layer of the stack through co-invests or funds with our venture partners — from the chip layer to the agentic & AI model layer.  

Where we think this goes next: Edge and ASIC at the application layer

First: AI moves onto the device. Today, most AI runs in a data center somewhere. Increasingly, it will run on your phone, in your car, on the factory floor, or inside a pair of glasses, at the edge, so to speak.  

The models have gotten small enough: well under 10 billion parameters, versus the hundreds of billions behind the frontier models running in the cloud to do real work locally, and running them, on-device is faster, cheaper, and more private than pinging the cloud. Running locally means data never leaves the machine and businesses own and customize their AI.  

The growth numbers are eye-watering. For example, the edge AI market is expected to grow from $29 billion in 2025 to $197 billion by 2034, at a 24% annual clip. Phones, laptops, cars, robots, medical devices, retail kiosks are all becoming AI devices in their own right.  

Second: Custom chips go mainstream. Until now, designing your own AI chip was something only the giants did: Google's TPU, Amazon's Trainium, Meta's MTIA, Microsoft's Maia.  

The next wave is much broader with companies designing chips for one specific job. Self-driving cars. Robots. Video. Real-time translation. Defense. Once a workload is high-volume and predictable, the same logic that pushed hyperscalers to build their own silicon applies to everyone else.  

The application layer is moving in the same direction. The next wave of AI demand will come from agentic AI, possibly monetised through ads and usage fees. Either way, the conclusion is the same: more inference, likely more of it running on device, and running on chips built for the job.  

Put those trends together: inference moving to the edge, custom chips moving into specific applications, and agentic AI driving a new monetization layer on top, and you get a much bigger and more distributed semiconductor market than the one we have today. Less of the value sits with one or two GPU vendors. More of it sits with the companies that own the device, the workload, and the chip designed for it.  

This is where it stops being a diagnosis and starts being a thesis. If inference is a forever cost, the most valuable place to sit is wherever that cost gets paid, the device, the workload, and the chip built for it. Four points outline the types of companies we are targeting for our investment exposure:  

  • Inference-first silicon. Chips built for the workload, not the benchmark; a 65% cost advantage means every high-volume job eventually gets its own chip. That's d-Matrix: in-memory accelerators purpose-built for inference. Run the same job billions of times, and purpose-built wins.
  • Edge-native silicon and devices. Models under 10 billion parameters now do real work locally with a $29 billion market headed to $197 billion. That's Hailo, putting inference chips in cameras, cars, and machines, and Figure AI, building the robots they'll power. On-device is faster, cheaper, more private, and the business owns its AI.
  • Inference infrastructure. Training built one Nvidia; inference needs a supply chain. Fluidstack is the Google-backed data center build-out anchored by frontier lab demand; Lambda is the wholesale AI factory for Microsoft, Nvidia, and the independent labs.  
  • Agentic applications that own the workload. A single agentic task fires off dozens of model calls, and the agent keeps working when nobody's watching the meter never stops. Poolside AI is the example: purpose-built for software engineering, sharper with every code-execution feedback loop, and deployable inside the customer's own security boundary. The workload and the data both staying home.

The common thread is the workload. The training era rewarded whoever built the biggest model. The inference era rewards whoever owns the device it runs on, the job it runs for, and the chip it runs best on. That value won’t sit with one or two vendors. It’s being built across the stack.  

Read more insights in the accompanying presentation here.

This article was jointly written by the investment team:

Jeffrey Pehl, Director, Investments

Gavin Ezekowitz – Chief Investment Officer

Matt Curtolo – Managing Director, Investments

Alexander Prater – Investment Associate

About the Author

Jeff Pehl contributes research and market analysis to the broader investment effort, drawing on more than a decade in Global Investment Research at Goldman Sachs across the United States, Australia & Asia Pacific, as well as his broader experience across markets and assets classes including time at a global EM fund. His background adds a public-markets, macro and thematic lens to manager evaluation and portfolio context.

General information only. Not financial advice. Wholesale clients only.

© BFA Global Investors 2026.