Artificial intelligence

Before the NPU Became the Future of AI, Kneron Was Already Building It

Today, the Neural Processing Unit, or NPU, is rapidly becoming one of the most important pieces of the AI computing landscape. NPUs are appearing across smartphones, PCs, automobiles, cameras and intelligent devices. The technology industry is beginning to recognize that the future of artificial intelligence cannot exist only inside massive data centers. Intelligence also needs to live where people actually use it.

But almost a decade before the NPU became part of the mainstream AI conversation, Kneron was already building for that future. Founded in 2015 by AI researcher and engineer Albert Liu, Kneron began with a surprisingly simple question: What if artificial intelligence did not have to travel to the cloud every time it needed to think?

At the time, that was far from the dominant vision of AI.

The industry was moving in the opposite direction. AI models were becoming larger. Computing infrastructure was becoming more centralized. GPUs were emerging as the workhorses of machine learning, and enormous amounts of computing power were beginning to concentrate inside increasingly powerful data centers.

Kneron saw another part of the equation. Training artificial intelligence requires enormous computing resources. But once intelligence has been created, does every decision really need to return to a distant server?

A camera recognizing a person should not necessarily need the cloud to understand what it sees. A car making a split second decision cannot afford to wait for a remote data center. A factory robot should not stop thinking because its internet connection disappears. And increasingly, consumers and enterprises may not want sensitive information leaving their devices in the first place.

The challenge was no longer simply how to create more powerful AI. It was how to put that intelligence everywhere. That problem helped shape Kneron’s early work on the Neural Processing Unit.

A Different Kind of AI Processor

An NPU is not simply a smaller GPU.

GPUs were originally built for highly parallel computing and proved remarkably effective at accelerating the mathematical operations behind machine learning. Their enormous computational capabilities helped ignite the modern AI revolution and remain essential to training frontier models.

The NPU approaches the problem from another direction. It is designed specifically around neural network computation, allowing AI inference to be performed with very different power, latency and deployment requirements. For Kneron, that distinction was fundamental.

The company says it introduced its first NPU intellectual property for endpoint devices in 2016, when today’s conversation around AI PCs and ubiquitous on device intelligence was still years away. By 2018, Kneron was publicly discussing ultra low power NPU architectures capable of running deep learning networks directly on devices. Then, in 2019, Kneron introduced the KL520, an AI system on chip powered by its NPU architecture and designed specifically for edge AI inference.

That moment is important to understanding Kneron’s story. Today, describing a company as an “Edge AI chip company” places it inside an existing market category. Kneron’s history suggests something more interesting. It was building the technology while the category itself was still being defined.

The Harder Problem Was Making the NPU Adapt

Kneron was also thinking about another problem that has become even more important with the explosion of modern AI: AI changes incredibly quickly. A processor optimized for today’s neural network can become far less useful when tomorrow’s models look completely different. So rather than designing an NPU around a single fixed workload, Kneron pursued a reconfigurable architecture designed to support different neural networks and evolving AI applications.

That technical work later received recognition from IEEE. In 2021, research connected to Kneron’s neural processing architecture received the IEEE Circuits and Systems Society’s Darlington Best Paper Award. In 2023, Kneron received the IEEE Consumer Technology Society’s Corporate Innovation and Leadership Award for its work in consumer technology.

AI Is Starting to Leave the Data Center

For the past several years, the AI race has largely been measured in scale. More parameters. More GPUs. Larger clusters. Larger data centers. More electricity. That race produced extraordinary advances. Generative AI would not exist in its current form without massive cloud computing infrastructure.

But the success of AI has created another problem. What happens when billions of devices want to use it continuously? Sending every interaction from every camera, vehicle, computer, robot and intelligent machine to the cloud introduces costs that are not purely financial. It can mean additional latency, bandwidth requirements, energy consumption and privacy considerations. The more ubiquitous AI becomes, the more important the location of computation becomes.

That is why the next chapter of AI may look very different from the first. The cloud will not disappear. Neither will the GPU. Instead, intelligence is beginning to spread outward. From the data center to the factory floor. From the cloud to the car. From enormous GPU clusters to computers, cameras, robots and machines capable of processing AI where the data is actually created. This is the world the NPU was built for.

And it is remarkably close to the world Kneron imagined when it started nearly a decade ago.

The GPU Era Created an Unexpected Problem

The extraordinary success of GPUs has demonstrated just how powerful AI can become when enormous amounts of computing power are concentrated together. It has also revealed the limits of relying exclusively on that model.

AI data centers require enormous amounts of electricity. Advanced chips require sophisticated cooling infrastructure. Cloud inference costs accumulate every time a model is used. And governments and enterprises are increasingly asking where their data goes when AI processes it.

The next great competition in AI may therefore be about more than who can build the largest model. It may be about who can deliver the most intelligence with the least amount of energy, latency, infrastructure and cost.

That changes the semiconductor equation. The question is no longer simply: How powerful is the processor? It becomes: How efficiently can intelligence be delivered? And that is precisely where the NPU becomes strategically important.

Kneron’s Bet Is Getting Bigger

Kneron’s original thesis was that AI should be able to operate directly on devices. Its next thesis is considerably larger.

The company increasingly sees the NPU not merely as a component inside an endpoint device, but as the foundation for an AI computing architecture that can scale across different layers of infrastructure. In other words, Kneron is attempting to take the principles that helped make on-device AI possible and push them upward into much larger AI workloads.

That could become an important shift. For years, the semiconductor industry largely divided AI computing into two worlds: enormous computing resources in the cloud and relatively small amounts of intelligence at the edge. Those boundaries are beginning to blur.

Edge systems are becoming dramatically more powerful. Models are becoming more efficient. Specialized processors can handle workloads that once required much larger infrastructure. And increasingly sophisticated AI can operate closer to where information is generated.

The edge is no longer simply the place where small AI models go. It is becoming another layer of AI infrastructure.

What Comes After the GPU Era?

The company made an early bet on specialized neural processing before NPUs became a mainstream technology category. It spent years developing reconfigurable NPU architectures designed for a world in which AI would need to operate beyond the data center.

Now that world is arriving.

The GPU will continue to be one of the most important technologies in artificial intelligence. It will power frontier model training, massive data centers and workloads requiring extraordinary computational scale. But the future of AI is unlikely to belong to one processor alone. It will require an ecosystem of computing architectures, with different processors performing different jobs at different points between the cloud and the physical world.

GPUs helped solve one of the defining problems of the first AI era: How do we create intelligence at unprecedented scale?

NPUs may help solve the defining problem of the next one: How do we put that intelligence everywhere? That distinction is what makes Kneron’s history worth paying attention to.

Technology companies often describe themselves as being ahead of the curve. Far fewer can point to a technological bet made years before the rest of the market began moving in the same direction. Kneron started building NPU technology when much of the world was still focused on centralizing AI computation.

Nearly a decade later, NPUs are moving into the center of the global AI conversation. Now Kneron is betting that the same architectural shift is only beginning. The GPU helped ignite the AI revolution. The NPU could determine how far that revolution ultimately reaches. And Kneron has been preparing for that future since before most of the world had a name for it.

Comments

TechBullion

FinTech News and Information

Copyright © 2026 TechBullion. All Rights Reserved.

To Top

Pin It on Pinterest

Share This