The Rise of Consumer AI PC's aka. NPUs: History, Technology, Devices, and Economics
Estimated read time: 9 minutes
A Neural Processing Unit, or NPU, is a dedicated chip built into modern laptops and desktops that handles AI calculations locally instead of sending them to a cloud server. Apple started the trend in November 2020 when the M1 chip shipped with a 16 core Neural Engine capable of 11 trillion operations per second.
Qualcomm, Intel, and AMD have since added their own NPUs, and by 2026 a 40 TOPS or faster NPU is the baseline requirement for a Microsoft Copilot+ PC. This guide covers how NPUs got here, how they actually work, which devices have them today, and whether running AI locally really is cheaper than paying for cloud subscriptions.
What is an NPU and how is it different from a CPU or a GPU
An NPU is silicon designed for one job: running the matrix multiplication behind machine learning inference. A CPU is built for general, low-latency tasks, a GPU is built for high-throughput parallel work like graphics and model training, and an NPU sits between them as a low-power accelerator dedicated to AI workloads.
NPUs typically run in lower-precision formats such as INT8 rather than the 32-bit math a CPU uses, which is what lets them hit TOPS figures in the tens or hundreds while drawing only a few watts. Consumer NPUs are built for inference, meaning they run an already-trained model. If you want to train or fine-tune a model from scratch, you still need a GPU or a cloud service.
When did NPUs first show up in personal computers
The idea started on phones. Apple's A11 Bionic chip brought a Neural Engine to the iPhone in 2017, and Qualcomm's Hexagon DSP was doing similar work in Android phones around the same time.
The shift to PCs came in November 2020, when Apple's M1 chip brought a 16 core, 11 TOPS Neural Engine to the MacBook Air, MacBook Pro, and Mac mini, making it the first mainstream computer with a built-in NPU. Every other major chipmaker followed within a few years, as the timeline below shows.
history of npus
Which laptops and desktops actually have NPUs in 2026
Nearly every new premium laptop released in 2026 ships with an NPU, but budget machines and most desktop towers still do not. Here is how the four major chipmakers stack up.
Apple's M1 through M5 chips all include a Neural Engine. The MacBook Air M1 launched at 999 dollars and the Mac mini M1 at 699 dollars in 2020, while the current MacBook Pro lineup with M5 Pro and M5 Max starts at 2,199 dollars and climbs to 3,899 dollars for the top 16 inch configuration.
Qualcomm's Snapdragon X Elite (2024) and Snapdragon X2 Elite (2026) power Windows on Arm laptops such as the Lenovo Yoga Slim 7x, HP OmniBook X, and ASUS Zenbook A16, with Hexagon NPUs rated at 45 and 80 TOPS respectively.
Intel's Meteor Lake chips brought the first Intel NPU to laptops like the Dell XPS 13 Plus in 2023, and 2026's Panther Lake generation (marketed as Core Ultra Series 3) raised that to a dedicated 50 TOPS NPU, with total platform AI performance reaching up to 180 TOPS once the GPU and CPU pitch in.
AMD brought its XDNA NPU to x86 laptops with the Ryzen 7040 "Phoenix" series in 2023, then to desktops for the first time in March 2026 with the Ryzen AI 400 series, built on the familiar AM5 socket.
consumer npu models
AMD's Ryzen AI 400 desktop chips, first previewed at CES 2026 and formally launched at Mobile World Congress in March 2026, are not sold as boxed retail chips. You only get one inside a prebuilt system from HP, Lenovo, or Dell, at least for this first wave.
Can on-device AI replace cloud AI
For a lot of everyday tasks, yes. For the largest and newest models, not yet.
Live camera effects, dictation, offline transcription, and image touch ups already run entirely on-device. Smaller open models like Llama and Mistral also run locally through free tools such as LM Studio, Ollama, and llama.cpp, and Apple's unified memory Macs handle this especially well: a 128GB M5 Max can reportedly run Llama 3 70B at around 48 tokens per second using Apple's MLX framework, while Qualcomm has claimed its original Snapdragon X Elite could run a 13 billion parameter model at roughly 30 tokens per second.
The largest frontier models, along with most image and video generation at top quality, still require datacenter GPUs and stay cloud only. If you want to see how much of that cloud infrastructure is running at any given moment, OpenAI's live server map gives a real-time look at global capacity.
On-device AI often needs a bit of setup too. Running Llama 3 locally means downloading model weights and a runtime, though some AI features come pre-installed, such as Apple's on-device dictation and Windows 11's Voice Typing.
Does running AI on an NPU actually protect your privacy
Yes, in the sense that raw data does not have to leave the device, though the operating system around the NPU still matters.
Apple processes most Siri voice requests entirely on-device through the Neural Engine, and Qualcomm markets its Hexagon NPU the same way: computation happens locally, so personal data is not automatically sent out. That said, an NPU only handles computation. The OS and installed apps can still collect diagnostics or usage telemetry regardless of where the AI processing happens, so it is worth checking privacy settings on any AI feature you turn on.
One side effect worth knowing about: NPU driver support on Linux still lags well behind Windows and macOS, according to reports from Framework laptop's own community forums. That is inconvenient for Linux users who want local AI, but it does mean an NPU without a working driver simply cannot run, which is its own quiet privacy floor.
Can you upgrade or add an NPU to an existing PC
No. NPUs are built into the CPU or SoC package rather than sold as a separate card, so there is no way to bolt one onto a machine that does not already have one.
In laptops, the NPU is soldered in place with the rest of the chip, so getting a better one means buying a new laptop. On AMD's AM5 desktop platform, a chip swap to the Ryzen AI 400 series adds an NPU the same way any CPU upgrade would, once those chips reach retail rather than OEM-only channels. There is no standard NPU expansion card for general computing; small devices like Google's Coral Edge TPU sticks exist for IoT and hobbyist projects, not mainstream PC upgrades.
Is it cheaper to run AI locally than to pay for cloud AI subscriptions
For heavy, sustained use, increasingly yes, and the math has gotten more favorable through 2026.
Lenovo's 2026 total cost of ownership analysis found that on-premises AI infrastructure can reach breakeven against cloud rental in as little as four months for high-utilization workloads. On a per-token basis, the same analysis found a large model like Llama 3.1 405B costs about 4.74 dollars per million tokens on owned hardware versus 29.09 dollars on a comparable AWS cloud instance, an 8 times difference against standard cloud pricing, and up to 18 times cheaper than paying for a fully hosted frontier API outright.
costs of local vs cloud ai
These specific figures come from a business running its own inference servers rather than a single laptop, but the underlying logic holds at consumer scale too. A Copilot+ PC with a capable NPU pays for itself faster the more you rely on local models instead of metered API calls.
What should you look for when buying an AI PC in 2026
● If you plan to run local AI models yourself rather than just using built-in Copilot or Apple Intelligence features, memory matters more than the TOPS number alone. A 128GB Mac or Snapdragon X2 laptop can run a 70B parameter model that will not fit on a 16 to 32GB Windows machine.
● Check for the 40 TOPS Copilot+ PC threshold if you want Windows features like Recall and Cocreator. Every 2026 Panther Lake, Ryzen AI 400, and Snapdragon X2 chip clears that bar comfortably.
● Remember NPUs cannot be upgraded later, so buy the chip you actually want on day one rather than planning to add AI hardware afterward.
● If your business is already paying for several AI subscriptions, or experimenting with running its own agents and chatbots instead of metered APIs, capable local hardware can start paying for itself within months rather than years.
If you are exploring how AI agents could take over some of the repetitive work currently costing your business a subscription fee, or whether an AI SEO service makes more sense than running search visibility tools yourself, both are worth a look before you commit budget to new hardware.
New to running any of these tools hands-on? Our free beginner's AI course walks through the best AI tools to use today, no coding required. And if your only real NPU-adjacent need right now is customer facing chat, it is worth comparing what a dedicated AI chatbot service can do for you versus a general purpose assistant.
Bottom line
Consumer NPUs went from a novelty in the 2020 M1 Mac to a baseline requirement across nearly every premium laptop sold in 2026. The chips are fixed once you buy them, cloud AI still wins for the very largest models, and local AI increasingly wins on cost and privacy for everything else. Evaluate whether the privacy, speed, and offline capability of an NPU actually match what you need, weigh that against your current cloud subscriptions, and keep an eye on the next wave of chips, since this hardware category is still moving fast.
Frequently asked questions
What TOPS rating do I need for an AI PC in 2026?
Microsoft's Copilot+ PC baseline is 40 TOPS. Most 2026 chips, including Intel Panther Lake, AMD Ryzen AI 400, and Qualcomm Snapdragon X2 Elite, clear that with NPUs rated between 50 and 80 TOPS.
Do I need an NPU to use ChatGPT or Google Gemini?
No. Cloud based chatbots run entirely on the provider's servers regardless of your hardware. An NPU only matters if you want to run a model locally, offline, on your own device.
Can I add an NPU to my existing desktop PC?
Not directly. NPUs ship built into the CPU package. On AMD's AM5 platform, upgrading to a Ryzen AI 400 series chip gets you an NPU the same way any CPU swap would, once retail chips are available. On Intel and most other platforms you need a new motherboard and CPU.
Is running AI locally actually cheaper than a ChatGPT subscription?
For light, occasional use, a monthly subscription is usually still cheaper. For heavy, constant use, owned hardware can break even against cloud costs in as little as four months, according to Lenovo's 2026 total cost of ownership analysis.
Which 2026 laptops have the fastest NPUs?
Qualcomm's Snapdragon X2 Elite Extreme currently leads at 80 TOPS, followed by Intel's Panther Lake and AMD's Ryzen AI 400 series, both rated around 50 to 60 TOPS on the NPU alone.