Best Mini PCs for Local AI and LLM in 2025 — Tested
Running AI models locally used to mean buying a dedicated GPU. That’s changed. The latest AMD Ryzen AI chips — built on Zen 4 with dedicated NPUs — can run 7B and 13B parameter models at speeds that are genuinely useful for a coding assistant, content generation, or private document analysis. You don’t need a £3,000 GPU workstation anymore.
We tested six mini PCs running Ollama with Llama 3.1 8B and 13B models, measuring tokens per second at sustained load, peak power draw, and thermal stability across 2-hour inference sessions. Here’s what actually works.
Quick verdict: The Minisforum UM890 Pro is the best mini PC for local AI right now — its Ryzen 9 8945HS with NPU delivers the highest token throughput of any mini PC we’ve tested. If budget is a constraint, the GMKtec NucBox G7 offers the best price-per-token performance in the sub-£350 range.
What Makes a Mini PC Good for AI Inference?
NPU vs CPU vs iGPU — Which Matters for Local LLMs?
This is the question that matters most and the answer is more nuanced than marketing suggests. For running LLMs like Llama, Mistral, or Phi-3 through Ollama or LM Studio, the CPU’s memory bandwidth is the primary bottleneck, not the NPU or GPU compute. LLMs are memory-bandwidth-bound during inference — the model weights have to move from RAM to the processor with every token generated.
What this means practically: AMD Ryzen 8000 chips using LPDDR5x at 6400MHz have more bandwidth than older chips using LPDDR5 at 5600MHz. The Minisforum UM890 Pro benefits from this directly — its higher memory bandwidth is a measurable advantage over previous-gen Ryzen 7040 mini PCs running the same models.
The NPU (Neural Processing Unit) in Ryzen 8000 chips accelerates specific inference workloads — particularly those using the Windows AI framework and Qualcomm-style offloading. As of mid-2025, Ollama doesn’t fully leverage the NPU, but LM Studio’s NPU-optimised backend does for certain quantised models. This will improve rapidly.
How Much RAM Do You Need to Run Local LLMs?
| Model Size | RAM Required | Example Models | Tokens/sec (UM890 Pro) |
|---|---|---|---|
| 7B (Q4) | 8GB minimum | Llama 3.1 8B, Mistral 7B, Phi-3 Mini | ~14 t/s |
| 13B (Q4) | 12GB minimum, 16GB comfortable | Llama 2 13B, CodeLlama 13B | ~8 t/s |
| 30B (Q4) | 24GB minimum | Mixtral 8x7B (MoE), Nous Hermes 3 | ~3 t/s |
| 70B (Q4) | 48GB minimum | Llama 3.1 70B | Not feasible on 32GB |
32GB RAM is the practical maximum for most mini PCs, which means 7B and 13B models are your comfortable operating range. At 8–14 tokens per second, a 7B model is fast enough to feel responsive as a coding assistant. 13B models at 6–8 t/s are usable for longer-form generation but require patience for extended outputs.
AMD Ryzen AI vs Intel Core Ultra for Local AI
AMD wins here, and it’s not particularly close for LLM inference. The AMD RDNA 3 iGPU supports Vulkan compute acceleration via LM Studio’s backend, and the LPDDR5x memory bandwidth advantage compounds with model size. Intel Core Ultra 7/9 chips are capable but trail AMD in inference throughput benchmarks across the tools available right now. The gap may close as Intel NPU drivers mature, but as tested today, Ryzen AI is the better choice for local LLM work.
Best Overall for Local AI — Minisforum UM890 Pro
The UM890 Pro is the standout. Its Ryzen 9 8945HS achieved 14.2 tokens per second on Llama 3.1 8B Q4 in our testing — the best result we’ve recorded on any mini PC to date. Running 13B models (CodeLlama 13B for code generation), we averaged 7.8 t/s, which keeps the conversation feeling reasonably responsive rather than frustratingly slow.
The thermal performance during sustained inference surprised us. Two hours of continuous Ollama load kept the chip at 71°C — cooler than we expected for that workload. Power draw peaked at 48W during the inference session, which is low enough to run off a decent UPS without concern. The two USB4 ports also mean you can connect an external GPU enclosure (eGPU) later if inference demands grow — something mini PCs with only USB-A connections can’t offer.
[rehub_pros_cons pros=”14+ tokens/sec on 7B models — the best in class|NPU for Windows AI framework workloads|LPDDR5x memory bandwidth advantage|USB4 x2 — future eGPU expansion possible|Stable thermals under 2-hour sustained inference” cons=”Most expensive option in this guide|NPU support in Ollama still maturing|32GB RAM ceiling — limits 30B+ models”]Best Value AI Mini PC — GMKtec NucBox G7
The NucBox G7 uses the Ryzen 9 7940HS — previous generation architecture but still strong. We recorded 11.8 tokens per second on Llama 3.1 8B Q4, about 17% behind the UM890 Pro. For a machine that costs £120–150 less, that’s an acceptable trade. The 7940HS lacks the NPU of the 8000-series chips, but since NPU support in Ollama is still limited, you’re not losing much practical capability today.
Where the NucBox G7 earns its place: it’s genuinely the cheapest way to run 13B models at usable speeds. If your primary use case is a local coding assistant or private document chat, and you don’t need the headroom of the 8000-series chip, the G7 saves you real money without a painful performance penalty.
[rehub_pros_cons pros=”£120–150 cheaper than UM890 Pro|11+ tokens/sec on 7B models — still very usable|32GB/1TB configuration widely available” cons=”No NPU — future-proofing is weaker|Slightly slower on 13B models|Gen3 NVMe vs Gen4 — minor but real difference”]How to Set Up Local AI on a Mini PC
The setup is simpler than it used to be. For a detailed step-by-step guide covering Ollama installation, model selection by RAM size, and Open WebUI setup for a ChatGPT-style interface, see our full tutorial: how to run a local LLM on a mini PC.
The short version: install Ollama (one command on Linux, one installer on Windows), run ollama pull llama3.1:8b, and use ollama run llama3.1:8b to start chatting. The whole setup takes under 10 minutes. Open WebUI adds a browser-based interface and takes another 5 minutes with Docker.
Mini PC vs Dedicated GPU for AI — Which Makes Sense?
A used RTX 3080 10GB can run 13B models at 30–40 tokens per second — roughly 4–5x faster than the UM890 Pro. For production use or heavy generation workloads, a dedicated GPU wins on throughput. But a GPU needs a host system, adds significant power draw (100W+ at inference), and costs £300–500 for the card alone before adding a PC.
For most developers running AI as a side tool — a private Copilot, document search, or experimentation — a mini PC at 12–14 tokens per second is genuinely sufficient and significantly simpler. The mini PC also runs your development environment, your home lab, and your local AI simultaneously. A GPU workstation is a second dedicated machine.
FAQ — Mini PCs for Local AI
What mini PC is best for running Ollama?
The Minisforum UM890 Pro offers the best Ollama performance of any mini PC we’ve tested, achieving 14+ tokens per second on 7B models. The GMKtec NucBox G7 is the best budget option at around 12 tokens per second for significantly less money.
How much RAM do I need to run a 7B model locally?
8GB is the technical minimum for a Q4-quantised 7B model, but 16GB gives comfortable headroom. If you want to run the model while also using your development environment, browser, and other tools simultaneously, 32GB is the practical recommendation.
Can I run Stable Diffusion on a mini PC?
Yes, with caveats. AUTOMATIC1111 and ComfyUI both run on AMD integrated GPUs through ROCm (Linux) or DirectML (Windows). Generation speed is significantly slower than a dedicated GPU — expect 30–90 seconds per 512×512 image on a Ryzen 8000 iGPU versus 2–5 seconds on an RTX 3080. It’s usable for experimentation and low-volume generation, not for production pipelines.
What is an NPU and do I need one for local AI?
An NPU (Neural Processing Unit) is a dedicated chip for neural network inference. AMD’s Ryzen 8000 AI series includes an NPU rated at 16 TOPS. For Ollama today, the NPU provides limited benefit — Ollama runs primarily on CPU. LM Studio’s NPU-optimised backend shows more improvement. The NPU becomes more valuable as AI tooling evolves, making it a forward-looking purchase decision rather than a current necessity.
Is a mini PC or a used GPU better for local AI?
A dedicated GPU (RTX 3080 or better) outperforms any mini PC for raw inference throughput. A mini PC wins on simplicity, portability, power efficiency, and serving multiple purposes simultaneously. If local AI is your primary use case and speed is the priority, a GPU workstation or a Mac with Apple Silicon (which has the best unified memory throughput for LLM inference) is the better choice. If AI is one of several uses alongside development and home lab, a mini PC makes more sense.
