The one thing that decides speed: memory bandwidth
Large language models are mostly a memory problem. To produce each token, the machine has to read essentially every active weight in the model from RAM. A 5 GB model on a system that can move 40 GB per second has a hard ceiling no matter how many cores you add. That is why a mini PC with dual-channel DDR5-5600 (about 89.6 GB/s theoretical) answers much faster than an older tiny with dual-channel DDR4-2666 (about 42.7 GB/s), and why a machine running a single stick, or an N100 with its single memory channel, is slower still. Always buy or install RAM as a matched pair.
Prompt processing, meaning reading your question and documents before the answer starts, is the part that benefits from compute: more cores, or a GPU. Token generation, the part you watch scroll, is the part bandwidth limits.
What fits in 16, 32 and 64 GB
Ollama's README no longer publishes a RAM-per-model rule of thumb, so use the file sizes from its model library and leave headroom for the OS, the context and anything else the box runs. Ollama's default context window is 4,096 tokens; longer contexts and parallel requests need more memory on top of the model.
- 16 GB: 3B to 8B models comfortably. Llama 3.1 8B is 4.9 GB, Qwen3 8B is 5.2 GB and Gemma 3 4B is 3.3 GB. Gemma 3 12B (8.1 GB) fits but leaves little room for anything else.
- 32 GB: adds 14B models (Qwen3 14B, 9.3 GB) and the 27B to 32B class: Gemma 3 27B at 17 GB, Qwen3 32B at 20 GB, and Qwen3 30B at 19 GB.
- 64 GB: Llama 3.1 70B at q4_K_M (43 GB) loads, but on a mini PC it is a batch-job model, very slow and best left to run while you do something else.
One thing changes the math in a mini PC's favor. Qwen3 30B is a mixture-of-experts model: Qwen describes it as 30B-A3B, meaning only about 3B parameters are active for each token. You need RAM for the full 19 GB, but each token only reads a fraction of it, so MoE models are the most practical large models for bandwidth-limited machines.
Measured speeds on mini PC hardware
These are the only mini-PC-class numbers on this page, and each comes from a published measurement:
- Radeon 780M through llama.cpp's Vulkan backend: on a Ryzen 9 7940HS, 19.91 tokens/sec generation and 281.62 tokens/sec prompt processing on a 7B Q4_0 model (3.56 GiB). A Ryzen 7 8840HS posted 20.10 tokens/sec generation on the same test, and a Ryzen 7 7840U laptop with two DDR5-5600 sticks posted 18.22.
- Intel N150 CPU-only through Ollama: Jeff Geerling's benchmarks measured a 16 GB GMKtec G3 Plus at 9.06 tokens/sec on Llama 3.2 3B and 2.13 tokens/sec on DeepSeek R1 14B.
A 780M machine running a 7B to 8B model feels conversational. An N-series box is pleasant with 3B models and painfully slow above that. For 8th to 10th-gen Intel tinies and DDR4 Ryzen machines we don't have a published measurement for this exact class, so the honest summary is slow but usable for 7B to 8B chat and for overnight batch jobs like summarizing documents, and too slow for interactive use above about 14B.
Software: Ollama, llama.cpp, LM Studio, Open WebUI
- Ollama is the easiest server: one install, `ollama run <model>`, and an API other apps can call. By default it keeps a model in memory for 5 minutes after the last request. On AMD, its docs say Vulkan is enabled by default when the backend is installed; ROCm support is limited to listed discrete Radeon cards, which don't include the 780M.
- llama.cpp is the engine under many tools and the source of the Vulkan numbers above. Use it directly when you want to tune threads, context and GPU offload yourself.
- LM Studio is a desktop app. On x64 it requires a CPU with AVX2 and recommends at least 16 GB of RAM; on Linux it needs Ubuntu 20.04 or newer.
- Open WebUI gives you a ChatGPT-style web interface for Ollama and OpenAI-compatible APIs, runs entirely offline, and installs as one Docker container.
When a mini PC is the wrong tool
If you want 30B to 70B models at conversational speed, a used mini PC won't get you there. Two better routes: a desktop with a discrete GPU, where the model lives in fast VRAM, or an Apple Silicon Mac, whose unified memory is shared by CPU and GPU at much higher bandwidth. Apple lists the current Mac mini at 153 to 170 GB/s on the base chip and 307 GB/s on the M5 Pro, roughly 1.7 to 3.4 times the theoretical 89.6 GB/s of dual-channel DDR5-5600. A mini PC is the right tool when you want a cheap, quiet, always-on box for small models, private document summaries, or an endpoint for Home Assistant voice or n8n workflows.
If you mainly want to run cloud-model agents like Claude Code around the clock, you don't need local inference at all; see the AI agent server page. Comparing brands for a Ryzen mini? See Beelink vs GMKtec vs Minisforum.





















