Skip to main content
used mini pcs

Local AI, honestly

Running local LLMs on a used mini PC

A mini PC can run local models with Ollama, llama.cpp or LM Studio, and the right one does it surprisingly well. Two numbers decide the experience: how much RAM you have (which models fit) and how fast that memory is (how quickly they answer). This page gives measured speeds where they exist and plain warnings where they don't.

What you need

SpecRecommendedWhy
CPU generationRyzen 7040/8040 (Phoenix/Hawk Point, Radeon 780M) for the best results; any AVX2-capable Intel 6th gen+ or Ryzen for small modelsPhoenix chips pair DDR5 with the Radeon 780M iGPU, which llama.cpp can use through Vulkan. Older chips run CPU-only. LM Studio requires AVX2 on x64, so check the exact CPU's spec page, especially on Pentium and Celeron units.
RAM32 GB minimum for useful models, 64 GB for 30B-class; always two matched sticksThe whole model file has to sit in memory, plus room for the context and the OS. Ollama's library lists Llama 3.1 8B at 4.9 GB, Qwen3 14B at 9.3 GB, Gemma 3 27B at 17 GB and Llama 3.1 70B (q4_K_M) at 43 GB.
Memory bandwidthDual-channel DDR5-5600 (about 89.6 GB/s theoretical) beats dual-channel DDR4-2666 (about 42.7 GB/s)Each generated token streams the whole active model through memory, so token speed is capped by memory bandwidth, not core count. Theoretical bandwidth is transfer rate × 8 bytes × channels. A single-channel machine halves it.
Storage typeNVMe SSD, 1 TB if you keep several modelsModel files run from about 5 GB to over 40 GB each. NVMe makes loading a model into RAM quick, which matters because Ollama unloads idle models after 5 minutes by default.
NICsOne 1 GbE port is fine; 2.5 GbE is a bonusModels download once. After that, chat traffic from Open WebUI or an API client is tiny.
Idle wattsAbout 6 to 10 W for a Ryzen 7840HS mini at idleServeTheHome measured 6 to 10 W idle on a Beelink SER7 (Ryzen 7 7840HS), with a peak around 77 to 79 W under load that settled to 44 to 46 W sustained. An LLM box idles most of the day, so idle draw dominates the yearly cost.
NoiseAudible under long generations; quiet at idleInference keeps the CPU or iGPU pinned for the length of every answer. ServeTheHome measured the SER7 at about 35 dBA under load in a 34 dBA room, but small chassis vary, so expect the fan to be noticeable during long jobs.
iGPURadeon 780M via Vulkan is worth having; older Intel UHD iGPUs mean CPU-onlyllama.cpp's Vulkan backend has been measured on 780M machines (see below). Ollama documents Vulkan as enabled by default when its backend is installed. The 780M is not on Ollama's ROCm-supported Radeon list, so don't plan on ROCm.

Which mini PCs fit

A model gets the “Local LLM” badge when it:

  • Intel 8th gen or newer, Intel N-series, or AMD Ryzen 2000 or newer (not Ryzen Embedded)
  • Takes at least 32 GB of RAM (manufacturer's stated maximum)
  • Has two or more memory slots, so RAM can run dual-channel (memory bandwidth sets LLM speed)
  • Has an M.2 slot that takes an NVMe SSD

Rules use manufacturer specs only. 33 of 58 catalog models qualify.

Best overall

Minisforum UM790 Pro (Ryzen 9 7940HS)

The UM790 Pro uses the Ryzen 9 7940HS with Radeon 780M, the exact chip in a llama.cpp Vulkan result of 19.91 tokens/sec on a 7B Q4_0 model. Minisforum lists dual-channel DDR5 SO-DIMMs up to 5600 MHz, 64 GB max, two M.2 2280 PCIe 4.0 slots and 2.5G Ethernet. That is enough for 8B to 14B models at conversational speed and 30B-class models when you can wait.

On eBay now: The lowest ask at $325 delivered with Best Offer from a 100% seller, but it's a barebones unit with no RAM or SSD per the title. Budget for two DDR5-5600 SO-DIMMs and an NVMe SSD, and confirm the 120 W adapter is included. $325.00

View on eBay ↗: MinisForum UM790 PRO Mini PC w/ AMD Ryzen 9 7940HS Barebone SFF PC NO

Best value

Lenovo ThinkCentre M75q Gen 2 Tiny

A common off-lease 1-liter Lenovo that its PSREF rates for 64 GB of DDR4-3200 across two slots, with 6- and 8-core Ryzen PRO GE chips. It is CPU-only and slower than DDR5 machines, but 64 GB fits 30B-class models such as Qwen3 30B (19 GB) that 16 GB boxes can't hold at all.

On eBay now: The lowest ask at about $140 delivered from a 99.8% seller, for a quad-core Ryzen 3 PRO 4350GE with 8 GB. The title cuts off before storage, so ask whether a drive and adapter are included before you count it as complete. $139.99

View on eBay ↗: Lenovo ThinkCentre M75q Gen 2 Tiny Ryzen 3 Pro 4350GE 3.50 GHz 8 GB DD

What it costs to leave on all year

An always-on box spends almost all its time near idle, so idle draw sets the bill. The math:

idle watts × 24 h × 365 days ÷ 1,000 = kWh per year · kWh × $0.1831/kWh = dollars per year

Electricity at 18.31¢/kWh, the U.S. average residential price, July 2026 (EIA). Your rate is on your bill; swap it into the formula. Real draw rises above idle whenever the box is working, and depends on drives, RAM, BIOS power settings and the OS.

The one thing that decides speed: memory bandwidth

Large language models are mostly a memory problem. To produce each token, the machine has to read essentially every active weight in the model from RAM. A 5 GB model on a system that can move 40 GB per second has a hard ceiling no matter how many cores you add. That is why a mini PC with dual-channel DDR5-5600 (about 89.6 GB/s theoretical) answers much faster than an older tiny with dual-channel DDR4-2666 (about 42.7 GB/s), and why a machine running a single stick, or an N100 with its single memory channel, is slower still. Always buy or install RAM as a matched pair.

Prompt processing, meaning reading your question and documents before the answer starts, is the part that benefits from compute: more cores, or a GPU. Token generation, the part you watch scroll, is the part bandwidth limits.

What fits in 16, 32 and 64 GB

Ollama's README no longer publishes a RAM-per-model rule of thumb, so use the file sizes from its model library and leave headroom for the OS, the context and anything else the box runs. Ollama's default context window is 4,096 tokens; longer contexts and parallel requests need more memory on top of the model.

  • 16 GB: 3B to 8B models comfortably. Llama 3.1 8B is 4.9 GB, Qwen3 8B is 5.2 GB and Gemma 3 4B is 3.3 GB. Gemma 3 12B (8.1 GB) fits but leaves little room for anything else.
  • 32 GB: adds 14B models (Qwen3 14B, 9.3 GB) and the 27B to 32B class: Gemma 3 27B at 17 GB, Qwen3 32B at 20 GB, and Qwen3 30B at 19 GB.
  • 64 GB: Llama 3.1 70B at q4_K_M (43 GB) loads, but on a mini PC it is a batch-job model, very slow and best left to run while you do something else.

One thing changes the math in a mini PC's favor. Qwen3 30B is a mixture-of-experts model: Qwen describes it as 30B-A3B, meaning only about 3B parameters are active for each token. You need RAM for the full 19 GB, but each token only reads a fraction of it, so MoE models are the most practical large models for bandwidth-limited machines.

Measured speeds on mini PC hardware

These are the only mini-PC-class numbers on this page, and each comes from a published measurement:

  • Radeon 780M through llama.cpp's Vulkan backend: on a Ryzen 9 7940HS, 19.91 tokens/sec generation and 281.62 tokens/sec prompt processing on a 7B Q4_0 model (3.56 GiB). A Ryzen 7 8840HS posted 20.10 tokens/sec generation on the same test, and a Ryzen 7 7840U laptop with two DDR5-5600 sticks posted 18.22.
  • Intel N150 CPU-only through Ollama: Jeff Geerling's benchmarks measured a 16 GB GMKtec G3 Plus at 9.06 tokens/sec on Llama 3.2 3B and 2.13 tokens/sec on DeepSeek R1 14B.

A 780M machine running a 7B to 8B model feels conversational. An N-series box is pleasant with 3B models and painfully slow above that. For 8th to 10th-gen Intel tinies and DDR4 Ryzen machines we don't have a published measurement for this exact class, so the honest summary is slow but usable for 7B to 8B chat and for overnight batch jobs like summarizing documents, and too slow for interactive use above about 14B.

Software: Ollama, llama.cpp, LM Studio, Open WebUI

  • Ollama is the easiest server: one install, `ollama run <model>`, and an API other apps can call. By default it keeps a model in memory for 5 minutes after the last request. On AMD, its docs say Vulkan is enabled by default when the backend is installed; ROCm support is limited to listed discrete Radeon cards, which don't include the 780M.
  • llama.cpp is the engine under many tools and the source of the Vulkan numbers above. Use it directly when you want to tune threads, context and GPU offload yourself.
  • LM Studio is a desktop app. On x64 it requires a CPU with AVX2 and recommends at least 16 GB of RAM; on Linux it needs Ubuntu 20.04 or newer.
  • Open WebUI gives you a ChatGPT-style web interface for Ollama and OpenAI-compatible APIs, runs entirely offline, and installs as one Docker container.

When a mini PC is the wrong tool

If you want 30B to 70B models at conversational speed, a used mini PC won't get you there. Two better routes: a desktop with a discrete GPU, where the model lives in fast VRAM, or an Apple Silicon Mac, whose unified memory is shared by CPU and GPU at much higher bandwidth. Apple lists the current Mac mini at 153 to 170 GB/s on the base chip and 307 GB/s on the M5 Pro, roughly 1.7 to 3.4 times the theoretical 89.6 GB/s of dual-channel DDR5-5600. A mini PC is the right tool when you want a cheap, quiet, always-on box for small models, private document summaries, or an endpoint for Home Assistant voice or n8n workflows.

If you mainly want to run cloud-model agents like Claude Code around the clock, you don't need local inference at all; see the AI agent server page. Comparing brands for a Ryzen mini? See Beelink vs GMKtec vs Minisforum.

Setup outline

  1. 1Before you pay, ask the seller whether a BIOS supervisor (administrator) password is set. With one in place you cannot change boot order, UEFI or Secure Boot settings, so get it removed or skip the unit.
  2. 2On first power-up, open BIOS setup and look on the security page for a Computrace or Absolute setting. Absolute's persistence module lives in the firmware of machines from many PC makers and can reinstall its agent after a wipe, so if it shows as activated, ask the seller for proof the previous owner released the device.
  3. 3Update the BIOS to the latest release from the maker's support site (look it up by serial number). Old firmware is the usual cause of odd memory and NVMe behavior on off-lease units.
  4. 4Install RAM as two matched SO-DIMMs, never a single stick: a single stick halves memory bandwidth, which directly halves token speed. On DDR5 Ryzen boxes, confirm the speed in BIOS or with `dmidecode` after first boot.
  5. 5Run a full pass of MemTest86+. Large models fill RAM end to end, so a marginal stick that never bothered a desktop workload will crash inference.
  6. 6Install Ubuntu 24.04 LTS (or Windows 11 if you want LM Studio as a desktop app), then install Ollama and pull a small model first, such as an 8B, to confirm everything works.
  7. 7On a Radeon 780M machine, run `ollama ps` while a model is loaded: the processor column shows whether the model landed on the GPU, the CPU, or split across both.
  8. 8Add Open WebUI in Docker for a browser chat interface, and keep it on your LAN or behind Tailscale rather than exposing it to the internet.

Local LLM: common questions

How much RAM do I need for Ollama on a mini PC?

16 GB runs 3B to 8B models (Llama 3.1 8B is 4.9 GB in Ollama's library). 32 GB adds 14B models and the 27B to 32B class (Gemma 3 27B is 17 GB, Qwen3 32B is 20 GB). 64 GB is needed for anything larger, such as Llama 3.1 70B at 43 GB, which runs very slowly on a mini PC.

How fast is a Radeon 780M mini PC for local LLMs?

Published llama.cpp Vulkan results on 780M machines show about 18 to 20 tokens/sec of generation on a 7B Q4_0 model: 19.91 on a Ryzen 9 7940HS and 20.10 on a Ryzen 7 8840HS. That is comfortably conversational for 7B to 8B models.

Can an Intel N100 or N150 run local AI?

Small models only. Jeff Geerling measured a 16 GB N150 mini PC at 9.06 tokens/sec on Llama 3.2 3B and 2.13 tokens/sec on DeepSeek R1 14B. Most N-series machines top out at 16 GB on a single memory channel.

Does Ollama support the Radeon 780M?

Not through ROCm; Ollama's supported Radeon list covers discrete RX cards. Ollama's docs say its Vulkan backend is enabled by default when installed, and llama.cpp's Vulkan backend has published 780M results. Run `ollama ps` to see whether a model loaded on the GPU or the CPU.

Should I buy a mini PC or a Mac mini for local LLMs?

For 30B-plus models at chat speed, a Mac with more unified memory is the stronger tool: Apple lists 153 to 307 GB/s of memory bandwidth on the current Mac mini, versus about 90 GB/s theoretical for dual-channel DDR5-5600. A used x86 mini PC wins for a quiet, always-on Linux box running small models next to other services.

Is DDR4 fine for local LLMs?

It works, at roughly half the bandwidth of dual-channel DDR5-5600, so expect noticeably slower generation. A 64 GB DDR4 box still has one real advantage over a 16 GB DDR5 one: it can load 30B-class models at all.

Sources

Run more on the same box

Upgrading? Sell the old one to us.

Tell us what you have (model, CPU, RAM, drive, whether the BIOS is unlocked), add a few photos and your price, and we’ll reply with an offer by email. No listing, no fees.

Get an offer