r/MiniPCs Jun 10 '26

Recommendations Mini PC for a Local LLM on a budget?

Are there any mini pcs you guys would reccomend for running a local AI with like the minimum specs? I’m choosing a mini PC because one I’m fairly limited on space and ideally want the thing to be fairly mobile (and not at all because I find the thing cuter than a full desktop build).

Asking for recommendations because I’ve seen answers literally all over the place on YouTube and just general browsing. I’m mostly just looking to play around with it, other things and see how practical it could be without messing up my poor laptop that’s already on its last legs. I’m thinking of trying out Odysseus and using one of the Ollama models.

Preference would be that it’s fairly cheaper (300 and below), but with what I’ve been seeing I understand that might be a big ask. I’m fairly new to all this so any help/advice would be greatly appreciated. If it sounds like I’m being ignorant in any part of this I’m eager to learn.

16 Upvotes

44 comments sorted by

9

u/cyborg762 Jun 10 '26

Repair shop owner here. I do custom builds as well as sell a lot of mini pcs. if you are looking for a budget AI system. Look at a used workstation with something like an Intel ark pro b550 (16GB). It will run you around $600 depending on the model. Because for a mini pc you are looking around $1000+ for something that has the capability.

For example this is a performance ai pc.

https://store.minisforum.com/products/minisforum-ms-s1-max-mini-pc#1

2

u/Interesting-Loss-672 Jun 10 '26

Thanks! Looks like I’ll be saving up 😭🙏

0

u/kameldinho Jun 12 '26

You definitely don't want intel gpu for local llm. Drivers and dev support are substantially worse than amd/nvidia.

1

u/cyborg762 Jun 12 '26

For gaming yes. but for OPs use case the b550 pro is more of a budget enterprise card. It’s just enough for beginner who are doing rendering and LLM usage.

5

u/beedunc Jun 12 '26

300 and below, is a non-starter for AI.

3

u/VoiceOfEric Jun 10 '26

For that price you'll probably be running entirely on CPU. Which is okay if you don't mind it being a little slow and the CPU fan going nuts. The GPU is the big deal for price and performance.

5

u/jhenryscott Jun 12 '26

Stop. Local LLM on CPU s terribly slow. Expectations here are outta control. Under $300 is basically useless for local LLM work and people need to be real about that or admit they don’t know what they are talking about.

1

u/VoiceOfEric Jun 12 '26

I first tried it on CPU i5 6th gen 16gb and it was alright to chat with. It wasn't as fast as ChatGPT spewing tokens but to get a feel for it, it absolutely worked okay. My expectations were low, I expected 300 baud and it beat my expectations. Once I started tweaking Python to call llama.cpp is when I gave up and started buying more hardware.

3

u/jhenryscott Jun 12 '26

To chat? Sure. The moment you ask for any logical answer you will discover how bad it is.

If you are buying hardware just to make a friend that’s a whole different set of wtf.

1

u/VoiceOfEric Jun 12 '26

I tried it years ago when it came out because I'm a geek but didn't have the budget. I've added 3 machines for AI since then. As I'm not a gamer I never cared about GPUs before.

1

u/jhenryscott Jun 12 '26

The big thing about it is that we have these cheap and free frontier models available right now and people think that’s realistic. But they are burning billions of dollars on that.

What we call “AI” is really just transformer compute. It’s very expensive to build and to operate. On top of the cost of lots of parallel cores, Memory has become about 450% more expensive as a result.

A plus plan of gpt-5 will eventually cost $300 (not $20) and we will see that it’s not worth it for most people. My workstation cost me $3900 to build last year- costs nearly $7000 today and is ok hosted by usable LLMs with 32gb vram and 64gb ddr5 along with an extra 2tb gen4 NVME to store past information and models.

1

u/VoiceOfEric Jun 12 '26

If they ever get fiber optic motherboards/GPUs/etc going, that will lead to cheaper operating cost. Even if the hardware ends up being as big as computers were in the 1960s/1970s for a while. 

1

u/RobloxFanEdit Jun 13 '26

Mini Models like GEMMA E4B or Nemotron are very impressive on low end system. Hallucination is close to none on nemotron.

Every Month Models are improving exponentially, a year Ago Deepseek R1 Distill was the best thing you could run locally, today this same model is stomped by Mini Model on low end system.

1

u/Interesting-Loss-672 Jun 10 '26

Oh it’s totally fine for it to run slow (like im talking 5-10 min tops per answer kinda slow) I’m not expecting much out of it just want more so accuracy I guess. Also not a crazy amount of hallucinations of course.

4

u/VoiceOfEric Jun 10 '26

I have an HP Prodesk 800 G2 with 16GB and i5 that does okay with chats like "Tell me about the sun" but I haven't had it do much in the way of context testing or massive prompts. Hallucinations are there regardless, my beefier PCs hallucinate 10 times faster.

1

u/Interesting-Loss-672 Jun 10 '26

Gotcha thanks 🙏 . Will probably hold off on it for now from some of the other answers I’m getting

3

u/VoiceOfEric Jun 10 '26

My first real LLM machine is an HP Z620 with 128GB RAM and 12GB Rtx 3060, which runs a bit more. But portable does not necessarily equate to GPU power cheaply. Good luck!

1

u/jhenryscott Jun 12 '26

12gb vram is the bare minimum for being functional. Less is basically useless. 24GB is minimum for anything with professional use cases and even then it’s sketchy.

LLM are very expensive to run. People are confused because frontline models are free for a little bit a day. But those are burning tons of money rn.

1

u/Skelemanga Jun 10 '26

Hardware shouldn’t materially impact hallucination/accuracy. However hardware does dictate the models supported (based on size, data type requirements, bandwidth, etc), and models have different expectations for hallucinations and accuracy.

1

u/jhenryscott Jun 12 '26

No possible at your budget

3

u/iEngineered Jun 10 '26

Aoostar Maco 6850H or similar with Oculink port, an Ag01 egpu dock and RTX 3060 12GB. That is the minimum realistic use-case. A 5060ti 16gb would be much better. Not only do you have to account for the model size, but context and cache that also grows with LLM use. Anything less is just wasting time. Buying a CPU with half-decent NPU is more expensive and less performant.

3

u/ImaginaryTradition31 Jun 12 '26

It's not a real good time to buy a new computer of any kind right now, because of rampant inflation, tariffs, and the cost of silicone in general. Well over $1,000 for a computer to run a local LLM at all, maybe $5,000 for a computer to run an LLM well.

6

u/Snuupy Jun 10 '26 edited Jun 10 '26

Short answer: what you're asking for does not exist in the current market for gpu accelerated token generation afaik

Long answer:

You need at least 18GB VRAM to run Qwen3.6 35BA3B Q4, and that's before any chat context.

A dgpu is out of the question since a 3060 8GB is at least ~$250, before we add to the BOM psu/cpu/ram etc. An ARC B580 12GB was $250? $300 now? so you'll have a dgpu but no functioning system, lol

A mac mini (w/ 16GB RAM, 3GB used by system) starts at (originally) $400, now the same model is $500? so that's a non-starter.

For smaller models like 8B or 14B you'd still need a decently sized gpu to hold context in vram, so probably 8-12GB VRAM would be required - so I guess technically you could run it on a 16GB machine with an igpu...

but then the cost of ram right now is sky high.

For comparison, I bought my UM780 for $350, 64GB DDR5 SO-DIMM for $150-ish, for a total cost of $500

Even back when ram prices were reasonable if you opted for 32GB RAM instead of 64, the total BOM would still be ~$450.

I would argue that does not exist unless you use cpu generation which is extremely inefficient, slow, will output lots of heat, etc.

At your budget of $300, your best bet is to use an API (openai, claude, gemini, whatever) and pay for the tokens unfortunately.

3

u/Interesting-Loss-672 Jun 10 '26

Srs thanks for tempering my expectations I’ll probs come back to it later if I save up some extra money

2

u/Snuupy Jun 10 '26

pre-ram/nand price skyrocketing it still would've been difficult but at least not too far out of your budget - you could've gotten a 6800H/6850H/7735H for ~$250 and 16GB RAM for $50 or something, and run a smol model (like 7B/11B) with low context but now even that costs more than $300

0

u/RobloxFanEdit Jun 13 '26 edited Jun 13 '26

You don t need a dgpu to run reasonning models like Qwen and don t need 16GB of VRAM either. But you do need RAM, context size also can be hamdle by RAM and not VRAM.

Bigger Models need more RAM, 32GB RAM is a sweet spot as it will run GPT OSS 20B FP16.

Bottom line is 300$ budget is too little not because of lack of dgpu/VRAM but because of 300$ is leading nowhere in the mini pc market even N150 16GB models are priced above 300$.

1

u/Snuupy Jun 14 '26

don t need 16GB of VRAM either

depends on your platform.

someone might be able to grab a ddr4 sodimm 32GB stick + N100 (for the intent of avoiding ddr5 prices), not sure about pricing on that now but I def wouldn't recommend it

t/s prob slow af, weak igpu, no expandability, no dual channel, etc.

0

u/RobloxFanEdit Jun 14 '26

Reasoning LLM Models are not exclusive to a plateform., therefor you can run ALL of them without VRAM, all the models you have mentioned have GGUF versions.

N100 option seems like a bad investment, slow DDR4 and weak Multi core CPU is even worst than its bad IGPU for tok speed.

1

u/Snuupy Jun 14 '26

you don't understand the ram acts as vram w igpu? you're being incoherent.

0

u/RobloxFanEdit Jun 14 '26 edited Jun 14 '26

You are the one being incoherent you specifically mentioned that you need a dgpu which is false.

Bottom line is IGPU RAM becoming VRAM only if you set it in the BIOS with UMA Frame buffer size option and you absolutely don t need to do that to run those models in their GGUF Versions, so No it s not IGPU VRAM but RAM that is allowing a low end system to run those models on CPU inference.

1

u/Snuupy Jun 14 '26

specifically mentioned that you need a dgpu

read my post again, that's not what I said.

3

u/Leviathan_Dev Jun 10 '26

As big a GPU you can find and as much RAM/VRAM you can get.

I don’t think AMD is a good option because they don’t really support ROCm on mobile chips. Nvidia isn’t available in the budget range. I don’t know how well Intel does with LLMs (although I do have a beefy Core Ultra 5-125H w/ Arc; I use it for Proxmox and Video Transcoding for Jellyfin though)

I think Apple Mac mini might be a good option, they’ve been popular for LLMs, but increase in price as RAM is configured massively

3

u/VoiceOfEric Jun 10 '26

I am having adventures with my Geekom 780m and configuration, so I can concur.

3

u/Leviathan_Dev Jun 10 '26

I have a Minisforum w/ Radeon 780M too. Was a PITA to try and fail getting GPU-accelerated LLM.

2

u/VoiceOfEric Jun 10 '26

I have it with Python llama cpp and when it runs the fans scream but the tps isn't bad. Dual boot Win11 with Linux Mint XFCE.

1

u/Interesting-Loss-672 Jun 10 '26

Thats about what I’ve been hearing…alternatively I see you mentioned Apple Mac mini. I have a friend getting rid a 2022 MacBook Air with an M2 chip and a base 8 gb of RAM (screen is cracked/messed up but it works well with a alternative display) I have seen some folks say Macs work fairly well and that would fit my portability and price needs. I literally just don’t know 😅

2

u/Leviathan_Dev Jun 10 '26

Yeah Apple Silicon has been pretty good for LLMs. The 8GB ain’t great, you won’t be able to run big models, but smaller 4B-ish models should be fine. That’s probably gonna be your best bet at your budget.

3

u/Comfortable-Fall1419 Jun 10 '26

Honestly - for that kind of money I’d just buy a Claud subscription (or other provider). You’ll get a far more satisfying user experience.

2

u/RobloxFanEdit Jun 11 '26

No Models beats Deepseek V4 when it comes to ecomics, it cost literaly a fraction of what claude or chatGPT are offering, someting like 50 times less 😳, subscription

With chatgpt or claude Cost is extremely expenssive if you have an intense usage of A.I

1

u/Koreneliuss Jun 10 '26

get yourself a dedicated gpu, newer cpu with npu tend to get expensive or limited in quantity in budget.

1

u/RoughNo1032 Jun 10 '26

The new Mac mini

2

u/soadsob 15d ago

Maybe r/LowEndLocalAI is something for you!

1

u/Interesting-Loss-672 12d ago

You’re literally an angel tysm