this post was submitted on 07 May 2026
643 points (99.5% liked)

Technology

84449 readers
4052 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 2 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] MagicShel@lemmy.zip 12 points 1 day ago* (last edited 1 day ago) (3 children)

It's a MacBook Pro. 36GB of ram. I am sure Macs have some kind of gpu and I understand it somehow combines GPU ram with system ram, but I don't really know Mac hardware very well.

It's beefy for a laptop, but the desktop I built for myself several years ago had 32 GB of ram and a GTX 1660, so I'm guessing they are similar in capability. I gave that to my daughter, so I can't run a comparison right now.

EDIT: After doing just a bit of research, I've learned the unified memory architecture that Macs use, while not ideal for many purposes, is actually a big advantage for running larger inference models. So it's possible that this particular model wouldn't run at all on my Linux box or would run much slower because the full model wouldn't fit in the 6GB of VRAM and create a lot of memory thrashing.

[–] boonhet@sopuli.xyz 3 points 1 day ago

Yup, you want memory accessible to the GPU for local AI. AMD Strix Point and Mac devices are popular options. CPU can run LLMs but very slowly. I've got 32 GB of RAM and 8 VRAM and it's borderline useless for models that don't fit in the VRAM.

[–] SabinStargem@lemmy.today 4 points 1 day ago (1 children)

You can use something like KoboldCPP on Linux, which allows both RAM and VRAM combined to run a model. O'course, not as fast when compared to pure VRAM or the Mac approach, but it is an option. I use my 128gb RAM with some GPUs for running models.

[–] boonhet@sopuli.xyz 1 points 1 day ago (1 children)

Ollama and llama.cpp allow it too but it's super slow in my experience.

[–] SabinStargem@lemmy.today 1 points 21 hours ago (1 children)

Speed depends on how much of the model is on VRAM, and the dense/MoE architecture of that model. The RAM's benefit is more about having the ability to run the model in the first place. In any case, a dense Qwen3.6 27b would take up about 27-33gb-ish of memory, plus whatever context size you set.

Upcoming implementation of MTP will increase the size of models, but in exchange, they will also run faster. About a 30%ish boost for dense models, a bit less for Mixture of Expert varieties, from the looks of it.

[–] boonhet@sopuli.xyz 1 points 21 hours ago (1 children)

When I've tried running a ~14 gigabyte distillation of whatever model it is I was trying to run, it would come out super slow at I believe 50/50 GPU to CPU. It gets so slow it was just more bearable to run a 7 or 8 b model that would actually fit entirely in VRAM and run entirely on GPU. Also made the rest of computer usage more bearable.

To be fair I do only have a 6 core 6 thread CPU though. It shot up to 600% usage so even the DDR4 memory wasn't really bottlenecking it. I suspect a 9950X would fare a lot better.

[–] SabinStargem@lemmy.today 1 points 20 hours ago

I am using a 5950x, with 128gb of DDR4 3600 memory. The GPUs are a 3060 and 4090, totaling 36gb of VRAM. IMO, being bottlenecked by the CPU is definitely a thing, it just comes third after the VRAM and RAM considerations.

With a 35b+3a MoE at Q8 with KV8, I get...

[11:54:32] CtxLimit:18858/262144, Init:0.18s, Processed:17294 in 7.66s (2259.18T/s), Generated:1564/32768 in 29.01s (53.91T/s), Total:36.85s

[–] humanspiral@lemmy.ca 1 points 1 day ago

decent performance on 6gb gpu without quantization: https://www.youtube.com/watch?v=8F_5pdcD3HY&t=9s