It's a MoE model and the A3B stands for 3 Billion active parameters, like the re...

abhikul0 · 2026-04-16T14:23:59 1776349439

Mac has unified memory, so 36GB is 36GB for everything- gpu,cpu.

zozbot234 · 2026-04-16T14:39:37 1776350377

CPU-MoE still helps with mmap. Should not overly hurt token-gen speed on the Mac since the CPU has access to most (though not all) of the unified memory bandwidth, which is the bottleneck.

abhikul0 · 2026-04-16T15:09:32 1776352172

I'll try to use that, but llama-server has mmap on by default and the model still takes up the size of the model in RAM, not sure what's going on.

zozbot234 · 2026-04-16T15:14:31 1776352471

Try running CPU-only inference to troubleshoot that. GPU layers will likely just ignore mmap.

mhitza · 2026-04-16T14:32:02 1776349922

For sure I was running on autopilot with that reply. Though in Q4 I would expect it to fit, as 24B-A4B Gemma model without CPU offloading got up to 18GB of VRAM usage

dgb23 · 2026-04-16T14:17:37 1776349057

Do I expect the same memory footprint from an N active parameters as from simply N total parameters?

daemonologist · 2026-04-16T14:35:53 1776350153

No - this model has the weights memory footprint of a 35B model (you do save a little bit on the KV cache, which will be smaller than the total size suggests). The lower number of active parameters gives you faster inference, including lower memory bandwidth utilization, which makes it viable to offload the weights for the experts onto slower memory. On a Mac, with unified memory, this doesn't really help you. (Unless you want to offload to nonvolatile storage, but it would still be painfully slow.)

All that said you could probably squeeze it onto a 36GB Mac. A lot of people run this size model on 24GB GPUs, at 4-5 bits per weight quantization and maybe with reduced context size.

pdyc · 2026-04-16T14:18:43 1776349123

i dont get it, mac has unified memory how would offloading experts to cpu help?

bee_rider · 2026-04-16T14:22:48 1776349368

I bet the poster just didn’t remember that important detail about Macs, it is kind of unusual from a normal computer point of view.

I wonder though, do Macs have swap, coupled unused experts be offloaded to swap?

abhikul0 · 2026-04-16T14:39:38 1776350378

Of course the swap is there for fallback but I hate using it lol as I don't want to degrade SSD longevity.