Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm disappointed that the smallest model size is 21B parameters, which strongly restricts how it can be run on personal hardware. Most competitors have released a 3B/7B model for that purpose.

For self-hosting, it's smart that they targeted a 16GB VRAM config for it since that's the size of the most cost-effective server GPUs, but I suspect "native MXFP4 quantization" has quality caveats.



Native FP4 quantization means it requires half as many bytes as parameters, and will have next to zero quality loss (on the order of 0.1%) compared to using twice the VRAM and exponentially more expensive hardware. FP3 and below gets messier.


A small part of me is considering going from a 4070 to a 16GB 5060 Ti just to avoid having to futz with offloading

I'd go for an ..80 card but I can't find any that fit in a mini-ITX case :(


I wouldn’t stop at 16GB right now.

24 is the lowest I would go. Buy a used 3090. Picked one up for $700 a few months back, but I think they were on the rise then.

The 3000 series can’t do FP8fast, but meh. It’s the OOM that’s tough, not the speed so much.


Are there any 24GB cards/3090s which fit in ~300mm without an angle grinder?



Oh nice, thank you :)

Admittedly a little tempting to see how the 5070 Ti Super shakes out!


I'm waiting too :)

50xx series supports MXFP4 format, but I'm not sure about 3090.


if you're going to get that kind of hardware, you need a larger case. IMHO this is not an unreasonable thing if you are doing heavy computing


Noted for my next build - I am aware this is a problem I've made for myself, otherwise I like the mini-ITX form factor a lot


Which do you like more OOM for local AI, or an itty bit case?


with quantization, 20B fits effortlessly in 24GB

with quantization + CPU offloading, non-thinking models run kind of fine (at about 2-5 tokens per second) even with 8 GB of VRAM

sure, it would be great if we could have models in all sizes imaginable (7/13/24/32/70/100+/1000+), but 20B and 120B are great.


I am not at all disappointed. I'm glad they decided to go for somewhat large but reasonable to run models on everything but phones.

Quite excited to give this a try


Eh 20B is pretty managable, 32GB of regular RAM and some VRAM will run you a 30B with partial offloading. After that it gets tricky.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: