Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Lossless model compression experiment: GLM-5.2 in 25% less memory (brianbell-x.github.io)
17 points by hambandit 27 days ago | hide | past | favorite | 1 comment


So, if I'm understanding correctly, this is just for lower bandwidth transfers of full BF16 weights over the wire, not for serving, correct? Have you benchmarked the performance against a SOTA general purpose compression algorithm like zstd?

Also, are all that many people handling the BF16 weights directly? GLM-5.2's reference deployment is FP8, and many vendors are even serving at NVFP4 which seems to offer negligible degradation over the FP8 reference deployments.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: