6 ms·
> A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision It actually runs fine at FP8 on this hardware too, with the full 1M context.
by nojs 15d ago
> A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision
It actually runs fine at FP8 on this hardware too, with the full 1M context.
- CamperBob2 15d agoFlash will run on 4x, but at the time I ran that test there were no 4-card quants for the full 744B-A40B 5.3 model. There are now, though, with KLD figures close to the FP8 level. I need to do some more benchmarking to see if they live up to the hype.
- nojs 15d agoOh right, I was referring to flash. I haven’t tried these either, but the ones for 5.2 looked interesting.