5 ms·
On the other hand, it used about the same tokens as GLM 5.2 and got 1 point lower score. The fact that we have a GLM 5.2-class model that can run on two 3090's
by 2001zhaozhao 1mo ago
On the other hand, it used about the same tokens as GLM 5.2 and got 1 point lower score.
The fact that we have a GLM 5.2-class model that can run on two 3090's comfortably at Q8 is absolutely insane. It wasn't long ago that GLM 5.2 was considered amazing for open weight models.