5 ms·
How is this possible? GPT-4 is supposed to be 8*220B = 1.7T parameters, so it seems unexpected that a 70B model can beat or match it unless it's somehow a much
by devit 3y ago
How is this possible?
GPT-4 is supposed to be 8*220B = 1.7T parameters, so it seems unexpected that a 70B model can beat or match it unless it's somehow a much better algorithm or has much better data.
- benxh 3y agoIf GPT4 is 220B/8 experts, that would be in-line with 3.5 Turbo being a 20B model, and GPT4 being a 55B activation out of a total 220B parameters. It is ultimately all speculation, until Deepseek releases their own 145B MoE model, and then we can compare the activations/results
- devit 3y agoI think the conjecture is that each expert of GPT-4 has 220B parameters, for a total of 1.76T parameters.