7 ms·
Then the article contradicts the slides because it states that OpenAI choose not to disaggregate prefill and decode. Idk you men with "predict"---conventional L
by bjourne 22d ago
Then the article contradicts the slides because it states that OpenAI choose not to disaggregate prefill and decode. Idk you men with "predict"---conventional LLM serving comprises only two phases.
- dist-epoch 22d agosorry, my mistake, I meant draft not predict > it states that OpenAI choose not to disaggregate prefill and decode They disaggregate INSIDE the chip, not by having separate machines for the 3 phases. the slides: https://x.com/beffjezos/status/2092416851737518190 https://x.com/beffjezos/status/2092416851737518190