5 ms·
Why do output tokens cost 5x more than input tokens?
- lazyMonkey69 5mo agoI think the paged attention part is a bit oversimplified. Nice read otherwise!
- ani17 5mo agoAuthor here. I wanted to understand what vLLM and llama.cpp are actually doing under the hood, but the codebases are massive. So I wrote a stripped down version from scratch to see the core ideas without the production complexity. Code: https://github.com/Anirudh171202/WhiteLotus https://github.com/Anirudh171202/WhiteLotus