6 ms·
not from a tech field at all but would it do the context window any good to use "think" mode but discard them once the llm gives the final answer/reply? is tha
by pomtato 1y ago
not from a tech field at all but would it do the context window any good to use "think" mode but discard them once the llm gives the final answer/reply?
is that even possible to disregard genrated token's selectively?