4 ms·Token-Count-Based Batching: Faster, Cheaper Embedding Inference for Queries1 points by fzliu 9mo ago