5 ms·
We use text-embedding-3-large, with both quantization and MRL reduction, plus oversampling on the search to compensate for the compression.
by pamelafox 1y ago
We use text-embedding-3-large, with both quantization and MRL reduction, plus oversampling on the search to compensate for the compression.