4 ms·
Just noticed in the model's metadata: > MiniLMv2 L6-H384 (30M params) progressively distilled L24→L12→L6 from RoBERTa-large-v8 teacher (neobert-v2 recipe, min_
by joouha 2mo ago
Just noticed in the model's metadata:
> MiniLMv2 L6-H384 (30M params) progressively distilled L24→L12→L6 from RoBERTa-large-v8 teacher (neobert-v2 recipe, min_words=0). Byte-level BPE tokenizer. INT8 quantized.