5 ms·
I wonder if the NCD metric says something about distillation too. Would you expect that a model that has been distilled/seen traces from other models would have
by tadkar 23d ago
I wonder if the NCD metric says something about distillation too. Would you expect that a model that has been distilled/seen traces from other models would have a smaller NCD? It would be really interesting to see if this holds up and provides evidence of distillation or certainly evidence of model outputs being used in the training mix.
- krackers 23d agohttps://eqbench.com/results/creative-writing-v3/hybrid_parsimony/charts/ox-alpha__phylo_tree_parsimony_rectangular.png https://eqbench.com/results/creative-writing-v3/hybrid_parsi...
- dejanseo 23d agoThis looks spot on, where is it from?
- dejanseo 23d agoHello! I wrote the above article and did the NCD on model outputs. The same thought crossed my mind when I saw Gemma misclassified as Gemini quite frequently. And GLM almost as Claude and not as Gemini at all. Gave me the feeling as if GLM didn't train on Gemini generated synthetic data at all but mainly on Claude and GPT.