8 ms·
Highly encourage you to read the blog post (https://usetokenless.com/blog/building-tokenless https://usetokenless.com/blog/building-tokenless). Essentially, we
by rohaga 2mo ago
Highly encourage you to read the blog post (https://usetokenless.com/blog/building-tokenless https://usetokenless.com/blog/building-tokenless). Essentially, we estimate the confidence of a specific model failing or succeeding on a specific task using our own foundation models.
A turn here is a tool call/user input, anything that causes the model to get some new input. We're working on adding Minimax M3 and other models. We think that people have some intuitions about which models are good when--we seek to quantify them scientifically.
- CuriouslyC 2mo agoIt will inspire more confidence if you say your classifier. Saying "your own foundation model" (charitably) suggests marketing hyperbole. If you're fine tuning a LLM to do this, at best it's going to be worse than a frontier model with suitably tuned skills (hence it should just be something in harness, not a service), and I wouldn't refer to that as "your own foundation model."