6 ms·
By post-train I presume you mean a finetune? Unless that's wrong (please correct me if so). I haven't looked into model architecture people are working with fo
by fennecfoxy 1mo ago
By post-train I presume you mean a finetune? Unless that's wrong (please correct me if so).
I haven't looked into model architecture people are working with for this stuff too deeply yet but I presume the core idea is fine-tuning a lightweight reasoning-enabled LLM specifically using search as a metric for training?
- Malp 1mo agoThat or providing a concrete RL env for $your_search_corpus_etc_here