7 ms·
Here is some more technical information on how this was trained, as well as a download link. https://huggingface.co/thomsonreuters/Thomson-1.0-Small https://hu
by johnnypangs 23d ago
Here is some more technical information on how this was trained, as well as a download link.
https://huggingface.co/thomsonreuters/Thomson-1.0-Small https://huggingface.co/thomsonreuters/Thomson-1.0-Small
(Full disclosure I’m a TR employee, although I had nothing to do with making this)
- deleted 23d ago[deleted]
- helloplanets 23d agoFull technical report PDF: https://huggingface.co/spaces/tri-fair-lab/publications/blob/main/Thomson_1_0_Technical_Report.pdf https://huggingface.co/spaces/tri-fair-lab/publications/blob... > In this report, we argue that frontier performance can be achieved by a wide range of institutions through Continual Learning on readily available open-weight models. > As opposed to existing limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation with a frozen model, our Continual Learning approach takes advantage of the effectiveness of a modern mid- & post-training stack while introducing safeguards preserving both plasticity and stability at each training stage and seeking to make the minimal number of high-impact interventions on the parameters. For the large model, Thomson is utilizing the fine tuning stack they describe in the article, running it on Snowdon 1.0-Large, which in turn is a fine tune of Qwen3.5 397B. Same thing for the small model, but it's a fine tune of Snowdon 1.1-Small, which is a fine tune of Qwen3.6 35B. As for the small version's run: > The full pipeline consumed approximately 1.63 × 10²³ FLOP over 35,207 B200 GPU-hours, showing that these results are achievable with compute and personnel budgets substantially lower than commonly thought. That would amount to around a quarter to half a million dollars of spend on that run. 100k minimum, if they got a great deal.
- kennywinker 23d agoI wonder where the other $39.5 million went?
- ustad 23d agoMen in the middle wages
- hypfer 23d agoSo it's a qwen fine-tune? I mean that's a reasonable thing to do, but then the press release shouldn't be written the way it is written. They're not as detached from the rest as the industry as the writing suggests. __ > It is obtained by repurposing the open-weight Qwen3.6-35B-A3B model and substantially improving it on a wide range of performance domains. nice wording on the HF page tho. "Repurposing". Lmao
- make3 23d agothe press release says this explicitly
- hypfer 23d agoYep. Explicitly enough to be legally safe for sure. Which is the thing with press releases. Would've been nice to not do the bare legal minimum tho
- dgellow 23d agoWhat isn’t clear in the press release? Reads fine to me
- mkl 23d agoNo it doesn't. There is no mention of Qwen at all.
- make3 22d agothey mention using an open weights model, which one it is doesn't really matter
- dgellow 23d agoThanks, that’s great, lots of details