Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Translationaut
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
Translationaut
7d ago
Still 10$/month with this referral code?!
2.
▲
by
Translationaut
1mo ago
There is the art of saying no: https://dl.acm.org/doi/10.5555/3737916.3739489 It is possible to create (subjective) reasoning traces like https://huggingface.co/datasets/Bachstelze/ethica
3.
▲
by
Translationaut
8mo ago
The idea of the ethical reasoning dataset is not to erase specific content. It is designed to present additional thinking traces with an ethical grounding. So far, it is only a fraction of the available data. This doesn't solve alignme
4.
▲
by
Translationaut
8mo ago
There is this ethical reasoning dataset to teach models stable and predictable values: https://huggingface.co/datasets/Bachstelze/ethical_coconot_6... An Olmo-3-7B-Think model is adapted with it. In theory, it sho
5.
▲
by
Translationaut
2y ago
The authors propose a novel approach where checklists are automatically generated to systematically assess and guide LLM outputs, ensuring more comprehensive and reliable evaluations by LLMs. E.g. it increases in the frequency of exact agre
6.
▲
by
Translationaut
2y ago
Why do we need vectors for search anyway? The results are often unrelated to the query. Aren't therefore exact matches better? One could also annotate the corpus with related tags and hypothetical questions, if we need more results.
7.
▲
by
Translationaut
3y ago
E.g. adapters, inference optimization and more (multilingual) models like https://huggingface.co/CohereForAI/aya-101
8.
▲
by
Translationaut
3y ago
What is the advantage of Ollama Python over huggingface?
9.
▲
by
Translationaut
3y ago
What is the point in using Ollama over huggingface if you use Python? Also, REST endpoints can be provided with huggingface transformer. Here with Go, it seems to make sense to use an abstraction. Though don't you lose a lot of flexibi
10.
▲
by
Translationaut
3y ago
This seems only to work cause large GPTs have redundant, undercomplex attentions. See this issue in BertViz about attention in Llama: https://github.com/jessevig/bertviz/issues/128
11.
▲
by
Translationaut
3y ago
Those minified models are still equal or bigger compared to the initial "attention is all you need" transformer.
12.
▲
by
Translationaut
3y ago
Have you also tried the bigger models? The smaller models are good for assisted generation: https://huggingface.co/blog/assisted-generation Those models of LaMini-Flan-T5 are trained to follow instructions and not to r