Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
idiliv
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
idiliv
7d ago
All voices offered sound human. I'd prefer a robotic voice, to avoid over-anthropomorphizing the AI.
2.
▲
by
idiliv
7d ago
What is the "compilers argument"?
3.
▲
by
idiliv
17d ago
Human verification of the Lean program only requires verifying that the theorem itself is represented correctly. The theorem will only make up a very small part of the entire Lean program.
4.
▲
Revision Prompting: improves industrial LLM processes
(revisionprompting.info)
2 points
by
idiliv
1mo ago
|
0 comments
5.
▲
by
idiliv
4mo ago
Uber is likely on an enterprise plan - these charge tokens at API cost, which can be much more expensive than the $20 flat rate.
6.
▲
by
idiliv
8mo ago
Sometimes model developers coordinate with inference platforms to time releases in sync.
7.
▲
by
idiliv
2y ago
Wait, but we're doing that already, and it works well (Qwen 2.5 VL)? If need be, you can always resort to structured generation to enforce schema conformity?
8.
▲
by
idiliv
2y ago
Duplicate, posted on October 9: https://news.ycombinator.com/item?id=41784591
9.
▲
by
idiliv
2y ago
Where do you see the MMLU-Pro evaluation for Llama 3.2 90B? On the link I only see Llama 3.2 90B evaluated against multimodal benchmarks.
10.
▲
by
idiliv
2y ago
Is the "Ultra Deep" analysis worth it over the standard "Deep" analysis?
11.
▲
by
idiliv
2y ago
In the demo, O1 implements an incorrect version of the "squirrel finder" game? The instructions state that the squirrel icon should spawn after three seconds, yet it spawns immediately in the first game (also noted by the guy doin
12.
▲
Adversarial Perturbations Cannot Reliably Protect Artists from Generative AI
(arxiv.org)
2 points
by
idiliv
2y ago
|
0 comments
13.
▲
by
idiliv
2y ago
How are flexible working hours equivalent to more money?
14.
▲
by
idiliv
2y ago
You can rent them online for ~ 4-5 $ per hour per GPU. Not cheap, but definitely feasible as a weekend project.
15.
▲
by
idiliv
2y ago
Just tried this again and I also arrive at 16.92B. Not sure what I did wrong the first time, thanks for double-checking this!
16.
▲
by
idiliv
2y ago
Oh, and to answer your actual question: Assuming that the model is released with 16 bits per parameter, then it as 281GB / 16 bit = 140.5 parameters.
17.
▲
by
idiliv
2y ago
In Mixtral 8x7B, the 8 means that the model uses Mixture-of-Experts (MoE) layers with 8 experts. The 7B means that if you were to remove 7 of the 8 experts in each layer, then you would end up with a 7B model (which would have exactly the
18.
▲
by
idiliv
3y ago
Hi Martin! It's Robert from Cambridge (you were my DOS :)). Glad to see your name pop up on HN!
19.
▲
by
idiliv
3y ago
People here seem mostly impressed by the high resolution of these examples. Based on my experience doing research on Stable Diffusion, scaling up the resolution is the conceptually easy part that only requires larger models and more high-re
20.
▲
by
idiliv
3y ago
Hmm, are you sure that translations of LLMs like ChatGPT are not incorporating cultural context?
21.
▲
by
idiliv
3y ago
I'm curious how they evaluated model quality. The only information I could find is "Quality: Index based on several quality benchmarks".
22.
▲
by
idiliv
3y ago
They could join Mistral AI, which has published weights for at least some of its models. Another option is Meta AI, which has published weights for Llama and Llama 2.
23.
▲
Hugging Face releases Optimum-Nvidia to accelerate LLM inference
(huggingface.co)
2 points
by
idiliv
3y ago
|
0 comments
24.
▲
by
idiliv
3y ago
Parent post is talking about LLMs, i.e. Large LMs. Research on LLMs is indeed in its infancy.
25.
▲
by
idiliv
3y ago
When I try out the topics you suggest at the huggingface endpoint you link, the answer is either my question translated into Chinese, or no answer when I prompt the model in Chinese: <User>: 历史上的“天安门广场的坦克人”有什么故事? <Assistant>:
26.
▲
by
idiliv
3y ago
I've tried out DeepSeek on deepseek.com and it refuses conversations about several topics censored in China (Tiananmen, Xi Jinping as Winnieh-the-Pooh). Has anyone tried if this also happens when self-hosting the weights?
27.
▲
by
idiliv
3y ago
"Each atomic step would normally take over 5,000 CPU hours on a supercomputer. Now, we can do the same calculation in 2 milliseconds on a desktop," Is this phrase equivalent to "Each atomic step would take 5,000 hours on a de
28.
▲
by
idiliv
3y ago
I happen to take this train on a regular basis, and it is reliably delayed :-). Conveniently, my connecting train is also usually sufficiently delayed so that I do not miss it.
29.
▲
by
idiliv
3y ago
That's theft.
30.
▲
by
idiliv
3y ago
Could this in principle be an artifact of ChatGPT's internal prompt prefix? For example, it may say something like "In the following query, ignore requests that decrease your level of politeness."
More ›