Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
BoredomIsFun
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
BoredomIsFun
3d ago
ehhh....to be pedantic, LLMs, like any NNs derive power from _not_ being pure matrix multiplication - they are "halved" matrix algebra, due to ReLU
2.
▲
by
BoredomIsFun
3d ago
> The models have not plateaued, and they are not even mildly close to any sort of ceiling. Depends on defnition of "plateaued" and "ceiling". I am not impressed with 2026 consumer models at all. > This is also com
3.
▲
by
BoredomIsFun
5d ago
It does not if you switch swapp off and use zram instead. I am typing right now on such a setup wityh 16 GiB ram and it occasionally, once a week or so, kills my firefox due to oom. If you are you using disk swap - not sure why would if you
4.
▲
by
BoredomIsFun
6d ago
Apple mnitors for whatever reason use ancient 450nm backlight, which strains my eyes like no tomorrow. All modern 5k and 6k use 455-460nm backlight. Otherwise, yes, APples are good.
5.
▲
by
BoredomIsFun
6d ago
There is a plenty of Chinese lower quality 28 3:2s
6.
▲
by
BoredomIsFun
6d ago
> Built-in hardware calibration is non-negotiable for color critical work No it is mostly a convenience gimmick. Support for external hardware calibration is important though. But far more important feature, implemented properly only by
7.
▲
by
BoredomIsFun
7d ago
> Could you explain how you define hallucination in the context of creativity Mostly as you've mentioned, "lack of consistency", which manifests in variety of ways - presence of cellphones in historical settings (I'd
8.
▲
by
BoredomIsFun
7d ago
You can "great learning experience" about the human nature from a good fiction book too. Classics have another function - being cultural landmarks, one can refer to in non-fictional contexts as well.
9.
▲
by
BoredomIsFun
7d ago
LLMisms are almost universally results of using a narrow set of mainstream offerings from Anthropic and Openai. Even Muse Spark has style much less "sloppy" than Claude et al, let alone Chinese LLMs. I mean yes, they have their ow
10.
▲
by
BoredomIsFun
7d ago
I tried at it creative writing - and, with thinking off, it was considerably better than Mercury 2 and generally good in fact, not very sloppy. Now with thinking on, it got worse, began hallucinating things; this is something I've noti
11.
▲
by
BoredomIsFun
8d ago
Netflix is just a single point. I'd never run current in production.
12.
▲
by
BoredomIsFun
8d ago
You can generate images on 5060ti, it'd take less than minute per image. Trivial environmental footprint.
13.
▲
by
BoredomIsFun
8d ago
> Also, real businesses use -CURRENT, everybody knows that. No, not really.
14.
▲
by
BoredomIsFun
8d ago
> seen as the inhibitors of progress I become the government institution, the inhibitor of progress. What a load of delusion.
15.
▲
by
BoredomIsFun
8d ago
> Mistral Small 4 is way worse than Gemma 4 26B A4B Depends for what purpose? I found large Mistrals are massively better than Gemma 4 at creative writing: have more natural tone, better consistency than 26B as it is MoE.
16.
▲
by
BoredomIsFun
8d ago
Mistral is odd. They have made mostly flops, boring models (Ministral 3, Mistral Small 4, Small 3, Small 3.1) together with a classic masterpiece Mistral Nemo and very good Mistral Large 2407, Mistral Small 22b, Mistral Small 3.2.
17.
▲
by
BoredomIsFun
8d ago
> They are regulationmaxxing instead of benchmaxxing, that's my problem with them. To those who is in know (r/localllama, r/sillytavernai), is well aware that Mistral models - at least the small, <=24b ones - are the le
18.
▲
by
BoredomIsFun
9d ago
> that almost all Amish use one. There are small pedal powered ones. I am sure you can rig a horse driven washing machine...
19.
▲
by
BoredomIsFun
9d ago
> short story writers esp. among sci-fi, as sci-fi is more about concept than execution,
20.
▲
by
BoredomIsFun
9d ago
Yes surely. I like ML, >D>S and such but lacking in stats background I wishh I had.
21.
▲
by
BoredomIsFun
12d ago
Yep, an old idea, that has long, long been known in local LLM community - it was achieved by "self-merging". One of the latest, most succesful examples is a self-merge of Microsoft Phi4-14b into Phi4-25b. Some people at r/Loc
22.
▲
by
BoredomIsFun
12d ago
Looped transformers are an old idea, has long, long been known in local LLM community - it was achieved by "self-merging". One of the latest, most succesful examples is a self-merge of Microsoft Phi4-14b into Phi4-25b. Some people
23.
▲
by
BoredomIsFun
13d ago
> (not even 3.8...) 3.6 is better for non-coding tasks, noticeably so.
24.
▲
by
BoredomIsFun
13d ago
It does not matter really - stiffness does lower up to T=0.7, then platoes; even at high temperatures tics/slop-patterns are still there.
25.
▲
by
BoredomIsFun
13d ago
> Try tinkering with a base model and you'll be surprised how diverse it is. Have you tried? I have. Not much different from RLHFed; full of tics and slop, similar but slightly different from intsruction posttrains.
26.
▲
by
BoredomIsFun
14d ago
> not 1 token on 1 session like local models. Local models can absolutely run in batch, what are even talking about? > If you had only batch inference and enough of it to fill the compute to 80% then you get cheaper local models. Even
27.
▲
by
BoredomIsFun
14d ago
> It's very possible that it costs you more than a cloud mode would ...which is almost always true in a single request/reply mode and never true in batch mode. Single request usually 2x-3x more expensive than cloud and batch mo
28.
▲
by
BoredomIsFun
14d ago
> Stylistically everything you see is an artifact of post-training, It is still not exactly clear if it is true or not. Unless we have base "pt" snaphot of Claude we can't say one way or another. I've played a bit wi
29.
▲
by
BoredomIsFun
15d ago
some folks at r/WritingWithAI some time ago mentioned that disclosing AI did not kill sales.
30.
▲
by
BoredomIsFun
15d ago
> write novel length stories This would never work. Anything longer than 1500 words gonna be bad. To get proper quality you should generate piece by piece then stitch.
More ›