Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
benxh
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
benxh
16d ago
NATO v Yugoslavia 1999 is not like the others
2.
▲
by
benxh
1mo ago
Depends on what software/hardware you'll run it. GGUFs from Unsloth can run on pretty much every single potato; full weights need beefy gpus
3.
▲
by
benxh
3mo ago
benchmark where gemini flash is better than fable btw.
4.
▲
by
benxh
8mo ago
Crazy calling sovereign states "US Puppets".
5.
▲
by
benxh
8mo ago
Minimax has been great for super high speed web/js/ts related work. It compares in my experience to Claude Sonnet, and at times gets stuff similar to Opus. Design wise it produces some of the most beautiful AI generated page I&#x
6.
▲
by
benxh
8mo ago
It is arguable that the new Minimax M2.1 and GLM4.7 are drastically above Sonnet 3.7 in capabilities.
7.
▲
by
benxh
11mo ago
cline is used by a lot of devs
8.
▲
by
benxh
1y ago
The longer "it" reasons, the more attention sinks are used to come to a "better" final output.
9.
▲
by
benxh
2y ago
To prove you right, you can read up on the incredible giga-brained countrywide experiments by Kardelj in Socialist Yugoslavia [0]. The result being a country where no-one wanted to work, and everyone had a great standard of living (while th
10.
▲
by
benxh
2y ago
My biggest gripe with Ollama is the badly named models, e.g. under deepseek-r1, it defaults to the distill models.
11.
▲
by
benxh
2y ago
I'm pretty sure that Neosync[0] does this to a pretty good degree, it is open source and YC funded too. [0] https://www.neosync.dev/
12.
▲
by
benxh
2y ago
I am assuming this will be solved this year.
13.
▲
by
benxh
3y ago
If GPT4 is 220B/8 experts, that would be in-line with 3.5 Turbo being a 20B model, and GPT4 being a 55B activation out of a total 220B parameters. It is ultimately all speculation, until Deepseek releases their own 145B MoE model, and
14.
▲
by
benxh
3y ago
I personally was affected by this fire, although I've always kept 3 month backups of production data, encrypted, on-site, just in case of emergencies like this. Haven't touched their services for anything production related ever s
15.
▲
by
benxh
3y ago
It's buried deep in the Gemini report, but goddamn are these incredible stats.
16.
▲
by
benxh
3y ago
The Albanian takeover of AI continues. It's incredibly exciting!
17.
▲
by
benxh
3y ago
I can't wait to see this open sourced, there's a lot of sampling strategies that help coding. And I also can't wait to see how much Phind will improve further if the Glaive dataset is added onto it. Edit: Contrastive search,
18.
▲
by
benxh
3y ago
It was known by Polynesians for at least 1000 years before Columbus. See sweet potatoes.
19.
▲
by
benxh
3y ago
It's missing a lot of crucial details. Nothing on the dataset used, nothing on the data mix, nothing on their data cleaning procedures, nothing on the tokens trained.
20.
▲
by
benxh
3y ago
Not all of them per se, take a look at something like Mistral. It's a 7B model displaying incredible performance. IMO, we still haven't even scratched the surface of what is possible with small LLMs. Especially not with pre-filter
21.
▲
by
benxh
3y ago
Added, and reached out on Twitter.
22.
▲
by
benxh
3y ago
I would like to get in touch with you related to books4. Do you happen to have discord? or would twitter be ok? There's currently multiple attempts at creating what you describe as books4.
23.
▲
by
benxh
3y ago
I've had some success using vast.ai[0] with the Oobabooga LLM WebUI (LLaMA2) instances. One click to start up, minimal editing in the interface settings to enable OpenAI compatible interface. [0] https://cloud.vast.ai/
24.
▲
by
benxh
3y ago
Yeah wildly inaccurate.
25.
▲
by
benxh
3y ago
This reads like a Serbian owned business from the North of Kosovo. But anybody following the local politics would know that: 1) Corruption as an issue is disappearing in Kosovo, especially compared to Albania, Macedonia, Montenegro and Serb
26.
▲
by
benxh
3y ago
Yes, but the acquisition of that data itself is illegal in almost all jurisdictions, since libgen is treated as a piracy website. Now if there were a pipeline to access books from Amazon or the Google Books project for training it would be
27.
▲
by
benxh
3y ago
How does one go about becoming a distributor of Quest products in countries which arent served at all by Meta? Considering I can leverage existing infrastructure and network connections to bring it to market?
28.
▲
by
benxh
3y ago
To be honest, I've been asking myself the same thing, technically the amount of "good quality" data in libgen is huge, way larger than the books3 dataset. However it would probably run afoul of copyright. Then again, a huge a
29.
▲
by
benxh
3y ago
So a model fine-tuned on libgen?
30.
▲
by
benxh
3y ago
Many of the currently ongoing reproductions of LLaMa for starters. (see: Red Pajama[0]). Or any of the OpenAssistant affiliated projects. [0] - https://www.together.xyz/blog/redpajama
More ›