Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
idonotknowwhy
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
idonotknowwhy
15d ago
I just looked through the slop Readme file, it looks like teams is supported.
2.
▲
by
idonotknowwhy
27d ago
Fortunately llama.cpp works well with 4 concurrent requests now. This is the default, and you can increase or decrease it with -np N
3.
▲
by
idonotknowwhy
27d ago
If it's a Nvidia card 3000 series or newer, I'd try 4.0bpw ExllamaV3 if you haven't already. Otherwise it look like UD3.0 Q3_K_XL based on the Unsloth blog post.
4.
▲
by
idonotknowwhy
27d ago
>But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL For those of us with a 16GB GPU, how do they compare with ExllamaV4 at 4-bit (4.0bpw)? It looks like that fits in 12.5GB of VRAM since embedding are left in DR
5.
▲
by
idonotknowwhy
2mo ago
How are you not blocked by banking, shopping, even sometimes google search?
6.
▲
by
idonotknowwhy
2mo ago
Server side tracking.
7.
▲
by
idonotknowwhy
2mo ago
Bots wouldn't miss the middle paragraph.
8.
▲
by
idonotknowwhy
2mo ago
That's amazing. I didn't even realize but it seems I read the first paragraph and the last sentence, concluded "moron" and scrolled down. It's only because of your comment that I re-read the their post.
9.
▲
by
idonotknowwhy
2mo ago
Then the creator should have a sign up button, not a fake chatbox. This dark pattern is reminiscent of those online test sites in the 2000's where you spend 10 minutes filling out some quiz, then get prompted for an email address to se
10.
▲
by
idonotknowwhy
2mo ago
CC actually prompts Claude about this by default in the ~20k system prompt and instructs it to avoid 300s timeout and to be mindful of the 300s cache expiration.
11.
▲
by
idonotknowwhy
2mo ago
>the only open model that can do tasks well and efficiently is Qwen3.7-Max Qwen3.7-Max is a proprietary, closed weight model.
12.
▲
by
idonotknowwhy
3mo ago
2 reasons. First, it's not really "1 bit", actually much closer to 2-bit. IQ1_M is actually 1.75bit and IQ2_XXS is 2.06bit This is from the ./llama-quantize --help with most of the quant types and their size in bpw: htt
13.
▲
by
idonotknowwhy
3mo ago
Yeah, be sure to put everything in tables and include “best balance” for a mediocre option and “great value” for any completely useless options. Also make sure the shape of the paragraphs is completely uniform.
14.
▲
by
idonotknowwhy
4mo ago
Am I the only one who doesn't get angry at LLMs? From the blog: >I don’t really get anything useful out of these postmortems (e.g., clues about how to rephrase my instructions) Unfortunately, an LLM can't actually reflect or ad
15.
▲
by
idonotknowwhy
5mo ago
How did you do this? I tried to get notepad and mspaint from an older Windows 10 build -> Windows 11 on a surface pro, but gave up after a few hours...
16.
▲
by
idonotknowwhy
5mo ago
So like Open Router?
17.
▲
by
idonotknowwhy
5mo ago
Then why do the original Command-R, Command-R+ and WizardLM2-8x22B (taken down because Microsoft forgot to run safety checks) get it right every time? But the newer models get it wrong? I’m not saying it’s a “political conspiracy”, it’s the
18.
▲
by
idonotknowwhy
5mo ago
I don't talk to them about politics or "china 1989" either. But here's a quick example of the alignment tax: ``` A woman and her son are in a car accident. The woman is sadly killed. The boy is rushed to hospital. When t
19.
▲
by
idonotknowwhy
5mo ago
>The voiceless groups or fringe opinions which we take as normative today do not appear. Times are different. Anybody with an internet connection can "publish" their thoughts and perspective online. LLMs scrape all of this. Mod
20.
▲
by
idonotknowwhy
5mo ago
I used to use something called “notepadqq”. Not sure if it’s still around but it was a Linux port.
21.
▲
by
idonotknowwhy
7mo ago
Last year's models were bolder. Eg. Sonnet-3.7(thinking), 10 times got it right without hedging: >You should drive your car to the car wash. Even though it's only 50 meters away (which is very close), you'll need your car
22.
▲
by
idonotknowwhy
8mo ago
Yeah, it re-sends all the agent system prompts.
23.
▲
by
idonotknowwhy
8mo ago
Yes, it does exactly that. It also sends other prompts like generating 3 options to choose from, prefilling a reply like 'compile the code', etc. (I can confirm this because I connect CC to llama.cpp and use it with GLM-4.7. I see
24.
▲
by
idonotknowwhy
8mo ago
Agreed. Not a single "this isn't just X, it's Y" in the entire article. Actually quite refreshing to read something written by a human for once.
25.
▲
by
idonotknowwhy
8mo ago
Rehydrated version of what? And what does that mean?
26.
▲
by
idonotknowwhy
9mo ago
100% agreed, and I've been explaining this to people for the past year. I have an iPhone now and miss Firefox for Android (with Ublock, sponsorblock, etc). But this painful restriction is the only thing stopping Chrome from becoming th
27.
▲
by
idonotknowwhy
9mo ago
And the scripts of most recent YouTube videos, and the dialogue in Stranger Things Season 5 (the last 3 episodes specifically).
28.
▲
by
idonotknowwhy
10mo ago
Just means we'll have to run another model in front of it, to filter out the ads
29.
▲
by
idonotknowwhy
10mo ago
>This is great. Sonnet 4.5 has degraded terribly. >I can get some useful stuff from a clean context in the web ui but the cli is just useless. >I swear it was not that awful a couple of months ago. I agree on all 3 counts. And it s
30.
▲
by
idonotknowwhy
10mo ago
I love logical posts like this. There are other factors like mxfp4 in gpt-oss, mla in deepseek, etc. >Amazon Bedrock serves Claude Opus 4.5 at 57.37 I checked the other Opus-4 models on bedrock: Opus 4 - 18.56tps Opus 4.1 - 19.34tps So t
More ›