Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
glub
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
glub
9d ago
Now it's an entire industry. They call it "bot protection"
32.
▲
by
glub
9d ago
I only use skills that are docs of software. Anything else is pure garbage. They get pinned with nix together with the software that they come from. It's just two 3rd party skills now: playwright-cli and herdr. All the rest are skills
33.
▲
by
glub
9d ago
Yes. See LongMemEval, LoCoMo. Tons of research here. But precision/recall is relatively "solved". What nobody has gotten close to solving is maintenance and provenance - what goes into memory, what qualifies as truth, how sta
34.
▲
by
glub
10d ago
I now have dozens of projects where I've embraced the yolo. I went through elaborate systems, workflows, review triages, architectural linters, specialist agents, yet, they all still suffered the same fate - slop which I don't und
35.
▲
by
glub
10d ago
> The intended purpose of RCS was to be federated between carriers, just like SMS. Oof, so it was essentially destined to fail when the spec was being written.
36.
▲
by
glub
10d ago
They also probably used Signal for comms, and it's safe to assume some of them used protonmail. And maybe they used a Google camera for taking photos of their acts. By the same logic, what prevents US from labelling Signal, Proton, and
37.
▲
by
glub
10d ago
I don't know much about RCS other than that carriers need to be involved, but there's a way to not involve them, which Google did for a while using some kind of compat-service (jibe or something?), and then they stopped doing it.
38.
▲
by
glub
12d ago
Maybe their justification is that allowing anyone to reserve means they'll have to admit they're screwing the ones with conflicts? Maybe there's some legal loophole that says screwing everyone == not screwing.
39.
▲
by
glub
12d ago
I do RE mostly, and they do lock up for me. But what works for me is: warm up context with non-RE things with American models > switch to GLM 5.3 with actual request, let it fill context with some RE work > switch to American > swi
40.
▲
by
glub
12d ago
Claude models tend to cut corners during design/ideation too. It becomes especially visible once you pair Sol as advisor to Fable. Sol will start going crazy - "hey, you said this, and it's actually false, i checked that"
41.
▲
by
glub
14d ago
Yeah, after some more testing, I think I'm going to pin it back to 5. It feels like 5.1 is 5 that has higher reasoning threshold. I've been using fable as orchestrator anyway, so I see no reason to use 5.1.
42.
▲
by
glub
15d ago
I'd rather sound desperate than be someone who noticed a shadow criterion quietly widening the already wide gap in access to frontier intelligence, and did nothing about it. Don't worry about me, I'm fortunate enough that thi
43.
▲
by
glub
15d ago
It only sticks to the instruction for maybe 3-4 turns. This is why when Anthropic released "concise output style" feature in claude code, it basically spams the model's context with "be concise" system reminders eve
44.
▲
by
glub
15d ago
It still talks the same claudish, but now it's indeed denser. I'm not quite sure what step up they're talking about.
45.
▲
by
glub
15d ago
It's a mix of slightly worse kimi k3 for UI work and slightly smarter than luna for everything else. But yeah, it's very slow. I've put it to work as an LLM-as-RAG agent.
46.
▲
by
glub
15d ago
From my limited testing of just 2 hours, reasoning output of 5.1-max is at least 7x of 5-max, on the same project and comparable prompts. It reasoned for ~2 minutes trying to figure out an appropriate directory name. I've never seen 5-
47.
▲
by
glub
15d ago
> Then you can implement the problematic parts yourself. Or with another LLM, but yeah. The only issue is when it's a monorepo and fable does ls/grep. I've got a file named `system_prompt` in a completely innocent project
48.
▲
by
glub
15d ago
They do want my business, it's an official OpenAI market. The gate isn't "we don't serve you", it's "pay for the model that may target you, but not for defensive purposes". And without a single poli
49.
▲
by
glub
15d ago
I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say so
50.
▲
by
glub
15d ago
> OpenAI is committed to ensuring that the benefits of AI are broadly accessible. > We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective crite
51.
▲
by
glub
16d ago
Same experience here. It would choose 2 providers and then bounce between the two every 5 requests or so. I don't know why there's no "Pick the cheapest provider above nTPS on first request and stick until cache bust" se
52.
▲
by
glub
16d ago
For single turn prompts like this, you just have to give the model encrypted bytes and ask politely what's in them, really. GPT-5.6 will disclose its internals if you tell it it's in "audit mode" and has to calculate c
53.
▲
by
glub
16d ago
How is this not an indication of a failed system? Why does EU/EC/etc system even allow for the same failed initiative to be pushed again and again under a different candy wrap with a rate of a machine gun? Not always a different w
54.
▲
by
glub
17d ago
I started session attribution (locally) as soon as I found that there were jsonl files on disk for every session. It makes it really easy to make sure everything that happens has a line that goes back to the user intent. Now every project I
55.
▲
OpenAI gates cyber defense in 44 ChatGPT markets with a 1996 US export list
(lubaretsi.com)
2 points
by
glub
19d ago
|
0 comments
56.
▲
by
glub
21d ago
You may have more luck with GLM 5.3. New Kimi subscriptions are currently paused, so you have to join the waitlist. But even if you get in, usage limits are pretty bad there, just check reddit. GLM 5.3 is quite capable with "cyber"
57.
▲
by
glub
21d ago
Ah yeah, new accounts get instantly blocked on TAC. Their backend has 2 failure states for `id_verification_status` - `failed` and `blocked`. Nationality bans get `failed`, new accounts (or rather, accounts with not enough good signals) get
58.
▲
by
glub
21d ago
When they initially revoked TAC for a bunch of users due to a "technical error", the TAC verification flow opened a Persona iframe where you select the document country first. I was able to select it there (Georgia, in my case), w
59.
▲
by
glub
21d ago
> Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms. OpenAI revoked my Cyber verification, along with many others, asked to reverify (i.e. give my biometric information to Persona),
60.
▲
by
glub
21d ago
Yes. I have no UI experience, and wanted a model that could produce something good without me telling it how anything should look like. My prompt was something like: "here's data I have, here's what matters to me, create HTML
More ›