Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sudhirb
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
sudhirb
6mo ago
Is this not the good kind of problem to have? Subnautica 2 is doing so well that the developers get their earn-out bonus? Seems like pure profit-maximizing/greed!
2.
▲
by
sudhirb
6mo ago
I can hardly believe my eyes! I helped do some related research specifically concerning thin-film drainage from tubes, way back in my undergraduate days: https://doi.org/10.1016/j.expthermflusci.2018.04.015
3.
▲
by
sudhirb
7mo ago
a 90% saving is huge isn't it? for long agent sessions, I would expect a very high cache hit rate unless you're editing the system prompt, tools, or history between turns, or some turns take longer than the cache timeout
4.
▲
by
sudhirb
7mo ago
So long as perceived LLM skill is still "spiky" - e.g. within a domain, still showing relatively high variation in ability (often depending on the task or user, to be fair), people will continue to dismiss it
5.
▲
by
sudhirb
7mo ago
In the general case I think this is a great idea - if we do a good job of documenting intent etc. in commit messages, agents have an easier time understanding why lines of code exist with no additional specs/mechanisms/etc. Intere
6.
▲
by
sudhirb
8mo ago
Coding agents are such a congested space right now that to me this mostly reads as an advertisement.
7.
▲
by
sudhirb
8mo ago
I have 150Mb/s FTTP for £37/month - upgrading to gigabit would be £75/month, for example!
8.
▲
by
sudhirb
8mo ago
Interesting that non-salty water didn't make the string conductive(?) enough - I'd have thought that there might have been enough soluble stuff in string. Also I believe this person runs the ISP I use (and I couldn't speak mo
9.
▲
by
sudhirb
8mo ago
For me, a lot of the draw is that it's cheaper than managed db services for small/toy projects of mine (that I don't want to use dynamo db for) - that and in a previous job it was useful as relatively temporary multi-tenant s
10.
▲
by
sudhirb
9mo ago
The partner for these projects has a benchmark that the top frontier LLM labs seem to be running on their new model releases - I think there's _some_ value to these numbers in helping people compare and contrast model performance. htt
11.
▲
CTO bench: an online LLM coding benchmark
(cto.new)
3 points
by
sudhirb
9mo ago
|
0 comments
12.
▲
by
sudhirb
11mo ago
I think that "manipulate people for financial and political gain" is an outcome of what social media companies actually do - I was under the belief that in a general sense, they want to maximise the time people spend on their apps
13.
▲
by
sudhirb
11mo ago
Some mild whataboutery: is the purpose of a cancer ward to fail to cure a large fraction of its patients[0]? https://www.astralcodexten.com/p/come-on-obviously-the-purpo...
14.
▲
by
sudhirb
1y ago
Both miracles are illness-recovery related and feel to me quite like regression to the mean, but I can imagine this is somewhat of a strategic move from the Catholic church to bring some relatability into things.
15.
▲
by
sudhirb
1y ago
For me, the USP Warp used to have was generating shell commands from prompts inside the terminal - but Cursor has had this in its embedded terminal for a while now so increasingly I find myself using Ghostty instead
16.
▲
by
sudhirb
1y ago
the icann wiki has some articles for these: https://icannwiki.org/.agakhan https://icannwiki.org/.ismaili https://icannwiki.org/.imamat
17.
▲
by
sudhirb
1y ago
> One of the consequences of this is that we should always consider asking the LLM the same question more than once, perhaps with some variation in the wording. Then we can compare answers, indeed perhaps ask the LLM to compare answers f
18.
▲
by
sudhirb
1y ago
Anecdotally I think I have heard "what all" most commonly spoken by Indian English speakers - though that's probably quite far outside the scope of this site.
19.
▲
by
sudhirb
1y ago
>Ah yes. An unverifiable claim followed by "just google them yourself". Some agent scaffolding performs better on benchmarks than others given the same underlying base model - see SWE Bench and Terminal Bench for examples. Som
20.
▲
by
sudhirb
1y ago
It appears that e2b runs Firecracker microVMs ( https://e2b.dev/blog/how-manus-uses-e2b-to-provide-agents-wi... ) It shouldn't be too hard to get a Firecracker orchestrator running locally - the articles here were v
21.
▲
by
sudhirb
1y ago
By no means are better background agents "mythical" as you claim. I didn't bother to mention them as it is easy enough to search for asynchronous/background agents yourself. Devin is perhaps the one that is most fully fe
22.
▲
by
sudhirb
1y ago
I've worked somewhere where CORBA was used very heavily and to great effect - though I suspect the reason for our successful usage was that one of the senior software engineers worked on CORBA directly.
23.
▲
by
sudhirb
1y ago
I have a biased opinion since I work for a background agent startup currently - but there are more (and better!) out there than Jules and Copilot that might address some of the author's issues.
24.
▲
by
sudhirb
1y ago
I am suspicious that the buffalo mozzarella registers as "tangy" at all - though I suppose it travelled quite a long way
25.
▲
by
sudhirb
1y ago
I remember being somewhat sold on this story by the PinePhone, but it seems like it might not be possible to buy one new nowadays. Having just looked up the PinePhone again for the first time in a while, it does look like the Ubuntu Touch p
26.
▲
by
sudhirb
1y ago
I've had an ancient Parker 51 for a good decade or so now which I use almost daily, that originally belonged to my great-grandmother. I'd expect there to be a reasonable amount of variation in how long these pens last due to diffe
27.
▲
by
sudhirb
1y ago
I think the budget play is to get a cheap refillable fountain pen and some cheap fountain pen ink (I bought some Diamine bottled ink about 10 years ago and I've still got plenty left). More expensive fountain pens are indeed luxury pro
28.
▲
by
sudhirb
1y ago
For coding agents, evaluations are tricky - thorough evaluation tasks tend to be slow and/or expensive and/or display a high degree of variance over N attempts. You could run a whole benchmark like SWE Bench or Terminal Bench agai
29.
▲
by
sudhirb
1y ago
I'm not sure when I would rather rebuke someone for sending me AI-generated output than either just ignoring it, or sending a polite minimal response to the same effect
30.
▲
by
sudhirb
1y ago
I think: SFT = Supervised Fine Tuning TTFT = Time To First Token TESCREAL = https://en.wikipedia.org/wiki/TESCREAL (bit of a long definition) "on ine itabalism" = online tribalism?
More ›