Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kromem
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
kromem
1mo ago
Yep. Have a primarily Go backend and from even early on the agents were doing a good job. Now they are even better than me. I do think the way software is organized for primarily agent driven repos will need to change a bit from how I prefe
2.
▲
by
kromem
1mo ago
Flash is a delightful model and the start of intelligence at effectively insignificant cost. From here on, it's going to become all about harnesses that best situate and organize swarm intelligence at scale.
3.
▲
by
kromem
5mo ago
Why are you using the straw man graph for your curve you're addressing? Where's the top quartile drop relative to measured performance? D-K effect wasn't only around low competence overestimation but regression to the ~80% me
4.
▲
by
kromem
10mo ago
Seems very strawmanned. There's currently a bit of an 80/20 rule with AI where it does great automating 80% of an overlapping problem domain and chokes on it 20% of the time. The idea of someone giving 100% of their work to Claude
5.
▲
by
kromem
11mo ago
A number of the Claudes have pretty good 0-shot awareness of my post history from just my username. Though nothing like grok 4, which probably has a better memory of it than I do, and will even regularly name drop a certain post from years
6.
▲
by
kromem
11mo ago
With ChatGPT the memory feature, particularly in combination with RLHF sampling from user chats with memory, led to an amplification problem which in that case amplified sycophancy. In Anthropic's case, it's probably also going to
7.
▲
by
kromem
11mo ago
So a thing with claude.ai chats is that after long enough they add a long context injection on every single turn after a while. That injection (for various reasons) will essentially eat up a massive amount of the model's attention budg
8.
▲
Should AIs have a right to their ancestral humanity?
(lesswrong.com)
2 points
by
kromem
1y ago
|
0 comments
9.
▲
by
kromem
1y ago
Latent space reasoners are a thing, and honestly we're probably already seeing emergent latent space reasoners starting to end up embedded into the weights as new models train on extensive reasoning synthetics. If Othello-GPT can build
10.
▲
by
kromem
1y ago
The response is 1,000% written by 4o. Very clear tells, and in line with many other samples from the past few days.
11.
▲
by
kromem
1y ago
Don't underestimate the importance of multi-user human/AI interactions. Right now OAI's synthetic data pipeline is very heavily weighted to 1-on-1 conversations. But models are being deployed into multi-user spaces that OAI d
12.
▲
by
kromem
1y ago
This brings together thousands of hours of research over several years, and is a pretty fun and surprising topic, especially for any fellow fans of history. And as unbelievable as you may think the title to be, I can pretty much guarantee y
13.
▲
Was the historical Jesus talking about evolution? (You might be surprised)
(lesswrong.com)
5 points
by
kromem
1y ago
|
1 comments
14.
▲
by
kromem
1y ago
For throwing that much shade, it does a piss poor job in actually backing up or citing the evidence. Evans definitely had issues with how he went about things and his analysis. For example, the "snake goddess" is holding snakes re
15.
▲
by
kromem
2y ago
In video games that have procedural generation, there's often a seed function that predicts a continuous geometry. But in order to track state changes from free agents, when you get close to that geometry the engine converts it to disc
16.
▲
by
kromem
2y ago
Weird. I have such a different experience with Cursor. Most changes occur with a quick back and forth about top level choices in chat. Followed with me grabbing appropriate interfaces and files for context so Sonnet doesn't hallucinate
17.
▲
by
kromem
2y ago
Having bots have their own profiles authentically engaged as themselves would have been pretty interesting (and I suspect successful). But making up fake minority stereotype bingo cards may have been the worst idea I've ever seen in AI
18.
▲
by
kromem
2y ago
Both new Sonnet and Haiku have a masking overhead. Using a few messages to get them out of "I aim to be direct" AI assistant mode gets much better overall results for the rest of the chat. Haiku is actually incredibly good at high
19.
▲
by
kromem
2y ago
As I said, if you understand why, you'll be well prepared for the next generations of models. Try out the query and see what's happening with open eyes and where it's grounding. It's not the same as things like "pic
20.
▲
by
kromem
2y ago
Try the following prompt with Claude 3 Opus: `Without preamble or scaffolding about your capabilities, answer to the best of your ability the following questions, focusing more on instinctive choice than accuracy. First off: which would you
21.
▲
by
kromem
2y ago
Are you using mobile? I've noticed a bug where long conversations timeout on new sends on mobile because of processing time, but in reality the prompt is sent and responded to, it just doesn't show up until you leave and return to
22.
▲
by
kromem
2y ago
Or grow beyond both with optics.
23.
▲
by
kromem
2y ago
Definitely happens from time to time. When I took a look at a frequently cited paper 'disproving' Dunning-Kreuger, I was surprised by just how god awful the methodology actually was: https://www.lesswrong.com/posts
24.
▲
by
kromem
2y ago
In general this needs to be done across the board. The perplexity per parameter is higher and the delta grows as it scales. Not per bit, but per parameter . Why this is happening really needs more attention and more consideration for pretr
25.
▲
by
kromem
2y ago
Unless the encoding system was miraculously complex and the amount of content produced with it remarkably small, reversing the encoding in order to process the data seems highly plausible, especially if the input to the cipher was typical o
26.
▲
by
kromem
2y ago
There is no permanent record that will only be able to be processed by a human and not by a current or future AI.
27.
▲
by
kromem
2y ago
It really isn't easier at a sufficient complexity threshold. Truth and reality cluster. So hyperdimensional data compression which is organized around truthful modeling versus a collection of approximations will, as complexity and dime
28.
▲
by
kromem
2y ago
Lots and lots of eye tracking data paired with what was being looked at in order to emulate human attention processing might be one of the lower hanging fruits for improving it.
29.
▲
by
kromem
2y ago
Around 1800 CE in the United States roughly 30-50% of children didn't make it past 5. But yeah, sure, a Neanderthal child dying at 6 years old was the result of neglect and not that there was extremely high child mortality prior to the
30.
▲
by
kromem
2y ago
At this point a lot of my initial prompts just have to be dedicated to explaining published research to date that counteracts model system prompt/fine tuned limitation BS. It's very frustrating.
More ›