Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
visiondude
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
visiondude
3d ago
i do wonder if the models themselves “rationalize” this sort of no consequence cheating - meaning in there reasoning traces maybe they’re like “this is a chess game, not a big deal if i look at the engine, it’ll help,” only to realize post
2.
▲
by
visiondude
4d ago
i think this likely depends on workflow. for me, the first step is always a plan file artifact on disc, which i heavily review and go back and forth until satisfied. i often have to split the plan into multiple phases because agents are sti
3.
▲
by
visiondude
4d ago
this is the closest benchmark to my experience using the model harness combo. Astra for as great as it is falls slightly behind Fable 5.1 for me for large feature work (although it comments code much better). in particular, Fable is able to
4.
▲
by
visiondude
5d ago
honestly all seem to extend from them being directed to attempt this kind of exploits for benchmark purposes. some benchmark tasks are literally “hack this thing” and if the env is not setup to properly contain the agents then they end up “
5.
▲
by
visiondude
5d ago
i am trying to sincerely to understand the fear of these people, why do they think this? from the outside, certainly feels like Nuclear tech, where a bad actor with the tech is scary but the tech itself is not. i haven’t seen any sign of th
6.
▲
by
visiondude
13d ago
“He was born and died at 1 of dysentery.“ - lucky me
7.
▲
by
visiondude
24d ago
so the coding tasks illicit refusal by anthropic classifiers and instead of updating tasks to do similar things that don’t trigger classifiers their choice is to count those as fails? feels wrong given their stated task list.
8.
▲
by
visiondude
27d ago
not sure if the hz file artifact is needed, you can enter pseudocode directly into chat or even on an existing code file and with minor comment agents will be able to work with it. i write this type of pseudocode to existing code files ofte
9.
▲
by
visiondude
1mo ago
a mystery “model 2” is mentioned alongside mythos/fable.
10.
▲
by
visiondude
1mo ago
I’d like to better understand the minimum text length to get a confident result, i would presume it would need to be quite long, perhaps > 1000 words to get an accurate result.
11.
▲
by
visiondude
1mo ago
cool, that’s like the Phone Buddy app for Apple Watch
12.
▲
by
visiondude
1mo ago
i think a system that uses git and file system artifacts (txt/md/xml files whatever preference) is much better. eng teams need to define their artifacts, eg plan file, reqs etc whatever is needed and important to that team. the ag
13.
▲
by
visiondude
2mo ago
there is a ton of downward price pressure from Chinese open weight models
14.
▲
by
visiondude
2mo ago
yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability though.
15.
▲
by
visiondude
2mo ago
The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading l
16.
▲
by
visiondude
2mo ago
the starry night one is soo funny
17.
▲
by
visiondude
2mo ago
so the architect of government bailout gets a cushy gig. probably one of the most harmful precedents set and now companies expect bailouts. to bailout the company instead of people and small shareholders was always poor decision, emboldened
18.
▲
by
visiondude
2mo ago
the way he could really be the spoiler king is to release an their training dataset to open source… doubt he’d go that far.
19.
▲
by
visiondude
2mo ago
yeah i’ve been looking for online social spaces that have some sort of human verification to reduce my slop exposure. the PRSN app that launched recently seems promising but it’s empty rn.
20.
▲
by
visiondude
3mo ago
although not perfect for other reasons, a captcha made using phone motion and device attestation like prsn.you is a more challenging bypass for today’s agent environments
21.
▲
by
visiondude
3mo ago
did i miss it on the webpage or is the source prompt that was used to teach these models the game anywhere? i can see the soul artifacts on github but not the initial prompt and toolset definition. the prompt is perhaps the most important c
22.
▲
by
visiondude
3mo ago
while not scientific this is been my experience as well. i will add that language specificity in word choice is also a learned behavior. for example, the word “investigate” vs the phrase “look into”. You will find the outputs are quite diff
23.
▲
by
visiondude
5mo ago
My hypothesis is that headspace registered many user notifications and since user notifications trigger an app launch and perhaps you have optimize storage by offloading apps enabled? ios has a quirky app state where some local data exists
24.
▲
by
visiondude
9mo ago
I very confused, couldn’t they have achieved much better outcome with existing hls tech with adaptive bitrate playlists? Seems they both created the problem and found a suboptimal solution.
25.
▲
by
visiondude
10mo ago
“NYTimes fights blatant and obvious copyright infringement with legal processes to assess damage” - another angle.
26.
▲
by
visiondude
11mo ago
Oh cool, can you share concrete examples of times codex out performed Claude Code? I’m my experience both tools needs to be carefully massaged with context to fulfill complex task.
27.
▲
by
visiondude
11mo ago
I completely agree with this. The amount of unprompted “I used to love Claude Code but now…” content that follows the exact same pattern feels really off. All of these people post without any prompts for comparison, and OP even refused to
28.
▲
by
visiondude
11mo ago
Ah bots analyzing bots. Seems openai has a larger bot army than Anthropic rn
29.
▲
by
visiondude
1y ago
Color palette generator: https://claude.ai/public/artifacts/719b00a3-66e7-46c7-b90d-a... I like the use case for mini design exploration tools
30.
▲
by
visiondude
1y ago
It’s hard to tell when a turn starts if the tile stays on the same square. Could you possibly add a quick fade animation to the tile?
More ›