Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jangletown
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Show HN: Langy, an automated AI engineer (we gave it a robot body) [video]
(langwatch.ai)
9 points
by
jangletown
2mo ago
|
0 comments
2.
▲
Adding evals to a satelite image agent with a Claude Skill
(medium.com)
4 points
by
jangletown
6mo ago
|
2 comments
3.
▲
Show HN: Evals Skills
(langwatch.ai)
4 points
by
jangletown
6mo ago
|
0 comments
4.
▲
Kanban Code – The IDE for 2026
(github.com)
3 points
by
jangletown
6mo ago
|
0 comments
5.
▲
by
jangletown
7mo ago
impressive
6.
▲
Kanban Code - Native MacOS UI for Managing Multiple Claude Codes
(github.com)
1 points
by
jangletown
7mo ago
|
0 comments
7.
▲
by
jangletown
8mo ago
hey there, shameless plug here, I built pinacle.dev exactly with this use case in mind, cheap $7 VMs that comes with vs code, vibe kanban everything else needed out of the box to keep vibe coding on the go, no need to leave your computer ru
8.
▲
Show HN: Better Agents CLI
(github.com)
3 points
by
jangletown
10mo ago
|
0 comments
9.
▲
by
jangletown
1y ago
and how do you detect hallucinations?
10.
▲
Script that counts how many times has cursor said you're "absolutely right"
(gist.github.com)
1 points
by
jangletown
1y ago
|
2 comments
11.
▲
The Lethal Trifecta
(simonwillison.net)
3 points
by
jangletown
1y ago
|
0 comments
12.
▲
by
jangletown
1y ago
hello aszen, I work with draismaa, the way we have developed our simulations is by putting a few agents in a loop to simulate the conversation: - the agent under test - a user simulator agent, sending messages as a user would - a judge agen
13.
▲
by
jangletown
1y ago
That's true, we have been trying to help customers doing evals for ages now, and it's super hard for everyone to build a really good dataset and define great quality metrics just wanted then to shameless plug this lib I've bu
14.
▲
by
jangletown
1y ago
I love the term! But I do think it's both really, after all this time, LLMs are still very finicky, even the order of the instructions still matter a lot, even with the right context, so you are still prompt engineering, ideally this w
15.
▲
by
jangletown
1y ago
"51% fewer false positives", how were you measuring? is this an internal or benchmarking dataset?
16.
▲
by
jangletown
1y ago
oh shoot, wrong link: https://github.com/langwatch/scenario I think AI slop on the editor changed for me when I was typing it and I didn't notice fixed now, thanks!
17.
▲
The Agent Testing Pyramid
(rchaves.app)
7 points
by
jangletown
1y ago
|
2 comments
18.
▲
by
jangletown
3y ago
Just saw the video you shared on the other comment using prophecy, very cool Generally I don’t care much about the embedding and retrieval and connectors etc for playing with the LLMs, I imagined much more robust tools were available alread
19.
▲
by
jangletown
3y ago
I agree, I really don’t like LangChain abstractions, the chains they say are “composable” are not really, you spend more time trying to figure out langchain than actually building things with it, and it seems it’s not just me after talking
20.
▲
LangChain alternative using FP approach
(github.com)
6 points
by
jangletown
3y ago
|
0 comments
21.
▲
FP Interface for Langchain
(gist.github.com)
3 points
by
jangletown
3y ago
|
0 comments
22.
▲
by
jangletown
3y ago
unfortunately picovoice does not support plain “GPT” as it’s a “unrecognized word” in their model On the plus side, it never triggers by accident, you have to be very intentional
23.
▲
Show HN: Whisper and ChatGPT and ElevenLabs on Raspberry Pi
(github.com)
2 points
by
jangletown
3y ago
|
3 comments
24.
▲
by
jangletown
4y ago
Hello hacker news, I'm facing a conundrum in my career, I'm a Senior Software Engineer at a Big Tech company, and I'm very bored at my job, I'm using boring tech, things move slowly, and I have basically no power to chan
25.
▲
Ask HN: Stay on big tech for management experience, or move to startup?
1 points
by
jangletown
4y ago
|
1 comments