Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
suninsight
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
suninsight
1y ago
Very cool and nicely executed ! Definitely see a lot of value in this. I was actually building a version of this using NonBioS.ai, but this is already pretty well done, so will just use this instead.
2.
▲
by
suninsight
1y ago
So I can attest to the fact that all of the things proposed in this article actually works. And you can try it out yourself on any arbitrary code base within few minutes. This is how: I work for a company called NonBioS.ai - we already impl
3.
▲
What I learned managing an AI developer while seeking enlightenment
(pocha.substack.com)
4 points
by
suninsight
1y ago
|
0 comments
4.
▲
by
suninsight
1y ago
how about halarax ...halucinate and paralax
5.
▲
by
suninsight
1y ago
So we tried that route - but problem is that these interfaces aren't suited for asynchronous updates. Like if the agent is working for the next hour or so - how do you communicate that in mediums like these. An Agent, unlike a human, i
6.
▲
by
suninsight
1y ago
I think if you use Cursor, using Claude Code is a huge upgrade. The problem is that Cursor was a huge upgrade from the IDE, so we are still getting used to it. The company I work for builds a similar tool - NonBioS.ai. It is in someways sim
7.
▲
by
suninsight
1y ago
1. Multi-Agent is divide a part into tasks and hand off each part to a different Agent. This is different in the sense that a task is not divided into parts aprior. When the agent gets to a roadblock - lets say it is unable to fix a softwar
8.
▲
by
suninsight
1y ago
He is NOT talking about multi-agent systems, which is exactly why he is calling it an Agency. The author goes to great length to explain why this is NOT a multi-agent system because it can be easily misunderstood to be that.
9.
▲
by
suninsight
1y ago
This isn't multi-agents at all. Infact if you read the article in detail, you will realize that the author goes in detail to explain how this system is different from multi-agents. And this is exactly why the author calls it "Agen
10.
▲
From AI to Agents to Agencies
(blog.nishantsoni.com)
10 points
by
suninsight
1y ago
|
10 comments
11.
▲
by
suninsight
1y ago
So what we do at NonBioS.ai is to use a cheaper model to do routine tasks, but switch to a higher thinking model seamlessly if the agent get stuck. Its most cost efficient, and we take that switching cost away from the engineer. But broadly
12.
▲
by
suninsight
1y ago
This will not end well.
13.
▲
by
suninsight
1y ago
Most of our stuff is built in house actually, simply because everything else is still kind of catching up. You can find a bunch of information on the blog ( https://www.nonbios.ai/blog ) The only software that we use is Langf
14.
▲
by
suninsight
1y ago
It is AI Software Dev called NonBioS.ai
15.
▲
by
suninsight
1y ago
As someone who works for a company having a real Agent in production, (not a workflow), I cannot disagree more than the very first statement here: Use Agent Frameworks like Langraph. We did exactly that, and had to throw everything away jus
16.
▲
by
suninsight
1y ago
Yes, but we dont believe that this is a 'fundamental' problem. We have learnt to guide their actions a lot better and they go down the rabbit a lot less now than when we started out.
17.
▲
by
suninsight
1y ago
That is very accurate with what we have found. <thinking> models do a lot better, but with huge speed drops. For now, we have chosen accuracy over speed. But speed drop is like 3-4x - so we might move to an architecture where we '
18.
▲
by
suninsight
1y ago
Yes it works really well. We do something like that at NonBioS.ai - longer post below. The agent self reflects if it is stuck or confused and calls out the human for help.
19.
▲
by
suninsight
1y ago
So managing context is what takes the maximum effort. We use a bunch of strategies to reduce it, including, but not limited to: 1. Custom MCP server to work on linux command line. This wasn't really a 'MCP' server because we
20.
▲
by
suninsight
1y ago
It only seems effective, unless you start using it for actual work. The biggest issue - context. All tool use creates context. Large code bases come with large context out of the bat. LLM's seem to work, unless they are hit with a size
21.
▲
by
suninsight
1y ago
https://nonbios.ai - [Disclosure: I am working on this.] - We are in public beta and free for now. - Fully Agentic. Controllable and Transparent. Agent does all the work, but keeps you in the loop. You can take back control anyt
22.
▲
by
suninsight
1y ago
No we dont use QEMU - never heard of them till now. We built our own software from scratch - using Ubuntu - for AI. We are completely on the cloud. Every user gets a full Ubuntu Cloud VM for his NonBioS AI Engineer to work on. We covered th
23.
▲
by
suninsight
1y ago
I also did not get it, but now I get it a bit, I think. Look at it this way. You have to get some work done - maybe book a flight ticket. So you go to two sites - first you go to flight fare comparison, then you book the ticket on the airli
24.
▲
by
suninsight
1y ago
Very cool product ! We, at NonBioS.ai [AI Software Dev], built something like this from scratch for Linux VM's, and it was a heavy lift. Could have used you guys if had known about it. But can see this being immediately useful at a ton
25.
▲
by
suninsight
2y ago
Key questions: 1. The key data point seems to be Figure 6a. Where it compares performance on BABILong and claims Titans performance is at ~62%, as compared to GPT-4o-mini at ~42% for 100k sequence length. However, GPT-4o and Claude are miss
26.
▲
by
suninsight
2y ago
> Doctors will often recommend exercise, but I find that these days even moderately strenuous exercise like riding a bicycle destroys my sleep quality for several days. There's something about it that appears to be too physiological
27.
▲
by
suninsight
2y ago
So we are still on V2.7 - works pretty good for us. Havent tried V3 yet, and not looking to upgrade. I think the next big feature set we are looking for is a prompt evaluation system. But we are coming around to the view that it is a big en
28.
▲
by
suninsight
2y ago
Thanks for the pointer ! We are actually toying with building out a prompt evaluation platform and were considering extending langfuse. Maybe just use this instead.
29.
▲
by
suninsight
2y ago
Bunch of them : Langsmith, Lunary, Phoenix Arize, Portkey, Datadog and Helicone. We also picked Langfuse - more details here: https://www.nonbios.ai/post/the-nonbios-llm-observability-pi...
30.
▲
by
suninsight
2y ago
I did a similar test and tried to pull up certain categories of individuals I am interested in, with their names and linkedin profile links. ChatGPT hallucinated the names and the links. I simply cannot move to a search, where there is rand
More ›