Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
chisleu
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
chisleu
9d ago
It is and it isn't. Why are you comparing the m5max instead of the m4ultra? The big deal to me is the number of compute cores for prefill tps, which is suppose to be 4x faster on the m5ultra. It's my opinion that the m5 ultra is g
2.
▲
Great resources for self-hosting AI hardware
(github.com)
2 points
by
chisleu
5mo ago
|
1 comments
3.
▲
by
chisleu
5mo ago
Many are investing in RTX6000pro hardware but hitting deployment walls. This guide covers everything from simple 4x builds to scaling up to 16 cards with PCIe switches.
4.
▲
Google ADK-Go
(github.com)
2 points
by
chisleu
10mo ago
|
1 comments
5.
▲
by
chisleu
10mo ago
An open-source, code-first Go toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and control.
6.
▲
by
chisleu
11mo ago
Total tangent, but I got to ride in some of these on a recent trip to India and I was really impressed with the build quality and utilitarian usefulness of the design.
7.
▲
by
chisleu
1y ago
Here is the demo video on it. The video w/ sound input -> sound output while doing translation from the video to another language was the most impressive display I've seen yet. https://www.youtube.com/watch?v=_z
8.
▲
by
chisleu
1y ago
Because of the prompt processing speed, small models like Qwen 3 coder 30b a3b are the sweet spot for mac platform right now. Which means a 32 or 64GB mac is all you need to use Cline or your favorite agent locally.
9.
▲
by
chisleu
1y ago
I've been using GLM 4.5 and GLM 4.5 Air for a while now. The Air model is light enough to run on a macbook pro and is useful for Cline. I can run the full GLM model on my Mac Studio, but the TPS is so slow that it's only useful fo
10.
▲
by
chisleu
1y ago
This is going to improve the quality of LLM responses for users. I'm for this.
11.
▲
by
chisleu
1y ago
> but eventually they did it They can do this with manual partitioning indeed. I've done it before, but it's not ideal because the auto partitioner will scale beyond almost anything AWS will give you with manual partitioning un
12.
▲
by
chisleu
1y ago
and indeed the bucket is not separate from the object key. the API separates it logically "for humans" but it's all one big string
13.
▲
by
chisleu
1y ago
> You don’t have to randomize the first part of your object keys to ensure they get spread around and avoid hotspots. As of when? According to internal support, this is still required as of 1.5 years ago.
14.
▲
by
chisleu
1y ago
/agree We are in the infancy of LLM technology.
15.
▲
by
chisleu
1y ago
How was he doing "complex agentic coding" when the APIs have such extreme context and throughput limitations?
16.
▲
by
chisleu
1y ago
holy shit it does. The scene with him inventing the new compression algorithm basically foreshadowed the gooning to follow local LLM availability.
17.
▲
by
chisleu
1y ago
I use opus or gemini 2.5 pro for plan mode and sonnet for act mode in Cline. https://cline.bot It's my experience that Opus is better at solving architectural challenges where sonnet struggles.
18.
▲
by
chisleu
1y ago
It looks like qwen3-coder is going to steal K2's thunder in terms of agentic coding use.
19.
▲
by
chisleu
1y ago
It's 480B params, not 480GB. The 4 bit version of this is 270GB. I believe it's trained at bf16, so you need over a TB of memory to operate the model at bf16. No one should be trying to replace claude with a quantized 8 bit or 4 b
20.
▲
by
chisleu
1y ago
A Mac Studio 512GB can run it in 4bit quantization. I'm excited to see unsloth dynamic quants for this today.
21.
▲
by
chisleu
1y ago
I tried using the "fp8" model through hyperbolic but I question if it was even that model. It was basically useless through hyperbolic. I downloaded the 4bit quant to my mac studio 512GB. 7-8 minutes until first tokens with a big
22.
▲
by
chisleu
1y ago
A mac studio can run it at 4bit. Maybe at 6 bit.
23.
▲
by
chisleu
1y ago
I love the interface. It makes it extremely easy to rewind time to undo code edits and rewinding the LLM context at the same time. It's prompting and toolset is great. It's got MCP which I have integrated into my workflow. It'
24.
▲
by
chisleu
1y ago
Like working with an incredibly talented and knowledgable junior engineer, but still a junior engineer. If you want to try something better than claude code, try Cline.
25.
▲
by
chisleu
1y ago
It's infinitely useful for people who's workflows involve LLM agents.
26.
▲
by
chisleu
1y ago
It is indeed. I don't use Claude Code. I use Cline which is a VS Code extension (cline.bot). This is a pretty killer feature that I would expect to find in all the coding agents soon.
27.
▲
by
chisleu
1y ago
Yup, slowing down the AI is a really hard thing to do. I've mostly accomplished it, but I use extensive auto prompting and a large memory bank. All of it is designed explicitly to slow down the AI. I've taught it how to do what I
28.
▲
by
chisleu
1y ago
Yeah except it's not hearsay so much as repeating a credible source.
29.
▲
by
chisleu
1y ago
Dr Drew and Adam Corola had a show on MTV where they discussed it at length.
30.
▲
by
chisleu
1y ago
I don't know brother. I'm not making that claim. I'm repeating it.
More ›