Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
scottcha
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
scottcha
16d ago
Neuralwatt | https://neuralwatt.com | REMOTE | Seattle or Denver/Boulder metros | Full-time Hiring: 2 Engineers + Director of Sales & Operations Energy is becoming one of the biggest constraints on AI infrastructure. Ne
2.
▲
by
scottcha
2mo ago
We use vllm as it generally has the best ecosystem support. Parameters are largely dependent on what type of requests you are serving (concurrency, input/output ratios, cached hit patterns). We've never had a limitation at the tok
3.
▲
by
scottcha
2mo ago
I run an AI platform and we need to tokenize fast and early to make a lot of decisions on the subsequent steps (things like routing, rate limiting and such). Its really important to do this efficiently even though its not a large % of tota
4.
▲
by
scottcha
2mo ago
Hi, I'm co-founder of Neuralwatt. While there aren't other providers selling by energy we also produce the same tokens stats you get at others (input, output, cached) and you can compare using those. The CO2 number is really jus
5.
▲
by
scottcha
2mo ago
They do have a growing amount of Scope 1 emissions (emissions from their on site sources) which originally was primarily on site diesel but due to grid interconnect delays have been growing number of on site gas turbines. This certainly wou
6.
▲
by
scottcha
3mo ago
Thanks for the feedback! Our primary focus is charging by energy, for token pricing we really just try to be close to the market. That being said I'll take a look at our token pricing to see if we need an update there https:/
7.
▲
by
scottcha
3mo ago
Hi I'm the CTO of neuralwatt, would love to hear your feedback on what your experience was. Feel free to email me scott@neuralwatt.com. Also for GLM5.2 we run the FP8 quantization at 1M context which is a common deployment target.
8.
▲
by
scottcha
3mo ago
Yes way better. We host both and while qwen3.6 is over 100tps we usually can do glm around that too.
9.
▲
by
scottcha
3mo ago
I use glm5.1 plus pi with a few customized skills and am very happy with it. I hadn’t touched my Claude 5x plan for a couple of weeks but opened it back up in Claude code when fable was released and did a few tasks and still was happy to re
10.
▲
by
scottcha
4mo ago
I use claude code and pi.dev side by side most days and i'm mostly choosing pi for most work in last couple of weeks.
11.
▲
by
scottcha
6mo ago
Pretty cool idea, but whats the stack behind this? As 15-25 tok/s seems a bit low as expected SoA for most providers is around 60 tok/s and quality of life dramatically improves above that.
12.
▲
by
scottcha
6mo ago
I think there was a clarification posted on Reddit that said Claude Agents SDK didn't apply for now.
13.
▲
by
scottcha
6mo ago
I use OpenCode and have just started using Nanoclaw with ClaudeCode (my coworker has a post coming on this) and sometimes ClaudeCode with Claude Code Router. I do a range of small to complex work with these but I also do drop back in to Cl
14.
▲
by
scottcha
6mo ago
Mine are pretty unique since we optimize the energy for and run an inference service api so forces me to dogfood alot of different options.
15.
▲
by
scottcha
6mo ago
Yes GLM5 and KimiK2.5 are pretty close replacements for sonnet.
16.
▲
by
scottcha
6mo ago
I switch between Claude Code (Opus/Sonnet) and Qwen (OpenCode, OpenClaw) multiple times throughout the day and Qwen 3.5 is really nice. I do also use KimiK2.5 and GLM5 pretty often too and I'm starting to get a sense that the age
17.
▲
by
scottcha
6mo ago
We offer multiple SOA models at https://portal.neuralwatt.com at very generous pricing since we have options to bill per kWh instead of per token. Recipes for your favorite tools here: https://github.com/neuralw
18.
▲
by
scottcha
6mo ago
I actually built this analysis while I worked at Microsoft so I 100% agree. Doing the work at the platform level is the way to go and you can actually make a significant impact with this kind of approach. The other value of this that'
19.
▲
by
scottcha
7mo ago
There have been a few questions about the state of Show HN lately. Was actually interested in this post but I see all the OPs responses to questions are Dead? I do see its a new account but I don't really see anything egregious or ag
20.
▲
by
scottcha
8mo ago
That is a pretty good article although the one factor not mentioned that we see that has a huge impact on energy is batch size but that would be hard to estimate with the data he has. We've only launched to friends and family but I
21.
▲
by
scottcha
1y ago
Neuralwatt | https://neuralwatt.com | REMOTE (US – Seattle/Denver/Boulder metros only) | Full-time | $180k–$220k DOE Energy is the #1 constraint in new datacenter buildouts. Neuralwatt is reshaping AI compute around en
22.
▲
by
scottcha
1y ago
Neuralwatt | https://neuralwatt.com | REMOTE (US – Seattle/Denver/Boulder metros only) | Full-time | $180k–$220k DOE Energy is the #1 constraint in new datacenter buildouts. Neuralwatt is reshaping AI compute around en
23.
▲
by
scottcha
1y ago
I’ve asked that question on linked in to the Cerebras team a couple times and haven’t ever received a response. There is system max tdp values posted online but I’m not sure you can assume the system is running in max tdp for these queries
24.
▲
by
scottcha
1y ago
Turns out there is multiple publications associating tachycardia and other heart symptoms with long covid. https://pmc.ncbi.nlm.nih.gov/articles/PMC8356730/
25.
▲
by
scottcha
1y ago
Maybe I’m a statistical anomaly or maybe I just don’t know the baseline occurrence rate for this stuff but I have 3 close acquaintances two of which are this persons age or younger with similar symptoms (tachycardia, though to a lesser degr
26.
▲
by
scottcha
1y ago
The are many great things about Aurora, here are a few as I've been using it since it came out. 1. Its open source & open weights and free to use non-commercially. 2. Its configurable to easily fit on my local gpu for development p
27.
▲
by
scottcha
1y ago
Yeah, that was my first thought. I actually wrote a blog post a few weeks ago modeling the point at which agent recursion really gets out of control. https://www.neuralwatt.com/blog/agent-bedlam-a-future-of-end...
28.
▲
by
scottcha
1y ago
Shameless plug . . . I run a startup who is working to help this https://neuralwatt.com We are starting with an os level (as in no model changes/no developer changes required) component which uses RL to run AI with a ~25%
29.
▲
Mixture of Experts: When Does It Deliver Energy Efficiency?
(neuralwatt.com)
2 points
by
scottcha
1y ago
|
0 comments
30.
▲
by
scottcha
1y ago
My Grandparents lived in a very small farming town (pop 500) and word would get around town when chicks had arrived and she would take us down there to see them.
More ›