Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
0xjunhao
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
End-to-End OCR with Vision Language Models
(ubicloud.com)
3 points
by
0xjunhao
10mo ago
|
0 comments
2.
▲
by
0xjunhao
1y ago
Thank you! We have incorporated your suggestion.
3.
▲
by
0xjunhao
1y ago
Yes, typically users send the newest user message and the full conversation history. These combined become the prompt. Our API endpoint will try to route requests that has the same prefix to the same vLLM instance (similar to longest prefix
4.
▲
by
0xjunhao
1y ago
Yes, decoding is very I/O heavy. It has to stream in the whole of the model weights from HBM for every token decoded. However, that cost can be shared between the requests in the same batch. So if the system has more GPU RAM to hold la
5.
▲
by
0xjunhao
1y ago
Hi, I'm the author of this post. Writing it was a great learning experience. I gained a lot of insight into vLLM. If you have any feedback or questions, feel free to drop a comment below!
6.
▲
by
0xjunhao
1y ago
FYI a literature review from a deep research AI ====== The effect of noise on sleep is multifaceted, involving various types of noise exposure, physiological mechanisms, and consequences on sleep quality and health. I. Introduction Environm
7.
▲
by
0xjunhao
1y ago
IMHO the most appropriate medical decision should depend on the patient's economic situation.
8.
▲
by
0xjunhao
1y ago
I remember trying to publish something of a similar style to arxiv and getting rejected. Seems that having an abstract and references is key :)
9.
▲
by
0xjunhao
1y ago
If you have ever lost someone due to other drivers not watching the road, you will start to appreciate that these cars are at least always watching the road and won’t kill people standing right in front of them. Nevertheless, a lot of impro
10.
▲
by
0xjunhao
1y ago
What's the difference between resign and layoff :)
11.
▲
by
0xjunhao
1y ago
Hi, I had a quick question. Would it be correct to say the following? 1. For long inputs and short outputs, the inference can be arbitrarily number of times faster, as it avoids repeated KV computation. 2. Conversely, for short inputs and l
12.
▲
by
0xjunhao
1y ago
Before I became a software engineer, I was a computational physicist. My days back then were pretty much tweaking some parameters, running a job, then reading papers and checking back after a few minutes or hours. Increasingly, I’m starting
13.
▲
by
0xjunhao
1y ago
With the rise of Agentic AI, this increasingly feels like the right move, unless AWS drastically lowers their prices.
14.
▲
by
0xjunhao
1y ago
In a world of LLMs, it's great to see classic NLP works like Harper. Both definitely have their own use cases.
15.
▲
SageAttention3: Microscaling FP4 Attention. 5x Speed up
(arxiv.org)
2 points
by
0xjunhao
1y ago
|
0 comments
16.
▲
Ilya Sutskever's SSI is raising $1B+ on $30B valuation
(siliconangle.com)
10 points
by
0xjunhao
2y ago
|
3 comments
17.
▲
by
0xjunhao
2y ago
Maybe we should build more railways. The ground is more stable than the air.