Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
helloericsf
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
helloericsf
10mo ago
Wow, didn't know there were talks about relaxing GDPR. Can you share a few links? Many thanks.
2.
▲
by
helloericsf
10mo ago
Thanks for sharing. I bet their DPO and EU customers are super interested in the findings. The CEO should have handled it better, IMO.
3.
▲
by
helloericsf
10mo ago
Not a DBA, how do you do DB permission rollout gating?
4.
▲
by
helloericsf
1y ago
Thanks for reaching out. Just reposted.
5.
▲
by
helloericsf
1y ago
If you're in SF, you don't want to miss this. The Qwen team is making their first public appearance in the United States, with the VP of Qwen Lab speaking at the meetup below during SF teach week. https://partiful.com&
6.
▲
Context Engineering for AI Agents: Lessons
(manus.im)
120 points
by
helloericsf
1y ago
|
4 comments
7.
▲
Context Engineering for AI Agents: Lessons
(manus.im)
3 points
by
helloericsf
1y ago
|
0 comments
8.
▲
by
helloericsf
1y ago
The base plan limit is not hard to hit. Then you're on the usage based rocket.
9.
▲
by
helloericsf
1y ago
Honestly depends on when they got in. Seed investors? They're probably fine with their preferences. Series B and beyond? That's where it gets messy. What round you thinking?
10.
▲
by
helloericsf
1y ago
How does it stack up against the new Grok 4 model?
11.
▲
Better than DeepSeek R1? MiniMax-M1:open-weight hybrid-attention reasoning model
(huggingface.co)
6 points
by
helloericsf
1y ago
|
0 comments
12.
▲
kit - Code Intelligence Toolkit
(github.com)
1 points
by
helloericsf
1y ago
|
0 comments
13.
▲
by
helloericsf
2y ago
3 repos DualPipe https://github.com/deepseek-ai/DualPipe EPLB https://github.com/deepseek-ai/eplb profile-data https://github.com/deepseek-ai/profile-data
14.
▲
DeepSeek Open Source Optimized Parallelism Strategies, 3 repos
(github.com)
103 points
by
helloericsf
2y ago
|
8 comments
15.
▲
by
helloericsf
2y ago
Github: https://github.com/deepseek-ai/DeepGEMM - Up to 1350+ FP8 TFLOPS on Hopper GPUs - No heavy dependency, as clean as a tutorial - Fully Just-In-Time compiled - Core logic at ~300 lines - yet outperforms expert-tu
16.
▲
DeepSeek Open Source DeepGEMM – FP8 GEMM Library(300 lines for 1350+ FP8 TFLOPS)
(twitter.com)
4 points
by
helloericsf
2y ago
|
1 comments
17.
▲
by
helloericsf
2y ago
HF: https://huggingface.co/Wan-AI/Wan2.1-T2V-14B Github: https://github.com/Wan-Video/Wan2.1 Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video gen
18.
▲
Alibaba Open Source Large-Scale Video Generative Models: Wan2.1
(twitter.com)
8 points
by
helloericsf
2y ago
|
2 comments
19.
▲
by
helloericsf
2y ago
this might help: https://x.com/main_horse/status/1894215779521794058/photo/1
20.
▲
by
helloericsf
2y ago
- Efficient and optimized all-to-all communication - Both intranode and internode support with NVLink and RDMA - High-throughput kernels for training and inference prefilling - Low-latency kernels for inference decoding - Native FP8 dispatc
21.
▲
DeepSeek open source DeepEP – library for MoE training and Inference
(github.com)
536 points
by
helloericsf
2y ago
|
71 comments
22.
▲
by
helloericsf
2y ago
They don't have h100. wink,wink.
23.
▲
by
helloericsf
2y ago
Don't think the decision is based on infra, or any technical reasons. It's more on the service support side. How a 200-person company supports 44M iPhone users in China?
24.
▲
by
helloericsf
2y ago
What do you mean by "lower"? To my understanding, they will open 5 infra related repos this week. Let's revisit your comparison question on Friday.
25.
▲
by
helloericsf
2y ago
X: https://x.com/deepseek_ai/status/1893836827574030466 BF16 support Paged KV cache (block size 64) 3000 GB/s memory-bound & 580 TFLOPS compute-bound on H800
26.
▲
DeepSeek Open Source FlashMLA – MLA Decoding Kernel for Hopper GPUs
(github.com)
441 points
by
helloericsf
2y ago
|
108 comments
27.
▲
by
helloericsf
2y ago
Blog: https://qwenlm.github.io/blog/qwen2.5-max/ Seems like the model is not available for download on HF/Github
28.
▲
New Qwen2.5-Max Outperforms DeepSeek V3 in Benchmarks
(twitter.com)
3 points
by
helloericsf
2y ago
|
2 comments
29.
▲
by
helloericsf
2y ago
Release blog: https://www.minimaxi.com/en/news/minimax-01-series-2 Report: https://filecdn.minimax.chat/_Arxiv_MiniMax_01_Report.pdf
30.
▲
Longest context up to 4M, MiniMax-01 hybrid 456B Open source model
(github.com)
19 points
by
helloericsf
2y ago
|
1 comments
More ›