Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
convexstrictly
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Retire the Abstractions
(hazyresearch.stanford.edu)
72 points
by
convexstrictly
1mo ago
|
64 comments
2.
▲
by
convexstrictly
2mo ago
Tokenization is done on the CPU. Models never see the raw characters. That's why you get trick questions like the number of r's in strawberry. There are many research papers on models using characters directly. One challenge is
3.
▲
by
convexstrictly
2mo ago
GitHub: https://github.com/marcelroed/gigatoken
4.
▲
Gigatoken: Fastest Tokenizer
(twitter.com)
26 points
by
convexstrictly
2mo ago
|
5 comments
5.
▲
by
convexstrictly
2y ago
Video Demo: https://x.com/OfficialLoganK/status/1869789822384255300
6.
▲
by
convexstrictly
2y ago
"Just when you thought it was over... we’re introducing Gemini 2.0 Flash Thinking, a new experimental model that unlocks stronger reasoning capabilities and shows its thoughts. The model plans (with thoughts visible), can solve comple
7.
▲
Gemini Flash 2.0 Thinking Experimental
(github.com)
4 points
by
convexstrictly
2y ago
|
3 comments
8.
▲
by
convexstrictly
2y ago
GB per second
9.
▲
by
convexstrictly
2y ago
Simran Arora: "Join us for a livestream this Thursday, Halloween/Diwali , and join our channel on the GPU Mode Discord server to hang out with us/get involved:" https://discord.com/login?redirect_to=%2
10.
▲
by
convexstrictly
2y ago
CUDA + ThunderKittens 4.5 hour tutorial https://www.youtube.com/watch?v=xcpEl0cGCC4
11.
▲
What Questions Are in the Chinese College Entrance Exam?
(cherylwu.substack.com)
1 points
by
convexstrictly
2y ago
|
0 comments
12.
▲
Generative AI Is Not Going to Build Your Engineering Team for You
(stackoverflow.blog)
1 points
by
convexstrictly
2y ago
|
0 comments
13.
▲
by
convexstrictly
2y ago
Results https://x.com/sbeastwindy/status/1801525876267372874
14.
▲
Building GPT2o – Part 1: Audio
(medium.com)
3 points
by
convexstrictly
2y ago
|
1 comments
15.
▲
by
convexstrictly
2y ago
Aider uses Treesitter to improve code generation. https://aider.chat/2023/10/22/repomap.html Aider: https://github.com/paul-gauthier/aider It is state of the art on SWE-Bench and SWE-Ben
16.
▲
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
(arxiv.org)
7 points
by
convexstrictly
2y ago
|
0 comments
17.
▲
OpenAI says it has begun training a new flagship A.I. model
(nytimes.com)
8 points
by
convexstrictly
2y ago
|
0 comments
18.
▲
California residents: call your legislators about AI bill SB 1047
(twitter.com)
22 points
by
convexstrictly
2y ago
|
11 comments
19.
▲
by
convexstrictly
2y ago
Candle is a minimalist ML framework for Rust with a focus on performance (including GPU support) and ease of use https://github.com/huggingface/candle
20.
▲
by
convexstrictly
2y ago
An author claims better performance than LoRA in 50% of the time. https://twitter.com/Rui45898440/status/1772996453557997924
21.
▲
LISA: Layerwise Importance Sampling for Memory-Efficient LLM Fine-Tuning
(arxiv.org)
3 points
by
convexstrictly
2y ago
|
1 comments
22.
▲
by
convexstrictly
2y ago
The federal government requests comments on regulation of AI models with openly available weights. The deadline is March 27, 2024. Earlier thread. https://news.ycombinator.com/item?id=39494760
23.
▲
NTIA AI Open Model Weights RFC
(regulations.gov)
1 points
by
convexstrictly
2y ago
|
1 comments
24.
▲
Mechanics of Next Token Prediction with Self-Attention
(arxiv.org)
1 points
by
convexstrictly
2y ago
|
0 comments
25.
▲
Dive Deeper into Yi-9B
(huggingface.co)
1 points
by
convexstrictly
3y ago
|
0 comments
26.
▲
by
convexstrictly
3y ago
It would be good to know on which listings Amazon is the seller. Filtering by that criterion may also be useful. For professional cards, I've noticed dihuni.com has good prices. I have never purchased from them and have no idea what
27.
▲
You can now train a 70B language model at home
(answer.ai)
4 points
by
convexstrictly
3y ago
|
1 comments
28.
▲
Shape Suffixes – Good Coding Style (2024)
(medium.com)
2 points
by
convexstrictly
3y ago
|
0 comments
29.
▲
by
convexstrictly
3y ago
The results are from the paper The Unreasonable Effectiveness of Eccentric Automatic Prompts https://arxiv.org/abs/2402.10949
30.
▲
Star Trek prompt optimal for grade school math on Llama-70B
(twitter.com)
2 points
by
convexstrictly
3y ago
|
1 comments
More ›