Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
forrestp
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
forrestp
1y ago
My understanding is that your levers are roughly better / more diverse embeddings or computing more embeddings (embed chunks / groups / etc) + aggregating more cosine similarities / scores. More flops = better search w&#
2.
▲
by
forrestp
2y ago
It's expensive in this field to verify other people's work. There are a few other papers in the last 3 years that have the same high-level idea but call the anchor tokens something different -- Gist tokens being the only one I per
3.
▲
by
forrestp
2y ago
Gradient.ai | SF Bay Area | Onsite/hybrid | Staff SWE | Senior SWE | Enterprise Account Executive Our vision is to power the future of enterprise automation. Gradient is a full stack AI platform that enables businesses to automate oper
4.
▲
by
forrestp
2y ago
Right. We are sleep deprived -- couldn't stop over the weekend. Please forgive the typos
5.
▲
by
forrestp
2y ago
All (training / evals / inference) performed on their L40s clusters. These machines are underrated but capable of serious work
6.
▲
by
forrestp
2y ago
We are training on top of llama 3. The 256k reasoning benchmarks are on the open LLM leaderboard. And re: token count: our copy was wrong -- it's pre-prepped copy for a model run that didn't pan out. Updating to correct number --
7.
▲
Gradient AI Releases 1M Context Llama 3 8B
(twitter.com)
80 points
by
forrestp
2y ago
|
29 comments
8.
▲
by
forrestp
2y ago
Our team @ https://gradient.ai/ has more checkpoints coming soon with longer context lengths.
9.
▲
Llama-3 8B Instruct 262k
(huggingface.co)
4 points
by
forrestp
2y ago
|
2 comments