Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lwarfield
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
lwarfield
7d ago
I'm kinda surprised that the mixture of depths paper didn't come up here. It approaches the other direction of sometimes dropping layers: https://arxiv.org/abs/2404.02258
2.
▲
by
lwarfield
9d ago
The using a J lens is super cheap compared to inference. You basically add a single matrix multiply per layer. You probably wouldn't even notice the overhead in a good implementation. You should even be able to create a j lens from scr
3.
▲
by
lwarfield
9d ago
If the author would like, I self computed a j lens for the 27b version of the qwen model. I used it for my own exploration in this area, and can share it if you want.
4.
▲
by
lwarfield
15d ago
Same for me. Every single time I tried it got flagged. I think this will be my litnus test for if the safeguards are good enough for benign requests.
5.
▲
by
lwarfield
21d ago
Best museum in Seattle! It also feels a lot more "hands on", than most museums.
6.
▲
by
lwarfield
23d ago
\s Take my angry upvote!
7.
▲
by
lwarfield
23d ago
Personally I'm curious to the point of doing borderline LLM archecture research. This come from genuine curiosity, and not a want to use LLMs better. I haven't gotten much out of for using LLMs though. It makes me understand the s
8.
▲
by
lwarfield
1mo ago
I currently have fable organize a bunch of 5.6 sol agents when working on my personal projects. This makes me wonder if I should add something along the lines of "For tasks that involve visual analysis, have gemini 3.7 look at images g
9.
▲
by
lwarfield
1mo ago
I've always wondered why the industry relies on the giant monolithic system prompt. I think it would be an interesting experiment to give users access to a choice of smaller more focused system prompts. You could have a common core for
10.
▲
by
lwarfield
1mo ago
Yes it is: > Same model weights as Mythos 5, deployed with higher-coverage safeguards (see Section 4.5.2.2)
11.
▲
by
lwarfield
1mo ago
> 6.2 [Appendix redacted] > This appendix describes the criteria for our blocking bioclassifier exemption policy, and has been redacted from the public version of this report for security reasons. >6.3 [Appendix redacted] > Thi
12.
▲
by
lwarfield
1mo ago
> More capable than Mythos 5 in some areas, less capable in others; overall slightly more capable. This sounds like it might be a Mythos finetune for some specific task. EDIT: After reading some more reading, it looks like model 2 might
13.
▲
Layer Scope: How I used $20 of compute to make a new way to look at LLMs
(blog.lwarfield.dev)
5 points
by
lwarfield
1mo ago
|
1 comments
14.
▲
by
lwarfield
1mo ago
Hey hackernews, I'm a long time lurker and first time poster. I've been doing a lot of LLM research on my own, and a friend of mine mentioned I should do a writeup of a small project of mine. I mainly want to show that you don’t n
15.
▲
by
lwarfield
2mo ago
Its interesting seeing how many of these researchers became the heads of frontier labs!
16.
▲
by
lwarfield
2mo ago
For beginners I'd recommend the Welch Labs Illustrated Guide To AI if your not well versed in reading papers. Its a beautiful book that I've enjoyed going through. I'd recommend going through these papers after reading that
17.
▲
by
lwarfield
3mo ago
Do they state if they used an API endpoint without a system prompt, or were these done via prompting the currently existing chatbots with a system prompt? Without a system prompt, I'd imagine there would be more variance in answers.
18.
▲
by
lwarfield
3mo ago
There's also lowering the number of experts you run in MoE models.
19.
▲
by
lwarfield
4mo ago
>Yeah, it's kind of mind boggling that Ted Chiang (of all people!) can't imagine intelligence without a body. and the whole thing just begs a lot of questions. Damn, what a line! Another thing that bothered me with his baseline
20.
▲
by
lwarfield
5mo ago
This is some real "There is no claw in ba sing se" stuff.
21.
▲
by
lwarfield
5mo ago
Well that didn't take long.
22.
▲
by
lwarfield
5mo ago
I do a lot of lindyhop Swing dancing as a hobby. It's not innovative work, but I find it rewarding.
23.
▲
by
lwarfield
6mo ago
Damn, I need to going with my embeddings project. I've currently got a prototype for using embeddings (not gemini in my case) for making a game that's kinda reverse connections: collections.lwarfield.dev
24.
▲
by
lwarfield
7mo ago
I've always thought the really low bandwidth support they added a few years ago was to support the french subs. It matched all the requirements of VLF/ELF communications.
25.
▲
by
lwarfield
1y ago
I ran into this guy nere Interlaken 2 days ago! Had a nice long talk at the post office over how he was going to present this in Geneva. I heard that this is result is with little or no rejection drugs as well!
26.
▲
by
lwarfield
2y ago
When I broke a joint in my pinky a few years ago it was pretty easy to tell. Early on the range of motion was the limiting factor, and I'd move it back and forth as much as I could without any pain. After that I worked on strength in a