Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tudelo
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
tudelo
11d ago
Writing is fluff. What do you want, a list of bullet points? There would be no books. Maybe you have a career in writing 2 page books?
2.
▲
by
tudelo
11d ago
You ever wonder why recent claude models speak in riddles? I dunno, maybe all those "rare" books? They may have been rare for a reason.
3.
▲
by
tudelo
11d ago
Well, I can tell what sort of angle you most enjoy. Anyways -- I think there is GOOD writing and BAD writing, but only subjectively. So if you enjoy it, power to you. It's certainly not random, but it is the sort of verbosity that turn
4.
▲
by
tudelo
11d ago
Like every downtown in the US. In short: not good.
5.
▲
by
tudelo
1mo ago
Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually representing? The best run? The average? It is all in some ways a falsehood
6.
▲
by
tudelo
1mo ago
I don't think it is vague in the slightest. Take the most simple examples, how many LLM's have you tested making them? There are stylistic choices pertaining to games that is well beyond a 0/1 reward. Even something as basic
7.
▲
by
tudelo
1mo ago
Your first failure was trying to get claude to do anything :)
8.
▲
by
tudelo
1mo ago
There is, in my personal opinion, a reason that reasoning is not front and center. Partially related to distillation.
9.
▲
by
tudelo
1mo ago
I have been in RDP land but I would never go back unless required (windows). While it could just be personal experience/failures, latency was not good and reliability was not good. Things are probably better as of late though I hesitat
10.
▲
by
tudelo
1mo ago
From what I understand pre-training is totally irrelevant to this and as far as post training goes there will be multiple steps, for claude and codex and the like that ship with a harness, the harness is definitely included in evaluation. H
11.
▲
by
tudelo
2mo ago
Tmux is good because: 1) sessions stay running if you disconnect 2) window management - you can split your screen, have an octobox a-la redzone style, and focus/unfocus etc. This is based on developing on a remote server - but even loc
12.
▲
by
tudelo
2mo ago
I just wanted to echo your comment. Most people, in my experience, pursue an undergrad program with the intention of being employable. This is my personal experience, but also a learned opinion from being a teaching assistant for some year
13.
▲
by
tudelo
2mo ago
It is RLVR, Not a puzzle, Not leetcode
14.
▲
by
tudelo
2mo ago
> Only systems which required less than $10,000 to run are shown. (Notes[1]) Am I lost or are their many models on this ranking (Opus 5 included) that clear this?
15.
▲
by
tudelo
2mo ago
I honestly just use GPT models nowadays, Claude models are too restrictive and more of a quitter and fable/whatever is just too expensive to be worth it.
16.
▲
by
tudelo
2mo ago
I find it useful for code reviews (spawn a subagent with minimal/no context to review X commit). Of course, this is more or less a shortcut that could be done with a seperate agent. Another use is multiple reviews at once if tokens are
17.
▲
by
tudelo
2mo ago
It would make sense. Massive distribution vector
18.
▲
by
tudelo
2mo ago
A degree is not a bad thing. This forum is pretty biased on startup culture but I bet the vast majority would say its not worth it personally but worth it on a career level. Even then, the space to explore things outside of your immediate i
19.
▲
by
tudelo
2mo ago
I echo the sentiment. Most work is described as basic and unimaginative, yet we still have every large company having outages despite employing "the best". Even worse, they game uptime and outages in a way that mirrors gerrymander
20.
▲
by
tudelo
3mo ago
Most likely, this (reverse engineering) is one of the numerous things these LLM companies target. You can also assume all of the internet has been slurped up in to any frontier model. That doesn't mean what you want will be a one shot
21.
▲
by
tudelo
3mo ago
> For a significantly shorter critque of the book, check out qntm's critique. I mostly agree with qntm assessment. But it's a bit too emotional and personal and doesn't cover the parts i find the most harmful. This page se
22.
▲
by
tudelo
3mo ago
It is extremely easy to burn tokens if that is required. Explore this codebase. Team x wants y feature, research and generate a full plan. What does feature x in codebase y actually mean? Analyze code coverage in x. Map out code flow and
23.
▲
by
tudelo
3mo ago
You can use images inline in Racket. Decidedly less esoteric :)
24.
▲
by
tudelo
3mo ago
This is tangential and offtopic but kirkland beef hotdogs are 10/10 for value.
25.
▲
by
tudelo
3mo ago
People will reply to you calling you crazy, but SF/bay is the only place I have ever experienced where many people will literally leave their cars unlocked because a broken window isn't worth the hassle. Yes, locking your parked c
26.
▲
by
tudelo
4mo ago
It's interesting... Opus seems horrible at keeping text aligned. Markdown it is I suppose
27.
▲
by
tudelo
4mo ago
The bolded quote "It’s harder to read code than to write it." is hilarious given todays context... it has only become more true :)
28.
▲
by
tudelo
4mo ago
Most of my work has been in core infra at large companies. Having the code written faster does not change rollout velocity all that much... It does help with signals and idiot proofing on bugs but when things break and cost real (very real)
29.
▲
by
tudelo
4mo ago
I seriously doubt it. Degradation would be in some part related to the conditions the painting was held in, which would be nearly impossible to backtrack outside of one-off case studies. Imagine a painting that was stuck in a room full of s
30.
▲
by
tudelo
6mo ago
I mean if you don't have your company paying for it I wouldn't bother... We are talking sessions of 500-1000 dollars in cost.
More ›