Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kimjune01
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
kimjune01
14d ago
deepswe is public and can be considered contaminated.
2.
▲
by
kimjune01
1mo ago
beam search works when you have enumerable branches or some predefined dimension
3.
▲
by
kimjune01
1mo ago
Location: Vancouver, Canada Remote: Yes Willing to relocate: Yes for research Technologies: Python, TypeScript, Go, Rust, C++; LLM and coding agents, evaluation harnesses and graders, benchmark design and auditing, tool use/MCP, RAG an
4.
▲
by
kimjune01
1mo ago
unfortunately, knowing about ethics doesn't compel you to ethical behavior
5.
▲
by
kimjune01
1mo ago
ah, the good old days where everything you needed to know was all signal no noise
6.
▲
by
kimjune01
1mo ago
if the recipient can't tell, does it matter?
7.
▲
by
kimjune01
1mo ago
would you prefer that they advertised? Or is your life good enough to not need to learn about new offers?
8.
▲
by
kimjune01
1mo ago
an audit for the bench for the curious: https://www.june.kim/terminal-bench-frame
9.
▲
by
kimjune01
1mo ago
working on building a bench for human ICs for hiring https://june.kim/human-signal
10.
▲
by
kimjune01
1mo ago
for throughput, you can use a stopwatch and a big movie transfer from one end to another
11.
▲
by
kimjune01
1mo ago
It only takes a small minority to ruin it for the rest of us. for example, bike theft
12.
▲
by
kimjune01
1mo ago
Vancouver dream, not achievable with a median income
13.
▲
by
kimjune01
1mo ago
would you let taste make withdrawls from your bank account
14.
▲
by
kimjune01
1mo ago
Thinking that harnesses and agents can improve without human involvement, even if it works, will yield a much lower growth rate than if a human gets involved in the loop.
15.
▲
by
kimjune01
2mo ago
it should be entirely acceptable to filter out factually incorrect or unverifiable submissions without human intervention
16.
▲
by
kimjune01
2mo ago
paper submissions require LLM disclosure
17.
▲
by
kimjune01
2mo ago
nowhere near plateau, but right at the inflection point of diminishing returns imo
18.
▲
by
kimjune01
2mo ago
i think it's a rite of passage to have attempted encoding thinking and the scientific process for AI/ML researchers
19.
▲
by
kimjune01
2mo ago
the issues were already validated by the maintainers, and the maintainers merged it into their repo voluntarily. When maintainers accepted the PRs, they were the ones who found it acceptable. I found that well-tested PRs are more likely to
20.
▲
by
kimjune01
2mo ago
actor paradigm is huge for clamping down on agents
21.
▲
by
kimjune01
2mo ago
as the proof of verification decreases, the value of credentials that act as shortcut proofs of human competence will decrease, too.
22.
▲
by
kimjune01
2mo ago
i learned that no matter how good i am at writing prose, it doesnt matter if nobody reads it. so i write for AI agents instead, hoping that it'll get picked up by an agent and find propagation that way
23.
▲
by
kimjune01
2mo ago
for small to medium sized bugs I managed to get a bit more than half of my PRs to get merged. ~90 PRs since May https://june.kim/speedrunning-open-source Verify yourself: { merged: search(query: "is:pr is:merged author
24.
▲
by
kimjune01
2mo ago
You can actually run these benches yourself, as Frontier-Bench is open source. Also have a look at these other coding benchmarks I audited. Frontier-Bench v0.1: all 74 tasks grade in a container brought up after the agent's is destroye
25.
▲
by
kimjune01
2mo ago
the bottleneck for these open source repos is maintainer attention. AI seemingly does not yet improve that throughput.
26.
▲
by
kimjune01
2mo ago
one time i told a google voice to kill itself and i could never get it to work again
27.
▲
by
kimjune01
2mo ago
someone's promotion depends on benching harder
28.
▲
by
kimjune01
2mo ago
it's not just social science, even AI papers that would be trivial to replicate dont
29.
▲
by
kimjune01
2mo ago
anyone else notice that the topline numbers are effort xhigh? anybody actually use the models at those levels?
30.
▲
by
kimjune01
2mo ago
end the fed?
More ›