Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mikeknoop
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
mikeknoop
5mo ago
Fun memory trip. Learned assembly on those old Z80s in middle school. I had to go re-dig up SafeGuard, a program I made by reverse engineering TI's TestGuard, to stop admins from wiping your calculator memory and all your games! https
2.
▲
Ndea (YC W26) is hiring a symbolic RL search guidance lead
(ndea.com)
1 points
by
mikeknoop
6mo ago
3.
▲
by
mikeknoop
2y ago
One must now ask whether research results are analyzing pure LLMs (eg. gpt-series) or LLM synthesis engines (eg. o-series, r-series). In this case, the headline is summarizing a paper originally published in 2023 and does not necessarily ha
4.
▲
by
mikeknoop
2y ago
I think we agree; to clarify, sharp messaging isn't inaccurate messaging. And I believe the story is not overhyped given the evidence: the benchmark resisted a $1M prize pool for ~6 months. But I concede we did obsess about the story t
5.
▲
by
mikeknoop
2y ago
Correct, fine-tuning is not new. It's long been used to augment foundational LLMs with private data. Eg. private enterprise data. We do this at Zapier, for instance. The new and surprising thing about test-time training (TTT) is how ef
6.
▲
by
mikeknoop
2y ago
> I'd heartily recommend maybe taking down the marketing vibrance down a notch and keep things a bit more measured, it's not entirely a meme, though some of the more-serious researchers don't take it as seriously as a resu
7.
▲
by
mikeknoop
2y ago
Author here -- six months ago we launched ARC Prize, a huge $1M experiment, to test if we need new ideas for AGI. The ARC-AGI benchmark remains unbeaten and I think we can now definitely say "yes". One big update since June is tha
8.
▲
by
mikeknoop
2y ago
Context: ARC Prize 2024 just wrapped up yesterday. ARC Prize's goal is to be a north star towards AGI. The two major categories of this year's progress seem to fall into "program synthesis" and "test-time fine tunin
9.
▲
by
mikeknoop
2y ago
I met my Zapier co-founder bryanh through HN 15 years ago when someone made a similar service to OP called "hacker newsers". We were the only two people in Missouri at the time which led to a meetup. https://news.ycombi
10.
▲
by
mikeknoop
2y ago
I personally am slightly surprised at o1's modest performance on ARC-AGI given the large leaps in performance on other objectively hard benchmarks like IOI and AIME. Curiosity is the first step towards new ideas. ARC Prize's whole
11.
▲
by
mikeknoop
2y ago
I bet pretty well! Someone should try this. It's likely expensive but sampling could give you confidence to keep going. Ryan's approach costs about $10k to run the full 400 public eval set at current 4o prices -- which is the arbi
12.
▲
by
mikeknoop
2y ago
Author here. Which aspects are misleading? How can it be improved?
13.
▲
by
mikeknoop
2y ago
High efficiency "search" is necessary to reach AGI. For example, humans don't search millions of potentially answers to beat ARC Prize puzzles. Instead, humans use our core experience to shrink the search space "intuitiv
14.
▲
by
mikeknoop
2y ago
ARC isn't perfect and I hope ARC is not the last AGI benchmark. I've spoken with a few other benchmark creators looking to emulate ARC's novelty in other domains, so I think we'll see more. The evolution of AGI benchmark
15.
▲
by
mikeknoop
2y ago
(ARC Prize co-founder here). Ryan's work is legitimately interesting and novel "LLM reasoning" research! The core idea: > get GPT-4o to generate around 8,000 python programs which attempt to implement the transformation, s
16.
▲
by
mikeknoop
2y ago
Yes there is a secondary leaderboard called ARC-AGI-Pub (in beta) with no limitations: https://arcprize.org/leaderboard
17.
▲
by
mikeknoop
2y ago
(You can direct link to a task like this: https://arcprize.org/play?task=009d5c81 in case you want to share!)
18.
▲
by
mikeknoop
2y ago
Here is some published research on the human difficulty of ARC-AGI: https://cims.nyu.edu/~brenden/papers/JohnsonEtAl2021CogSci.p... > We found that humans were able to infer the underlying program and generate
19.
▲
by
mikeknoop
2y ago
That is correct for ARC Prize: limited Kaggle compute (to target efficiency) and no internet (to reduce cheating). We are also trialing a secondary leaderboard called ARC-AGI-Pub that imposes no limits or constraints. Not part of the prize
20.
▲
by
mikeknoop
2y ago
I agree, $1M is ~trivial in AI. The primary goal with the prize is to raise public awareness about how close (or far today) we are from AGI: https://arcprize.org/leaderboard and we hope that understanding will shift more wo
21.
▲
ARC Prize – a $1M+ competition towards open AGI progress
(arcprize.org)
588 points
by
mikeknoop
2y ago
|
337 comments
22.
▲
by
mikeknoop
3y ago
(Zapier co-founder) Perhaps the least known feature of Zapier is that you can use the dev platform ( https://zapier.com/developer ) to make your own private apps for any API (including internal/private services) and then
23.
▲
by
mikeknoop
3y ago
Couldn't find your contact info. Email me?
24.
▲
by
mikeknoop
3y ago
Very cool. Congrats on getting this launched @gogwilt!
25.
▲
by
mikeknoop
3y ago
We have a joke about this at Zapier -- don't be an oauth butt! "We support standard oauth butttt..."
26.
▲
by
mikeknoop
3y ago
(Zapier cofounder) Super excited for this. Tool use for LLMs goes way beyond just search. Zapier is a launch partner here -- you can access any of the 5k+ apps / 20k+ actions on Zapier directly from within ChatGPT. We are eager to see
27.
▲
by
mikeknoop
3y ago
Agree, will make the landing page better. Would recommend checking out this LangChain notebook: https://github.com/hwchase17/langchain/blob/master/docs/modu...
28.
▲
by
mikeknoop
3y ago
We just added a "Gmail: Create Draft Reply" action and by default the NLA API will auto-guess the appropriate thread. Share feedback if you see this not working as expected as it's new!
29.
▲
Show HN: Zapier's first API
235 points
by
mikeknoop
3y ago
|
32 comments
30.
▲
by
mikeknoop
4y ago
This seems targeted at adding friction for alternative app stores eg. https://altstore.io/
More ›