Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
michaelbuckbee
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
michaelbuckbee
22d ago
Movie production (editors, sound design, and a portion of fx work).
2.
▲
by
michaelbuckbee
23d ago
I feel like the closest we've gotten to that so far has been the flight sims.
3.
▲
by
michaelbuckbee
24d ago
It also distorts the testing as it encourages non-typical behavior.
4.
▲
by
michaelbuckbee
25d ago
I'm very bullish on MCP (or at least MCP "like" implementations), as they solve a lot of problems for non-devs as they're much easier and safer to add into ChatGPT, Claude and other desktop + web apps. They provide a set
5.
▲
by
michaelbuckbee
28d ago
Nice. Happy to be mistaken on that point as that seems very useful.
6.
▲
by
michaelbuckbee
28d ago
I've been using OpenRouter since relatively early (I build evvl.ai - an eval platform on top of it) and think that they're selling at a good time. An underpinning of their model is that API calls / inference are similar acros
7.
▲
by
michaelbuckbee
29d ago
Mostly it seems like greatly increased usage pushing their systems to the limit.
8.
▲
by
michaelbuckbee
29d ago
https://www.githubstatus.com/history
9.
▲
by
michaelbuckbee
1mo ago
The eval world is split into: 1. Long form task based examinations like this that test the ability of the model+harness to remain on task, tool calling, overall effectiveness and taste. 2. More direct 1:1 and qualitative comparisons that yo
10.
▲
by
michaelbuckbee
1mo ago
Since most analytics is done with JS (Google Analytics, etc.) very little of this shows up in site visit stats.
11.
▲
by
michaelbuckbee
1mo ago
Parent of a middle and a high schooler and we try very hard to not pressure the kids, but a lot of it is coming from their peers and the school.
12.
▲
by
michaelbuckbee
1mo ago
The 2 vs 3 days makes less of an impact on personal decision making but has massive benefits for decision making at the country wide response level.
13.
▲
by
michaelbuckbee
2mo ago
To your slow point: I did a quick eval for a data viz task to do a qualitative comparison and Fable was much faster (by 4x), but I struggle with the "token inefficiency" as a sort of whatever metric. Kimi was half the cost and pro
14.
▲
by
michaelbuckbee
2mo ago
It's kind of ridiculous how good these are getting. 3.5 Flash lite is pretty comparable to Opus 4.8 (at least for the couple tests I did) while simultaneously being 6x faster and 19x cheaper. https://fy2zp1ri90.evvl.io/
15.
▲
by
michaelbuckbee
2mo ago
It's a free site, so I was trying to limit both the privacy and risk exposure. Making the content auto expire after a short period of time greatly decreases the attractiveness of the site to lots of SEO spammers and other types of abus
16.
▲
by
michaelbuckbee
2mo ago
Like Simon concludes the article, the main use of this isn't to say which model is "better", but to try and poke at the model to sort out things like quality vs cost vs speed. So I put together a quick comparison of the last
17.
▲
by
michaelbuckbee
3mo ago
I built a simple (free) eval tool for my own uses (Github Gists + Model Outputs) after not being able to find a suitable one in the market. The market's being split into 1. Longitudinal LLM observability tooling Most eval startups have
18.
▲
by
michaelbuckbee
3mo ago
Vesting schedule?
19.
▲
by
michaelbuckbee
3mo ago
Something that's improved my life has been buying a sticker sheet of those LED darkening dots. They're only a couple bucks and look much cleaner than other solutions I've tried while still allowing for _some_ light to come th
20.
▲
by
michaelbuckbee
3mo ago
This, more than anything else I'd ever read about Inform, really makes me want to give it a try.
21.
▲
by
michaelbuckbee
3mo ago
I ran a quick eval to see what this looks like qualitatively vs just calling Opus 4.7 or GPT 5.5 directly. As expected, Fusion was 7x slower and 4x the cost. This isn't a knock against it, just that it I think this places Fusion into a
22.
▲
by
michaelbuckbee
3mo ago
A counterpoint to this is that we have some real different definitions of AI. If you consider things like the machine learning filters in your smartphone camera and Google's AI Overviews for searches it's entirely plausible that t
23.
▲
by
michaelbuckbee
3mo ago
I think it's your last point that's actually the strongest. There's always gaps between theoretical and practical, but to see China investing so hard in the future while the US digs in it's heels is infuriating.
24.
▲
by
michaelbuckbee
3mo ago
I thought it was more implied, but let me be more explicit: - This is something I made for myself without a lot of commercial thought, so I still haven't thought through pricing + usage + limits + operational limits. In it's curre
25.
▲
by
michaelbuckbee
3mo ago
This is all very fair criticisms. This thread asked: "What are tools you have made for yourself?" and that's genuinely what this is. I wanted it so I made it and then AI makes it so easy to just throw up a marketing page. I&#
26.
▲
by
michaelbuckbee
3mo ago
The funniest thing I've made is a free utility called "Moniker" that contextually renames files based on their contents. Uses local AI models and I was able to snag this great domain name. https://finalfinalreallyf
27.
▲
by
michaelbuckbee
4mo ago
It's not just comparing all the models, it's also comparing all the providers and configurations of those models. If you're doing any kind of production AI work you'll end up with outages caused by calling a single provi
28.
▲
by
michaelbuckbee
4mo ago
WRT the native grammar, consider adjective order. Few native english speakers (me included) can off the cuff name the proper order, but everyone knows the "right" order. https://dictionary.cambridge.org/us/gra
29.
▲
by
michaelbuckbee
4mo ago
Assassin's Creed Brotherhood is kind of like that for the architecture and period it covers. There's an interesting small YT channel that did a series on ACB + History https://www.youtube.com/watch?v=hebq-fObdhY
30.
▲
by
michaelbuckbee
4mo ago
That's a a lot less than I expected. Is it difficult to get coverage?
More ›