Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aliljet
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
aliljet
13d ago
To be clear, this is a hybrid aircraft, but it's still an awesome step forward!
2.
▲
by
aliljet
13d ago
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
3.
▲
by
aliljet
13d ago
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
4.
▲
by
aliljet
14d ago
This is probably Astra getting ready for release.
5.
▲
by
aliljet
1mo ago
Does diagnosis offer a better prognosis for those that are infected?
6.
▲
by
aliljet
1mo ago
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels l
7.
▲
by
aliljet
1mo ago
Accuracy is truly what people die for in the OCR game. Price isn't the primary function here.. it's an equation of price, accuracy, speed, and in mayn cases regulation.
8.
▲
by
aliljet
1mo ago
Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...
9.
▲
by
aliljet
1mo ago
What genuinely disappointing result. Long time pixel user here and I've been routinely buying these phones with the argument that you're getting the most value of any modern smart phone. Now? I'm just waiting for the pixel 10
10.
▲
by
aliljet
1mo ago
Is there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
11.
▲
by
aliljet
1mo ago
There is a more serious question in here that's not being answered. How effective is the retrieval in finding buried needles in larger and larger haystacks. And there's a correlary question, how effective could you be in finding p
12.
▲
by
aliljet
2mo ago
I'm trying and failing to find value running a potential Qwen 3.8 27b dense model on a 16 core, 128 GB of ram, 2080ti box. Yes, the GPU yells for help, but the problem is that no math works to upgrade this machine even when pouring $20
13.
▲
by
aliljet
2mo ago
Honestly, I have a 2080ti that I use to play and I can tell you the math isn't there to upgrade it. It's much easier to just find a 3090/4090/5090 and keep pace with the software and hardware simultaneously.
14.
▲
by
aliljet
2mo ago
I'd be curious about an.option that would allow glm use with a low end GPU like a 2080 ti...
15.
▲
by
aliljet
3mo ago
The benchmarks here are confusing at best. Am I reading correctly that this model is essentially as good or better than all frontier models right now?
16.
▲
by
aliljet
3mo ago
I was just using infinity parser 2 (flash, to be fair) for pennies self-hosted to run through thousands of pages of documents with remarkable confidence. I decided to use https://huggingface.co/datasets/allenai/olm
17.
▲
by
aliljet
3mo ago
I'm curious about this. What models/tools have you been using?
18.
▲
by
aliljet
3mo ago
How does this compare with infinty parser 2 which seemed to be running the table on every other OCR tool ( https://huggingface.co/datasets/allenai/olmOCR-bench ). To be fair, there's no single winning OCR bench
19.
▲
by
aliljet
3mo ago
This sounds incredible. Have these models effectively solved the problem of trying to use a fast-processing network to predict the world's state ahead? For example, to catch a ball?
20.
▲
by
aliljet
3mo ago
The problem here is always the cost-benefit. For $200/mo, you're receiving subsidized best of breed access. There's no model competing for that price anywhere. If a 27B param model is what you choose, show me your hardware! I
21.
▲
by
aliljet
4mo ago
Is this just one giant marketing plot?
22.
▲
by
aliljet
4mo ago
Where can a user reasonably host this in an affordable way to access the local LLM revolution?
23.
▲
by
aliljet
4mo ago
I'm really running into this deep at the edges of content creation. Take, for example, a need to general some kind of legal work. The cost of painstakingly checking and rechecking each case cited is reducing the value of these frontier
24.
▲
by
aliljet
4mo ago
Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context len
25.
▲
by
aliljet
5mo ago
This is so cool. I would love to revitalize a generation of great, but perhaps boring older cars with FSD. Just so much work...
26.
▲
by
aliljet
5mo ago
Why did Spirit die? Was there any last of this that had to do with their abysmal customer service?
27.
▲
by
aliljet
5mo ago
What systems are you actively using? And what systems have you tried? It seems like law, generally, may be hitting a tipping point on LLM use...
28.
▲
by
aliljet
5mo ago
This is a tough moment. Claude is simultaneously becoming substantially more expensive, substantially less reliable (single 9 of reliability), and substantially less performant. It's really hard to justify the cost of a subscription ov
29.
▲
by
aliljet
5mo ago
I wonder how this kind of response from Anthropic is actually being read by the community at large. If you consider the rough sentiment of the r/ClaudeCode subreddit against the r/Codex subreddit, you can see that there is a defin
30.
▲
by
aliljet
5mo ago
Why is this being made public?
More ›