Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ammar_x
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
ammar_x
4mo ago
Is there some sort of a leaderboard for this test? Like if you'd give each of Opus 4.8 and GPT 5.5 a score out of 100, what would the scores be?
2.
▲
by
ammar_x
4mo ago
Absolutely! We need new and better benchmarks like this. I have a question: why not use the maximum available reasoning on each LLM? For example, I see that Opus 4.7 at `max` reasoning but Sonnet 4.6 at `high`. Wouldn't it be a fairer
3.
▲
by
ammar_x
4mo ago
https://x.com/serenaa_ge/status/2059308400866111692
4.
▲
DeepSWE: A contamination-free benchmark for long-horizon coding agents
(deepswe.datacurve.ai)
67 points
by
ammar_x
4mo ago
|
20 comments
5.
▲
Xiaomi Mimo-v2.5 pricing is now permanently reduced
(twitter.com)
5 points
by
ammar_x
4mo ago
|
0 comments
6.
▲
by
ammar_x
4mo ago
I usually do this for complex features: - Opus 4.7 writes the code - I make GPT-5.5 in Codex to review it (given context) - I provide the review back to Opus and ask it to verify the review findings - Make Opus plan the fixes then execute t
7.
▲
by
ammar_x
4mo ago
Cool, but body font size is too small for comfortable reading!
8.
▲
by
ammar_x
4mo ago
I've used Brave Search and found it better than Google's in some cases
9.
▲
by
ammar_x
4mo ago
This looks great for quick audio operations without the need to use heavy apps. One question: I tried the "Fade In" effect; is there a way to control its timing (i.e. the part of the clip where the effect is applied) ?
10.
▲
by
ammar_x
4mo ago
You can use V4 Pro with Claude Code [1]. I tried it and it's impressive. [1]: https://api-docs.deepseek.com/quick_start/agent_integrations...
11.
▲
by
ammar_x
11mo ago
My "trick" was to divide things into batches (which can be big with LLMs with larger context sizes) and classify the items in each batch, then take the resulting categories from each batch and feed them into an LLM to group semant
12.
▲
by
ammar_x
11mo ago
Language support is not mentioned in the repo. But from the paper, it offers extensive multilingual support (nearly 100 languages) which is good, but I need to test it to see how it compares to Gemini and Mistral OCR.
13.
▲
by
ammar_x
11mo ago
Claude Skills seem to be the option that offers highest flexibility to add more capabilities at most simplicity. Better than MCP in my opinion. Hope it becomes a standard and get adopted by OpenAI and the rest of labs.
14.
▲
by
ammar_x
2y ago
Good question! I selected the edition with the smallest Goodreads ID¹ that has the publication date and cover photo available. If all editions don't have publication date nor cover photo, then we get the one with the smallest ID. And y
15.
▲
Exploring Goodreads data: Analysis of 10M books
(ammar-alyousfi.com)
4 points
by
ammar_x
2y ago
|
2 comments
16.
▲
by
ammar_x
2y ago
Hi Jeremy, congratulations for the launch. How does this compare to Dash? I've used Dash for many applications, so I'm wondering what are the advantages of FastHTML?
17.
▲
by
ammar_x
2y ago
Been looking for such a website to show weather for the whole year like this. Thanks for sharing.
18.
▲
by
ammar_x
2y ago
I have Raycast extensions for GPT and Claude models. Whenever I have a question, the most powerful LLMs in the world are two key strokes away. This way is easier than going to the browser then ChatGPT tab for example then creating a new cha
19.
▲
by
ammar_x
2y ago
The article compares GPT-4o to Sonnet from Anthropic. I'm wondering how Opus would perform at this test?
20.
▲
by
ammar_x
3y ago
How does it compare to Plotly Dash?
21.
▲
A study confirms: Big changes in GPT-4 performance since its launch
(twitter.com)
2 points
by
ammar_x
3y ago
|
0 comments
22.
▲
What's your favorite interface for GPT API?
1 points
by
ammar_x
3y ago
|
0 comments
23.
▲
by
ammar_x
3y ago
Can you explain more? Like which tool do you use for this wiki page? Or is it an internal tool? And do you use it to write meeting notes and then discuss on the same page?
24.
▲
Ask HN: What's your go-to platform for written discussions, and why?
10 points
by
ammar_x
3y ago
|
4 comments
25.
▲
by
ammar_x
3y ago
How is this different or better than Dash or Streamlit?
26.
▲
by
ammar_x
3y ago
Like many people here have noticed, it's definitely less quality now than before. It's annoying to be honest to reduce the quality significantly without a notice while we are paying the same amount. I'm willing to pay $40 for
27.
▲
by
ammar_x
3y ago
I couldn't find it on the app store, can you post the link?
28.
▲
Ask HN: Have you noticed decreased quality in GPT-4 reasoning recently?
4 points
by
ammar_x
3y ago
|
5 comments
29.
▲
Ask HN: What is the main communication channel at your remote company?
1 points
by
ammar_x
4y ago
|
0 comments
30.
▲
by
ammar_x
4y ago
Well, we have less than 2 TB of data, and although we are running MySQL on a large instance with ~120 GB of RAM, it's extremely slow when dealing with big tables (like a 25 GB table) and that's why we need "big data" too
More ›