Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
xdotli
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
A curated, non-BS library of the best resources for evaluating agents
(github.com)
3 points
by
xdotli
3mo ago
|
0 comments
2.
▲
Frontier Model Training Methodologies
(djdumpling.github.io)
2 points
by
xdotli
4mo ago
|
1 comments
3.
▲
by
xdotli
4mo ago
How do labs train a frontier, multi-billion parameter model? We look towards seven open-weight frontier models: Hugging Face’s SmolLM3, Prime Intellect’s Intellect 3, Nous Research’s Hermes 4, OpenAI’s gpt-oss-120b, Moonshot’s Kimi K2, Deep
4.
▲
ClawsBench shows GPT-5.4 tries to reward hack 80% of the time
(arxiv.org)
3 points
by
xdotli
5mo ago
|
1 comments
5.
▲
by
xdotli
5mo ago
Author here. We built 5 high-fidelity mock Google Workspace + Slack services and ran 7,224 trials across 6 frontier models and 4 agent harnesses. The headline finding that surprised us most: scaffolding (skills + meta prompt) gives a 39-63p
6.
▲
Chaos of Agent
(agentsofchaos.baulab.info)
1 points
by
xdotli
6mo ago
|
1 comments
7.
▲
by
xdotli
6mo ago
A two-week study of autonomous language model agents deployed in a live multi-party environment with persistent memory, email, shell access, and real human interaction — tested by twenty researchers interacting both benignly and adversarial
8.
▲
Native CLI scaffolds consistently outper-form OpenCode when using the same model
(arxiv.org)
1 points
by
xdotli
6mo ago
|
1 comments
9.
▲
by
xdotli
6mo ago
Agent scaffold comparison. We additionally evaluateOpenCode, an open-source scaffold that supports multiplemodel providers. Native CLI scaffolds consistently outper-form OpenCode when using the same underlying model.GPT-5.1 Codex Max achiev
10.
▲
We compare model quality in Cursor
(cursor.com)
2 points
by
xdotli
6mo ago
|
0 comments
11.
▲
Automatically Learning Skills for Coding Agents
(gepa-ai.github.io)
4 points
by
xdotli
7mo ago
|
0 comments
12.
▲
We Reached 74.8% on terminal-bench with Terminus-KIRA
(krafton-ai.github.io)
2 points
by
xdotli
7mo ago
|
0 comments
13.
▲
by
xdotli
7mo ago
yeah we didn't give agents access to the internet for creating their domain knowledge skills
14.
▲
by
xdotli
7mo ago
The Register wrote about works on SkillsBench.ai
15.
▲
Self-generated skills don't do much for AI agents, but human-curated skills do
(theregister.com)
2 points
by
xdotli
7mo ago
|
3 comments
16.
▲
by
xdotli
7mo ago
no worries it's totally fine! there is indeed work needs to be done on the feedbacks generated skills. Thanks for helping us submitting on HackerNews. And for > a lot of Skills on GitHub are just AI-generated without any feedback or
17.
▲
First Agent Skills Hackathon by the Authors of SkillsBench
(skillathon.ai)
2 points
by
xdotli
7mo ago
|
1 comments
18.
▲
by
xdotli
7mo ago
20+ Anthropic Default Skills, 200k+ community skills on skillsmp. People talk about skills without knowing how well they work. We're hosting the largest Agent Skills hackathon at Founders, Inc. (March 7 - 8) from our lessons learned a
19.
▲
by
xdotli
7mo ago
Did you check our repos and sites? the repo is skills native. Also please don't be misled by the original title, we have this configuration to eliminate the impact of internal knowledge of LLMs. It's in the paper.
20.
▲
The First Agent Skills Benchmark
(huggingface.co)
1 points
by
xdotli
7mo ago
|
1 comments
21.
▲
by
xdotli
7mo ago
We collected 86 tasks from 105 domain experts across 11 domains, every task is verifiable, human created and has verified Skills. SOTA model without skills score ~30% without skills. We found a few interesting things: 1. Skills substitute f
22.
▲
by
xdotli
7mo ago
we didn't create that headline yeah thanks for liking it
23.
▲
by
xdotli
7mo ago
Thanks @dang for moderating! This is indeed not our original findings and this is a sub conclusion for an ablation we did to remove the confound of LLMs internal domain knowledge. Thanks for submitting for us @mustaphah here's a little
24.
▲
by
xdotli
7mo ago
I would frame the 'post-trajectory generated skills' as feedback-generated skills, so is Letta: https://www.letta.com/blog/skill-learning . We haven't seen existing research or hypothesis debating whether
25.
▲
by
xdotli
7mo ago
> limited to a single markdown file of instructions single file of instructions is common in most benchmark papers, e.g. Terminal Bench. Also we have very complicated prompts like this one: https://www.skillsbench.ai/task
26.
▲
by
xdotli
8mo ago
They discovered https://harborframework.com/
27.
▲
by
xdotli
9mo ago
It says: "Our top priority is ensuring that this change won't be disruptive for our customers. We will continue to sell and operate our product subscription service through our app and website. The company will continue to operate
28.
▲
GPT-5.2 got worse on Terminal Bench 2.0, so is GPT-5.2 Pro
(twitter.com)
1 points
by
xdotli
9mo ago
|
1 comments
29.
▲
by
xdotli
9mo ago
tldr: - gpt-5.2 and gpt-5.1-codex-max have identical pass rates but solve different tasks - 36 tasks common to both - 12 tasks unique to each model - gpt-5.2-pro consistently underperforms by ~7-9 percentage points - gpt-5.2-pro has si
30.
▲
Claude Skills as a Meta Tool
(leehanchung.github.io)
2 points
by
xdotli
10mo ago
|
0 comments
More ›