Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sjmaplesec
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
Haiku 4.5 + skills outperforms Opus 4.7. 9 models tested with and without skills
(tessl.io)
4 points
by
sjmaplesec
5mo ago
|
2 comments
2.
▲
by
sjmaplesec
5mo ago
Full set of models tested: claude-opus-4-7 claude-opus-4-6 claude-sonnet-4-6 claude-haiku-4-5 gpt-5.4 gpt-5.3-codex gpt-5-codex cursor-composer-2 11 Skills used were here: https://github.com/mcollina/skills
3.
▲
Roast My Skill
(roast-my-skill.vercel.app)
3 points
by
sjmaplesec
6mo ago
|
1 comments
4.
▲
by
sjmaplesec
6mo ago
Pass in a skill and it'll roast the contents: "This is not just useless - it is an insult to the very concept of functionality."
5.
▲
by
sjmaplesec
6mo ago
I ran some evals to see which Anthropic models use skills the best, between Opus, Sonnet and Haiku. I was pretty impressed how good Haiku was with skills at completing various tasks
6.
▲
Claude Code model comparison: Skill usage
(tessl.io)
1 points
by
sjmaplesec
6mo ago
|
1 comments
7.
▲
by
sjmaplesec
7mo ago
Link to all the review scans is here - mostly in the 50-70% range https://tessl.io/registry/skills/github/googleworkspace/cli
8.
▲
Googleworkspace/CLI isn't optimized – Test your skills
(tessl.io)
1 points
by
sjmaplesec
7mo ago
|
2 comments
9.
▲
by
sjmaplesec
7mo ago
There's so much more we can do around activation and skills creation. Looking at the eval results, there are even cases where the context makes the agent worse. Scenario 5, test 1 72% -> 22% https://tessl.io/eval-run
10.
▲
by
sjmaplesec
7mo ago
The review eval tests language, activation etc of skills. I guess you could move it all to a skill quick and then run an eval on that if using Tessl. This checks if the way you write the instructions etc are being well understood by the age
11.
▲
by
sjmaplesec
7mo ago
An eval is to an LLM as a test is to code.
12.
▲
by
sjmaplesec
7mo ago
Tessl can generate the evals, both to test anthropic best practices as well as running scenarios with and without the skill to check if it's helping
13.
▲
by
sjmaplesec
7mo ago
Can add this as a skill or as part of a skill, and so you don't need to keep prompting the same things.
14.
▲
by
sjmaplesec
7mo ago
No, the context can be human created as much as it could be llm generated. The suggestions are based on Anthropic best practices and allow the agents to activate, and use the skills better, make the text clearer for the agent etc.
15.
▲
Agents.md file isn't the problem. Your lack of Evals is
(tessl.io)
42 points
by
sjmaplesec
7mo ago
|
17 comments
16.
▲
by
sjmaplesec
7mo ago
This resonates with my experience: we have dozens of internal “playbooks” and prompt snippets floating around, and nobody knows which ones still work after model changes. If you can make “skill quality” visible over time (regressions, drift
17.
▲
by
sjmaplesec
1y ago
This is so true!
18.
▲
Datadog CEO on AI and Observability
(ainativedev.io)
2 points
by
sjmaplesec
1y ago
|
4 comments
19.
▲
by
sjmaplesec
1y ago
I'm pretty sure Olivier Pomel rarely does podcasts, but this was a pretty good one. Some of my thoughts: - Customers "lie to themselves" saying they prefer noise to missed issues, when in practice 2 false alarms make them los
20.
▲
by
sjmaplesec
2y ago
Page title: Armon Dadgar, Hashicorp co-founder, on AI Native DevOps: Can AI shape the future of Autonomous DevOps workloads? Link is to an interesting podcast episode about Gen AI being used in infra
21.
▲
AI with IaC – "80% of the value is in codifying 20% of the assumptions"
(tessl.io)
2 points
by
sjmaplesec
2y ago
|
1 comments
22.
▲
Snyk founder, Guy Podjarny, creates new AI Native startup, Tessl
(tessl.io)
2 points
by
sjmaplesec
2y ago
|
1 comments
23.
▲
SourMint Malicious SDK
(snyk.io)
102 points
by
sjmaplesec
6y ago
|
44 comments
24.
▲
by
sjmaplesec
6y ago
Totally - Sounds like they started as a legitimate Ad SDK too!
25.
▲
by
sjmaplesec
6y ago
Amazing that this has been going for a year now - Let's see how Apple deal with the existing apps on the AppStore.
26.
▲
Snyk Closes $150M to Accelerate Developer-First Security
(snyk.io)
7 points
by
sjmaplesec
7y ago
|
0 comments
27.
▲
npm passes the 1 Millionth package milestone!
(snyk.io)
5 points
by
sjmaplesec
7y ago
|
0 comments
28.
▲
Largest JVM Ecosystem Survey Report
(snyk.io)
1 points
by
sjmaplesec
8y ago
|
0 comments
29.
▲
O11ycast Ep. #3, Distributed Systems with Paul Biggar of Dark
(heavybit.com)
2 points
by
sjmaplesec
8y ago
|
0 comments