Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dial481
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
dial481
3mo ago
This is badly needed. Jira is horrendous. Can your plugin system support custom workflow triggers yet?
2.
▲
Show HN: Allelix – Annotate your DNA against 7 public databases offline
(github.com)
3 points
by
dial481
3mo ago
|
0 comments
3.
▲
by
dial481
5mo ago
MemPalace's benchmark claims have been picked apart and debunked in the project's GitHub issues: github.com/milla-jovovich/mempalace/issues/27 github.com/milla-jovovich/mempalace/issues/29 g
4.
▲
Show HN: Proposal for a real long-term AI memory benchmark
(penfieldlabs.substack.com)
4 points
by
dial481
5mo ago
|
0 comments
5.
▲
Milla Jovovich's MemPalace Claims 100% on LoCoMo. Its Benchmarks.md Disagrees
(penfieldlabs.substack.com)
4 points
by
dial481
5mo ago
|
0 comments
6.
▲
by
dial481
6mo ago
That's encouraging to hear from someone with IR experience, thanks. Agree completely.
7.
▲
LoCoMo AI Benchmark: 6.4% of answer key wrong, judge accepts 63% of fake answers
(github.com)
3 points
by
dial481
6mo ago
|
3 comments
8.
▲
by
dial481
6mo ago
We audited the LoCoMo benchmark (one of the most cited eval for LLM agent memory) and found 99 score-corrupting errors in 1,540 questions (6.4%). Separately, we tested the LLM judge with adversarially generated wrong answers, it accepted 62