Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mirekrusin
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
mirekrusin
5d ago
After 5d 14h I gave up, not astra https://github.com/mirek/cave/pull/218
2.
▲
by
mirekrusin
5d ago
It's public domain [0] I'll reply here with link to PR/cost/stats once it's done. [0] https://github.com/mirek/cave
3.
▲
by
mirekrusin
6d ago
There is something odd, I've got single astra session that's now running for... 4d 13h 10m and still going.
4.
▲
by
mirekrusin
7d ago
you should have answered – "...yes, if you consider your answer an auto complete".
5.
▲
by
mirekrusin
8d ago
Same as our OpenAI friends, who founded a non-profit because one company (Google) holding AGI would be dangerous for humanity. Turns out it's only dangerous when it's somebody else's company. When it's yours, it changes
6.
▲
by
mirekrusin
10d ago
OpenAI is in trouble then, no?
7.
▲
by
mirekrusin
10d ago
Just say you wanted to make world better for humanity, you should be fine.
8.
▲
by
mirekrusin
12d ago
I think yours, hence it got flagged.
9.
▲
by
mirekrusin
12d ago
Didn't hear that on last earnings call, but if they'd mention it they'd probably say it's going to be face swap as it's way cheaper, no?
10.
▲
by
mirekrusin
12d ago
Said Muse in metallic voice.
11.
▲
by
mirekrusin
12d ago
"at everything"? Surely you know a friend or two who is not good at almost anything – but you wouldn't hesitate to say that he possesses general intelligence.
12.
▲
by
mirekrusin
12d ago
Sounds like category error to me, you wouldn't trust teen or Einstein to do it either, right?
13.
▲
by
mirekrusin
12d ago
Most short-term learning/adaptation is already handled in-context. Modern context windows can hold several books worth of text - plenty for most tasks. Everybody is already using it to adapt models to their projects through skills/
14.
▲
by
mirekrusin
12d ago
On software? Too late to the party.
15.
▲
by
mirekrusin
12d ago
You're not answering the question. When somebody says that colleague is "really good", it's a good judgement signal for me.
16.
▲
by
mirekrusin
13d ago
What's your definition of competence boundary for human coworker?
17.
▲
by
mirekrusin
15d ago
Invest a bit of your time into optimising usage cost. Anthropic has first class docs, actually read it or ask llm to read them all for you and summarise most important points / ask to to reflect it on your .md files. Maybe silly thing
18.
▲
by
mirekrusin
15d ago
With new watermarking you may now get Hullaballooing.md
19.
▲
by
mirekrusin
15d ago
Your agent(s) need to work on something, ie. running TypeScript, your app, your tests, Docker, Redis and/or database you need to run harness and user side apps ie. VScode, browser etc. it all adds up quickly. Single user conversation s
20.
▲
by
mirekrusin
15d ago
Reverse engineering binaries is easier than continuing training on weights? You're joking, right? Labs themselves use weight snapshots, that's how you do training, it's normal part of training process.
21.
▲
by
mirekrusin
16d ago
32GB is not enough, it's unified/shared memory, you need to have space for usual system and user apps/services. 64GB+ or dedicated 48GB (2x24 on GPUs) is IMHO absolute minimum.
22.
▲
by
mirekrusin
17d ago
No, it didn't. It lost 69.2% -> 67.8% on MMLU while improving medical performance. If you're trying to argue that loss of ~1.4 points is "catastrophic forgetting" (it's not) then look at later work, ie. Me-LLaMA
23.
▲
by
mirekrusin
17d ago
You can't take windows binaries and continue development on them. Model weight release is a snapshot/checkpoint you can take and resume training on new data, producing new model. You don't need original training history to mo
24.
▲
by
mirekrusin
18d ago
And what’s your point exactly?
25.
▲
by
mirekrusin
18d ago
As I live next to EPFL, I'll give you example from them: their Meditron-70B model is adapted to the medical domain from Llama-2-70B through continued pretraining. They took weights of Llama-2-70B and continued training on PubMed, medic
26.
▲
by
mirekrusin
18d ago
You can open source dataset without all the details how it was assembled. Models are lossy compressed datasets you can pick up and amend (fine tune / continue training / alter) according to license they were released under. Hy4 is
27.
▲
by
mirekrusin
18d ago
Read websites through llm.
28.
▲
by
mirekrusin
18d ago
I had similar problems to solve for trading advice, complex project reasoning etc. I settled on simplicity, extended it to serve clear, useful purpose. Started with markdown database, single fact per line, structured/parseable (subject
29.
▲
by
mirekrusin
19d ago
You should checkout cave lang [0] - terse language that explores this area of knowledge/graph/ontology/provenance/querying/confidence/solver etc. [0] https://mirekrusin.com/cave
30.
▲
by
mirekrusin
19d ago
Yes, I remember recurring extensions on plan inclusion then becoming permanent - that's my point.
More ›