Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jafitc
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
jafitc
2mo ago
"thanks to lobbying". if you use "for" it sounds like you are adressing the OP and he did the lobbying that you don't like.
2.
▲
by
jafitc
5mo ago
bigger change here might not be model quality, but debuggability. once you hide the reasoning, remove the knobs, and let the model choose its own effort, it gets much harder to tell whether the model got worse or just got harder to inspect.
3.
▲
by
jafitc
6mo ago
subprime mortgages sprinkled on top of prime ones, treated as prime ones. because they were printing money. subprime code sprinkled on the backbone of software we use everyday. because they are printing code. reckoning
4.
▲
by
jafitc
6mo ago
that's a great story!
5.
▲
by
jafitc
2y ago
I think you should consider trimming that file. Exclude movies with very low number of rating or potentially very low scores too. The long tail reduction would be significant
6.
▲
by
jafitc
3y ago
"People are really bad at understanding just how big LLM's actually are. I think this is partly why they belittle them as 'just' next-word predictors" https://nitter.net/jam3scampbell/status
7.
▲
by
jafitc
3y ago
Deepinfra Mixtral is $0.27 / M tokens as per their website
8.
▲
ChatGPT Website Updated With Long-Term Memory Feature (RAG)
1 points
by
jafitc
3y ago
|
0 comments
9.
▲
Which models are the Most Actively Liked on HuggingFace since inception?
(twitter.com)
1 points
by
jafitc
3y ago
|
0 comments
10.
▲
by
jafitc
3y ago
Important to note that this model excels in reasoning capabilities. But it was on purpose not trained on the big “web crawled” datasets to not learn how to build bombs etc, or be naughty. So it is the “smartest thinking” model in weight cla
11.
▲
by
jafitc
3y ago
Do you think the ISIS is bound by the words “non-commercial” in a license file when they have the source anyway? It was available even before this, all they changed is that law abiding citizens can put apps in the App Store and charge money
12.
▲
Small 100M transformer model nails 12x12 digit multiplication without COT
(twitter.com)
2 points
by
jafitc
3y ago
|
0 comments
13.
▲
by
jafitc
3y ago
This "vibe" check that it's even better than GPT-4 Turbo is not what its Elo rating shows on the Chatbot Arena based on not 1 but thousands of user votes. GPT-4 (Turbo) is in a league of its own still.
14.
▲
Mixtral 8x7B Above Gemini Pro – Chatbot Arena Leaderboard Updated
(huggingface.co)
2 points
by
jafitc
3y ago
|
1 comments
15.
▲
by
jafitc
3y ago
This is based on users choosing the better from 2 models at a time, and calculating an ELO rating from who-beats-who. BYOT - bring your own tests style. Gives a better picture of real-world performance and more robust against contamination.
16.
▲
by
jafitc
3y ago
from the announcement tweet: https://twitter.com/rasbt/status/1735293149965062476 --- So, we've been quietly building something new for running AI experiments and deploying models ... Our Lightning AI Studios
17.
▲
Lightning AI Studios – A persistent GPU cloud environment
(lightning.ai)
1 points
by
jafitc
3y ago
|
1 comments
18.
▲
My thoughts on Mistral's stellar rise
(twitter.com)
3 points
by
jafitc
3y ago
|
0 comments
19.
▲
by
jafitc
3y ago
Important note: Bing balanced mode (default) uses GPT 3.5 Only Precise and Creative modes use GPT-4 https://twitter.com/emollick/status/1732495030143549541 Also see: An Opinionated Guide to Which AI to Use: ChatGP
20.
▲
by
jafitc
3y ago
All I can say is it’s really fast
21.
▲
by
jafitc
3y ago
It’ll never be completely gone. But you’ll need it in less and less everyday scenarios and time goes on Just like we need to write less and less assembly by hand
22.
▲
by
jafitc
3y ago
OpenAI provides “instruct” version of their models (Not optimized for chat)
23.
▲
by
jafitc
3y ago
Isn’t that the first sentence?
24.
▲
by
jafitc
3y ago
Brain is just neurons and synapses at the end of the day. The whole universe might just be a stochastic swirl of milk in a shaken up mug of coffee. Looking at something under a microscope might make you miss its big-picture emergent behavio
25.
▲
by
jafitc
3y ago
The fact that the makers of such LLM make a post about it shows that they have incentive to cater to even these kind of use cases
26.
▲
by
jafitc
3y ago
These are not actual tests they used for themselves. Some third party did these tests first (in article and spread on social) to which the makers of Claude are responding. I knew it’s a weird test right when I first encountered it. Interest
27.
▲
by
jafitc
3y ago
Language can be ambiguous. But these LLMs were fine tuned on realistic human question and answer pairs to make them user friendly . I’m pretty sure the average person wouldn’t prefer an LLM whose output is always playing grammar Nazi or
28.
▲
by
jafitc
3y ago
We already know LLMs are good at summarizing. Question is how good they are are retaining minute details from extremely long context, say 200k tokens. That’s the frontier Claude and now GPT-4 Turbo are pushing
29.
▲
by
jafitc
3y ago
Interestingly human memory works the other way. We tend to remember out of place things more often. E.g. if there was a kid in a pink hat and blue mustache at a suit and tie business party, everybody is going to remember the outlier.
30.
▲
by
jafitc
3y ago
Imagine an assembly that you didn’t make, but was passed down to you by aliens. Now we have to tinker with it to learn instead of read Textbooks
More ›