Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
skyde
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
skyde
28d ago
Would converting those quant to MLX preserve the accuracy/size ? Or this only work with GGUF?
2.
▲
by
skyde
6mo ago
Actually not surprised. I guess this is for the same reason “say it twice” [1] is working. Because LLm are trained as causal language model, past token cannot attend to future token. One copy of the layer set solve this. [1] https:/&#
3.
▲
by
skyde
1y ago
this is CUDA backend to MLX not MLX backend for CUDA!
4.
▲
by
skyde
2y ago
Can you give more detail on what you mean by it can be a valuable experience with the right people around to help. My son (7 years old) is gifted in Math and as a parent I find it extremely hard to decide how much I should push him (registe
5.
▲
by
skyde
2y ago
is it only me or this completely miss all the recent research on causal inference using causal graphical model ?
6.
▲
by
skyde
2y ago
Des os is the best graphic calculator ever built. And its amazing it has un directly in your browser or without internet on your phone. Just wish it was open source :-) Anyone know of an open source library like 3blue1brown Manim library t
7.
▲
by
skyde
2y ago
how does this compare to ESP32-S3-BOX-3B ?
8.
▲
by
skyde
2y ago
Thanks a lot for writing that. I agree 100% with you. But I always wondered how polymath like Leonard davinci and Isaac newton that are excellent in many area are possible.
9.
▲
by
skyde
2y ago
Is MIT class also taught by outsourced instructor instead of MIT instructor?
10.
▲
by
skyde
2y ago
Redis Sentinel provides high availability and monitoring for Redis, but it does not guarantee strong consistency. Linearizability requires that once a write is acknowledged, all subsequent reads should reflect that write. if min-replicas-
11.
▲
by
skyde
2y ago
Paxos and Raft are consensus algorithms that provide certain guarantees and capabilities that a master-slave system with synchronous replication, such as PostgreSQL, cannot offer. These algorithms ensure that a majority of nodes (a quorum)
12.
▲
by
skyde
2y ago
Redis is a very bad store for a distributed lock but Postgres is only slightly better. What you truly need is something like ZooKeeper and etcd that are designed to achieve distributed consensus using algorithms like Paxos or Raft. This ens
13.
▲
by
skyde
2y ago
But inside on epoch there is a lot of duplication already. By duplication I mean if context length is N there is many sequence of N word that are not unique.
14.
▲
by
skyde
2y ago
Could not try it. Saying valid institutional or company email address. It doesn’t recognize my university.
15.
▲
by
skyde
2y ago
It “work” but the LLM having to use the calculator mean the LLM doesn’t understand arithmetic enough and doesn’t know how to use an follow a set of step (algorithm ) natively to find the answer for bug numbers. I believe this could be fixed
16.
▲
by
skyde
2y ago
Just discovered e-graph recently and I have a good understanding of compiler from taking compiler class at university. I would like to understand why you say e-graph would need control-flow to be revamped. Do you have anything I could read
17.
▲
by
skyde
2y ago
https://github.com/uwplse/tensat
18.
▲
by
skyde
2y ago
What do you mean by close to CNN? What is your architecture? Is it just a fully connected layer of chebyshev?
19.
▲
by
skyde
2y ago
Given that Alice has 13 brothers and 31 sisters, we can update the Prolog program with this information. We need to adjust the fact about Alice's siblings and then use the rule to calculate the number of sisters her brothers have. Here
20.
▲
by
skyde
2y ago
Asking gpt to first output prolog program seem to 100% fix it! Given that Alice has 13 brothers and 31 sisters, we can update the Prolog program with this information. We need to adjust the fact about Alice's siblings and then use the
21.
▲
by
skyde
2y ago
What do you mean by simplest in term of optimization? I get it find solution that are easy for SGD or Adam optimizer to find. But why would such solution be less simple than other ?
22.
▲
by
skyde
2y ago
Where is the code for it ?
23.
▲
by
skyde
2y ago
It seems it has been done before: "Syntax-Aware Transformer Models for Neural Machine Translation" by Yang et al. (2019). This model enhances the transformer architecture with syntax-aware attention mechanisms that consider depend
24.
▲
by
skyde
2y ago
Why not apply same concept every time a word is split into more than one token? Basically if a word contain a Prefix, suffix or root word. We could have a token position relative to the start of the word in the embedding.
25.
▲
by
skyde
2y ago
Forest are good at classification but they cannot leverage pre-training on unclassified data.
26.
▲
by
skyde
2y ago
As a parent myself I see this happening daily. Teacher public shaming my kids for having pretzel in his lunch saying it’s not healthy. Then later the same day the school give all the kids crackers as snack :-) Or teachers spending 2 hours t
27.
▲
by
skyde
3y ago
Author claim " you can’t reinforce understanding.". But I can think of 2 counterexample: -A student that does't understand statistics or calculus then do a bunch of exercise at home and suddenly understand the concept. -Or a
28.
▲
by
skyde
3y ago
In EF line expression you can easily specify which navigation should be Eager loaded. var customersWithOrderDetail = context.Customers.Include("Orders").ToList(); Would generate : SELECT * FROM Customers JOIN Orders ON Customers.I
29.
▲
by
skyde
3y ago
Please Enlighten me because this sound exactly like what sales is. Like being real estate agent are just like car salesmen that happen to sell house instead of car. They show you the product try to use your emotion to manipulate you into bu
30.
▲
by
skyde
3y ago
We are talking about university teacher. They are definitely the one writing and grading the exams.
More ›