Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kotach
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
kotach
10y ago
It would be innovative if they had some super fast on-line optimization of delivery routes. Optimized routes would allow them to chain several pickups from restaurants with a chain of deliveries. Given the fact that the optimization is glob
2.
▲
by
kotach
10y ago
Vowpal Wabbit is IO limited. Meaning that there's no way it is slower than anything else on a single machine. On multiple machines it glides faster than light. So, the benchmark is probably incorrect for VW.
3.
▲
by
kotach
11y ago
Checkout Dagger [2], SEARN [3] and LOLS [1] (LOLS is available in vowpal wabbit search capabilities). A lot of interesting stuff on mimicking optimal policies, local optimality, joint learning and similar stuff :D The whole point of playing
4.
▲
by
kotach
11y ago
If it works for chess, it'll work for Go. Chess has lots of games that you can learn from, Komodo wins any grandmaster or draws. The problem with Go was lack of evaluation function that would guide the policy. So it had to be learned s
5.
▲
by
kotach
11y ago
LSTM would converge even faster. A K-level breadth first search mimicking the optimal policy and a simple learning to search algorithm with a cost sensitive binary linear classifier would work well too. After training it would be a constant
6.
▲
by
kotach
11y ago
chess grandmaster can easily be beat by a smartphone app (komodo), is smartphone using more energy than a human brain? the problem with Go is that there's little data and the game is more complex - but given the small sizes of the deep
7.
▲
by
kotach
11y ago
Studies on twins, especially those that observe separated ones, pretty much show how much bodies behave in a deterministic way, from diseases, to relationships, names, jobs, wishes.
8.
▲
by
kotach
11y ago
Now, train the network jointly over the game sequence. Or even better, when given a chance to take action rollout on each action and learn jointly on that rest of gameplay. Reinforcement learning is very hard. Especially when you create mea
9.
▲
by
kotach
11y ago
Given Langford's locality vs globality argument this also gets quite obvious for the 4th game mistake and overconfidence that AlphaGo had. The rate of growth of the compounding error for local decision maker is going to cause these kin
10.
▲
by
kotach
11y ago
Yes, it is true. In the case of Super Mario he does the learning by simulating level-K BFS from positions that resulted in errors (unseen states) and thus minimizes the regret for the next K moves. Although, if you checkout his papers, the
11.
▲
by
kotach
11y ago
That's not really a problem. Given a large enough dataset you want to generalize from it - there are always states not present in the dataset - the whole point is now to extract features out of your dataset to allow generalization on u
12.
▲
by
kotach
11y ago
They train using trajectories but train them to guess the trajectory locally, not globally. Discounted long-term rewards are just a hack, they aren't joint learning. The concept of label bias, or decision bias is a joint/structure
13.
▲
by
kotach
11y ago
Yes, the "label bias" is more of a structured learning / joint learning term that is present in natural language processing. But reinforcement learning suffers only if you do the learning to minimize local loss of the decisio
14.
▲
by
kotach
11y ago
What you are talking about here is called "label bias". [2] It is present only if training is done badly. When you have a game of Go, or Super Mario level. You don't want to make your decisions by just checking the local feat
15.
▲
by
kotach
11y ago
I believe the whole point of pretraining on reference policies, which a collection of "optimally" played human games is, is just avoidance of bad local optimum. It can be a case that training and learning on just a learned policy
16.
▲
by
kotach
11y ago
The questions you pose require solving the game, at least (ii). https://en.wikipedia.org/wiki/Solved_game
17.
▲
by
kotach
11y ago
7.5-point komi variant played by AlphaGo and Lee has a win or lose outcome. There's no draw. But yes, a more formal definition of global optimality does not include victory as a necessary outcome.
18.
▲
by
kotach
11y ago
AlphaGo is approximating global optimality by finding local optimality. Local optimality is already computationally very hard, but it is exactly what AlphaGo is doing. The rollouts they are doing, evaluating every probable move, it is a sea
19.
▲
by
kotach
11y ago
Chess can be played godlike on a smartphone. Result of years of refining algorithms. Same could probably be accomplished with Go.
20.
▲
by
kotach
11y ago
AlphaGo is certainly controllable.
21.
▲
by
kotach
11y ago
I'd say I could do without concepts, modules and coroutines but ranges ! Ranges were so nice and would finally allow for easier stream handling.
22.
▲
by
kotach
11y ago
AlphaGo has a learned evaluation function for each move. Evaluation function exists but it is not as simple as it can be for chess.
23.
▲
by
kotach
11y ago
This comment is on-topic. Everything else is off-topic. Improve that precision!
24.
▲
by
kotach
11y ago
Yeah, that's exactly the food I was thinking of. Cheeseburger + chocolate milk shake. Quinoa with kale. What would you say is more healthy? The comparison is idiotic. If a person suffers anorexia their whole diet is unhealthy and has t
25.
▲
by
kotach
11y ago
All food is healthy. Diets can be unhealthy. It really is interesting that the whole science is concentrating on a single ingredient.
26.
▲
by
kotach
11y ago
Comparison on a task that is generating text. It also depends on how the HMMs are trained. Are they trained by reducing the loss over each document, or are they trained by just taking the frequency of all the needed transitions? Training us
27.
▲
by
kotach
11y ago
There's a lot of money in agriculture. Increased demand of food due to exponential growth of the population. Increased demand for luxurious food like meat due to exponential economic growth of the population. Long term, it is the most
28.
▲
by
kotach
11y ago
not really, because AI isn't really our species, and given all of the other creation we've been doing - having to do with other animals, taking care of them, nurturing them, and eventually supporting 150 billion of them dying each
29.
▲
by
kotach
11y ago
assuming you see an image every 400ms which is given blinking and activation of neural pathways a good approximation. Billion images per that rate is equivalent to 12 years of never stopping to watch there have been systems that learned to