Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
irthomasthomas
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
irthomasthomas
1mo ago
Have you seen the news about decrypting the hidden COT in U.S. models? [0] The decoded logs revealed instances where Claude memorized answers to test questions beforehand while making its final output look like it had derived the answer ste
62.
▲
by
irthomasthomas
1mo ago
That is more an indictment of AA than DS
63.
▲
by
irthomasthomas
1mo ago
rumor is that is what Ilya has done at SSI.
64.
▲
by
irthomasthomas
1mo ago
then what is the point of using operouter for this model? Just use the deepseek API and save the 5% fee on top of the better caching rate.
65.
▲
by
irthomasthomas
1mo ago
But they don't perform the same test on other models as far as I could tell? So we don't lnow if this is peculiar to kimi models or not.
66.
▲
by
irthomasthomas
1mo ago
From: Che Chang <redacted> Date: Feb 23, 2026, at 8:04 PM Subject: Re: Former Apple Employees at OpenAI Retaining Non-public, Confidential, and Proprietary Information To: <redacted> Hi [Apple in-house legal counsel] and [Apple
67.
▲
by
irthomasthomas
2mo ago
I think it's interesting to see them visibly struggling to improve. Claude pelicans aren't much better today than they where 18 months.
68.
▲
by
irthomasthomas
2mo ago
I guess expert+chatgpt beats chatgpt alone, so why not hire top experts to drive the search?
69.
▲
by
irthomasthomas
2mo ago
Why you think that?
70.
▲
by
irthomasthomas
2mo ago
"Senators don't have the luxury that a research scientist has of waiting for evidence." - 1977 https://www.youtube.com/watch?v=xbFQc2kxm9c FYI obesity 1977 ~14% 2026 40%+ T2 Diabetes 1977 ~3% 2026 ~12%
71.
▲
by
irthomasthomas
2mo ago
It can still generate a wrong reference and select the wrong verse.
72.
▲
by
irthomasthomas
2mo ago
Two world wars and one world cup...
73.
▲
by
irthomasthomas
2mo ago
Does x help you avoid that deadspot in the middle, where the flaps cant reach?
74.
▲
by
irthomasthomas
2mo ago
But you can weigh up the evidence. A crime has been commited afterall.
75.
▲
by
irthomasthomas
2mo ago
interesting... thanks.
76.
▲
by
irthomasthomas
2mo ago
But zero evidence provided that this was an unsupervised agent attack. I still find it incredible that a company who protect their IP so much would allow these dangerous experiments to run unsupervised and risk leaking their secrets. Why
77.
▲
by
irthomasthomas
2mo ago
If it's true that they run agents like this unsupervised, it is only a matter of time before an openai agent leaks its model weights.
78.
▲
by
irthomasthomas
2mo ago
Unless openai release the logs we have only their word that this was done fully autonomously and without their knowledge by an agent running their newest super powerful model. For all we know they could have bought zero days and left them l
79.
▲
by
irthomasthomas
2mo ago
No. Lora is for tuning behaviour and how the model applies what it learned in training. Teaching a model new facts is still expensive.
80.
▲
by
irthomasthomas
2mo ago
Benchmarks show a large drop in quality as context grows. Your opus will be acting like haiku above 300k tokens.
81.
▲
by
irthomasthomas
2mo ago
Yeah something is up. I have the same problem with K3 as I had with earlier kimis. I ask it to write <bash>code</bash> every turn, that does not seem very difficult, but kimi gets this wrong a large percentage of the time.
82.
▲
by
irthomasthomas
2mo ago
I use dvorak on PC but qwerty on phone
83.
▲
by
irthomasthomas
2mo ago
Because they do not know the name of the model before they train it. There is also distillation, where multiple models will be trained from a larger one. E.G. Sonnet was promoted to Opus at one point after it surpassed expectations.
84.
▲
by
irthomasthomas
2mo ago
Cooking with lard IS healthy. There was never any evidence to the contrary. The USDA promoted the low fat diet to sell cheap commodity crops, ultra-processed foods, and alternative vegetable oils. Those things ARE bad for you. The low smoke
85.
▲
by
irthomasthomas
2mo ago
Why do OpenAI never release logs to prove their claims? Why should we believe them when they write extraordinary anecdotes about the power of their products without ever providing proof?
86.
▲
by
irthomasthomas
2mo ago
Aaron Schwartz was being prosecuted and threatened with 35 years in jail for the crime of saving research papers to a thumb drive. What damage did he cause?
87.
▲
by
irthomasthomas
2mo ago
Changelog - fixed issue where model acts like qwen when prompted in chinese
88.
▲
by
irthomasthomas
2mo ago
Openai hacked HF with a zero-day. Definitely an interesting story! I just find their explanations hard to believe. They're admitting to a great deal of incompetence. I know, don't ascribe to malice... but still, the more extraordi
89.
▲
by
irthomasthomas
2mo ago
Simon says not to dismiss this as a publicity stunt, but I think we should reserve judgement, and not treat this as true until they publish the logs. A company that uses industrial espionage against Apple does not deserve the benefit of the
90.
▲
by
irthomasthomas
2mo ago
A piece of the frame is missing between pedals and back wheel. The frame of the bike passes through the bird. It also puts a cap on the bird's head, and a fish in it's mouth. The fish and the cap where always added when I asked an
More ›