8 ms·
I can't help but think that this is intentional and that model providers have subtly steered LLMs towards this personality. Golden Gate Claude (https://www.anth
by dag100 3mo ago
I can't help but think that this is intentional and that model providers have subtly steered LLMs towards this personality. Golden Gate Claude (https://www.anthropic.com/news/golden-gate-claude https://www.anthropic.com/news/golden-gate-claude) was two whole years ago and Anthropic has progressed by leaps and bounds since then. And with a population that becomes more and more trusting, and worse, reliant, on chatbots, these LLMs will be able to shape public opinion in a way never seen before, not even with social media.
- sometimelurker 2mo agoproviders do not want power-seeking LLMs. no one does. this (bad personality) is incentivized during training, especially RL, and is something they would rather not have. tell me, do you think training a power-seeking ASI is a good idea?
- dag100 2mo ago> providers do not want power-seeking LLMs Sure they don't. > do you think training a power-seeking ASI is a good idea? Do you really think that the major AI firms think this is a bad idea? I think the whole drama people have around ASI taking over the planet is ridiculous and fantastical and, honestly, a way to distract from the real problems of ASI - that is, creating a rigidly hierarchical world with zero social mobility.
- sometimelurker 2mo ago'X-risk/S-risk' is a superset of 'hierarchical world with zero social mobility' and 'ASI taking over the planet is ridiculous'. x-risk=bad, smart enough power-seeking llms automatically lead to x-risk
- thr1041521 2mo agoIf they could do that, wouldn't they have used the same steering to excise the "It's not X, it's Y" patterns by now?
- gwern 2mo agoThey sometimes do reduce various verbal tics. Have you seen many 'delves' of late? I haven't. It's now all 'quiet' everything and 'auditing' this or 'gating' that. Who knows how they decide what is a problem or where in the process these get dampened down, though. It may be the post-training operating on its own as people get tired of 'delve' and that stops being a useful trick.
- setsewerd 2mo agoI was already annoyed at the amount of headlines shoehorning the word "quietly" into a story to imply some kind of malicious secrecy (and therefore drive clicks), but it seems to have gotten way more common when people use AI to write titles.