9 ms·
The Future of Everything Is Lies, I Guess: Safety
- Cynddl 5mo ago> "Unavailable Due to the UK Online Safety Act" Anyone outside the UK can share what this is about?
- jazzpush2 5mo ago[flagged]
- jazzpush2 5mo agoTo be clear, that's not the full article, just the intro (though the whole thing isn't too long)
- 0x3444ac53 5mo agohttps://web.archive.org/web/20260413164025/https://aphyr.com/posts/417-the-future-of-everything-is-lies-i-guess-safety https://web.archive.org/web/20260413164025/https://aphyr.com...
- sieabahlpark 5mo ago[dead]
- satvikpendem 5mo agoIronic.
- starik36 5mo agoWhat specifically is unsafe in this article?
- onei 5mo agoIt's not that the article is inherently unsafe, it's that the UK law imposes a liability the author is unwilling to shoulder.
- cs02rm0 5mo agoAlthough Ofcom doesn't think geo blocking is sufficient to absolve them of that liability. Crazy as that is.
- aphyr 5mo agoI actually wound up geoblocking the UK based on Ofcom's February 2025 presentation for small services providers--they said that they intended to target "one-man bands" who (e.g.) failed to perform a child risk assessment or age verification, but that a geoblock would be considered compliant. I don't like doing this, but as someone who visits the UK regularly (and has been regularly pushing Ofcom on this matter) I figure better safe than sorry. https://player.vimeo.com/video/1053842235?app_id=122963 https://player.vimeo.com/video/1053842235?app_id=122963
- munksbeer 5mo agoI'm glad you have done this and I wish more would follow the same course. The more content that becomes unavailable in the UK, the more people might start to pay attention to the stupidity of the law. I doubt it, but even from an irrational anger perspective, I hate that these idiots can do idiotic (and worse, counter productive) stuff, and get no comeback on themselves.
- Jtarii 5mo ago>I'm glad you have done this and I wish more would follow the same course. The more content that becomes unavailable in the UK, the more people might start to pay attention to the stupidity of the law. The law isn't going to be repealed because a bunch of nerds geoblocked their personal blog.
- munksbeer 5mo agoThat is a weirdly aggressive reply.
- tristramb 5mo agoUse the Tor browser
- throwway120385 5mo agoAt scale I think our society is slowly inching closer and closer to building HM.
- nine_k 5mo agoWhat is HM here?
- zackmorris 5mo agoHacker Mews
- throwaway27448 5mo agoLooksmaxxing really has gone mainstream huh
- bitwize 5mo agoThought it was all the Rust catgirls.
- throw4847285 5mo agoSounds like a lovely co-op building, or perhaps a retirement community for aging hackers.
- derektank 5mo agoMaybe they meant AM (Allied Mastercomputer) from “I Have No Mouth, and I Must Scream“
- Sardtok 5mo agoHennes & Mauritz is a Swedish clothing retailer. On a serious note, I think they meant TN, as in Torment Nexus, but I could be wrong.
- throw4847285 5mo agoA Hidden Machine. That's right, a being that can cut, fly, surf, strength, and flash! Terrifying.
- jazzpush2 5mo agoEvery one of these posts is immediately pushed to the front page, this one within 4 minutes.
- acdha 5mo agoThat’s unsurprising given the author’s long history in the tech community. A ton of people see that domain and upvote.
- jazzpush2 5mo agoSure, but 4 front-page posts from the same url in 4 days surely sits at the tail of the distribution. (I guess they all capitalize on the same 'LLM-is-bad' sentiment).
- zdragnar 5mo agoIt's also aphyr, who is incredibly popular. Take one very popular author, have him write a series of posts on the zeitgeist everyone can't help but talk about, and yes, the outcome is that his posts are extremely popular. I still remember his takedown of mongodb's claims with the call me maybe post years and years ago filling me with a good bit of awe.
- macintux 5mo agoWhen I worked for Basho, aphyr was highly respected by some of the smartest people I’d ever worked with. Definitely no slouch.
- borski 5mo agoIt’s because it’s aphyr. If ‘tptacek posts a blog post, I bet it similarly does well, on average, because they’re a “known quantity” around these parts, for example.
- acdha 5mo agoDifferent URL, same domain, and exactly the kind of thing I’d expect a fair number of HN readers to have in a feed reader where they’d see it shortly after publication and decide to share it. Also, if you think this is just “LLM is bad”, I highly suggest reading the series first. The social impacts they talked about at the start of the series should resonate with a lot of people here and are exactly the kind of thing which people building systems should talk about. If you’re selling LLMs, you still want to think about how what you’re building will affect the larger society you live in and the ways that could go wrong—even if we posit sociopath/MBA-levels of disregard for impacts on other people, you still want to think about how LLMs change the fraud and security landscape, how the tools you build can be misused, how all of this is likely to lead to regulatory changes.
- macintux 5mo agoPrevious discussions from earlier posts on the topic: * https://news.ycombinator.com/item?id=47703528 https://news.ycombinator.com/item?id=47703528 * https://news.ycombinator.com/item?id=47730981 https://news.ycombinator.com/item?id=47730981
- imbus 5mo ago[dead]
- ibrahimhossain 5mo ago[flagged]
- Imnimo 5mo ago>Unlike human brains, which are biologically predisposed to acquire prosocial behavior, there is nothing intrinsic in the mathematics or hardware that ensures models are nice. How did brains acquire this predisposition if there is nothing intrinsic in the mathematics or hardware? The answer is "through evolution" which is just an alternative optimization procedure.
- cowpig 5mo agoThere are also many biological examples of evolution producing "anti-social" outcomes. Many creatures are not social. Most creatures are not social with respect to human goals.
- nyrikki 5mo agoThere is a reason we don’t allow corvids to choose if a person gets a medical treatment or not.
- b00ty4breakfast 5mo agoLuckily, this is a discussion of humans.
- fmbb 5mo agoThis is a discussion about large language models.
- Terr_ 5mo ago> just an alternative optimization procedure This "just" is... not-incorrect, but also not really actionable/relevant. 1. LLMs aren't a fully genetic algorithm exploring the space of all possible "neuron" architectures. The "social" capabilities we want may not be possible to acquire through the weight-based stuff going on now. 2. In biological life, a big part of that is detecting "thing like me", for finding a mate, kin-selection, etc. We do not want our LLM-driven systems to discriminate against actual humans in favor of similar systems. (In practice, this problem already exists.) 3. The humans involved making/selling them will never spend the necessary money to do it. 4. Even with investment, the number of iterations and years involved to get the same "optimization" result may be excessive.
- dgfl 5mo agoThe issue with most of these articles is that they seem to demonize the technology, and systematically use demeaning language about all of its facets. This one raises a lot of important points about LLMs, but the only real conclusion it seems to make is "LLMs are bad! We should never build them!". This is obviously unrealistic. The cat is out of the bag. And we're not _actually_ talking about nuclear weapons here. This technology is useful, and coding agents are just the first example of it. I can easily see a near future where everyone has a Jarvis-like secretary always available; it's only a cost and harness problem. And since this vision is very clear to most who have spent enough time with the latest agents, millions of people across the globe are trying to work towards this. I do think that safety is important. I'm particularly concerned about vulnerable people and sycophantic behavior. But I think it's better not to be a luddite. I will give a positively biased view because the article already presents a strongly negative stance. Two remarks: > Alignment is a Joke True, but for a different reason. Modern LLMs clearly don't have a strong sense of direction or intrinsic goals. That's perfect for what we need to do with them! But when a group of people aligns one to their own interest, they may imprint a stance which other groups may not like (which this article confusingly calls "unaligned model", even though it's perfectly aligned with its creators' intent). People unaligned with your values have always existed and will always exist. This is just another tool they can use. If they're truly against you, they'll develop it whether you want it or not. I guess I'm in the camp of people that have decided that those harmful capabilities are inevitable, as the article directly addresses. > LLMs change the cost balance for malicious attackers, enabling new scales of sophisticated, targeted security attacks, fraud, and harassment. Models can produce text and imagery that is difficult for humans to bear; I expect an increased burden to fall on moderators. What about the new scales of sophisticated defenses that they will enable? And for a simple solution to avoid the produced text and imagery: don't go online so much? We already all sort of agree that social media is bad for society. If we make it completely unusable, I think we will all have to gain for it. If digital stops having any value, perhaps we'll finally go back to valuing local communities and offline hobbies for children. What if this is our wakeup call?
- throw4847285 5mo agoThanks LLM!
- cowpig 5mo ago> I think it’s likely (at least in the short term) that we all pay the burden of increased fraud: higher credit card fees, higher insurance premiums, a less accurate court system, more dangerous roads, lower wages, and so on. I think the author is brushing against some larger system issues that are already in motion, and that the way AI is being rolled out are exacerbating, as opposed to a root cause of. There's a felony fraudster running the executive branch of the US, and it takes a lot of political resources to get someone elected president.
- nzoschke 5mo agoExcellent articles as expected from aphyr. I'm seeing that these tools are extremely powerful the hands of experts that already understand software engineering, security, observability, and system reliability / safety. And extremely dangerous in the hands of people that don't understand any of this. Perhaps reality of economics and safety will kick in, and inexperienced people will stop making expensive and dangerous mistakes.
- mursu 5mo agoThe future is happening. Instead of trying to raise awareness about evil AI... I think it would be more healthy if we could direct this energy to ways of improving the situation without condemning the unknown of AI evolution. As with anything.. there will be a bad side.. The bad guys will always be there.. be it AI or soccer matches.. should we stop developing nuclear energy because nuclear weapons are developed?
- fmbb 5mo agoThere is no natural law saying the good sides of any kind of tech will outweigh any bad sides. ”The future” is happening because it is allowed in our current legal framework and because investors want to make it happen. It is not ”happening” because it is good or desirable or unavoidable.
- philipkglass 5mo agoIn short, the ML industry is creating the conditions under which anyone with sufficient funds can train an unaligned model. Rather than raise the bar against malicious AI, ML companies have lowered it. This is true, and I believe that the "sufficient funds" threshold will keep dropping too. It's a relief more than a concern, because I don't trust that big models from American or Chinese labs will always be aligned with what I need. There are probably a lot of people in the world whose interests are not especially aligned with the interests of the current AI research leaders. "Don't turn the visible universe into paperclips" is a practically universal "good alignment" but the models we have can't do that anyhow. The actual refusal-guards that frontier models come with are a lot more culturally/historically contingent and less universal. Lumping them all under "safety" presupposes the outcome of a debate that has been philosophically unresolved forever. If we get hundreds of strong models from different groups all over the world, I think that it will improve the net utility of AI and disarm the possibility of one lab or a small cartel using it to control the rest of us.
- pixl97 5mo agoI mean that does partially reduce the chances of a cartel, but not really near as likely as you think. Most countries have a pretty strong ban on most kinds of weapons, the US is one of the few that lets everyone run around with their rooty tooty point and shooty, but most countries have implemented bans. Some because the government doesn't want the people having them, and in others the citizens call for the bans because they don't like the idea of getting shot by their fellow citizens. It won't be long before citizens and governments get tired of models being used for criminal activities and will eventually lay down laws around this. Models will have to be registered and safety tested, strict criminal prosecution will happen if you don't. And the big model companies will back their favorite politicians to ensure this will happen to. Now, that in general will be helpful as there will still be more models, but it will still not be a free for all.
- deleted 5mo ago[deleted]
- alfalfasprout 5mo ago
- conquera_ai 5mo agoFeels like we’re repeating classic distributed systems lessons: assume failure, constrain blast radiusand never trust components that can’t explain themselves reliably
- ibrahimhossain 5mo agoExactly assuming failure and constraining the blast radius feels like the only reliable path when the models themselves are black boxes. Patch based alignment starts looking fragile pretty quickly
- simianwords 5mo agoThe author is still grieving by watching a civilisation changing technology just passing by. Every single one of the problems they note applies to any technology that existed. The internet produced 4chan. Produced scammers. Produced fraud. Instrumental in spreading child porn. Caused suicides. Many people lost their lives due to bullying on the internet. Many develop have addictions to gaming. To anyone who has given it some thought, any sufficiently advanced technology usually affects both in good and bad ways. Its obvious that something that increases degrees of freedom in one direction will do so in others. Humans come in and align it. There's some social credit to gain by being cynical and by signalling this cynicism. In the current social dynamics - being cynical gives you an edge and makes you look savvy. The optimistic appear naive but the pessimists appear as if they truly understand the situation. But the optimists are usually correct in hindsight. We know how the internet turned out despite pessimists flagging potential problems with it. I know how AI will turn out. These kind of articles will be a dime a dozen and we will look at it the same way as we look at now at bygone internet-pessimists. This is response not just to this article, but a few others.
- raincole 5mo agoI think you underestimate people's grievance with technology. If you make a poll my guess is more than 50% of people will say the world was a better place pre-social media. If the AI tech keeps going at the direction it's going now, more and more people will start believing the world would be better if the internet and computer had never been invented. You talk like the internet being a net positive is a given. It really isn't, especially after it's proven that it doesn't democratize power (see Arab Spring, and China, and the US, and everywhere.)
- simianwords 5mo agoIts usually the educated and elite PMC types who have grievance with technology. They secured their status and have lucrative jobs mostly with the help of technology and they are too scared to have anything threaten their position in society. It is highly hypocritical to behave this way but they don't seem to have the self awareness to observe it objectively. Ask any poor person in India what their sentiment is with tech - it is usually optimism. > You talk like the internet being a net positive is a given. It really isn't, especially after it's proven that it doesn't democratize power (see Arab Spring, and China, and the US, and everywhere.) The world is far more democratic now than before and I attribute it to technology because it reduces information asymmetry.
- jagged-chisel 5mo ago"Alignment" In what world would I ever expect a commercial (or governmental) entity to have precise alignment with me personally, or even with my own business? I argue those relationships are necessarily adversarial, and trusting anyone else to align their "AI" tool to my goals, needs, and/or desires is a recipe for having my livelihood completely reassigned into someone else's wallet.
- sigbottle 5mo agoInteresting you single out commercial and government entities but not people. What defines the difference? Bureaucracy? Concentration of resources? Legal theory? I guess I'm trying to wonder why this line of thinking (in theory) doesn't turn to paranoia about everybody. I don't know much ethics or political theory or anything.
- jagged-chisel 5mo ago> … paranoia about everybody It does. People drive these entities. People hide behind the liability shields and authority of these entities. Also notice that I generalized with the phrase “…and trusting anyone…”
- robot-wrangler 5mo agoYou can tell that broad alignment between people is natural just by looking at the effort that corporations and governments make to undermine it. Alignment between people is perhaps not a state of nature, but it really is a pretty normal consequence of a fairly small amount of education and of middle-class existence that is left to itself (i.e. without brain-washing and deliberately working to create out-groups). If you're eating enough and have a few brain cells to rub together, then you definitely want that for your neighbors too because it promotes stability.
- zozbot234 5mo ago> You can tell that broad alignment between people is natural It really isn't. The whole point of the market system is to collectively align people's actions towards a shared target of "Pareto-optimized total welfare". And even then the alignment is approximate and heavily constrained due to a combination of transaction costs (which also account for e.g. externalities) and information asymmetries. But transaction costs and information asymmetries apply to any system of alignment, including non-market ones. The market (augmented with some pre-determined legal assignment of property rights, potentially including quite complex bundles of rules and regulations) is still your best bet.
- deleted 5mo ago[deleted]
- themafia 5mo ago> They also build secondary LLMs which double-check that the core LLM is not telling people how to build pipe bombs Such a fear mongering position. You can learn to build pipe bombs already. Take any chemical reaction that produces gas and heat and contain it. Congratulations, you have a pipe bomb. Meanwhile.. just.. ask an LLM if you can mix certain cleaning chemicals safely. > I see four moats that could prevent this from happening. Really? Because you just said: > human brains, which are biologically predisposed to acquire prosocial behavior You think you're going to constrain _human_ behavior by twiddling with the language models? This is foolishly naive to an extreme. If you put basic and well understood human considerations before corporate ones then reality is far easier to predict.
- bigfishrunning 5mo ago> Meanwhile.. just.. ask an LLM if you can mix certain cleaning chemicals safely. the cost of the wrong answer to this question is so incredibly high that I hope nobody is sincerely asking an LLM for this information. The things people trust to "machine that gives convincing answers that are correct 90% of the time" continue to shock me
- themafia 5mo ago> is so incredibly high that I hope nobody is sincerely asking an LLM for this information Google trumps the search results with it's LLM box. There's only one reason to do that. They know their audience is not engaging in discretion. > The things people trust to "machine that gives convincing answers that are correct 90% of the time" continue to shock me People are having intimate relationships with chat bots. There's a deeper sociological problem here.
- bigfishrunning 5mo agoThe liability of google's search box saying "Ammonia and bleach mix to make a great cleaning agent!" (disclaimer: please don't do that it will kill you) seems really high. I feel like we're all living in crazy world.
- atleastoptimal 5mo agoThere really are only 3 options that don't involve human destruction: 1. AI becomes a highly protected technology, a totalitarian world government retains a monopoly on its powers and enforces use, and offers it to those with preexisting connections: permanent underclass outcome 2. Somehow the world agrees to stop building AI and keep tech in many fields at a permanent pre-2026 level: soft butlerian jihad 3. Futurama: somehow we get ASI and a magical balance of weirdness and dance of continual disruption keeps apocalypse in check and we accept a constant steady-state transformation without paperclipocalypse
- deleted 5mo ago[deleted]
- raincole 5mo agoIn other words, only one option.
- tomjen3 5mo agoThis makes the assumption that AI will lead to the apocalypse. That's unfalsifiable, predicted about plenty of things in the past, and frankly annoying to keep seeing pop up. Its like listening to Christians talking about the rapture.
- atleastoptimal 5mo agoThe problem is that if someone is right about an existential disaster caused by AI, by the time they're proven right it would be too late. Frontier AI models get smarter every year, humans but humans don't get any smarter year over year. If you don't believe that somehow AI will just suddenly stop getting better (which is as much a faith-based gamble as assuming some rapturous outcome for AI by default), then you'd have to assume that at some point AI will surpass human intelligence in all fields, and the keep going. In that case human minds and overall will will be onconsequential compared to that of AI.
- zozbot234 5mo ago
- amarant 5mo agoThere's really only one thing we need to do to avoid the apocalypse, and that is to not hand over the launch codes to a LLM. Seems easy enough, I'm actually pretty confident in even the most incompetent of current world leaders in this particular task.
- anon35 5mo agoYou don't think a human using an LLM to generate content that convinces another human to press the launch button is a concern? Sure seems like there's more than one thing we need to do.
- amarant 5mo agoHonestly? I really don't! What kind of content do you think would trigger that? If humans were launching nukes based on Facebook posts we'd all be long dead! A good deep fake might trick your grandma, but it's not very likely to fool military intelligence.
- deleted 5mo ago[deleted]
- munificent 5mo ago> What kind of content do you think would trigger that? The kind of political propaganda that leads to the US reelecting a convicted rapist whose selects another rapist to lead the Department of Defense who then renames it to the Department of War and, true to the name, starts unilaterally attacking other countries.
- amarant 5mo agoIf trump getting elected was due to AI, I wonder why every nation isn't electing similarly awful politicians? Hungary just elected a new president who seems a lot better than his predecessor, and a lot better than trump. The Canadian prime minister is genuinely one of the best politicians I've seen in my lifetime! The list goes on and on. No blaming trump on anything other than the people who voted for him is like blaming school shootings on anything other than guns:a popular American passtime, and complete and utter nonsense.
- weinzierl 5mo agoOh boy, that’s a very generous view of human nature. The cynic in me agrees with the article’s premise, but not because I believe "alignment is a joke", but because I doubt that humans are "biologically predisposed to acquire prosocial behavior."
- goatlover 5mo agoHuman cooperation is the norm not the exception.
- weinzierl 5mo agoThe norm is competition and cooperation is the tool we invented to compete more effectively. Cooperation is only competition’s favorite strategy.
- jaeh 5mo agoThe norm is cooperation and competition is the tool we invented to cooperate more effectively. Competition is only cooperation's favorite strategy. (By choosing from competing groups we select more favorable cooperation partners, because there are too many to choose from.) Both of our statements are true. darned doublethink.
- catcowcostume 5mo agoIt's ok, you are allowed to start from wrong premises. It's nice that you acknowledge your shortcomings.
- ramoz 5mo agoAside from the sentiment and arguments made– You don't need to train new models. Every single frontier model is susceptible to the same jailbreaks they were 3 years ago. Only now, an agent reading the CEOs email is much more dangerous because it is more capable than it was 3 years ago.
- achierius 5mo agoAre they? I'm sure they're vulnerable to certain jailbreaks, but many common ones were demonstrably fixed.
- ramoz 5mo agoI retract that. I think what I meant to say was, they're as simple to jailbreak as they were three years ago. Different methods, still simple. Working with researchers that are able to get very explicit things out of them. Again, it feels much worse than before, given the capability of these models. There's basically guardrails encoded into the fine-tuned layers that you can essentially weave through (prompting). These 'guardrails' are where they work hard for benevolent alignment, yet where it falls short (but enables exceptional capability alignment). Again, nothing really different than it was three years ago.
- quantified 5mo agoThe Garden of Eden story is an apocryphal fable. But it sort of has a relevant twang to it. Geoffrey Hinton will not have his liver pecked out every day like Prometheus does.
- throwanem 5mo agoAre you sure? In some mythologies, the basilisk is notably birdlike, I believe.
- krishna3145 5mo agohttps://www.researchgate.net/publication/403780821_Adversarial_Threats_to_AI-Augmented_Security_Operations_A_Systematic_Survey https://www.researchgate.net/publication/403780821_Adversari...
- dredmorbius 5mo agoOther articles in this series discussed over the past five days: 1. Introduction: <https://news.ycombinator.com/item?id=47689648 https://news.ycombinator.com/item?id=47689648> (619 comments) 2. Dynamics: <https://news.ycombinator.com/item?id=47693678 https://news.ycombinator.com/item?id=47693678> (0 comments) 3. Culture: <https://news.ycombinator.com/item?id=47703528 https://news.ycombinator.com/item?id=47703528> 4. Information Ecology: <https://news.ycombinator.com/item?id=47718502 https://news.ycombinator.com/item?id=47718502> (106 comments) 5. Annoyances: <https://news.ycombinator.com/item?id=47730981 https://news.ycombinator.com/item?id=47730981> (171 comments) 6. Psychological Hazards: <https://news.ycombinator.com/item?id=47747936 https://news.ycombinator.com/item?id=47747936> (0 comments) And this submission makes: 7. Safety: <https://news.ycombinator.com/item?id=47754379 https://news.ycombinator.com/item?id=47754379> (89 comments, presently). There's also a comprehensive PDF version for those who prefer that kind of thing: <https://aphyr.com/data/posts/411/the-future-of-everything-is-lies.pdf https://aphyr.com/data/posts/411/the-future-of-everything-is...> (PDF) 26 pp. (Derived from aphyr's comment: <https://news.ycombinator.com/item?id=47754834 https://news.ycombinator.com/item?id=47754834>.)
- cold_tom 5mo agoFeels like people are mixing two different things here-alignment in small groups (family,teams) vs alignment at scale. The first happens naturally, the second almost always needs structure, incentives, and enforcement
- agentic_lawyer 5mo agoIf lies are our future, we have the tools necessary to deal with them. Frankly, this question was answered over a century ago by Dostoyevsky in Crime and Punishment, and every experienced criminal lawyer, prosecutor, and judge I've met already understood this very basic fact to be true: even lies point to the truth. What is unacceptable, and what I've used my entire life as a deliberate strategy to obfuscate personal affairs, deflect unpleasant conversations, and deal with fools I come across, is to mix of a small amount of truth within a complex web of lies and misdirection. This approach deals with two main challenges of lying effectively: lying in a consistent way and resisting the urge to be caught out in the lie. The truth is an abyss, and it frequently finds its most trenchant opponents flinging themselves willingly into it. The most important, revealing truths can be disclosed without any risk of being discovered, hiding in plain sight. The philosophers knew this and applied these lessons judiciously since the times of Plato. Sometimes speaking the truth is dangerous. I sometimes wish LLMs displayed that cautious refrain when discussing difficult matters. In my estimation, AGI will not have been reached until the models can produce works as mischievous as Plato, Averroes, Rousseau, or Derrida. We are a long way from that. The vanilla brand of lies put out today by LLMs are barely worth mentioning, even if troublesome. It's when the lies mask a deeper and profound truth that we'll know the game is up.
- deleted 5mo ago[deleted]
- GistNoesis 5mo agoThere is also the fact that it's very easy to plant backdoors in LLMs with plausible deniability : - You can just use the same tools you use to train them to make them behave in some specific ways if some specific preconditions are met. - You can also poison the training data, so that the LLMs are writing flawed code they are convinced is right because they saw it on some obscure blog but in fact it had some subtle flaw you planted. - You can poison the prompts as they are automatically injected from "skills" found online. You couple that with long running agents which may drift very var from the conditions where they were tested during the safety tests. You add the fact that in this AI race war, there is some premium to run agents capable of advanced offensive security with full permission, pushed using yolo dark-pattern. The training process is obscure and expensive so only really doable by big actors non replicable and non verifiable. And of course, now safe developers (aka those not taking the insane risk of running what really is and should be called malware), can't get jobs, get no visibility for any of their work, drown into a sea of AI slop made using a prompt and a credit card, and therefore they must sell their soul.md and hype for the madness.
- BloondAndDoom 5mo agoI don’t even see the pint of alignment or anything about security in LLMs. I feel like this is how “some people” reacted to the internet when I was young (lots of censorship), how hackers don’t let it happen, then how we are back to that world in the hand of corporations and governments who “think of the children). LLMs are out of the bottle and not going back there, only option is building for the new world on the defender side, everything else is politics. LLMs can hack, but also nmap made hacking easier do we make nmap illegal? We already have drones who kills people, now there is less human involvement, results are same. LLM can also make defending easier (at least for cyber security) but I guess real world security is not that different. Now evil things can be done faster, easier and at more scale. Also good things have the properties. It’s another tool in the toolbox, the idea that some entity will able to censor or align it as naive as thinking internet can be controlled. Some will do and manage anyway, but it’s not any different china’s firewall. Alignment is sold to us by companies like OpenAI and Anthropic , not because they care, because that gives them power and more control. When was the last time a big corporation actually cared about soft topics like this? Yes, never.
- intended 5mo agoTech changes do not impact attackers and defenders equally. Good things do not all have the same properties - That’s mistaking an incomplete assertion for a complete one. Cyber security is an attackers domain. Your security is typically because you are (were) not valuable enough to earn the attention of an attacker. When LLMs make targeting you cost effective, you will have to spend more energy defending yourself. This means that you have less time to do other useful things, reducing your net utility, while increasing attackers utility. Also - teams in these companies DO care, I have worked with them. The decision makers are regulated by the cadence of the quarterly share holders meeting. At that point things like safety are a cost center. Reducing safety spend while minimizing reduced time on site is rewarded by markets.
- kwar13 5mo agoI did not know about this: https://en.wikipedia.org/wiki/Saudi_infiltration_of_Twitter https://en.wikipedia.org/wiki/Saudi_infiltration_of_Twitter
- intended 5mo ago> I know this because a part of my work as a moderator of a Mastodon instance is to respond to user reports, and occasionally those reports are for CSAM, and I am legally obligated to review and submit that content to the NCMEC. Oh ** that. I have moderated all sorts of crap, and I am grateful that my worst has only been murders, hate speech, NCII, assaults, gore, and other forms of violence. > I sometimes wish that the engineers working at OpenAI etc. had to see these images too. Perhaps it would make them reflect on the technology they are ushering into the world, and how “alignment” is working out in practice This is a great idea. I’ve heard of new leaders being dropped in, and being sure they have a better handle on safety than the T&S teams. Only after they engage with the issues, and have their assumptions challenged by uncaring reality, did they listen to the T&S teams. There are a lot of assumptions on speech online that do not translate into operational reality. On HN and Reddit, everyone complains about moderation and janitors, but I highly recommend coders take it as civic service and volunteer. How can you meaningfully fix a mess, if you do not actually know what the mess is about?
- rupayanc 5mo agoThe power asymmetry point is what gets missed in most alignment debates. An AI model doesn't need to be misaligned to cause harm. It just needs to be misaligned with users while aligned with whoever's paying for it. That's not a future risk. That's how every enterprise SaaS product works already.
- schnitzelstoat 5mo agoIt's a tool, some people use the tool to do bad things. But they already did bad things before. Virtually all of the arguments here could also be applied against the Internet itself.
- jkman 5mo agoThat's a lazy argument. Obviously tools are tools. But if tool A revolutionized human society and has massively advanced technology (and CAN be used for harm), where tool B's positive impact is a drop in the bucket by comparison and has the potential for an outsized amount of harm, obviously tool B is comparatively a bad tool.
- ozozozd 5mo agoThank you for such high quality thinking and writing. Very impressive.