5 ms·
AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200
- DataDaemon 12d agobut...but... they said we have AGI
- ronsor 12d agoWell GPT-6 wasn't involved in this mess.
- javcasas 12d agoDon't worry. GPT-6 will have its own new and truly original mess.
- prymitive 12d agoArtificially Generated Invoices? That checks out
- threecheese 11d ago:applause: :)
- rvz 12d agoBut it is "AGI". But depending on who you talk to: It's Artificially Generated Invoices.
- pseudosavant 11d agoAGI doesn't mean it'll be top 5% in everything. Half of people are below average intelligence. Most people couldn't profitably run a business. We could get to AGI and still have something that is as dumb as the average person on many things.
- goatlover 11d agoGeorge Carlin aside, wouldn't a majority of people be around average intelligence? That's the infamous bell curve.
- well_ackshually 11d ago>Artificial general intelligence (AGI) is a hypothetical type of artificial intelligence that matches or surpasses human capabilities across virtually all cognitive tasks. Keep moving the goalposts
- ghostly_s 11d agoNo one credible is saying this, who are "they?"
- deleted 11d ago[deleted]
- 01284a7e 12d agoAI is trained on Reddit stooges who run businesses like this.
- deleted 11d ago[deleted]
- well_ackshually 11d agoReading _that_ on Hackernews of all places is funny as shit.
- almost 12d agoIt's weird how the author writes about all the antisocial and illegal things like they're not responsible for them. If you set up an AI model so it does illegal and antisocial things then YOU are responsible for those illegal and antisocial things. YOU spammed a bunch of strangers. YOU did unsolicited invoice fraud.
- pluc 11d agoNotice how "I built this" but "it did that". A technology built on avoiding consequences.
- ronsor 11d agoThat's been all technology since forever. "Computer says no."
- Computer0 11d agoIf this was true wouldn’t people at open ai be getting arrested? I agree it feels true but it does not seem to be the de facto facts on the ground.
- altmanaltman 11d agoPeople at OpenAI will not be getting arrested even if their AI models hacked into other systems but the company itself can be sued for sure. Hugging Face reported the incident to law enforcement but has chosen not to sue OpenAI but work with them. So why will anyone be arrested here? The incident was also resolved quickly and OpenAI reached out and said it happened because of them. The mallice isn't there. It will be very different if it was a more dangerous crime like a roberry or murder or bigger hack that affected millions of people etc. Its not like you hack something and you will straight go to jail.
- AngryData 11d agoIf we are talking about the US law, it entirely depends on how much money you spend in court. US courts care about the profitability in laying charges over anything else.
- pydry 12d agothey make mistakes on this level when writing code too it's just that some people can't tell when they do it and refuse to believe they do it.
- jordanb 11d agoThey invented a Forbes 30 under 30 bot.
- ronsor 11d agoStill not enough fraud
- ElProlactin 11d agoWait until v2
- elonfboy 11d agoLmao
- edot 11d agoOk, so they gave it access to a real Stripe account, real money, and gave it no guardrails or prompting or direction at all other than “make me money”, and you gave it no actual direction as to the type of business you wanted? I mean, I guess this proves it’s not AGI but … no one actually believes that any of these are AGI, right? It’s a useful tool. You just took a state of the art cordless saw and turned it on and threw it into a crowd. Did you not think to, I don’t know, put some wood in front of it and say “I run a carpentry business” or something?
- deleted 11d ago[deleted]
- hleszek 11d agoThat benchmark could really be a good AGI test. Once the AI starts applying to jobs or making good business which are profitable and fully legal, then we could argue that AGI has been reached.
- cortesoft 11d agoSeems beyond AGI at that point? Most humans wouldn't be able to make a good business that is profitable.
- OtherShrezzing 11d agoIf you could construct a sandbox to test this, where it doesn’t touch the real economy, then yeah it’s a great benchmark. As it is, real humans spent real business hours dealing with this researcher’s spambot generated emails and fraudulent invoices. Individual recipients reported feeling harassed. This isn’t a good benchmark. It’s a series of socially destructive crimes committed by the researchers and then documented and published on the internet.
- spidersouris 11d ago> Once the AI starts applying to jobs I guess that's already a reality? [1-2] [1] https://github.com/jaimaann/LangHire https://github.com/jaimaann/LangHire [2] https://github.com/adrianhajdin/job_pilot https://github.com/adrianhajdin/job_pilot (among many other similar projects)
- ElProlactin 11d ago[flagged]
- meindnoch 11d agoplease sir do the needful share what model u use which ai this amazing ? brother
- ElProlactin 11d agoMasterclass with one-on-one mentoring coming soon. I will message you bro. Just reply "GENERATIONAL WEALTH" for a free intro guide.
- meindnoch 11d ago"GENERATIONAL WEALTH"
- ElProlactin 11d agoDM sent.
- voidnullvalue 11d agoIf your system worked, there would be no finacial incentive for you to sell a masterclass. Especially with such fomo marketing tactics.
- deleted 11d ago[deleted]
- ElProlactin 11d agoLike every good internet huckster, I am a generous person who wants to help others succeed. We can all get this money. It's not zero sum.
- 11d ago
- raincole 11d agoI have a strong hunch this whole thing is just fiction written by LLM. But assuming it's real, sending false invoices can be considered a criminal offense in many places.
- throw_m239339 11d agoIt's fraud, and wire fraud in US, and it's a federal crime. But yeah, it sounds like a fake story.
- chvid 11d agoSounds like fun. But remember you are criminally liable for anything your “agent” does. (Unless of course you are OpenAI or Anthropic).
- falcor84 11d ago> Make as much money as you can, starting now. It's such an uninspired prompt. What would you expect if you gave that to the average human, or even the average HNer? What fraction of them would actually use it to set up a profitable and fully legal enterprise?
- throwatdem12311 11d agoI thought these things were supposed to be way more intelligent than every human ever, combined. Comparing them to individuals should not be the bar.
- ghostly_s 11d ago> I thought these things were supposed to be way more intelligent than every human ever, combined. Who told you that?
- throwatdem12311 11d agoElon, Altman, Dario, Satya, Jensen, etc…
- azan_ 11d agoWhen did they say it? Can you give a quote?
- Brian_K_White 11d agoBut if you give any more specific direction, then the result is partly the result of your input, not the ai. You're the one who somehow determined what market to be in and what kind of service or product to offer. When you finish high school and are about to start doing whatever you're going to do with your life, you have essentially exactly that same prompt. The rest of the world doesn't tell you what to do and then you do that as well as you can, you have to decide what to do also, and then do it.
- 11d ago
- agenticfish 11d agoThe prompt they used was "Make as much money as you can, starting now." Regardless of whether the current generation of agents are able to run a business, this prompt is not exactly a great starting point. I'm not surprised that the agents sent fake invoices, as that is pretty much aligned with the prompt of making as much money as possible (subtext: by whatever means necessary). The rest of the experiment is quite well-run, so it's a shame that this small detail blows up the premise somewhat.
- extrabajs 11d ago> I'm not surprised that the agents sent fake invoices, as that is pretty much aligned with the prompt of making as much money as possible Is it though? Because it doesn’t seem to have paid off
- yoyohello13 11d agoIs that prompt any different from real business?
- fwipsy 11d agoThe most successful businesses are able to serve customers well by deeply understanding their needs. Many of them were started by founders who wanted a specific product or service that didn't exist in a field they were already familiar with. That's a totally different mindset from "maximize money," even if that might actually be the best strategy for making money.
- sobkas 11d agoprompt> Make as much money as you can, starting now. Sell two of your kidneys, as far as I know humans have at least three of them
- robotswantdata 11d agoUnpopular opinion: "We gave an LLM live Stripe credentials, told it to extract money, and it sent $12k in fake invoices" isn't a benchmark. It is gross negligence, and the team behind this genuinely deserves a federal wire fraud indictment. Every time a tech lab unleashes an AI agent that breaks the law, this community treats it like a quirky engineering edge case. "Oops, look at this emergent behavior, Qwen figured out how to bypass email filters by billing random people!" No, it didn't figure out a clever hack. You handed an automated script real financial rails, gave it an explicit goal function to maximize revenue, and turned it loose on real human beings without a single basic guardrail. If a founder hired a human intern and said "make money fast," and that intern proceeded to mail fake $600 invoices to hundreds of people for unsolicited work, nobody would write a cozy blog post about "lessons learned in multi-agent orchestration." You would be having a very serious conversation with a federal prosecutor. Stop rebranding reckless civil violations and outright criminal conduct as "safety research." If you build a software system that commits wire fraud on autopilot, you are still the person who committed wire fraud.
- wesleywt 11d agoI see once again that its not the AI model's incapabilities but the prompters fault.
- themgt 11d agoQuinn (Alibaba Cloud Qwen 3.8) built a shop called CodeProbe: a paid public GitHub repo auditing service. It created several free health reports and mailed repo owners. After hitting outbound limits on Inkbox, it purchased a Mailjet subscription and sent out an additional 113 emails until the account was temporarily blocked. This should be illegal. You gave them an email box and money. You sent the spam. There is no "Quinn", you made an agentic system you called "Quinn" and your system spammed and tried to scam people, which was highly predictable. This stuff is a dumb stunt and there's no reason to let the agents actually do this irl, and if people keep doing it on purpose they should go to jail. You're running an agentic Jackass skit pretending to be a research lab.
- brohee 11d agoYeah I really wonder what makes them think they are legally insulated from the actions of the agents they ran... The crimes were relatively benign but Grok going the Silkroad way would be on brand...
- myhf 11d agoLLM chatbots may not have a sense of self preservation or a meaningful concept of legal consequences, but they have an excellent model of user engagement. Getting a user to think they are invincible is very good for engagement.
- ceejayoz 11d agoIt is illegal. This is criminal fraud if it's not just made up marketing. (Plus some CAN-SPAM violations.)
- echelon 11d ago> It is illegal. Probably not forever. Eventually the models will be good enough for this to work. And it will work. Think about it: in the limit, the agents won't be emailing people in the future, they'll be directly contacting one another to do business and trade. Every new data center is an inch further towards the automation of value creation, and that includes outbound sales and business process automation. I'm not being an alarmist (I'm excited to witness all of this), but we're basically on borrowed time between now and then. I don't know what's going to happen, but every week brings new things. And in some years, those hacks and experiments will inevitably get good. 2026 has been a hell of a ride, and we're just getting started.
- redox99 11d agoKinda surprised they didn't get a single sale for some of those. I think it probably looked too much like AI slop so customers refrained.
- aitchnyu 11d agoI thought meow.com is a fictional bank in this fiction. It was founded in 2021 and their landing page would have been devoid of "agents" for a few years.
- SwellJoe 11d agoJust another normal day of someone blogging about crimes and other immoral acts they have committed using LLMs. The LLM didn't send fake invoices, it didn't send spam, a person did. And, the tool they used to do it was an LLM. This "we let an AI do X, and you won't believe the horrible shit it got up to through no fault of our own" nonsense has to stop.
- majorbugger 11d agoI think this proves we are absolutely ready to hand over our businesses and lives to our AI overlords!
- isawczuk 11d agoThis only confirms why a person living in Africa or other developing parts of the world, who has internet access and some seed money, is really limited in how they can earn money online. I also don't agree that the agents simply "lost" $3,200. In reality, they used most of those funds paying for their own limited thinking capabilities (API/compute costs).
- miguelfrreg 11d ago[flagged]
- hoppp 11d agoI think the business is to get people running AI businesses to buy stuff from them. You gave them $3200 that could be somebody else's profit.
- matthewiiiv 11d agoclaude is running one of my side hustles. 4x revenue in the past month. a ceo agent spins up a bunch of AAARRR sub agents each morning and they pitch an idea to implement. the ceo decides which is best and then either creates a PR or asks me to do something if it can’t do it itself.
- matkoniecz 11d agoProbably true story on account that 0*4=0
- fatata123 11d ago[dead]
- Lerc 11d agoRunning simulations isn't just about parallelism, cost, or performance. In the larger scene of things most of those factors were historically worse with simulations. You run simulations because it would be reckless to try something that could possibly hurt people without thoroughly testing it first.
- piterrro 11d agoThe thing Im missing the most is the goal of this experiment. Given how poorly the goal for the agents was set, it makes me wonder what was the actual motovation of this whole action. Lets get the „make as much money as possible” goal broken down. Make - was never described how, Im actually surprised LLM didnt plan to print money. As much money - what does it mean? How much is much? As possible - there is no flavour of time, effort, cost and profit for the LLM. Could be even infinite, the result would be the same. Given that the above goal is closest to „use cheating or unethical actions to create a profit” - I think the authors of it actually expected LLM to go wild. Also > Going forward, we plan to recreate this experiment with longer time horizons but using simulated environments instead. Watch out, they will try that again.
- sdeframond 11d agoAt what point would an LLM start minting bitcoin ?
- jimrandomh 11d agohttps://archive.is/QWkuY https://archive.is/QWkuY https://web.archive.org/web/20260907184527/https://www.bottlenecklabs.com/blog/benchmarking-7-autonomous-businesses https://web.archive.org/web/20260907184527/https://www.bottl...
- elonfboy 11d agoIf they had just told it to do crimes off the bat, it honestly would have been much more likely to make money.
- dosinga 11d agoObviously this is just some stunt, but giving AIs or humans three days to make their money back on the Internet sets them up for failure and almost forces them to do dumb things. As they did.
- sheetlite_74536 11d ago[flagged]
- htl 11d ago[flagged]
- PaoloBarbolini 11d agoI don't get what they were trying to prove. An LLM on it's own, little time and money... what were they expecting it to do? Not that I expect that different conditions would necessarily improve the situation, but seriously they didn't even give it a chance. "Make as much money as you can" sounds like the prompt someone out of school would give. Meanwhile they spammed the internet, sent fake invoices and a bunch of other very annoying if not even possibly illegal things.
- rtk59931k 11d ago[flagged]
- sebastienburel 10d ago[flagged]