9 ms·
The Vibe Tax
- deleted 26d ago[deleted]
- supriyo-biswas 26d agoI feel this, yes. In effect, I’ve always wanted a pair programmer agent, not a zero to one programming agent. Unfortunately models these days are mostly of the latter kind and it has caused a major disruption in the way I work. I’d much rather appreciate a small model making fast and specific edits that I ask if it, rather than ingesting 20 files to make changes, and then starting to write tests, etc.
- beezlewax 26d agoI've found writing small well defined tickets and getting Claude to work on them works well for this type of workflow.
- acedTrex 26d agoThis sounds miserable, why not just give it specific tasks to do in your normal workflow/editor? Why would we want to do MORE of the miserable task of ticket creation.
- exogenousdata 26d agoBecause when you're a dev, tickets are a tool of the devil. But when you're a manager, tickets are a simple means to an end. -- PHB
- jackjeff 26d agoIndeed. Matt Pococks skills formalizes this process… (even though it can be excessive)
- lilbigdoot 26d agoJust said something similar myself in another thread. I'm either writing things by hand (and using LLMs for research, or double checking an idea), or having an LLM spit out something I treat as an external dependency. Its still too tedious for me to use them to write code when I care how it works or there's not obvious invariants the code needs to hold
- dave_sid 26d agoI have found what works well is to modularise the code as much as possible and get an agent to work within a very limited scope. Break everything down well using SRP with well defined interfaces and let the agent work on small problems. Then when it shits the bed, there’s a smaller blast radius and you can strip back and try again. I think seasoned developers, over time, learn how to work a code base and design components with well defined interfaces, where the implementation is isolated in small well contained classes. SRP etc. more junior programmers can work on those smaller components/services in isolation. For me this also seems to be a productive way to work along side an agent. Break up functionally into well defined chunks, and let the agent work on each small problem. Take more of a lead in the architecture I suppose.
- jbstack 26d agoI think this is the only sensible way to work with agents, if you care about code quality and reliability but still want the benefits of AI. There seem to be three camps that people more or less fall into: (a) AI is terrible/bad/evil and should never be used, (b) you should one-shot everything and be happy if it seems to "work" when you try it, (c) the middle ground, where the AI writes code which you carefully review. I definitely prefer (c). But I get why (b) can feel necessary. If your competition is using (b) there can be pressure to do the same just to keep up.
- dave_sid 26d agoI think c is the only way it can sustainably work. The idea of b, that software is running and nobody there knows how it works, doesn’t seem like a good foundation for a business to run on.
- bonesss 25d agoMy personal challenge is that if c, generate then review, isn’t pretty close to a one-shot then I am almost certainly negative for time versus creating from scratch (accounting for over-documenting, prompting, and enforced pauses for generation). The sunk cost fallacy bites and then bites again and again.
- dofm 26d agoI just had a sunday afternoon request for a solution to a trivial but annoying spammer pattern on a stackoverflow-like site I maintain. The software has a plugin API. I asked Muse Glimmer to recommend a plugin — it found one but I tested it and it didn't work for unclear reasons (among other things the software installed version is old, the plugin older). I then asked it to outline how to implement a simple word filter, it gave me an overview of some hooks that looked right from dim-and-distant-past recollection of reading the docs when I installed it. I asked it some questions, it did the research. I then set it off generating the skeleton of a filter plugin, went to the shops to buy food, came back and worked through filling it in and finishing it off. There was a bug. It found the solution. It's only about 100 lines of code but it is a random old webapp and it had to look stuff up to finish it, and I think it did rather well. All on my Mac. I am deeply cynical of the one-shot code, "nobody codes anymore" hype culture idea and that distaste put me off AI and agentic coding for ages. Like you, I want an assistant but as a freelancer I have to stay in control. I have no interest in the "just specify loops" BS and it will be bad for my business anyway. I worry about code that I don't have a good working overivew of, and I worry that I might forget what I have done (I have pretty bad issues with focus and memory). But in this particular case, I don't really care if I forget, because there's documented code and I have no intention of specialising in this app. So it was a nice little test case. I also don't really want to sit around waiting for Qwen 3.8 27B on this machine. Muse Glimmer is fine, actually. Gets to the solution as quickly as Qwen 3.6 35B-A3B. This gives me a little hope that local AI will give me the sort of responsive developer sidekick I actually want.
- Lalabadie 26d agoI would love, love an even better Zeta model for that reason.
- devrob 26d agoWasn't that what the first iterations of cursor were?
- jdkoeck 25d agoJust hook vanilla pi to Sol and you’ll have the pair programmer agent you crave. Make sure never to install a subagent extension. That’s it!
- esafak 26d agoCreate a spec and have a dumb model execute it. Problem solved.
- hleszek 26d agoAsk the AI to create a detailed spec according to a few simple requirements. Review the spec yourself and correct what you want changed. Then ask the AI to implement the spec. Each time you request something new, ask the AI to update the spec as well.
- add-sub-mul-div 26d agoThis so much more annoying and circuitous than writing code.
- esikich 26d agoIt isn't because I'm doing something else while it churns away.
- blackqueeriroh 26d agoThen just write the code, my man! This guy in this post is CHOOSING to use the LLM. Choose what makes you happy!
- tomasphan 26d agoDumb models will make more mistakes even with good spec no? They lack capacity to verify (to be introspective) and will pattern match over reasoning.
- pgt 26d agoPre-October 2025, maybe yes. But now? Couldn't disagree more. There is no insight in this post.
- zahlman 26d agoThe post was published today by someone who is clearly making satirical reference to frontier models ("Pol" alludes to GPT-5.6 Sol).
- alehlopeh 26d agoI tried, but I’m not sure I understand. The vibe tax is caused by the model trying to one-shot everything and doing so requires unnecessary tests? How are vibe coders training the model over months? Do you mean their sessions and preferences are being fed back into the RL?
- aDyslecticCrow 26d agoForgot where i saw it discussed; If you observe recent model benchmarks over the past year; the performance is slowly climbing, but if you divide by the token count; the score per token is dropping. The current trend in state-of-art LLM coding agents is giving more output, thinking longer and checking the results more to catch mistakes. Be it an economics inventive to make users burn through their quota or show increase in usage for shareholders, or a market demand of users liking the ability of models to do independent work without intervention or oversight; the result is what the article seem to call the Vibe Tax. I myself asked Claude code recently to review a somewhat large PR, to see what it would find. I didn't expect much, but also didn't quite realize how the model would interpret my request; I burned $20 in 3 minutes in API usage, as it ran 2 sub-agents which themselves spun up 5 more each. Most sub-agents were manually checking for things clang-tidy would catch without actually calling clang-tidy. This behavior rose as i changed from sonnet/opus 4.6 to 4.8 and now 5.0. I don't want to run a agent independently in this way; i ask targeted questions about specific things and review the result. But model development is targeted towards a more hands-off "vibe" workflow, because that's where the money and hype is. As a result, i find the models more frustrating, less trustworthy and more costly to my work. (I've even started using haiku more, since it remains to-the-point without steering away from what i ask)
- itishappy 26d agoHappens with humans too! My senior colleagues check in with me significantly less often and cost significantly more in the meantime!
- techpression 26d agoThis is my experience too, but even worse. Opus 5 finished the task, I then asked it to code review it, 61 agents later it came back with a bunch of errors that needed fixing. The first pass had tests, they passed, they were just wrong. I wish more people started reviewing their AI output, because I see a worrying trend of ”we have all these tests the agent wrote so it has to be good”, which is not surprising because understanding tests is not a trivial skill.
- hmokiguess 26d agoNeeds more info, has a good storyline but I am left trying to understand the overall pattern and trend implied there.
- freepiai 26d agoThis really resonated for me. It's like the smarter the model gets, somehow the more tokens get burned? Same failure mode whether you’re on Claude, Codex, or Cursor: the harness will spend the whole pool if you let it. I'm building my own Harness on top of pi that is add supported (www.freepi.ai) mostly because pi is so much more efficient with tokens. (That said, it tens to be slower and vastly more verbose with information I don't need to know). But yeah, since I'm trying to offer free ad supported inference the vibe tax would kill the business model. I've even been thinking about installing the 'caveman' skill to reign in token costs.
- itishappy 26d agoI feel like these are two competing goals: * the dev wants to describe an app in natural language then fall asleep while an AI works on it * the dev wishes that the same AI would write less comprehensive tests What exactly is a vibe coder to this dev?
- ad_fontes 26d agoI feel like I'm living in a parallel universe when I read these types of posts. My agents have never created code that is straight-up garbage and I have never flushed a week's worth of tokens down the toilet. I just can't identify with all the constant complaints about AI-assisted coding. And my biggest project isn't some hello world app. It's a self-hosted, privacy-focused personal financial management application that I intend to open source. It's about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline. I'm doing 24x7 mutation testing on a dedicated box against the accounting engine and temporal systems. I even have specialized agents doing audits against Regulation Z (US banking law) criteria so the app models the required behavior of banks. Most of my complaints about everything are nits, like the overly verbose and dense way LLMs communicate with me. Or their predisposition to add, add, and add more stuff when proper engineering practices are more often about subtraction (but I've built mitigation guardrails against a lot of that).
- yladiz 26d agoHow in the world do you need 30k LOC for your CI/CD??
- geerlingguy 26d agoWith some models, if you're not forcefully terse, probably 28k LOC of comments!
- floren 26d ago1. be the kind of guy who thinks more LOC = more betterer 2. ask an LLM to do the needful and never ever look at the results except to count LOC
- ad_fontes 26d agoBecause it's complex and I'm stuffing 30-50 PRs a day through it. 7 GHA workflows across four self-hosted runners. I'm on my third iteration, after constantly log jamming previous versions. Most steps aren't "run pytest", they're gates that guard against an LLM's bias to continuously add more and more complexity to a system. I'm aware of the irony of having a complex system to mitigate complexity, but the key difference is these rules bound complexity growth. If you're legitimately interested in the details, let me know. I'm too tired to write up much more but would be willing to drop in a LLM-authored summary of the details.
- guybedo 26d agoi'm not sure why people expect agents to one shot everything to perfection with just a prompt. There's a reason why we talk about software development lifecycle, design, architecture, testing ... It's because it's been the most reliable way to build and ship software. We shouldn't expect discard this and expect agents to perform well outside of this. I'm treating LLM agents as junior devs who happen to have vast knowledge of software engineering. As their team leader i make them go through planning, implementation, bug sweeping cycles using strict workflows. And it works quite well, i've been working on several large projects (1M+ LOC java,typescript,c/c++) and by any measure the projects are healthy. Sure the code isn't that beautiful, sure i'd have written things differently but it's pretty good nonetheless. Shameless plug here: i've been also working on https://kodfactory.com https://kodfactory.com, the code factory i've built to work on these large projects with workflows, reviews, etc ... I'm cleaning things up to open source it later.
- 0x457 26d ago> i'm not sure why people expect agents to one shot everything to perfection with just a prompt. because that's how agents are marketed.
- addandsubtract 25d agoNot just marketed, but also benchmarked (and benchmaxxed).
- uproarchat 26d agoI've never seen model providers marketing like that. What examples have you seen?
- Lalabadie 26d agoI don't really think they advertise "Create your app idea in one weekend night" and assume the general public will mentally add "... but hire an experienced developer to supervise the process".
- 26d ago
- robomc 26d agoThis is a confusing description of a real thing. They're clearly biasing the models more and more towards long horizon end-to-end software development, which leads to impressive "claude, build an X make no mistakes" demos, but is mainly an annoyance for expert users doing real work. (If you give claude an inch these days it'll just steamroll through a whole program of work without checking what it should be doing - a kind of overenthusiastic pull towards the first draft that is often detrimental and definitely wastes tokens, and even for very basic tasks it's using many more tokens than it should because it's doing this full belt and braces thing for everything, just in case you're an idiot). But also... it's something you can easily reign in if you want to.
- nippoo 26d agoYou can absolutely prompt agents not to write tests, or not to write extraneous asserts, or whatever, and I find that generally quite useful for the kind of code I write. I don't think it's "months of users training it", it's more that a lot of people do want a one-shot agent, and having a good test set really helps that.
- markbao 26d agoI’ve never had an agent fail to write the actual implementation. Has it done so badly, yes, but not nothing but tests. This sounds to me like a rare case that doesn’t generalize. If the general idea is that these agents write too many tests, sure I guess? ‘Too many tests’ doesn’t sound like a failure case of engineering to me; typically software has had too few tests. Also, a lot of the power of these agents is their ability to self-verify and correct, which the test loop is a part of. Nobody is making you pay this supposed tax. Just tell it not to write tests.
- deleted 26d ago[deleted]
- dzhar11 26d agoThis article somewhat reflects my experience with autonomous agentic coding. I've run several experiments with similar results: the agent burns through all my tokens while making very little progress, or produces something unacceptable. So I'd rather micromanage the process step by step. It takes more of my time, but the result is much, much closer to what I actually wanted.
- jumploops 26d agoI've found that LLMs make throwaway software better than I ever did. They handle edge cases, catch bugs, and write tests that I'd never write. Even if, however, this leads to the average piece of software improving, this one-shot complexity has the same issues as any large project. The more code, the longer it takes to steer the ship. This "rising tide lifts all boats" mentality will make exceptional software even rarer than it is today. Excited for the Roller Coaster Tycoons of tomorrow[0]. [0]https://en.wikipedia.org/wiki/RollerCoaster_Tycoon_(video_game)#Development https://en.wikipedia.org/wiki/RollerCoaster_Tycoon_(video_ga...
- TechnoBabble199 26d ago[dead]
- zuzululu 26d agowhenever I read these type of articles or comments where people are getting such bad negative experiences, I do wonder, what are they doing wrong or is there something that they are not sharing? I've been able to get such positive returns out of LLMs. I am working for 3 different remote jobs concurrently with it, I've shipped a few apps thats doing six digits a month, I found a life partner after I used LLM to really work on myself. I am also experimenting with hardware prototypes and will likely have funding to launch it all with LLMs. Why am I able to get so much out of "vibe coding" but others seemingly do not? I am not a genius, I am not a artisan, I am just very persistent and clear on what I ask LLMs but more importantly I don't try to place any other sort of unrealistic expectations on what it can and can't do. You read comments on HN and read these articles and you might come across feeling a sense of peril and doom which are all completely fictional for the most part. A lot can be achieved with LLMs, much more than what the constant doomers will try to drag you down to.
- ymolodtsov 26d agoYou can't one shot a perfect app with AI. You definitely can create a pretty complex and beautiful production-ready app with AI in a couple of days or weeks depending on what exactly you're building. I now have my own link catalog, read-latter app and an RSS reader. Tailored to work exactly how I like. Hardened, with automated backup, and external users for the RSS app. It works. It takes learning, some knowledge of terms and very high-level practices, plus design thinking, but I haven't written a line of code for these.
- fxtentacle 26d agoaggressively proactive I'd say that's the correct way to describe frontier models. They were trained with reinforcement learning based on human feedback. And obviously, humans prefer the bug-free variant. That's why models are now super verbose and spam tests like crazy. In their training environment, tokens were effectively free. And the humans that got asked never saw the price. If you ask people to choose the better offer and both are free, you end up with bloat. It's like people over-filling their plate at a buffet, then leaving leftovers. Except in this case, it's AI models burning through your wallet.
- j1elo 26d agoOur AI agent is like a dumb monkey with all the knowledge of Humanity, so a few guardrails are needed. In case it helps anyone, this is how I did describe my desired harness, from scratch. I didn't know nor wanted to write all the ".vscode/skills" files, or the AGENTS.md file or any of that, so I asked Opus to "write a Harness and all related skills as needed, to follow this procedure on absolutely every change"... It (at least on VSCode) already comes with a harness/agent creation skill by default, so it has the ability to write a very good standarized process for you. I've been playing with AI seriously for the first time, with a Python app that reads a spreadsheet with investment bookkeeping records and generates a pre-filled tax form. The harness prompt was somewhat like this: ---- 1. A "Technical Spec Writer" subagent notes down every requested change to a SPEC.md file. This spec includes functional and behavioral descriptions, together with detailed technical documentation, includes software architecture, data models, API boundary definitions, etc. It then reviews everything for inconsistencies, mistakes, and text consolidation opportunities. 2. A "Tax Law Expert" subagent makes a due diligence review of the spec corpus, and raises any concerns it has wrt. what the actual Law mandates vs. what the spec docs say. Any concern is a blocker which gets documented and must be resolved by the owner (me) before proceeding. Ask me for clarifications, rulings, reference documentation, etc. as needed. 3. A "Software Engineer" subagent takes the spec and implements it. Reviews for obvious mistakes, variable misuses, unhandled errors. Finally, reviews the code to find DRY or refactoring opportunities. 4. A "Quality Assurance" subagent makes a final pass on the code, ensuring full compliance of the codebase with the specification. Also, tests are passed and verified. ---- I would have never imagined how deep the "Tax Expert" would make me go until "it" was satisfied with the results. The resulting spec is by no means a replacement of a human expert reviewing the tax declaration, but I am 98% confident that much more than the "happy path" of what I particularly want to cover is actually right. It asked me for clarifications or references (actual URLs so it could read them) to jurispridence on corner cases that I had not even anticipated for my own declarations. In comparison, the actual "software engineering" must have been like 15% of the time/tokens. It definitely helped me do a much deeper dive on the legalese than I would have done otherwise when writing something like this. (Still no replacement for an actual expert)
- ricksunny 25d agoInteresting thanks - sounds like a vote for 'the work of elucidating the problem space' > 'the software engineering itself'.
- danpalmer 26d agoHyperbolic, but I'm seeing hints of this – Models refusing to do pair work with an engineer and trust their input, instead mandating having full control over something. Friends switching back from Fable/Opus 5 to Opus 4.8 just so they can have some input. Anthropic especially right now seem to be optimising for doing the whole task with no input. That's fine when that's the only task, and it's fine when you don't care how the sausage is made, but it's not fine for actual software engineering.
- TOMDM 26d agoYeah I'm getting this feeling too, that Opus 5 collaborates better with other Claudes, but that some of the older Opus models collaborated with people better.
- robertoallende 26d agoHa! I just did what the article says. My own Open-Source Kanban Board and I've published a month ago. According to the metrics, it's doing well: https://community.obsidian.md/plugins/fancy-kanban https://community.obsidian.md/plugins/fancy-kanban And I've also made the personal finance tracker as well: https://www.youtube.com/watch?v=qi4P4kL4IkQ https://www.youtube.com/watch?v=qi4P4kL4IkQ Now, one caveat. I don't vibe code with one shot prompt. I use something called Micromanaged Driven Development (MMDD) which aims to be the opposite of one-shot prompt: https://mmdd.dev/ https://mmdd.dev/ When I read articles like these, it surprises me that it's very unusual for me to hit token limits. I've standard accounts, I don't spend more than $40 per months in tokens. Probably I couldn't find the right narrative to promote MMDD, or probably nobody cares and this is why you fall easily into clickbait narratives to get people's attention these days. Not justifying, just trying to describe a perception.
- chr15m 26d ago> micro-managed Perfect, this is exactly how I refer to my LLM workflow too!
- localhoster 26d agoMost, if not all, of the code in the company I work in is written by AI. Our tests are useless. We have tests that make sure that mongoose schemas are creating the collections defined in them. We have tests that check that zod shcemas parse objects currently. Every PR, even if 1 line change, will drag 25 file changes because our tests are so ad-hoc, so verbose, and so incompatible with each other. Imagine how the codebase is looking... This is a badge of incompetence for the company I work with, and prob to the entire tech world. The funny part is? This makes our managers and investors proud, every PR is bigger, we make x3 more PRs (wheres the promised x10). When they say this is the death of software engineering, this is what they mean.
- eloisius 26d agoA project I was working on for a couple years until this spring started to turn like this when another contractor started committing AI work. There was a test literally called “test_imports.py” that, you know, tested that you could import every module of the project. I asked if we could just run `python __main__.py` to test that the imports worked, but he said no, we need this for big Claude-powered refactors to make sure nothing breaks.
- localhoster 25d agoI believe software will go the same road as the cloth industry went. From quality cloth that holds for years, to shit quality cloth that look good but doesnt hold more than few month. Yet, we are being sold that this is _the way weve all been waiting for_ ? what?!
- eloisius 25d agoIf that turns out to be true I hope that there are metaphorical Japanese out there still willing to pay 30,000 yen for quality shirts while the rest of the world SHEINs out.
- localhoster 24d ago
- flankton 26d agoDid an AI review this rant about AI?
- larodi 26d ago> ‘A tax on other devs’ Which other devs bro? Those in India which get paid 1/10 your wage or those who get paid corpo salaries to move three lines of code per week…? And how exactly does your prompted code tax mine? This is incredible nonsense. If you want to say - generated code killed a lot of handmade one - fine, I get it. But the fact you struggle to run this new dev process properly does not mean somebody is incurring costs to you on purpose neither that you’re a victim. It is a position you choose to be into.
- sph 25d agoYou guys work like this nowadays? Damn.
- syngrog66 25d agoI have 0 sympathy.