7 ms·
Asking AI to build scrapers should be easy right?
- bigiain 11mo agoSo I wonder how to fingerprint and block this? While I cans see _some_ good uses for it, there are clearly abusive uses for it, including in their examples. I mean jesus fuck, who wants cheap/free automation out there to "Skyvern can be instructed to navigate to job application websites like Lever.co and automatically generate answers, fill out and submit the job application."? I already have to deal with enough totally unsuitable scattergun job applications every time we advertise an open position. This is just asking to be used for abuse.
- Ldorigo 11mo agoI wonder why the focus on replaying UI interactions, rather than just skipping one step ahead to the underlying network/API calls? I've been playing around with similar ideas a lot recently, and I indeed started out in a similar approach as what is described in the article - but then I realized that you can get much more robust (and faster-executing) automation scripts by having the agents figure out the exact network calls to replay, rather than clicking around in a headless browser.
- fsckboy 11mo ago>Asking AI to build scrapers should be easy right? AI, build me a scraper what do you want to scrape [lists sites to scrape] oh, I've already scraped those relentlessly, here ya go
- showerst 11mo agoA point orthogonal to this; consider whether you need browser automation at all. If a website isn't using Cloudflare or a JS-only design, it's generally better to skip playwright. All the major AIs understand beautifulsoup pretty well, and they're likely to write you a faster, less brittle scraper.
- Etheryte 11mo agoThe vast majority of the modern internet falls into one of those two buckets though, no?
- showerst 11mo agoI mostly scrape government data so the sites are a little 'behind' on that trend, but no. Even JS heavy sites are almost always pulling from a JSON or graphql source under the hood. At scale, dropping the heavier dependencies and network traffic of a browser is meaningful.
- suchintan 11mo agoYeah, reverse engineering APIs is another fantastic approach. They aren't enough if you are dealing with wizards (eg typeform), but they can work really well
- pavel_lishin 11mo agoIf.
- suchintan 11mo agoIF you can use crawlers, definitely do. They aren't enough for anything that's login-protected, or requires interacting with wizards (eg JS, downloading files, etc)
- philipbjorge 11mo agoWe had a similar realization here at Thoughtful and pivoted towards code generation approaches as well. I know the authors of Skyvern are around here sometimes -- How do you think about code generation with vision based approaches to agentic browser use like OpenAI's Operator, Claude Computer Use and Magnitude? From my POV, I think the vision based approaches are superior, but they are less amenable to codegen IMO.
- suchintan 11mo agoI think they're complementary, and that's the direction we're headed. We can ask the vision based models to output why they are doing what they are doing, and fallback to code-based approaches for subsequent runs
- suchintan 11mo agoUnrelated, but thoughtful gave us some very very helpful feedback early in our journey. We are big fans!
- ahstilde 11mo agothis matches our personal experience, too
- franze 11mo agoIn AI First workshops. By now I tell them for the last exercise "no scrappers". the learning is to separate reasoning (AI) from data (that you have to bring.) and ai coded scrappers seem a logical, but always fail. scrapping is a scaling issue, not reasoning challenge. also the most interesting websites are not keen for new scrappers.
- pyuser583 11mo agoOver the past few days I've spent a lot of time dealing with terribly designed UIs. Some legitimate and desired use cases are impossible because poor logic excludes them. Is AI capable of saying, "This website sucks, and doesn't work - file a complaint with the webmaster?" I once had similar problems with the CIA's World Factbook. I shudder to think what an I would do there.
- suchintan 11mo agoIt's funny, one time we had a customer that wanted to use us to test their website for bugs.. Skyvern kept suggesting improvements unrelated to the issue they were testing for
- pyuser583 11mo agoSo how do clients process this sort of feedback? As a dev, “negative user feedback” gives me scares that “failed behavior testing” does not. The AI isn’t mad, and won’t refuse to renew. Unless it’s being run by the client of course. Are clients using your platform to assess vendors?
- suchintan 11mo agoNo, we don't have a lot of usage in that direction. People mainly use us to log into websites and either fill out forms or download files!
- nithril 11mo agoThe same day, a post on reddit was about: "We built 3B and 8B models that rival GPT-5 at HTML extraction while costing 40-80x less - fully open source" [1]. Not fully equivalent to what is doing Skyvern, but still an interesting approach. [1] https://www.reddit.com/r/LocalLLaMA/comments/1o8m0ti/we_built_3b_and_8b_models_that_rival_gpt5_at_html/ https://www.reddit.com/r/LocalLLaMA/comments/1o8m0ti/we_buil...
- suchintan 11mo agoThis is really cool. We might integrate this into Skyvern actually - we've been looking for a faster HTML extraction engine Thanks for sharing!
- guluarte 11mo agothe hardest part of scrapping is bypassing Cloudflare/captchas/fingerprinting etc
- suchintan 11mo agoDefinitely. What are your thoughts on the CloudFlare agent identity
- fragmede 11mo agoThe hardest part is not telling anyone how you're bypassing it!
- ThatPlayer 11mo agoI can talk about this bypass because they've fixed it: a site I was scraping rolled their own custom captcha that was just multiple choice. But they didn't have a nonce, so I would just attempt all the choices, and one of them would let me in.
- jimrandomh 11mo agoThe captcha put you on notice that your scraping wasn't authorized. Depending on the details and circumstances, bypassing it and scraping anyways may have been a crime.
- deleted 11mo ago[deleted]
- herpdyderp 11mo agoI'd be all over Skyvern if only they had enterprise compliance agreements available.
- suchintan 11mo agoWe do have them! We are HIPAA compliant, have soc-2 type 2 and offer self hosted deployments
- herpdyderp 11mo agoThank you for responding! Where is your compliance information? How do I sign a BAA?
- suchintan 11mo agoSend me an email suchintan@skyvern.com - we can get you started
- whinvik 11mo agoI feel like this is how normal work is. When I have to figure out how to use a new app/api etc, I go through an initial period where I am just clicking around, shouting in the ether etc until I get the hang of it. And then the third or fourth time its automatic. Its weird but sometimes I feel like the best way to make agents work is to metathink about how I myself work.
- suchintan 11mo agoI have a 2yo and it's been surreal watching her learn the world. It deeply resembles how LLMs learn and think. Crazy
- Retric 11mo agoOdd, I've been stuck by how different LLMs and kids learn the world. You don’t get that whole uncanny valley disconnect do you?
- goatlover 11mo agoHow so? Your kid has a body that interacts with the physical world. An LLM is trained on terabytes of text, then modified by human feedback and rules to be a useful chatbot for all sorts of tasks. I don't see the similarity.
- crazygringo 11mo agoIf you watch how agents attempt a task, fail, try to figure out what went wrong, try again, repeat a couple more times, then finally succeed -- you don't see the similarity?
- dingnuts 11mo agono I see something resembling gradient descent which is fine but it's hardly a child
- balder1991 11mo agoNo, because an agent doesn’t learn, it’s just continuing a story. A kid will learn from the experience and at the end will be a different person.
- _pdp_ 11mo agoThis is exactly the direction I am seeing agent go. They should be able to write their own tools and we are soon launching something about that. That being said... LLMS are amazing for some coding tasks and fail miserably at others. My hypothesis is that there is some sort of practical limit to how many concepts an LLM can hold into account no matter the context window given the current model architectures. For a long time I wanted to find some sort of litmus test to measure this and I think I found one that is an easy to understand programming problem, can be done in a single file, yet complex enough. I have not found a single LLM to be able to build a solution without careful guidance. I wrote more about this here if you are interested: https://chatbotkit.com/reflections/where-ai-coding-agents-go-to-die https://chatbotkit.com/reflections/where-ai-coding-agents-go...
- meowface 11mo agoWith the upcoming release of Gemini 3.0 Pro, we might see a breakthrough for that particular issue. (Those are the rumors, at least.) I'm sure not fully solved, but possibly greatly improved.
- Groxx 11mo agoalso training data quality. they are horrifyingly bad at concurrent code in general in my experience, and looking at most concurrent code in existence.... yeah I can see why.
- Grimblewald 11mo agoOr when code is fully vectorizable they default to using loops even if explicitly told not to yse loops. Code I got a LLM to solve for a fairly straightforward problem took 18 minutes to run. my own solution? 1.56 seconds. I consider myself to be at an intermediate skill level, and while LLMs are useful, they likely wont replace any but the least talented programmers. Even then i'd value human with critial thinking paired with an LLM over an even more competent LLM.
- disgruntledphd2 11mo agoThe really depressing part about LLMs (and the limitations of ML more generally) is that humans are really bad at formal logic (which is what programming basically is), and instead of continuing the path of making machines that made it harder for us to get it wrong, we instead decided to toss every open piece of code/text in existence into a big machine that then reproduces those patterns non-deterministically and use that to build more programs. One can see the results in a place where most code is terrible (data science is the place I see this most, as it's what I do mostly) but most people don't realise this. I assume this also happens for stuff like frontend, where I don't see the badness because I'm not an expert.
- pennaMan 11mo agoYes, it is easy. LLMs have reduced my maintenance work on scraping tasks I manage (lots of specialized high-traffic adfield sites) by 99% What used to be a constant almost daily chore with them breaking all the time at random intervals is now a self-healing system that rarely ever fails.
- suchintan 11mo agoThat's the dream
- ACCount37 11mo agoOne of the uses for AI I'm excited about - maintaining systems, keeping up with the moving targets.
- silver_sun 11mo agoInteresting. Could you elaborate? Is there a specific reason that it doesn't do 100% of the work already?
- TheTaytay 11mo agoCould you elaborate on your setup please?
- claysmithr 11mo agoI misread this as 'sky scrapers'
- pcblues 11mo agoYou gain experience getting interactions with other agencies optimised by dealing with them yourself. If the AI you rely on fails, you are dead in the water. And I'm speaking as a fairly resilient 50 year old with plenty of hands-on experience, but concerned for the next generation. I know generational concern has existed since the invention of writing, and the world hasn't fallen apart, so what do I know? :)
- hamasho 11mo agoOff topic, but because the article mentioned improper usage of DOM, I put down the UK government's design system/accessibility. It's well documented, and I hope all governments have the same standard. I guess they paid a huge amount of money to consultants and vendors. [1] https://design-system.service.gov.uk/components/radios/ https://design-system.service.gov.uk/components/radios/
- pu_pu 11mo agoNot at all in my opinion. Its a zero sum game against anti bot technologies also employing AI to block scrapers.
- moomoo11 11mo agoI tried skyvern like 6 mo ago and it didn’t work for scraping a site that sounds like welp. Ended up doing it myself. Was trying to scrape data across Bay Area. That said I’d try it again but I don’t want to spend money again.
- randunel 11mo agoThat 'welp' probably has a tonne of bot detection going on, given its popularity and the sheer amount of data it makes available without an account.
- jimrandomh 11mo agoYour example use case is automatically filling out an IRS form, operated by the sort of IRC department that makes a webform that's only up during business hours? Do you realize how legally risky that is to create, and how legally risky that will be to operate?
- suchintan 11mo agoWhat are some of the risks? This is a public web form available on the IRS website
- jimrandomh 11mo agoIf you're automating filling out the form, you aren't reading the instructions and you aren't checking what you're putting into it as much as you should be. And if you put in incorrect information, it tends to be considered fraud, even if it's downstream of a sloppy LLM rather than downstream of a particular fraudulent scheme.
- suchintan 11mo agoYou're right, but this is where the LLMs are especially useful. Our customers all prompt it to terminate if it doesn't have the right information / the pre submission confirmation doesn't match