6 ms·
You can't parse [X]HTML with regex. Because HTML can't be parsed by regex. Regex is not a tool that can be used to correctly parse HTML.
by danlitt 2mo ago
You can't parse [X]HTML with regex. Because HTML can't be parsed by regex. Regex is not a tool that can be used to correctly parse HTML.
- zarzavat 2mo agoParsing HTML with a regex is never a good option, but it's sometimes the only option.
- pkal 2mo agoIn the example from the article it certainly is an option. In Python you could either use a "soup" library or you could play around with a tool like https://www.w3.org/Tools/HTML-XML-utils/man1/hxpipe.html https://www.w3.org/Tools/HTML-XML-utils/man1/hxpipe.html. The more fundamental question for me is why the author didn't decide to make make code blocks non-breaking by default, or just add the class annotations when he writes the HTML?
- rokkamokka 2mo agoYou can parse a subset of it though, like if you're in control of the html yourself and avoid certain structures
- matheusmoreira 2mo agoTony the pony, he comes.
- timedude 2mo agoWhat do you mean you can't. I do it all the time
- throw1234567891 2mo agoYou do some parts all the time.
- gpvos 2mo agoAs the saying goes: you can fool all people some of the time, and some people all of the time. Actually, depending on what you're doing, it can be totally fine. Just be aware that you're losing some generality.
- tyho 2mo agoYou can absolutely parse HTML with regex, so long as the document is finite in length. Every finite language is regular, hence can be parsed with regex's.
- gpvos 2mo agoDo a web search for the parent of your comment, read the Stackoverflow answer. It's a classic. Learn about Zalgo and Tony the pony, he comes.
- magicalist 2mo agoZalgo and situational subset parsing aside: > You can absolutely parse HTML with regex, so long as the document is finite in length This isn't sufficient, unless I'm misinterpreting what you're saying. It's not enough to have documents of finite length (all documents are finite in length), you need documents with a max length, so you have a finite number of possible documents to parse.
- deleted 2mo ago[deleted]
- vrighter 2mo agoand you write your regex specifically for that document length, handling all possible nesting combinations. combinatoric explosion
- layer8 2mo agoSome commenters are missing that this is a reference to https://stackoverflow.com/a/1732454 https://stackoverflow.com/a/1732454.
- pwdisswordfishq 2mo agoGood ol' appeal to emotion devoid of technical argument.
- chrisandchris 2mo ago> Locked. There are disputes about this answer’s content being resolved at this time [sic] And also > Nov, 2020 And then, StackOverflow asks itself why it looses users.
- layer8 2mo agoThe answer is actually from 2009. As long as HTML/XML doesn’t suddenly become a regular language, I think it’s pretty timeless.
- magicalist 2mo agoIt's a 17 year old joke, maybe one of the most famous on stackoverflow for the aging engineers out there. No one wants new "funny" edits on it.
- prmoustache 2mo agoI do not count substitution as "parsing"