6 ms·
while it's mentioned in the post, it seems to me a bit burried: isn't this more like a port of `html5ever` from rust to python using LLM, as opposed to creatin
by minusf 9mo ago
while it's mentioned in the post, it seems to me a bit burried:
isn't this more like a port of `html5ever` from rust to python using LLM, as opposed to creating something "new" based on the test suite alone?
if yes, wouldn't be the distinction rather important?
- EmilStenstrom 9mo agoDepending on your perspective, you can take away any of the two points. The first iteration of the project created a library from scratch, from the tests all the way to 100% test coverage. So even without the second iteration, it's still possible to create something new. In an attempt to speed it up, I (with coding agent) rewrote it again based on html5ever's code structure. It's far from a clean port, because it's heavily optimized Rust code, that isn't possible to port to Python (Rust marcos). And it still depended on a lot of iteration and rerunning tests to get it anywhere. I'm not pushing any agenda here, you're free to take what you want from it!
- minusf 9mo agoThank you for the clarification, that was not entirely clear to me from the post. You also mention that the current "optimised" version is "good enough" for every-day use (I use `bs4` for working with html), was the first iteration also usable in that way? Did you look at `html5ever` because the LLM hit a wall trying to speed it up?
- EmilStenstrom 9mo agoIt was usable! Yeah, the handler based architecture that I had built on was very dependent on object lookups and method calls, and my hunch was that I had hit a wall trying to optimize the speed. I was slower than html5lib still, so decided to go with another "code architecture" (html5ever) that was closer to the metal. Worked out in getting me ~60% faster than html5lib. As for bs4, if you don't change the default, you get the stdlib html.parser, which doesn't implement html5. Only works for valid HTML.
- simonw 9mo agoI just had Codex CLI figure out where that first version ended and the new one began. It looks to me like this is the last commit before the rewrite: https://github.com/EmilStenstrom/justhtml/tree/989b70818874d895b99bf539a7e796464386867f https://github.com/EmilStenstrom/justhtml/tree/989b70818874d... The commit after that is https://github.com/EmilStenstrom/justhtml/commit/7bab3d2 https://github.com/EmilStenstrom/justhtml/commit/7bab3d2 "radical: replace legacy TurboHTML tree/handler stack with new tokenizer + treebuilder scaffold" It also adds this document called html5ever_port_plan.md: https://github.com/EmilStenstrom/justhtml/blob/7bab3d22c0da06b16bcd0d34cb91864298e721ba/docs/html5ever_port_plan.md https://github.com/EmilStenstrom/justhtml/blob/7bab3d22c0da0... Here's the Codex CLI transcript I used to figure this out: https://gistpreview.github.io/?53202706d137c82dce87d729263dfa2d https://gistpreview.github.io/?53202706d137c82dce87d729263df...