11 ms·
So you want to parse a PDF?
- BenGosub 1y agoDocling* works pretty well in PDF hell, but is terribly slow. *https://docling-project.github.io/docling/ https://docling-project.github.io/docling/
- JKCalhoun 1y agoYeah, PDF didn't anticipate streaming. That pesky trailer dictionary at the end means you have to wait for the file to fully load to parse it. Having said that, I believe there are "streamable" PDF's where there is enough info up front to render the first page (but only the first page). (But I have been out of the PDF loop for over a decade now so keep that in mind.)
- UglyToad 1y agoYes, you're right there are Linearized PDFs which are organized to enable parsing and display of the first page(s) without having to download the full file. I skipped those from the summary for now because they have a whole chunk of an appendix to themselves.
- jeroenhd 1y agoStreaming with a footer should still be possible if your website is capable of processing range requests and sets the content length header. A streaming PDF reader can start with a HEAD request, send a second request for the last few hundred bytes to get the pointers and another request to get the tables, and then continue parsing the rest as normal. Not great for PDFs generated at request time, but any file stored on a competent web server made after 2000 should permit streaming with only 1-2 RTT of additional overhead. Unfortunately, nobody seems to care for file type specific streaming parsers using ranged requests, but I don't believe there's a strong technical boundary with footers.
- wackget 1y ago> So you want to parse a PDF? Absolutely not. For the reasons in the article.
- ponooqjoqo 1y agoWould be nice if my banks provided records in a more digestible format, but until then, I have no choice.
- vander_elst 1y agoI find it pretty sad that for some banks the CSV export is behind a paywall.
- ponooqjoqo 1y agoMine offer CSV exports but the data in the CSV file is a small fraction of the data in the PDF statement. It's just a list of transactions, not a reconcilable "end of month" balance with all the data.
- Hackbraten 1y agoIn Germany, traditional banks and credit unions offer a financial API called FinTS [0]. A couple of desktop banking apps support FinTS, and consumers can typically use it free of charge. The API has been around since 1998 and is one of the best pieces of software ever produced in Germany imho (if we ignore for a second that that bar is pretty low to begin with). Unfortunately, it’s mostly traditional German banks and credit unions that offer FinTS. From a neobank’s point of view, chances are you’re catering to a global audience, so you just cobble together a questionable smartphone app and call it a day. That’s probably cheaper and makes more sense than offering a protocol that only works in Germany. I wish FinTS had caught on internationally though! [0]: https://en.wikipedia.org/wiki/FinTS https://en.wikipedia.org/wiki/FinTS
- Paul-Craft 1y agoNo shit. I've made that mistake before, not gonna try it again.
- yoyohello13 1y agoOne of the very first programming projects I tried, after learning Python, was a PDF parser to try to automate grabbing maps for one of my DnD campaigns. It did not go well lol.
- simonw 1y agoI convert the PDF into an image per page, then dump those images into either an OCR program (if the PDF is a single column) or a vision-LLM (for double columns or more complex layouts). Some vision LLMs can accept PDF inputs directly too, but you need to check that they're going to convert to images and process those rather than attempting and failing to extract the text some other way. I think OpenAI, Anthropic and Gemini all do the images-version of this now, thankfully.
- trebligdivad 1y agoSadly this makes some sense; pdf represents characters in the text as offsets into it's fonts, and often the fonts are incomplete fonts; so an 'A' in the pdf is often not good old ASCII 65. In theory there's two optional systems that should tell you it's an 'A' - except when they don't; so the only way to know is to use the font to draw it.
- UglyToad 1y agoIf you don't have a known set of PDF producers this is really the only way to safely consume PDF content. Type 3 fonts alone make pulling text content out unreliable or impossible, before even getting to PDFs containing images of scans. I expect the current LLMs significantly improve upon the previous ways of doing this, e.g. Tesseract, when given an image input? Is there any test you're aware of for model capabilities when it comes to ingesting PDFs?
- simonw 1y agoI've been trying it informally and noting that it's getting really good now - Claude 4 and Gemini 2.5 seem to do a perfect job now, though I'm still paranoid that some rogue instruction in the scanned text (accidental or deliberate) might result in an inaccurate result.
- throwaway840932 1y agoAs a matter of urgency PDF needs to go the way of Flash, same goes for TTF. Those that know, know why.
- internetter 1y agoI think a PDF 2.0 would just be an extension of a single file HTML page with a fixed viewport
- mdaniel 1y agoI presume you meant that as "PDF next generation" because PDF 2.0 already exists https://en.wikipedia.org/wiki/History_of_PDF#ISO_32000-2:_2020_(PDF_2.0) https://en.wikipedia.org/wiki/History_of_PDF#ISO_32000-2:_20... Also, absolutely not to your "single file HTML" theory: it would still allow javascript, random image formats (via data: URIs), conversely I don't _think_ that one can embed fonts in a single file HTML (e.g. not using the same data: URI trick), and to the best of my knowledge there's no cryptographic signing for HTML at all It would also suffer from the linearization problem mentioned elsewhere in that one could not display the document if it were streaming in (the browsers work around this problem by just janking items around as the various .css and .js files resolve and parse) I'd offer Open XPS as an alternative even given its Empire of Evil origins because I'll take XML over a pseudo-text-pseudo-binary file format all day every day https://en.wikipedia.org/wiki/Open_XML_Paper_Specification#Comparison_with_PDF https://en.wikipedia.org/wiki/Open_XML_Paper_Specification#C... I've also heard people cite DjVu https://en.wikipedia.org/wiki/DjVu https://en.wikipedia.org/wiki/DjVu as an alternative but I've never had good experience with it, its format doesn't appear to be an ECMA standard, and (lol) its linked reference file is a .pdf
- LegionMammal978 1y agoAs it happens, we already have "HTML as a document format". It's the EPUB format for ebooks, and it's just a zip file filled with an HTML document, images, and XML metadata. The only limitation is that all viewers I know of are geared toward rewrapping the content according to the viewport (which makes sense for ebooks), though the newer specifications include an option for fixed-layout content.
- farkin88 1y agoGreat rundown. One thing you didn't mention that I thought was interesting to note is incremental-save chains: the first startxref offset is fine, but the /Prev links that Acrobat appends on successive edits may point a few bytes short of the next xref. Most viewers (PDF.js, MuPDF, even Adobe Reader in "repair" mode) fall back to a brute-force scan for obj tokens and reconstruct a fresh table so they work fine while a spec-accurate parser explodes. Building a similar salvage path is pretty much necessary if you want to work with real-world documents that have been edited multiple times by different applications.
- UglyToad 1y agoYou're right, this was a fairly common failure state seen in the sample set. The previous reference or one in the reference chain would point to offset of 0 or outside the bounds of the file, or just be plain wrong. What prompted this post was trying to rewrite the initial parse logic for my project PdfPig[0]. I had originally ported the Java PDFBox code but felt like it should be 'simple' to rewrite more performantly. The new logic falls back to a brute-force scan of the entire file if a single xref table or stream is missed and just relies on those offsets in the recovery path. However it is considerably slower than the code before it and it's hard to have confidence in the changes. I'm currently running through a 10,000 file test-set trying to identify edge-cases. [0]: https://github.com/UglyToad/PdfPig/pull/1102 https://github.com/UglyToad/PdfPig/pull/1102
- farkin88 1y agoThat robustness-vs-throughput trade-off is such a staple of PDF parsing. My guess is that the new path is slower because the recovery scan now always walks the whole byte range and has to inflate any object streams it meets before it can trust the offsets even when the first startxref would have been fine. The 10k-file test set sounds great for confidence-building. Are the failures clustering around certain producer apps like Word, InDesign, scanners, etc.? Or is it just long-tail randomness? Reading the PR, I like the recovery-first mindset. If the common real-world case is that offsets lie, treating salvage as the default is arguably the most spec-conformant thing you can do. Slow-and-correct beats fast-and-brittle for PDFs any day.
- coldcode 1y agoI parsed the original Illustrator format in 1988 or 1989, which is a precursor to PDF. It was simpler than today's PDF, but of course I had zero documentation to guide me. I was mostly interested in writing Illustrator files, not importing them, so it was easier than this.
- sergiotapia 1y agoI did some exploration using LLMs to parse, understand then fill in PDFs. It was brutal but doable. I don't think I could build a "generalized" solution like this without LLMs. The internals are spaghetti! Also, god bless the open source developers. Without them also impossible to do this in a timely fashion. pymupdf is incredible. https://www.linkedin.com/posts/sergiotapia_completed-a-really-tricky-ai-pipeline-today-activity-7351668437078687744-OaNG https://www.linkedin.com/posts/sergiotapia_completed-a-reall...
- diptanu 1y agoDisclaimer - Founder of Tensorlake, we built a Document Parsing API for developers. This is exactly the reason why Computer Vision approaches for parsing PDFs works so well in the real world. Relying on metadata in files just doesn't scale across different source of PDFs. We convert PDFs to images, run a layout understanding model on them first, and then apply specialized models like text recognition and table recognition models on them, stitch them back together to get acceptable results for domains where accuracy is table stakes.
- rkagerer 1y agoSo you've outsourced the parsing to whatever software you're using to render the PDF as an image.
- bee_rider 1y agoSeems like a fairly reasonable decision given all the high quality implementations out there.
- throwaway4496 1y agoHow is it reasonable to render the PDF, rasterize it, OCR it, use AI, instead of just using the "quality implementation" to actually get structured data out? Sounds like "I don't know programming, so I will just use AI".
- do_not_redeem 1y agoPDFs don't always lay out characters in sequence, sometimes they have absolutely positioned individual characters instead. PDFs don't always use UTF-8, sometimes they assign random-seeming numbers to individual glyphs (this is common if unused glyphs are stripped from an embedded font, for example) etc etc
- throwaway4496 1y agoBut all those problems exist when rendering into a surface or rastering. I just don't understand how one thinks, this is a hard problem, let me make it harder by solving the problem into another kind of problem that is just as hard as solving it in the first place (PDF to structured data vs PDF to raster). And then solve the new problem, which is also hard. It is absurd.
- HocusLocus 1y agoThanks kindly for this well done and brave introduction. There are few people these days who'd even recognize the bare ASCII 'Postscript' form of a PDF at first sight. First step is to unroll into ASCII of course and remove the first wrapper of Flate/ZIP,LZW,RLE. I recently teased Gemini for accepting .PDF and not .EPUB (chapterized html inna zip basically, with almost-guaranteed paragraph streams of UTF-8) and it lamented apologetically that its pdf support was opaque and library oriented. That was very human of it. Aside from a quick recap of the most likely LZW wrapper format, a deep dive into Lineariziation and reordering the objects by 'first use on page X' and writing them out again preceding each page would be a good pain project. UglyToad is a good name for someone who likes pain. ;-)
- userbinator 1y agoAs someone who has written a PDF parser - it's definitely one of the weirdest formats I've seen, and IMHO much of it is caused by attempting to be a mix of both binary and text; and I suspect at least some of these weird cases of bad "incorrect but close" xref offsets may be caused by buggy code that's dealing with LF/CR conversions. What the article doesn't mention is a lot of newer PDFs (v1.5+) don't even have a regular textual xref table, but the xref table is itself inside an "xref stream", and I believe v1.6+ can have the option of putting objects inside "object streams" too.
- robmccoll 1y agoYeah I was a little surprised that this didn't go beyond the simplest xref table and get into streams and compression. Things don't seem that bad until you realize the object you want is inside a stream that's using a weird riff on PNG compression and its offset is in an xref stream that's flate compressed that's a later addition to the document so you need to start with a plain one at the end of the file and then consider which versions of which objects are where. Then there's that you can find documentation on 1.7 pretty easily, but up until 2 years ago, 2.0 doc was pay-walled.
- kragen 1y agoYeah, I was really surprised to learn that Paeth prediction really improves the compression ratio of xref tables a lot!
- leeter 1y agoI remember having a prior boss of mine asked if the application the company I was working for made could use PDF as an input. His response was to laugh then say "No, there is no coming back from chaos." The article has only reinforced that he was right.
- brentm 1y agoThis is one of those things that seems like it shouldn't be that hard until you start to dig in.
- ccgreg 1y agoSee https://digitalcorpora.org/corpora/file-corpora/cc-main-2021-31-pdf-untruncated/ https://digitalcorpora.org/corpora/file-corpora/cc-main-2021... for a set of 8 million PDF files from the web, as seen by a single crawl of Common Crawl.
- deleted 1y ago[deleted]
- Beefin 1y agofounder of mixpeek here, we fine-tune late interaction models on pdfs based on domain https://mixpeek.com/extractors https://mixpeek.com/extractors
- sgt 1y agoDo you offer local or on-premise models? There are certain PDF's we cannot send to an API.
- anon-3988 1y agoLast weekend I was trying to convert some PDF of Upanishads which contains some Sanskrit and English word. By god its so annoying, I don't think I would be able to without the help of Claude Code with it just reiterating different libraries and methods over and over again. Can we just write things in markdown from now on? I really, really, really, don't care that the images you put is nicely aligned to the right side and every is boxed together nicely. Just give me the text and let me render it however I want on my end.
- sgt 1y agoWhole point of PDF is that it's digital paper. It's up to the author how he wants to design it, just like a written note or something printed out and handed to you in person.
- jeroenhd 1y agoThe point of PDFs is that you design them once and they look the same everywhere. I do care very much that the heading in my CV doesn't split the paragraph below it. Automatically parsing and extracting text contents from PDFs is not a main feature of the file format, it's an optional addition. PDFs don't compete with Markdown. They're more like PNGs with optional support for screen readers and digital signatures. Maybe SVGs if you go for some of the fancier features. You can turn a PDF into a PNG quite easily with readily available tools, so an alternative file format wouldn't have saved you much work.
- Animats 1y agoCan you just ignore the index and read the entire file to find all the objects?
- UglyToad 1y agoYes this is generally the fallback approach if finding the objects via the index (xref) fails. It is slightly slower but it's a one time cost, though I imagine it was a lot slower back when PDFs were first used on the machines of the time.
- gcanyon 1y agoThe answer seems obvious to me: 1. PDFs support arbitrary attached/included metadata in whatever format you like. 2. So everything that produces PDFs should attach the same information in a machine-friendly format. 3. Then everyone who wants to "parse" the PDF can refer to the metadata instead. From a practical standpoint: my first name is Geoff. Half the resume parsers out there interpret my name as "Geo" and "ff" separately. Because that's how the text gets placed into the PDF. This happens out of multiple source applications.
- jiveturkey 1y agoprobably because ff is rendered as a ligature
- philipwhiuk 1y agoOr could be so is treated as special.
- jeroenhd 1y agoThere's a huge difference between parsing a PDF and parsing the contents of a PDF. Parsing PDF files is its own hell, but because PDFs are basically "stuff at a given position" and often not "well-formed text within boundary boxes", you have to guess what letters belong together if you want to parse the text as a word. If you're interested in helping out the resume parsers, take a look at the accessibility tree. Not every PDF renderer generates accessible PDFs, but accessible PDFs can help shitty AI parsers get their names right. As for the ff problem, that's probably the resume analyzer not being able to cope with non-ASCII text such as the ff ligature. You may be able to influence the PDF renderer not to generate ligatures like that (at the expense of often creating uglier text).
- Aardwolf 1y agoHow would that work for a scan of a handwritten document or similar, assuming scanners / consumer computers don't have perfect OCR?
- 1y ago
- pss314 1y agopdfgrep (as a command line utility) is pretty great if one simply needs to search text in PDF files https://pdfgrep.org/ https://pdfgrep.org/
- constantinum 1y agoOther PDF parsing woes include: 1. Identifying form elements like check boxes and radio buttons. 2. Badly oriented PDF scans 3. Text rendered as bezier curves 4. Images embedded in a PDF 5. Background watermarks 6. Handwritten documents PDF parsing is hell indeed: https://unstract.com/blog/pdf-hell-and-practical-rag-applications/ https://unstract.com/blog/pdf-hell-and-practical-rag-applica...
- v5v3 1y agoThose of you saying OCR and Vision LLM are missing the point. This is an article by a geek for other geeks. Not aimed at solution developers.
- jupin 1y ago> Assuming everything is well behaved and you have a reasonable parser for PDF objects this is fairly simple. But you cannot assume everything is well behaved. That would be very foolish, foolish indeed. You're in PDF hell now. PDF isn't a specification, it's a social construct, it's a vibe. The more you struggle the deeper you sink. You live in the bog now, with the rest of us, far from the sight of God. This put a smile on my face:)
- beng-nl 1y agoCould’ve been written by the great James Mickens.
- AtNightWeCode 1y agoPDF is a format for preserving layouts across different platforms when viewing and printing. It is not intended for data processing and so on. I don't see why a structured document format can't exist that simplifies processing and increases accessibility while still preserving the layouts.
- neuroelectron 1y agoWhat about open office docs? (ODF – OpenDocument Format, like .odt, .ods, .odp) JavaScript in particular is actively hostile to stability and determinism.
- AtNightWeCode 1y agoI have not looked at those formats but take docx for example. That structure is complicated because the layout needs to be described and editable.
- mft_ 1y agoI've been pondering for a while that we need to move away from layout-based written communication. As in, the need to make things look professionally laid out is an anachronism, and is (very) rarely related to comprehension of the actual content. For example, submissions to regulatory agencies are huge documents; we spend lots of time in (typically) Microsoft Word creating documents that follow a layout tradition. Aside from this time spent (wasted), the downside is that to guarantee that layout for the recipient, the file must be submitted in DOCX or PDF. These formats are then unfriendly if you want to do anything programatically with them, extract raw data, etc. And of course, while LLMs can read such files, there's likely a significant computational overhead vs. a file in a simple machine-readable format (e.g. text, markdown, XML, JSON). --- An alternative approach would be to adopt a very simple 'machine first', or 'content first' format - for example, based on JSON, XML, even HTML - with minimum metadata to support strurcture, intra-document links, and embedding of images. For human comsumption, a simple viewer app would reconstitute the file into something more readable; for machine consumption, the content is already directly available. I'm well aware that such formats already exist - HTML/browsers, or EPUB/readers, for example - the issue is to take the rational step towards adopting such a format in place of the legacy alternatives. I'm hoping that the LLM revolutoion will drive us in just this direction, and that in time, expensive parsing of PDFs is a thing of the past.
- xp84 1y agoI’m with you on PDF, but is docx really that bad in practice? I have not implemented a parser for it so I’m not pushing one answer to that. But it seems like it’s an XML-based format that isn’t about absolutely positioning everything unless you explicitly decide to, and intuitively it seems like it should be like an 80 on the parsing easiness scale if a JPEG is a 0, a PDF is a 15, and a markdown is 100.
- Zardoz84 1y agoDocx it's a proprietary format. So it's a direct no
- Anon_troll 1y agoExtracting text from DOCX is easy. Anything related to layout is non-trivial and extremely brittle. To get the layout correct, you need to reverse engineer details down to Word's numerical accuracy so that content appears at the correct position in more complex cases. People like creating brittle documents where a pixel of difference can break the layout and cause content to misalign and appear on separate pages. This will be a major problem for cases like the text saying "look at the above picture" but the picture was not anchored properly and floated to the next page due to rendering differences compared to a specific version of Word.
- ChrisMarshallNY 1y agoI've written TIFF readers. Same sort of deal. It's really easy to write a TIFF; not so easy to read one. Looks like PDF is much the same.
- bjoli 1y agoThe correct answer is, and has always been: Haha. What? Of course I don't. Are you insane?
- sychou 1y agoAmusing, cringey, and also painful that two of our most common formats - PDF and HTML/CSS/JS - are such a challenge to parse and display. Probably a quarter of AI compute power seems to go into understanding just those two.
- pcunite 1y agoBe sure and talk to Derek Noonburg, he knows PDF!
- ulrischa 1y agoParsing a pdf is the most painful thing you can do
- butlike 1y agoParsing PDFs is filed under 'might make me quit on the spot,' depending on the severity of the ask.
- csours 1y agoThe subsequent article "So you want to PRINT a PDF" is stuck in a queue somewhere. Well, I say 'stuck' - it actually got timed out of the queue, but that doesn't raise an error so no one knows about it.
- gethly 1y agoIf microsoft was able to push their docx garbage into being a standard, nothing surprise me any more.
- akaybk 1y ago[dead]