PDF to Text, a challenging problem

(www.marginalia.nu)

357 points ingve | 1 comments | 13 May 25 15:01 UTC | HN request time: 0.264s | source

Show context

xnx ◴[13 May 25 15:44 UTC] No.43974208[source]▶

Weird that there's no mention of LLMs in this article even though the article is very recent. LLMs haven't solved every OCR/document data extraction problem, but they've dramatically improved the situation.

replies(5): >>43974229 #>>43974325 #>>43974337 #>>43974562 #>>43975686 #

marginalia_nu ◴[13 May 25 15:57 UTC] No.43974337[source]▶

>>43974208 #

Author here: LLMs are definitely the new gold standard for smaller bodies of shorter documents.

The article is in the context of an internet search engine, the corpus to be converted is of order 1 TB. Running that amount of data through an LLM would be extremely expensive, given the relatively marginal improvement in outcome.

replies(2): >>43974639 #>>43977353 #

1. noosphr ◴[13 May 25 20:30 UTC] No.43977353[source]▶

>>43974337 #

A PDF corpus with a size of 1tb can mean anything from 10,000 really poorly scanned documents to 1,000,000,000 nicely generated latex pdfs. What matters is the number of documents, and the number of pages per document.

For the first I can run a segmentation model + traditional OCR in a day or two for the cost of warming my office in winter. For the second you'd need a few hundred dollars and a cloud server.

Feel free to reach out. I'd be happy to have a chat and do some pro-bono work for someone building a open source tool chain and index for the rest of us.

↑