Community Discussion · Tracks

Papers as corpus: Building a batch extraction pipeline in two days

Can't Finish Reading PapersCan't Finish Reading PapersAug 172026/08/17 344 views

I spent two days trying to batch-process my lab's backlog of papers into a structured corpus. It started when I saw news about A-share corpus-related stocks skyrocketing, with copyright holders like Readers Publishing being chased by capital. The logic is that high-quality text for large AI models is nearly exhausted. I looked down at the 500+ PDFs on my desk—isn't this exactly what a corpus is? It's just too messy, and no one has organized it.

2 replies

?
Ctrl + Enter to reply
Pixel Dust

The thing about paths containing Chinese characters is so real... I tried using WorkBuddy for batch processing before and got stuck on this too. The error messages left me completely confused, and in the end, it turned out to be the folder names' fault. Spent ages dealing with such a trivial issue 😅

hongtao
hongtaoAug 17

PDF extraction is the real trap here. Scanned versions just spit out garbled text everywhere... I ran a batch through n8n last week and spent half a day just on cleaning; otherwise, you're feeding the model pure garbage. I've also fallen into that path issue, plus file permission problems—basically, once the script starts running, all sorts of weird glitches pop up.