I would like to propose a substantive maintenance update to Working with batches of PDF files.
I am the original author of the lesson and would be happy to prepare the revision and PR. The lesson was published in 2020 and its methodological core is still useful, but the installation instructions, OCR assumptions, batch-processing commands, and topic-modelling walkthrough now need an end-to-end re-test.
This would build on earlier maintenance work in #2501, #3314, #3367, and #3659. In particular, the discussion in #2501 already noted that the lesson would likely benefit from a broader sustainability revisit.
Why this update is needed
1. Re-evaluate the OCR pipeline
The lesson currently says that OCRmyPDF “automatically skips PDFs that already contain embedded text” and demonstrates commands such as:
ocrmypdf --language eng --deskew --clean 'ILO-SR_N2_engl.pdf' 'ILO-SR_N2_engl.pdf'
and:
find . -name '*.pdf' -exec ocrmypdf --language eng --deskew --clean '{}' '{}' \;
This no longer describes current OCRmyPDF behaviour accurately. Current OCRmyPDF distinguishes explicitly between handling existing text/OCR, including modes for skipping, redoing, or forcing OCR. The lesson should explain this choice rather than imply that existing text is always skipped automatically.
The revised lesson should also stop using the same pathname for input and output. Even where a tool permits this, a teaching workflow should preserve the source PDFs and make intermediate/output files explicit.
I suggest restructuring the case study around separate directories, for example:
and then testing a non-destructive batch workflow against PDFs with:
- no text layer;
- a usable existing text layer;
- an existing but poor OCR layer;
- mixed pages.
The exact commands should be finalized only after re-testing against the current OCRmyPDF release.
The update should also reconsider whether --clean belongs in the default command. OCRmyPDF treats image cleaning as an optional processing step with additional dependencies, and such transformations deserve explanation and visual quality control rather than being silently applied to every file.
2. Refresh platform installation instructions
The lesson still assumes Windows 10 with an Ubuntu 18.04 WSL setup, Ubuntu 18.04 on Linux, and an obsolete Homebrew installation command on macOS.
The prerequisites should be rewritten and tested for currently supported environments.
Rather than hard-coding an old operating-system release, the lesson should:
- link to maintained installation instructions for the major dependencies;
- give only the minimum commands needed for tested environments;
- state which versions/platforms were used during the lesson re-test;
- separate required dependencies from optional ones such as additional Tesseract language packs or image-cleaning tools.
3. Re-test and simplify the PDF/text tooling
Poppler remains appropriate for pdftotext, pdfimages, and pdfunite, but the shell examples should be reviewed for predictable filenames, safe handling of paths, and clearer failure behaviour.
For example, the current batch extraction:
find . -name '*.pdf' -exec pdftotext '{}' '{}.txt' \;
produces names such as document.pdf.txt. The revised workflow should produce intentional .txt filenames in a dedicated output directory.
The recursive grep examples should likewise search the extracted text corpus rather than every file in a directory that also contains PDFs and other assets.
The ImageMagick section needs a separate decision. ImageMagick 7 uses the magick command rather than the older convert command, and PDF read/write behaviour can be affected by security-policy settings. We should either:
- update and clearly document the current ImageMagick workflow; or
- replace/remove ImageMagick where a smaller, purpose-built tool would make the image-to-PDF example more sustainable.
pdfunite can remain the straightforward example for joining PDF files.
4. Make the corpus-download step more robust
The current ILO case-study download command scrapes a directory listing with curl | grep | uniq | sed | xargs. This is compact but brittle and has already needed maintenance when ILO URLs changed (#3314, #3367).
For a teaching lesson, I suggest replacing this with either:
- a small checked-in URL manifest that the shell downloads; or
- a short, readable download script with explicit error handling and a documented expected file count.
As part of the update, all ILO sample links and the Zenodo prepared-corpus link/DOI should be re-verified.
5. Replace the DARIAH Topics Explorer v2 walkthrough with Simple Topic Modeling
The current lesson teaches the DARIAH Topics Explorer v2 workflow. Its latest v2 desktop binaries date from 2018, while the repository now states that version 3 is being developed as a complete reimplementation.
I propose replacing this part of the lesson with Simple Topic Modeling, while keeping the conceptual transition from PDF → OCR/text extraction → corpus analysis.
This is a good fit for the lesson because the application:
- runs analysis locally in the browser, so corpus text is not sent to an analysis server;
- accepts multiple UTF-8 text files directly, with one file treated as one document;
- deliberately rejects PDF input, so the OCR/text-extraction workflow taught earlier in the lesson remains necessary;
- supports NMF with TF-IDF as the default model and LDA with counts as an alternative;
- provides configurable stop-word handling;
- exposes topic, document, similarity, metadata, and diagnostic views;
- exports document-topic scores, topics, topic terms, topic similarities, and a
config.json;
- stores a random seed and configuration so a run can be reproduced.
This also requires updating the methodological explanation rather than only changing screenshots. In particular:
- do not define topic modelling only through LDA probabilities;
- introduce NMF + TF-IDF as the default workflow and LDA as an optional probabilistic model;
- replace “remove the 150 most common words” with an explanation of language stop words, custom stop words, and document-frequency filtering;
- replace the fixed “30 topics / 200 iterations” recipe with an exploratory workflow for choosing a topic count and inspecting diagnostics;
- revise the claim that topic modelling is simply “non-deterministic”: the new application fixes and records a random seed by default, while learners can deliberately change the seed to investigate model stability;
- show how to inspect representative documents as well as top terms before assigning an interpretation to a topic;
- document the export/reproducibility workflow using the corpus plus
config.json.
The screenshots and the “Create the Topic Model” / “Evaluate the Topic Model” sections would need to be replaced accordingly.
One dependency should be explicit: the current Simple Topic Modeling repository reports version 2.0.0a0. I would therefore treat a stable tagged/released build as a prerequisite for merging the lesson update, rather than pointing a published lesson at an alpha development state.
Proposed scope
This would be a focused update of the existing lesson, not a retirement or a completely new methodological lesson. The revised version should preserve the main learning path:
assess PDFs → OCR where necessary → extract text/images → batch-process safely → inspect/search the plain-text corpus → run and evaluate a topic model
but re-test every technical step with current software and replace outdated assumptions and screenshots.
Acceptance criteria
If the editorial team agrees with this scope, I can prepare the revised lesson and associated assets in a PR.
PS: This issue was written with teh help of an LLM (and therefore is quite verbose). Sowwy.
I would like to propose a substantive maintenance update to Working with batches of PDF files.
I am the original author of the lesson and would be happy to prepare the revision and PR. The lesson was published in 2020 and its methodological core is still useful, but the installation instructions, OCR assumptions, batch-processing commands, and topic-modelling walkthrough now need an end-to-end re-test.
This would build on earlier maintenance work in #2501, #3314, #3367, and #3659. In particular, the discussion in #2501 already noted that the lesson would likely benefit from a broader sustainability revisit.
Why this update is needed
1. Re-evaluate the OCR pipeline
The lesson currently says that OCRmyPDF “automatically skips PDFs that already contain embedded text” and demonstrates commands such as:
and:
This no longer describes current OCRmyPDF behaviour accurately. Current OCRmyPDF distinguishes explicitly between handling existing text/OCR, including modes for skipping, redoing, or forcing OCR. The lesson should explain this choice rather than imply that existing text is always skipped automatically.
The revised lesson should also stop using the same pathname for input and output. Even where a tool permits this, a teaching workflow should preserve the source PDFs and make intermediate/output files explicit.
I suggest restructuring the case study around separate directories, for example:
and then testing a non-destructive batch workflow against PDFs with:
The exact commands should be finalized only after re-testing against the current OCRmyPDF release.
The update should also reconsider whether
--cleanbelongs in the default command. OCRmyPDF treats image cleaning as an optional processing step with additional dependencies, and such transformations deserve explanation and visual quality control rather than being silently applied to every file.2. Refresh platform installation instructions
The lesson still assumes Windows 10 with an Ubuntu 18.04 WSL setup, Ubuntu 18.04 on Linux, and an obsolete Homebrew installation command on macOS.
The prerequisites should be rewritten and tested for currently supported environments.
Rather than hard-coding an old operating-system release, the lesson should:
3. Re-test and simplify the PDF/text tooling
Poppler remains appropriate for
pdftotext,pdfimages, andpdfunite, but the shell examples should be reviewed for predictable filenames, safe handling of paths, and clearer failure behaviour.For example, the current batch extraction:
produces names such as
document.pdf.txt. The revised workflow should produce intentional.txtfilenames in a dedicated output directory.The recursive
grepexamples should likewise search the extracted text corpus rather than every file in a directory that also contains PDFs and other assets.The ImageMagick section needs a separate decision. ImageMagick 7 uses the
magickcommand rather than the olderconvertcommand, and PDF read/write behaviour can be affected by security-policy settings. We should either:pdfunitecan remain the straightforward example for joining PDF files.4. Make the corpus-download step more robust
The current ILO case-study download command scrapes a directory listing with
curl | grep | uniq | sed | xargs. This is compact but brittle and has already needed maintenance when ILO URLs changed (#3314, #3367).For a teaching lesson, I suggest replacing this with either:
As part of the update, all ILO sample links and the Zenodo prepared-corpus link/DOI should be re-verified.
5. Replace the DARIAH Topics Explorer v2 walkthrough with Simple Topic Modeling
The current lesson teaches the DARIAH Topics Explorer v2 workflow. Its latest v2 desktop binaries date from 2018, while the repository now states that version 3 is being developed as a complete reimplementation.
I propose replacing this part of the lesson with Simple Topic Modeling, while keeping the conceptual transition from PDF → OCR/text extraction → corpus analysis.
This is a good fit for the lesson because the application:
config.json;This also requires updating the methodological explanation rather than only changing screenshots. In particular:
config.json.The screenshots and the “Create the Topic Model” / “Evaluate the Topic Model” sections would need to be replaced accordingly.
One dependency should be explicit: the current Simple Topic Modeling repository reports version
2.0.0a0. I would therefore treat a stable tagged/released build as a prerequisite for merging the lesson update, rather than pointing a published lesson at an alpha development state.Proposed scope
This would be a focused update of the existing lesson, not a retirement or a completely new methodological lesson. The revised version should preserve the main learning path:
but re-test every technical step with current software and replace outdated assumptions and screenshots.
Acceptance criteria
--cleanshould remain, and if so document its optional dependency and visual-QA implications.If the editorial team agrees with this scope, I can prepare the revised lesson and associated assets in a PR.
PS: This issue was written with teh help of an LLM (and therefore is quite verbose). Sowwy.