Skip to content

Modernize “Working with batches of PDF files”: OCR pipeline, batch scripts, and topic modelling #3860

Description

@maehr

I would like to propose a substantive maintenance update to Working with batches of PDF files.

I am the original author of the lesson and would be happy to prepare the revision and PR. The lesson was published in 2020 and its methodological core is still useful, but the installation instructions, OCR assumptions, batch-processing commands, and topic-modelling walkthrough now need an end-to-end re-test.

This would build on earlier maintenance work in #2501, #3314, #3367, and #3659. In particular, the discussion in #2501 already noted that the lesson would likely benefit from a broader sustainability revisit.

Why this update is needed

1. Re-evaluate the OCR pipeline

The lesson currently says that OCRmyPDF “automatically skips PDFs that already contain embedded text” and demonstrates commands such as:

ocrmypdf --language eng --deskew --clean 'ILO-SR_N2_engl.pdf' 'ILO-SR_N2_engl.pdf'

and:

find . -name '*.pdf' -exec ocrmypdf --language eng --deskew --clean '{}' '{}' \;

This no longer describes current OCRmyPDF behaviour accurately. Current OCRmyPDF distinguishes explicitly between handling existing text/OCR, including modes for skipping, redoing, or forcing OCR. The lesson should explain this choice rather than imply that existing text is always skipped automatically.

The revised lesson should also stop using the same pathname for input and output. Even where a tool permits this, a teaching workflow should preserve the source PDFs and make intermediate/output files explicit.

I suggest restructuring the case study around separate directories, for example:

input/
ocr/
text/

and then testing a non-destructive batch workflow against PDFs with:

  • no text layer;
  • a usable existing text layer;
  • an existing but poor OCR layer;
  • mixed pages.

The exact commands should be finalized only after re-testing against the current OCRmyPDF release.

The update should also reconsider whether --clean belongs in the default command. OCRmyPDF treats image cleaning as an optional processing step with additional dependencies, and such transformations deserve explanation and visual quality control rather than being silently applied to every file.

2. Refresh platform installation instructions

The lesson still assumes Windows 10 with an Ubuntu 18.04 WSL setup, Ubuntu 18.04 on Linux, and an obsolete Homebrew installation command on macOS.

The prerequisites should be rewritten and tested for currently supported environments.

Rather than hard-coding an old operating-system release, the lesson should:

  • link to maintained installation instructions for the major dependencies;
  • give only the minimum commands needed for tested environments;
  • state which versions/platforms were used during the lesson re-test;
  • separate required dependencies from optional ones such as additional Tesseract language packs or image-cleaning tools.

3. Re-test and simplify the PDF/text tooling

Poppler remains appropriate for pdftotext, pdfimages, and pdfunite, but the shell examples should be reviewed for predictable filenames, safe handling of paths, and clearer failure behaviour.

For example, the current batch extraction:

find . -name '*.pdf' -exec pdftotext '{}' '{}.txt' \;

produces names such as document.pdf.txt. The revised workflow should produce intentional .txt filenames in a dedicated output directory.

The recursive grep examples should likewise search the extracted text corpus rather than every file in a directory that also contains PDFs and other assets.

The ImageMagick section needs a separate decision. ImageMagick 7 uses the magick command rather than the older convert command, and PDF read/write behaviour can be affected by security-policy settings. We should either:

  1. update and clearly document the current ImageMagick workflow; or
  2. replace/remove ImageMagick where a smaller, purpose-built tool would make the image-to-PDF example more sustainable.

pdfunite can remain the straightforward example for joining PDF files.

4. Make the corpus-download step more robust

The current ILO case-study download command scrapes a directory listing with curl | grep | uniq | sed | xargs. This is compact but brittle and has already needed maintenance when ILO URLs changed (#3314, #3367).

For a teaching lesson, I suggest replacing this with either:

  • a small checked-in URL manifest that the shell downloads; or
  • a short, readable download script with explicit error handling and a documented expected file count.

As part of the update, all ILO sample links and the Zenodo prepared-corpus link/DOI should be re-verified.

5. Replace the DARIAH Topics Explorer v2 walkthrough with Simple Topic Modeling

The current lesson teaches the DARIAH Topics Explorer v2 workflow. Its latest v2 desktop binaries date from 2018, while the repository now states that version 3 is being developed as a complete reimplementation.

I propose replacing this part of the lesson with Simple Topic Modeling, while keeping the conceptual transition from PDF → OCR/text extraction → corpus analysis.

This is a good fit for the lesson because the application:

  • runs analysis locally in the browser, so corpus text is not sent to an analysis server;
  • accepts multiple UTF-8 text files directly, with one file treated as one document;
  • deliberately rejects PDF input, so the OCR/text-extraction workflow taught earlier in the lesson remains necessary;
  • supports NMF with TF-IDF as the default model and LDA with counts as an alternative;
  • provides configurable stop-word handling;
  • exposes topic, document, similarity, metadata, and diagnostic views;
  • exports document-topic scores, topics, topic terms, topic similarities, and a config.json;
  • stores a random seed and configuration so a run can be reproduced.

This also requires updating the methodological explanation rather than only changing screenshots. In particular:

  • do not define topic modelling only through LDA probabilities;
  • introduce NMF + TF-IDF as the default workflow and LDA as an optional probabilistic model;
  • replace “remove the 150 most common words” with an explanation of language stop words, custom stop words, and document-frequency filtering;
  • replace the fixed “30 topics / 200 iterations” recipe with an exploratory workflow for choosing a topic count and inspecting diagnostics;
  • revise the claim that topic modelling is simply “non-deterministic”: the new application fixes and records a random seed by default, while learners can deliberately change the seed to investigate model stability;
  • show how to inspect representative documents as well as top terms before assigning an interpretation to a topic;
  • document the export/reproducibility workflow using the corpus plus config.json.

The screenshots and the “Create the Topic Model” / “Evaluate the Topic Model” sections would need to be replaced accordingly.

One dependency should be explicit: the current Simple Topic Modeling repository reports version 2.0.0a0. I would therefore treat a stable tagged/released build as a prerequisite for merging the lesson update, rather than pointing a published lesson at an alpha development state.

Proposed scope

This would be a focused update of the existing lesson, not a retirement or a completely new methodological lesson. The revised version should preserve the main learning path:

assess PDFs → OCR where necessary → extract text/images → batch-process safely → inspect/search the plain-text corpus → run and evaluate a topic model

but re-test every technical step with current software and replace outdated assumptions and screenshots.

Acceptance criteria

  • Re-test the full lesson from a clean setup on current macOS and Linux, plus a documented Windows route (native and/or WSL).
  • Record the software versions used for the re-test.
  • Do not overwrite source PDFs in any example command.
  • Explain and test OCRmyPDF behaviour for PDFs with no text, existing text, and existing OCR.
  • Decide whether --clean should remain, and if so document its optional dependency and visual-QA implications.
  • Re-test every Poppler/ImageMagick (or replacement) command and update obsolete syntax.
  • Make batch commands produce predictable output filenames/directories and handle failures visibly.
  • Restrict text-search examples to the extracted text corpus.
  • Replace or harden the ILO bulk-download script and verify the expected corpus size.
  • Verify all ILO and Zenodo data links.
  • Replace the DARIAH Topics Explorer v2 instructions and screenshots with a stable release of Simple Topic Modeling.
  • Update the topic-modelling explanation for NMF/LDA, preprocessing, random seeds, diagnostics, interpretation, and reproducibility.
  • Confirm that the lesson's generated text corpus can be loaded directly into Simple Topic Modeling and that its result/config exports work as described.
  • Update figures/screenshots and any associated lesson assets.
  • Re-test the complete case study end to end before merge.

If the editorial team agrees with this scope, I can prepare the revised lesson and associated assets in a PR.

PS: This issue was written with teh help of an LLM (and therefore is quite verbose). Sowwy.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions