Skip to content

Counting and mining research data with Unix: suggested code fixes (LLM-assisted review) #3867

Description

@drjwbaker

I'm testing how useful language models are for supporting lesson maintenance. The fixes below were suggested by an LLM (Claude) after it reviewed this lesson's code: commands that no longer work or that behave differently from what the text describes, plus the text changes that follow from them. The commands were run against the Figshare dataset in bash 5.2 with GNU grep 3.11 and GNU coreutils 9.4 on Linux, and the counts match the lesson. The Windows points are based on how Git Bash behaves by default and weren't tested on Windows. Please treat these as suggestions. Feedback on whether this kind of review is useful is welcome.

Lesson: [Counting and mining research data with Unix](https://programminghistorian.org/en/lessons/research-data-with-unix). Line numbers refer to en/lessons/research-data-with-unix.md at 5bdece6.

  1. L51–57: the Figshare zip unpacks to data/ with no proghist/ folder. Tell readers to create ~/proghist and unzip the file into it.
  2. L51–55, L67, L77, L93, L111: c:\proghist\… fails in Git Bash because \ is an escape character. The replacement in the April 2025 note (c/Users/…) is missing its leading slash. Use ~/proghist/… on every platform and remove the note.
  3. L57, L67, L77, L93: /user/USERNAME/… and ~/users/USERNAME/… don't exist. Use ~/proghist/….
  4. L87–89: wc -c counts bytes and wc -m counts characters on every platform, so the "OS X/Linux use -m" note is inaccurate. Use -m in the text and note that -c counts bytes. The two give the same number for this ASCII data.
  5. L107, L111, L113: when grep searches several files, it puts the filename at the start of each line, which corrupts the first column of the saved subsets. Add -h (grep -ih …, grep -ivh …), explain it in one line, and add it to the Summary (L122).
  6. L111–113: saving the output to a .csv file doesn't convert it, and the contents are still tab-separated. Reword the text: spreadsheet programs use the file extension to decide how to split columns.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions