Skip to content

feat: v1.1.0 - harden library, add CLI, types, and 8 new styles - #8

Open
brianfunk wants to merge 8 commits into
devfrom
feat/v1.1.0-hardening
Open

brianfunk wants to merge 8 commits into
devfrom
feat/v1.1.0-hardening

Conversation

@brianfunk

Copy link
Copy Markdown
Owner

Summary

Makes capstring serious and robust while staying fun. Stays on 1.x (1.1.0). No style names removed.

Library

  • Smart tokenizer for all code styles: XMLHttpRequest hello_world → xml-http-request-hello-world
  • Real slug (diacritics folded, ASCII only): Crème Brûlée & Co. → creme-brulee-co. kebab keeps Unicode letters.
  • Unicode-aware title, sentence (capitalizes after . ! ?), alternate
  • Grapheme-safe reverse / flip: emoji, flags, combining accents survive
  • Conventional ASCII leet map, case preserved
  • Empty string returns '' (was false); { strict: true } throws instead of silently returning
  • New exports capstringAll() and CATEGORIES
  • 8 new styles: smallcaps, bubble, wide, strike, clap, morse, binary, piglatin (37 total)

Tooling

  • CLI: npx capstring kebab "Hello World", --all, --list, --json, reads stdin
  • Hand-written index.d.ts with a Style union; a test fails if it drifts from STYLES
  • files whitelist replaces .npmignore; sideEffects: false; prepublishOnly runs lint + coverage
  • 100% coverage thresholds enforced; Node 24 added to CI; CLI smoke test and npm pack --dry-run in CI

All behavior changes are listed in CHANGELOG under "Behavior changes".

Test plan

  • npm run lint clean
  • npm run test:coverage 94 tests, 100% statements / branches / functions / lines
  • npm pack --dry-run lists exactly index.js, index.d.ts, cli.js, bin/, CHANGELOG.md, README.md, LICENSE, package.json
  • node bin/capstring.js --all "XMLHttpRequest 😀" and echo "hello. world" | node bin/capstring.js sentence

Library:
- Smart tokenizer for code styles (camelCase, _, -, ., / boundaries)
- Real slug (diacritics folded, ASCII only); kebab keeps Unicode
- Unicode-aware title, sentence, alternate; grapheme-safe reverse/flip
- Conventional ASCII leet map, case preserved
- Empty string returns '' instead of false
- strict option throws TypeError/RangeError
- capstringAll() and CATEGORIES exports
- New styles: smallcaps, bubble, wide, strike, clap, morse, binary, piglatin

Tooling:
- CLI (cli.js + bin/capstring.js) with --all, --list, --json, stdin
- Hand-written index.d.ts with Style union; drift test
- files whitelist (drop .npmignore), sideEffects: false, prepublishOnly
- 100% coverage thresholds, Node 24 in CI, CLI smoke test

Docs: README regrouped with behavior notes, CHANGELOG 1.1.0, CLAUDE.md, AGENTS.md
@brianfunk

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-06T04:59:36.839120Z 9921cd7 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: da7e603570

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread index.js Outdated
.replace(/['’]/gu, '')
.replace(/([\p{Ll}\p{N}])(\p{Lu})/gu, '$1 $2')
.replace(/(\p{Lu})(\p{Lu}\p{Ll})/gu, '$1 $2')
.replace(/[^\p{L}\p{N}]+/gu, ' ')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve combining marks in the tokenizer

When a code style receives canonically decomposed text, this replacement treats every combining mark as a separator. For example, NFD Crème Brûlée becomes cre-me-bru-le-e under kebab, while the canonically equivalent NFC input becomes crème-brûlée; camel, acronym, and the other word-based styles are similarly corrupted. Include marks as part of their adjacent word or normalize before tokenizing so the library remains Unicode-safe.

AGENTS.md reference: AGENTS.md:L5-L6

Useful? React with 👍 / 👎.

Comment thread index.js
Comment on lines +364 to +366
case 'slug':
// url-slug (ASCII only, diacritics folded)
return slugify(str);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Route slug through camel-case tokenization

For camelCase input, this branch bypasses the shared code-style tokenizer and lowercases before inserting any boundaries: capstring('helloWorld', 'slug') returns helloworld and XMLHttpRequest returns xmlhttprequest. This contradicts the documented behavior that all code styles tokenize camelCase, so preserve word boundaries before applying the slug-specific ASCII folding.

AGENTS.md reference: AGENTS.md:L50-L50

Useful? React with 👍 / 👎.

Comment thread index.js Outdated

case 'strike':
// s̶t̶r̶i̶k̶e̶
return chars(str).map((ch) => (/\s/u.test(ch) ? ch : ch + '̶')).join('');

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Apply strike marks per grapheme

For multi-code-point emoji, adding the combining strike after every code point breaks the grapheme: 👨‍👩‍👧 becomes three graphemes and 🇺🇸 becomes two regional indicators instead of a flag. Iterate the existing grapheme helper rather than chars so the new style does not violate the library's emoji-safe guarantee.

AGENTS.md reference: AGENTS.md:L5-L6

Useful? React with 👍 / 👎.

Comment thread index.js Outdated
result = lower + 'way';
} else {
// Leading consonant cluster: `qu` moves as a unit; `y` is a vowel unless it leads the word
const cluster = /^(qu|[^aeiouy]+|[^aeiou]+)/.exec(lower)[0];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep qu with a preceding consonant cluster

For words where qu follows another leading consonant, the alternation stops before u, so squeal, square, and squid become uealsqay, uaresqay, and uidsqay instead of moving the full squ cluster (ealsquay, etc.). The qu alternative only handles words beginning directly with qu; extend the cluster match so qu remains a unit after preceding consonants too.

Useful? React with 👍 / 👎.

Comment thread index.js Outdated
Comment on lines +329 to +330
return str.toLowerCase().replace(/(^["'([]*|[.!?]+["')\]]*\s+)(\p{L})/gu, (_, lead, letter) =>
lead + letter.toUpperCase()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Skip opening quotes after sentence boundaries

When a sentence starts with a quote after punctuation, such as hello. "world", the pattern cannot reach the next letter because it expects the letter immediately after whitespace; the result is Hello. "world" rather than Hello. "World". Account for opening quotes or brackets after the boundary whitespace as well as closing punctuation before it.

Useful? React with 👍 / 👎.

- Tokenizer NFC-normalizes and keeps combining marks (NFD 'Crème' no longer
  splits into cre-me)
- slug goes through the shared tokenizer so camelCase splits
  (helloWorld -> hello-world, XMLHttpRequest -> xml-http-request)
- strike applies the combining mark per grapheme so ZWJ emoji and flags survive
- piglatin keeps 'qu' with a preceding consonant cluster (squeal -> ealsquay)
- sentence capitalizes after an opening quote or bracket that follows a boundary
@brianfunk

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 894f410ccc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread index.js Outdated
// Sentence case - capitalize first letter only
return str.charAt(0).toUpperCase() + str.substring(1).toLowerCase();
// Sentence case - capitalize the first letter and the first letter after . ! ?
return str.toLowerCase().replace(/(^["'([]*|[.!?]+["')\]]*\s+["'([]*)(\p{L})/gu, (_, lead, letter) =>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Recognize Unicode quotes at sentence boundaries

When sentence punctuation is followed by a typographic quote, capstring('hello. “world”', 'sentence') still returns Hello. “world” because the boundary classes skip only ASCII quotes and brackets. The fresh evidence in this revision is that the newly added quote-handling pattern on this line explicitly excludes common Unicode quotation marks such as “”‘’; include those marks so sentence casing remains Unicode-aware.

AGENTS.md reference: AGENTS.md:L5-L6

Useful? React with 👍 / 👎.

@brianfunk

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b0f54f6acc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread index.js
.replace(/['’]/gu, '')
.replace(/([\p{Ll}\p{N}])(\p{Lu})/gu, '$1 $2')
.replace(/(\p{Lu})(\p{Lu}\p{Ll})/gu, '$1 $2')
.replace(/[^\p{L}\p{N}\p{M}]+/gu, ' ')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Remove marks orphaned by discarded emoji

When a code style receives an emoji containing a variation selector, such as ❤️ or ☕️, this replacement discards the base symbol but retains U+FE0F because it is a \p{M} character. Consequently, capstring('❤️', 'kebab') returns an invisible non-empty string and hashtag returns #️, despite the tokenizer promising no words for input without letters or digits. The fresh case beyond the prior combining-mark report is specifically marks whose base was removed; retain marks only when attached to a retained letter or number.

AGENTS.md reference: AGENTS.md:L5-L6

Useful? React with 👍 / 👎.

Comment thread index.js Outdated

case 'strike':
// s̶t̶r̶i̶k̶e̶
return graphemes(str).map((g) => (/^\s$/u.test(g) ? g : g + '\u0336')).join('');

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Treat CRLF graphemes as whitespace

When strike processes Windows-style multiline text, Intl.Segmenter returns \r\n as one grapheme, but ^\s$ only matches a single whitespace code point. Thus capstring('a\r\nb', 'strike') produces a̶\r\n̶b̶, inserting an orphan combining strike at the start of the next line. The fresh case after switching this style to grapheme iteration is the multi-code-point CRLF grapheme; test the whole grapheme as one or more whitespace characters.

Useful? React with 👍 / 👎.

@brianfunk

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d52be1ee04

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread index.js Outdated
* @param {(cp: number) => number|undefined} shift - Returns the new code point, or undefined to keep
* @returns {string} Transformed string
*/
const mapCodePoints = (str, shift) => chars(str).map((ch) => {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep keycap emoji atomic in art styles

When bubble or wide receives a keycap emoji, this code maps its ASCII base independently from the variation selector and enclosing-keycap marks: for example, capstring('1️⃣', 'bubble') produces ①️⃣, while capstring('#️⃣', 'wide') produces #️⃣. Those are no longer standard Unicode keycap emoji sequences, so these styles violate the library's emoji-safe guarantee; iterate graphemes and preserve multi-code-point emoji rather than mapping their component code points.

AGENTS.md reference: AGENTS.md:L5-L6

Useful? React with 👍 / 👎.

Comment thread cli.js Outdated
io.stdout(`${JSON.stringify(results, null, 2)}\n`);
} else {
const width = Math.max(...STYLES.map((s) => s.length));
io.stdout(STYLES.map((s) => `${s.padEnd(width)} ${results[s]}`).join('\n') + '\n');

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Escape multiline results in plain --all output

When --all reads multiline stdin, styles such as same, title, and upper retain the embedded newlines, but this formatter inserts each raw result into a supposedly single output row. For input like hello\nworld, the command emits extra unlabeled lines interleaved with subsequent style rows, so users cannot reliably associate output with styles; escape embedded line breaks or otherwise delimit each result structurally.

Useful? React with 👍 / 👎.

@brianfunk

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9f0114dd32

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread index.js Outdated

// Normalize words (split on any whitespace, filter empty)
const words = str.toLowerCase().split(/\s+/).filter(w => w);
const words = toWords(str);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Defer tokenization until a code style needs it

For every nonempty input, this eagerly performs normalization, multiple Unicode-regex passes, splitting, filtering, and lowercasing before the selected style is known. Consequently, passthrough styles such as the default same and proper, as well as constant-output none, now scan and allocate several copies of large strings even though they immediately return the original string or ''; slug also repeats the entire tokenization inside slugify. Compute words only within the code-style paths and reuse them for slug to avoid this substantial performance regression.

Useful? React with 👍 / 👎.

@brianfunk

Copy link
Copy Markdown
Owner Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9921cd7366

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread index.js
.replace(/([\p{Ll}\p{N}])(\p{Lu})/gu, '$1 $2')
.replace(/(\p{Lu})(\p{Lu}\p{Ll})/gu, '$1 $2')
.replace(/(^|[^\p{L}\p{N}\p{M}])\p{M}+/gu, '$1') // drop marks orphaned by a removed base (❤️ -> U+FE0F)
.replace(/[^\p{L}\p{N}\p{M}]+/gu, ' ')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Treat keycap emoji as tokenizer separators

When a code style receives a keycap such as 1️⃣, this allow-list retains the ASCII digit together with U+FE0F and U+20E3, so snake returns the emoji unchanged and hashtag returns #1️⃣ even though toWords documents emoji as separators. Beyond the previously covered orphaned-selector case, the retained digit is fresh evidence that these marks evade the cleanup because they are attached to an allowed code point; identify keycap graphemes before admitting letters and digits.

AGENTS.md reference: AGENTS.md:L6-L6

Useful? React with 👍 / 👎.

Comment thread cli.js Outdated
} else {
const width = Math.max(...STYLES.map((s) => s.length));
// One row per style, so line breaks inside a result are shown escaped
const oneLine = (v) => v.replace(/\r/g, '\\r').replace(/\n/g, '\\n');

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Escape every Unicode line separator in --all output

When stdin contains U+2028 or U+2029, styles such as same preserve it and this formatter emits it verbatim, splitting the supposedly single style row; for example, a\u2028b places b on an unlabeled line. The revised formatter now handles CR and LF, but this fresh Unicode multiline case remains, so escape Unicode line/paragraph separators as well.

AGENTS.md reference: AGENTS.md:L6-L6

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant