Conversation
Library: - Smart tokenizer for code styles (camelCase, _, -, ., / boundaries) - Real slug (diacritics folded, ASCII only); kebab keeps Unicode - Unicode-aware title, sentence, alternate; grapheme-safe reverse/flip - Conventional ASCII leet map, case preserved - Empty string returns '' instead of false - strict option throws TypeError/RangeError - capstringAll() and CATEGORIES exports - New styles: smallcaps, bubble, wide, strike, clap, morse, binary, piglatin Tooling: - CLI (cli.js + bin/capstring.js) with --all, --list, --json, stdin - Hand-written index.d.ts with Style union; drift test - files whitelist (drop .npmignore), sideEffects: false, prepublishOnly - 100% coverage thresholds, Node 24 in CI, CLI smoke test Docs: README regrouped with behavior notes, CHANGELOG 1.1.0, CLAUDE.md, AGENTS.md
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: da7e603570
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| .replace(/['’]/gu, '') | ||
| .replace(/([\p{Ll}\p{N}])(\p{Lu})/gu, '$1 $2') | ||
| .replace(/(\p{Lu})(\p{Lu}\p{Ll})/gu, '$1 $2') | ||
| .replace(/[^\p{L}\p{N}]+/gu, ' ') |
There was a problem hiding this comment.
Preserve combining marks in the tokenizer
When a code style receives canonically decomposed text, this replacement treats every combining mark as a separator. For example, NFD Crème Brûlée becomes cre-me-bru-le-e under kebab, while the canonically equivalent NFC input becomes crème-brûlée; camel, acronym, and the other word-based styles are similarly corrupted. Include marks as part of their adjacent word or normalize before tokenizing so the library remains Unicode-safe.
AGENTS.md reference: AGENTS.md:L5-L6
Useful? React with 👍 / 👎.
| case 'slug': | ||
| // url-slug (ASCII only, diacritics folded) | ||
| return slugify(str); |
There was a problem hiding this comment.
Route slug through camel-case tokenization
For camelCase input, this branch bypasses the shared code-style tokenizer and lowercases before inserting any boundaries: capstring('helloWorld', 'slug') returns helloworld and XMLHttpRequest returns xmlhttprequest. This contradicts the documented behavior that all code styles tokenize camelCase, so preserve word boundaries before applying the slug-specific ASCII folding.
AGENTS.md reference: AGENTS.md:L50-L50
Useful? React with 👍 / 👎.
|
|
||
| case 'strike': | ||
| // s̶t̶r̶i̶k̶e̶ | ||
| return chars(str).map((ch) => (/\s/u.test(ch) ? ch : ch + '̶')).join(''); |
There was a problem hiding this comment.
Apply strike marks per grapheme
For multi-code-point emoji, adding the combining strike after every code point breaks the grapheme: 👨👩👧 becomes three graphemes and 🇺🇸 becomes two regional indicators instead of a flag. Iterate the existing grapheme helper rather than chars so the new style does not violate the library's emoji-safe guarantee.
AGENTS.md reference: AGENTS.md:L5-L6
Useful? React with 👍 / 👎.
| result = lower + 'way'; | ||
| } else { | ||
| // Leading consonant cluster: `qu` moves as a unit; `y` is a vowel unless it leads the word | ||
| const cluster = /^(qu|[^aeiouy]+|[^aeiou]+)/.exec(lower)[0]; |
There was a problem hiding this comment.
Keep qu with a preceding consonant cluster
For words where qu follows another leading consonant, the alternation stops before u, so squeal, square, and squid become uealsqay, uaresqay, and uidsqay instead of moving the full squ cluster (ealsquay, etc.). The qu alternative only handles words beginning directly with qu; extend the cluster match so qu remains a unit after preceding consonants too.
Useful? React with 👍 / 👎.
| return str.toLowerCase().replace(/(^["'([]*|[.!?]+["')\]]*\s+)(\p{L})/gu, (_, lead, letter) => | ||
| lead + letter.toUpperCase() |
There was a problem hiding this comment.
Skip opening quotes after sentence boundaries
When a sentence starts with a quote after punctuation, such as hello. "world", the pattern cannot reach the next letter because it expects the letter immediately after whitespace; the result is Hello. "world" rather than Hello. "World". Account for opening quotes or brackets after the boundary whitespace as well as closing punctuation before it.
Useful? React with 👍 / 👎.
- Tokenizer NFC-normalizes and keeps combining marks (NFD 'Crème' no longer splits into cre-me) - slug goes through the shared tokenizer so camelCase splits (helloWorld -> hello-world, XMLHttpRequest -> xml-http-request) - strike applies the combining mark per grapheme so ZWJ emoji and flags survive - piglatin keeps 'qu' with a preceding consonant cluster (squeal -> ealsquay) - sentence capitalizes after an opening quote or bracket that follows a boundary
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 894f410ccc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // Sentence case - capitalize first letter only | ||
| return str.charAt(0).toUpperCase() + str.substring(1).toLowerCase(); | ||
| // Sentence case - capitalize the first letter and the first letter after . ! ? | ||
| return str.toLowerCase().replace(/(^["'([]*|[.!?]+["')\]]*\s+["'([]*)(\p{L})/gu, (_, lead, letter) => |
There was a problem hiding this comment.
Recognize Unicode quotes at sentence boundaries
When sentence punctuation is followed by a typographic quote, capstring('hello. “world”', 'sentence') still returns Hello. “world” because the boundary classes skip only ASCII quotes and brackets. The fresh evidence in this revision is that the newly added quote-handling pattern on this line explicitly excludes common Unicode quotation marks such as “”‘’; include those marks so sentence casing remains Unicode-aware.
AGENTS.md reference: AGENTS.md:L5-L6
Useful? React with 👍 / 👎.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b0f54f6acc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| .replace(/['’]/gu, '') | ||
| .replace(/([\p{Ll}\p{N}])(\p{Lu})/gu, '$1 $2') | ||
| .replace(/(\p{Lu})(\p{Lu}\p{Ll})/gu, '$1 $2') | ||
| .replace(/[^\p{L}\p{N}\p{M}]+/gu, ' ') |
There was a problem hiding this comment.
Remove marks orphaned by discarded emoji
When a code style receives an emoji containing a variation selector, such as ❤️ or ☕️, this replacement discards the base symbol but retains U+FE0F because it is a \p{M} character. Consequently, capstring('❤️', 'kebab') returns an invisible non-empty string and hashtag returns #️, despite the tokenizer promising no words for input without letters or digits. The fresh case beyond the prior combining-mark report is specifically marks whose base was removed; retain marks only when attached to a retained letter or number.
AGENTS.md reference: AGENTS.md:L5-L6
Useful? React with 👍 / 👎.
|
|
||
| case 'strike': | ||
| // s̶t̶r̶i̶k̶e̶ | ||
| return graphemes(str).map((g) => (/^\s$/u.test(g) ? g : g + '\u0336')).join(''); |
There was a problem hiding this comment.
Treat CRLF graphemes as whitespace
When strike processes Windows-style multiline text, Intl.Segmenter returns \r\n as one grapheme, but ^\s$ only matches a single whitespace code point. Thus capstring('a\r\nb', 'strike') produces a̶\r\n̶b̶, inserting an orphan combining strike at the start of the next line. The fresh case after switching this style to grapheme iteration is the multi-code-point CRLF grapheme; test the whole grapheme as one or more whitespace characters.
Useful? React with 👍 / 👎.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d52be1ee04
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| * @param {(cp: number) => number|undefined} shift - Returns the new code point, or undefined to keep | ||
| * @returns {string} Transformed string | ||
| */ | ||
| const mapCodePoints = (str, shift) => chars(str).map((ch) => { |
There was a problem hiding this comment.
Keep keycap emoji atomic in art styles
When bubble or wide receives a keycap emoji, this code maps its ASCII base independently from the variation selector and enclosing-keycap marks: for example, capstring('1️⃣', 'bubble') produces ①️⃣, while capstring('#️⃣', 'wide') produces #️⃣. Those are no longer standard Unicode keycap emoji sequences, so these styles violate the library's emoji-safe guarantee; iterate graphemes and preserve multi-code-point emoji rather than mapping their component code points.
AGENTS.md reference: AGENTS.md:L5-L6
Useful? React with 👍 / 👎.
| io.stdout(`${JSON.stringify(results, null, 2)}\n`); | ||
| } else { | ||
| const width = Math.max(...STYLES.map((s) => s.length)); | ||
| io.stdout(STYLES.map((s) => `${s.padEnd(width)} ${results[s]}`).join('\n') + '\n'); |
There was a problem hiding this comment.
Escape multiline results in plain --all output
When --all reads multiline stdin, styles such as same, title, and upper retain the embedded newlines, but this formatter inserts each raw result into a supposedly single output row. For input like hello\nworld, the command emits extra unlabeled lines interleaved with subsequent style rows, so users cannot reliably associate output with styles; escape embedded line breaks or otherwise delimit each result structurally.
Useful? React with 👍 / 👎.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9f0114dd32
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| // Normalize words (split on any whitespace, filter empty) | ||
| const words = str.toLowerCase().split(/\s+/).filter(w => w); | ||
| const words = toWords(str); |
There was a problem hiding this comment.
Defer tokenization until a code style needs it
For every nonempty input, this eagerly performs normalization, multiple Unicode-regex passes, splitting, filtering, and lowercasing before the selected style is known. Consequently, passthrough styles such as the default same and proper, as well as constant-output none, now scan and allocate several copies of large strings even though they immediately return the original string or ''; slug also repeats the entire tokenization inside slugify. Compute words only within the code-style paths and reuse them for slug to avoid this substantial performance regression.
Useful? React with 👍 / 👎.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9921cd7366
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| .replace(/([\p{Ll}\p{N}])(\p{Lu})/gu, '$1 $2') | ||
| .replace(/(\p{Lu})(\p{Lu}\p{Ll})/gu, '$1 $2') | ||
| .replace(/(^|[^\p{L}\p{N}\p{M}])\p{M}+/gu, '$1') // drop marks orphaned by a removed base (❤️ -> U+FE0F) | ||
| .replace(/[^\p{L}\p{N}\p{M}]+/gu, ' ') |
There was a problem hiding this comment.
Treat keycap emoji as tokenizer separators
When a code style receives a keycap such as 1️⃣, this allow-list retains the ASCII digit together with U+FE0F and U+20E3, so snake returns the emoji unchanged and hashtag returns #1️⃣ even though toWords documents emoji as separators. Beyond the previously covered orphaned-selector case, the retained digit is fresh evidence that these marks evade the cleanup because they are attached to an allowed code point; identify keycap graphemes before admitting letters and digits.
AGENTS.md reference: AGENTS.md:L6-L6
Useful? React with 👍 / 👎.
| } else { | ||
| const width = Math.max(...STYLES.map((s) => s.length)); | ||
| // One row per style, so line breaks inside a result are shown escaped | ||
| const oneLine = (v) => v.replace(/\r/g, '\\r').replace(/\n/g, '\\n'); |
There was a problem hiding this comment.
Escape every Unicode line separator in --all output
When stdin contains U+2028 or U+2029, styles such as same preserve it and this formatter emits it verbatim, splitting the supposedly single style row; for example, a\u2028b places b on an unlabeled line. The revised formatter now handles CR and LF, but this fresh Unicode multiline case remains, so escape Unicode line/paragraph separators as well.
AGENTS.md reference: AGENTS.md:L6-L6
Useful? React with 👍 / 👎.
Summary
Makes capstring serious and robust while staying fun. Stays on 1.x (1.1.0). No style names removed.
Library
XMLHttpRequest hello_world→xml-http-request-hello-worldslug(diacritics folded, ASCII only):Crème Brûlée & Co.→creme-brulee-co.kebabkeeps Unicode letters.title,sentence(capitalizes after. ! ?),alternatereverse/flip: emoji, flags, combining accents surviveleetmap, case preserved''(wasfalse);{ strict: true }throws instead of silently returningcapstringAll()andCATEGORIESsmallcaps,bubble,wide,strike,clap,morse,binary,piglatin(37 total)Tooling
npx capstring kebab "Hello World",--all,--list,--json, reads stdinindex.d.tswith aStyleunion; a test fails if it drifts fromSTYLESfileswhitelist replaces.npmignore;sideEffects: false;prepublishOnlyruns lint + coveragenpm pack --dry-runin CIAll behavior changes are listed in CHANGELOG under "Behavior changes".
Test plan
npm run lintcleannpm run test:coverage94 tests, 100% statements / branches / functions / linesnpm pack --dry-runlists exactly index.js, index.d.ts, cli.js, bin/, CHANGELOG.md, README.md, LICENSE, package.jsonnode bin/capstring.js --all "XMLHttpRequest 😀"andecho "hello. world" | node bin/capstring.js sentence