Skip to content

Add Monospace Layout toggle for position-preserving text output - #1

Open
luw2007 wants to merge 2 commits into
batu3384:mainfrom
luw2007:feature/monospace-output-preset
Open

luw2007 wants to merge 2 commits into
batu3384:mainfrom
luw2007:feature/monospace-output-preset

Conversation

@luw2007

@luw2007 luw2007 commented Sep 10, 2026

Copy link
Copy Markdown

Summary

Adds a Monospace Layout toggle that preserves the exact horizontal and vertical positions of OCR text blocks using space padding. When enabled, captured text is rendered as monospace-aligned output — ideal for terminal screenshots, table layouts, and any UI where spatial alignment matters.

This is implemented as a toggle (Settings → General → Monospace Layout), not a separate output preset, so it composes with any existing output format.

How it works

The algorithm uses each OCR block's normalized boundingBox:

  1. Row clustering: blocks sorted by y (top-to-bottom), grouped into rows with a 0.02 threshold
  2. Horizontal alignment: each block's x position maps to a fixed column grid (col = Int(x * columns)), with space padding to reach that column
  3. Vertical spacing: median row height estimates blank lines between rows (capped at 6)

Default column count is 120, adjustable via a slider (40–200).

Changes

File Change
Models/AppModels.swift MonospaceLayoutSettings struct + UserDefaults store
AppState.swift @Published monospaceLayoutSettings, setters, persist
Models/OCRResult.swift monospaceAlignedText(columns:) method
Services/CaptureCoordinator.swift resolvedRawText(for:) applies monospace when enabled; works for normal capture and watch mode
Views/Settings/SettingsTabViews.swift Toggle + column slider in General tab
Views/SettingsView.swift Bindings and callbacks
ScreenTextGrabTests/OCRResultTests.swift 9 new tests

Design decisions

  • Toggle, not preset: The monospace transformation needs OCR block boundingBox data, which isn't available in the formatter layer (it only receives raw text). Applying it at the coordinator level keeps the formatter/writer interface unchanged and lets users combine monospace alignment with any output preset.
  • Fixed columns, not char-width estimation: Using a fixed column grid (col = Int(x * columns)) produces more predictable alignment than estimating per-character width from block sizes. The column count is user-adjustable.
  • No interface changes to CaptureOutputFormatter/CaptureOutputWriter: The toggle is applied before text reaches the formatter.

Testing

  • 9 unit tests covering: horizontal position preservation, multi-line output, blank line insertion, empty/single block, x-sorting within rows, explicit column counts, and column clamping.
  • All existing tests remain unchanged.
  • swiftc -typecheck passes with zero errors on the full project.

Notes for reviewer

  • The algorithm is a Swift port of a standalone OCR layout script; the core logic is OCR-engine-agnostic and only requires (x0, y0, x1, y1) per block.
  • A follow-up PR will add Umi-OCR-style gap-tree multi-column sorting as an optional enhancement, which composes with this monospace alignment.

Add a new .monospace CaptureOutputPreset that uses OCR block
boundingBox positions to produce monospace-aligned text:
- y-clustering groups blocks into lines
- x-sorting orders blocks within each line
- x-differences convert to spaces to preserve horizontal position
- y-differences convert to blank lines using median line height

Interface changes:
- CaptureOutputFormatter.format() / clipboardPayload() accept optional
  OCRResult parameter
- CaptureOutputWriter.copyCapturedText accepts optional OCRResult
- CaptureCoordinator protocol and implementations pass OCRResult through
- Main capture flow and watch mode pass the actual OCRResult
- History repaste / menu bar / table review pass nil (no block data)

UI:
- Monospace appears automatically in output preset menus (allCases)
- Menu bar icon: textformat
- Localized titles in Turkish and English

Tests:
- OCRResultTests: 8 new tests covering horizontal alignment, multi-line,
  blank line insertion, empty/single block, x-sorting, explicit columns
- CaptureOutputFormatterTests: 2 new tests for monospace preset with and
  without OCRResult
Replace the standalone .monospace output preset with a toggle-based
feature. When enabled, the capture pipeline uses OCR block boundingBox
positions to produce monospace-aligned text that preserves horizontal
and vertical layout.

Algorithm (ported from ocr-layout.swift):
- Fixed-column horizontal alignment: col = Int(x * columns), pad to col
- y-clustering into rows (threshold 0.02)
- Median row height for blank line estimation
- Vertical gaps convert to blank lines (capped at 6)

Changes:
- AppModels: MonospaceLayoutSettings struct + UserDefaults store
  (isEnabled Bool, columns Int default 120, clamped 20-300)
- AppState: @published monospaceLayoutSettings, setters, persist
- OCRResult: monospaceAlignedText(columns:) using fixed-column algorithm
- CaptureCoordinator: resolvedRawText(for:) applies monospace when toggle ON;
  works for both normal capture and watch mode
- Settings UI: Toggle + column slider (40-200) in General tab,
  under Output Format section
- Tests: 9 new OCRResultTests covering horizontal position, multi-line,
  blank lines, empty/single block, x-sorting, explicit columns, clamping

No changes to CaptureOutputFormatter/CaptureOutputWriter interfaces —
the toggle is applied at the coordinator level before text reaches
the formatter, so it composes with any existing output preset.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant