A DOCX and DOC engine written from scratch in Rust: parse, extract text, Markdown and HTML, lay out and render pages to PNG, PPM, BMP or JPEG, write DOCX. One core, a CLI, a terminal explorer and pythonic bindings.
Docs · Benchmarks · Conformance ledger · Failure modes
Reading a Word document should not require LibreOffice, a JVM or a C library. docboss is a clean-room reader built from the specifications (ECMA-376 for DOCX, Microsoft's [MS-DOC], [MS-CFB], [MS-OLEPS], [MS-ODRAW] and [MS-OFFCRYPTO] for the binary formats, PKWARE's APPNOTE for ZIP): safe Rust, no C dependencies, its own ZIP, XML, compound-file, font and raster code, one core behind the CLI, the terminal explorer and the Python extension. It is a lenient reader: a ZIP with a broken central directory is read from its local headers, a compound file with a bad FAT chain is followed as far as it goes, and every approximated or dropped item is reported instead of lost.
- Both formats, one model: DOCX, DOCM, DOTX and DOTM through ECMA-376 WordprocessingML (transitional and strict namespaces), DOC and DOT through [MS-DOC], including Word 97-2003 formatting, lists, tables, sections, notes, comments, fields, bookmarks and pictures, and the formatting, sections and tables of Word 6 and 95 files. The format is detected from the bytes, never from the file name.
- Fast: DOCX text at 3,304 files a second over 80 real-world files, about 4× docx2txt and 28× mammoth; DOC text at 3,130 files a second, about 19× antiword and catdoc; 13,839 DOCX files a second through
docboss.extract_textson 12 cores (benchmarks). - Complete text: word recall of 0.992 against LibreOffice's text export on DOCX, the highest of the six engines measured, and 0.985 on DOC: tables, headers, footers, footnotes, comments and text boxes included.
- Survives damaged files: text from 140 of 150 damaged DOCX files, where the other engines return text for 1 to 8, and from 142 of 150 damaged DOC files, where catdoc manages 100; no crash and no hang on 300 damaged DOCX and DOC files.
- Markdown and HTML: headings from styles, nested lists with the document's own labels, GFM tables with merged cells, footnotes, links and images, in one pass.
- Layout and rendering: Word-like line breaking, the Unicode bidirectional algorithm for right-to-left paragraphs, runs and tables, OpenType shaping (GSUB and GPOS, Arabic joining) for Arabic and Hebrew, tab stops, list labels, keep and widow rules, tables split across pages with repeated header rows, sections and columns, headers and footers with page numbers, footnotes, inline and floating images; an anti-aliased rasterizer over its own TrueType and CFF parsers, with metric-compatible font substitution (Carlito for Calibri, Liberation for Arial and Times New Roman). PNG, PPM, BMP and JPEG output, pages rendered on all cores.
- DOCX writing: model to DOCX, a builder API, Markdown to DOCX, and DOC to DOCX conversion, with deterministic output: the same input gives identical bytes.
- Encrypted files: password-protected DOCX (Agile and Standard encryption, [MS-OFFCRYPTO]) and DOC (RC4, RC4 CryptoAPI and, for Word 97 and later files, XOR obfuscation).
- Range-fetching I/O: documents open over files or
http(s)://URLs and fetch only the byte ranges they need: the ZIP central directory and the XML parts of a DOCX, the sectors of the Word streams of a DOC. - Conformance ledger checked by Lean: every clause of the specifications docboss reads has a row with a status, and a Lean 4 gate fails the build when a row claims more than the code and tests cite (conformance).
- Terminal explorer (
docboss tui): element tree, resolved-style inspector, ZIP or compound-file container view with hex and XML, Markdown and page previews.
pip install docboss # abi3 wheels (CPython 3.12+)
cargo install docboss-cli # the `docboss` binaryCoding agents get a bundled skill: docboss skill install writes it into ./.claude/skills/docboss (--global for ~/.claude/skills/docboss).
docboss info report.docx # format, metadata, page setup, page count, content counts
docboss text report.docx # plain text with list labels; --headers-footers, --comments, --no-notes
docboss text legacy.doc # the same for Word 97-2003 files
docboss md report.docx # GitHub-flavored Markdown; --images DIR exports and links images
docboss html report.docx --standalone # semantic HTML, images embedded as data: URIs
docboss json report.docx --blocks # the compact per-block view; omit --blocks for the whole model
docboss q report.docx '.metadata.title' -r # jq over the JSON model
docboss render report.docx --page 1 -o page.png --scale 2 # or .ppm / .bmp / .jpg; all pages with -o 'page-%d.png'
docboss images report.docx -o images # the document's images at their stored size and format
docboss convert legacy.doc -o legacy.docx # DOC or DOCX to a fresh DOCX, or .md/.html/.txt/.json by extension
docboss create md notes.md -o notes.docx # CommonMark + GFM composed into a DOCX
docboss fonts report.docx # each font the document names and the face it resolves to here
docboss diagnostics report.docx # everything the lenient reader approximated or dropped
docboss parts report.docx # ZIP entries or compound-file streams, with sizes
docboss hex legacy.doc WordDocument --length 64 # hexdump a part
docboss xml report.docx word/styles.xml # a part, indented
docboss text locked.docx --password secret # encrypted DOCX or DOC
docboss tui report.docx # interactive terminal explorerEvery read command accepts a local path or an http(s):// URL; a URL is read with range requests, not downloaded whole.
import docboss
doc = docboss.Document("report.docx") # or Document(data=raw_bytes), password="..."
print(doc.format, doc.metadata.title) # "docx" or "doc"
text = doc.extract_text(headers_footers=True)
md = doc.extract_markdown() # images="reference" | "embed" | "omit"
html = doc.extract_html(standalone=True)
for block in doc.blocks(): # kind, style, heading_level, list_label, text
print(block.kind, block.text)
png = doc.render(0, scale=2.0) # PNG bytes; format="ppm" | "bmp" | "jpeg"
pngs = doc.render_pages() # every page, rendered on all cores
imgs = doc.images() # .name, .content_type, .data
data = doc.to_docx() # a fresh DOCX, also for .doc input
texts = docboss.extract_texts(["report.docx", "legacy.doc"], threads=8) # batch, on Rust threads
docx = docboss.md.to_docx("# Notes\n\n- one\n- two\n") # Markdown -> DOCX bytesHeavy calls release the GIL, so documents extract and render in parallel from Python threads. extract_texts is the throughput path for pipelines: one call, a Rust thread pool, a list of strings back.
docboss.AsyncDocument reads files, bytes or http(s):// URLs asynchronously, fetching only the byte ranges each call needs:
import asyncio
import docboss
async def main() -> None:
doc = await docboss.AsyncDocument.open_url("https://example.com/report.docx")
text = await doc.extract_text()
print(doc.format, doc.bytes_fetched, "bytes fetched in", len(doc.requests), "requests")
full = await doc.document() # the synchronous Document, for rendering and the rest
asyncio.run(main())Rust
The library crates: docboss-core (open and detect), docboss-model (the document model), docboss-output (text, Markdown, HTML, JSON), docboss-layout and docboss-render (pages), docboss-write (DOCX), docboss-aio (async and remote reads), plus the readers underneath (docboss-docx, docboss-doc, docboss-zip, docboss-xml, docboss-cfb, docboss-crypt, docboss-font, docboss-metafile, docboss-mtef).
use docboss_output::{to_markdown, to_text, MarkdownOptions, TextOptions};
fn main() -> Result<(), Box<dyn std::error::Error>> {
let doc = docboss_core::open("report.docx")?; // DOCX or DOC, detected from the bytes
println!("{}", to_text(&doc, &TextOptions::default()));
println!("{}", to_markdown(&doc, &MarkdownOptions::default()));
let layout = docboss_layout::layout_document(&doc);
println!("{} pages", layout.pages.len());
docboss_render::render_page(&layout, 0, 2.0)?.save("page.png")?;
docboss_write::save(&doc, "copy.docx")?; // deterministic DOCX
Ok(())
}Composing a document:
use docboss_write::{DocumentBuilder, ListKind, Para, TableBuilder};
let mut doc = DocumentBuilder::new();
doc.title("Q3 Report").heading(1, "Summary");
doc.paragraph(Para::new().text("Revenue grew ").bold("12%").text("."));
doc.items(ListKind::Bullet, &["North", "South"]);
let width = doc.text_width();
doc.block(TableBuilder::new(2, width).header(&["Region", "Growth"]).row(&["North", "14%"]));
let bytes = doc.to_bytes()?;
# Ok::<(), docboss_write::Error>(())Reading a remote DOCX with range requests (docboss-aio with the http feature):
use docboss_aio::{AsyncDocument, ReadOptions};
# async fn run() -> docboss_aio::Result<()> {
let remote = AsyncDocument::open_url("https://example.com/report.docx").await?;
let doc = remote.read(&ReadOptions::default()).await?;
println!("{} of {} bytes fetched", remote.bytes_fetched(), remote.len());
# Ok(()) }Over 80 real-world DOCX files from the LibreOffice, Apache POI and python-docx test suites (word recall against LibreOffice's text export of each file):
| Engine | Text files/s | Word recall | Markdown files/s | HTML files/s |
|---|---|---|---|---|
| docboss | 3,304 | 0.992 | 3,324 | 3,355 |
| docx2txt | 757 | 0.974 | ||
| python-docx | 725 | 0.848 | ||
| docx2python | 127 | 0.974 | ||
| mammoth | 117 | 0.956 | 109 | 111 |
| pandoc | 7.4 | 0.894 | 7.6 | 7.5 |
Over 80 DOC files: docboss 3,130 files/s (recall 0.985), catdoc 161 (0.888), antiword 158 (0.948). LibreOffice converts either sample at 15 to 16 files/s in one batch, or about 0.7 to 0.8 s per file one process at a time.
Parallel, large documents, memory, damaged files, and where the others win
- Parallel:
docboss.extract_textsreads 300 DOCX files at 13,839 files/s on 12 cores (5.1× its sequential speed); the fastest other route is docx2txt in a process pool at 1,262. On DOC, docboss reaches 5,685 files/s against catdoc's 490. - One large document (3,000 sections, 1,188 pages): docboss extracts the DOCX in 85 ms against docx2txt's 808 ms and python-docx's 1,490 ms. catdoc wins the DOC version, 54 ms against docboss's 105 ms: it streams the text out of the piece table, where docboss builds its whole model first.
- Memory: on DOCX docboss peaks at 33.2 MB over 100 files, docx2txt slightly lower at 30.1 MB, the rest 55 to 82 MB. On DOC antiword and catdoc win, at about 24 MB with 2 to 3 MB of their own, against docboss's 40 MB (64 MB on a 6.2 MB file whose pictures take 12 MB of it).
- Damaged files (150 per format): docboss returns text for 140 DOCX files, the other engines for 1 to 8, and for 142 DOC files against catdoc's 100 and antiword's 71; catdoc hung once and antiword crashed 3 times. docboss neither crashed nor hung.
Measured on an Apple M3 Pro, one session, best of 3 after a warm-up. Method, versions and every table: benchmarks/README.md.
Pages rendered by docboss against LibreOffice's rendering of the same documents (LibreOffice to PDF, rasterized by pdfboss), 60 DOCX and 30 DOC files from the same corpora (two of the DOC files are docboss test fixtures), up to 5 pages each, windowed SSIM at scale 1.5:
| Format | SSIM median | SSIM p10 | Page counts equal |
|---|---|---|---|
| DOCX | 0.985 | 0.879 | 56 of 60 |
| DOC | 0.966 | 0.634 | 27 of 30 |
The p10 is the order statistic (the score a tenth of the files fall to or below); page counts compare docboss's full page count with LibreOffice's, and a file LibreOffice could not convert counts as a mismatch. The low tail is fonts LibreOffice substitutes differently, tables and shapes LibreOffice places differently, text in table cells that LibreOffice wraps around pictures and docboss does not, and tracked changes that LibreOffice shows with revision marks. Method: benchmarks/README.md.
Crate map
| Crate | Responsibility |
|---|---|
docboss-model |
The document model both readers produce: sections, paragraphs, runs, tables, styles with property resolution, numbering with list labels, notes, comments, media, diagnostics |
docboss-zip |
ZIP reader and deterministic writer: central directory, ZIP64, data descriptors, slicing-by-8 CRC-32, recovery from local headers, zip-bomb limits (APPNOTE) |
docboss-xml |
Zero-copy pull XML tokenizer with namespace resolution; transitional and strict OOXML namespaces map to the same ids |
docboss-docx |
OPC packages and WordprocessingML (ECMA-376 Parts 1 to 3), parts parsed on separate threads |
docboss-cfb |
Compound files ([MS-CFB]) and property sets ([MS-OLEPS]) |
docboss-doc |
Word binary documents ([MS-DOC], [MS-ODRAW] pictures), Word 6 and 95 formatting, RC4 and RC4 CryptoAPI decryption |
docboss-crypt |
Password-protected DOCX ([MS-OFFCRYPTO] Agile and Standard) and the XOR obfuscation of DOC |
docboss-core |
Format detection and one open/read over both readers |
docboss-output |
Text, Markdown, HTML, JSON and the per-block view |
docboss-font |
TrueType, OpenType CFF and collections, GSUB and GPOS layout features; system font discovery and metric-compatible substitution |
docboss-layout |
Pages from the model: line breaking, tabs, lists, tables, sections, headers and footers, footnotes, images |
docboss-metafile |
WMF ([MS-WMF]) and EMF ([MS-EMF]) pictures played into paths, text and bitmaps; DIB decoding |
docboss-mtef |
Equation Editor 3 equations: the MTEF of an OLE object's Equation Native stream read into the math model |
docboss-render |
Anti-aliased rasterizer and the PNG, PPM, BMP and JPEG encoders |
docboss-write |
DOCX from the model, a builder API, Markdown to DOCX |
docboss-aio |
Async reads over files or HTTP, fetching only the byte ranges needed |
docboss-cli |
The docboss command-line tool |
docboss-tui |
The terminal explorer |
docboss-py |
The PyO3 extension module (docboss._docboss) |
The reader is lenient and it says so: docboss diagnostics (the diagnostics property in Python, the diagnostics field of Document in Rust) lists every item that was approximated or dropped. The largest gaps:
- Layout: shaping covers Arabic, Hebrew and combining marks, not Indic, Thai or other scripts that reorder glyphs; right-to-left and vertical sections lay out left to right; no column balancing; text in table cells and text frames does not wrap around the floats anchored there, header and footer floats do not push body text aside, tables beside a float are not narrowed, and turned drawings wrap as if unturned (each reported by
diagnostics --layout); picture and pattern fills and shape effects are not drawn; 3-D, radar, stock, surface, bubble and bar-of-pie charts and SmartArt without a saved drawing show an empty frame; conditional formatting from table styles is not applied. CFF2 fonts are refused and TIFF images draw a placeholder. - DOCX reading:
w:altChunkcontent is skipped and reported; ActiveX controls are not read as content; ruby text is ignored. - DOC reading: password-protected Word 6 and 95 files are refused with an error; Word 6 and 95 files lose their comments, endnotes, drawing objects and equations, and Word 2 files are read as text only; freeforms and WordArt are reported as dropped.
- Equations (DOCX and DOC): MathType (MTEF 4 and 5) objects keep their picture but not their text, and are reported.
- Writing: embedded fonts are not written (the font table lists names only), no theme part is written, and raw HTML in Markdown is dropped except
<br>.
The conformance ledger lists every clause with its status and a note saying what is missing.
ledger/ is a Lean 4 project. lake exe ledger-index scans the Rust and Python sources for specification citations (ECMA-376 Part 1 §17.3.1.29, [MS-DOC] §2.5.1, APPNOTE §4.3.7, ...), and the gate's theorems hold each ledger row to them: an implemented row needs a code citation and a test citation, an incomplete one a code citation and a note, every required clause of each standard needs a row, every row must name a real clause by its real title, and a citation of a clause the standard does not have fails the build. The outlines of the nine standards are generated from their published texts. make ledger-gate runs it; make ledger-report regenerates the conformance table. The gate also checks Lean reference implementations of CRC-32 and the list number formats, whose differential vectors the Rust tests are held to.
failure-modes/ keeps a before and after render for each rendering bug fixed, named after the commit that fixed it, with the failure described in measured terms.
make ci # rustfmt check, clippy -D warnings, tests, rustdoc
make test-py # build the extension and run the Python tests
make ledger-gate # the Lean conformance gate
make help # everything elseLicensed under either of Apache License, Version 2.0 or MIT license at your option.