Skip to content

[Target: Wasm & TS] WebAssembly Streaming Runtime (polyxml-wasm) for Node.js, Bun, and Browsers #47

Description

@nth-bailey

🎯 Motivation & Background

Currently, PolyXML supports TypeScript interface and schema generation (Zod, Valibot, TypeBox). However, at runtime in JavaScript environments (Node.js, Bun, Cloudflare Workers, modern browsers), developers still rely on pure JavaScript XML parsers (e.g. fast-xml-parser, xml2js, or DOMParser).

In high-throughput microservices and data-intensive browser apps (financial reconciliation, geospatial GIS, telemetry):

  1. Allocation pressure: Compare the heap and GC effects of JS parsers with Wasm byte transfer and object conversion.
  2. Throughput: Measure end-to-end parsing on representative document sizes.
  3. Engine parity: Make PolyXML's Rust transcoder available in Node, Bun, and browsers while measuring Wasm-specific costs.

🛠️ Proposed Solution

Create a new workspace crate: crates/polyxml-wasm utilizing wasm-bindgen and wasm-pack:

1. Incremental Streaming API

Accept byte input as Uint8Array / ArrayBuffer, and process a web ReadableStream incrementally for repeated records:

import { createPolyXml } from '@polyxml/wasm';

const polyxml = await createPolyXml();
const completeDocument = polyxml.xmlToJson(xmlBytes);
for await (const record of polyxml.parseStream(readableStream)) {
  consume(record);
}

2. Fast Path Serialization / Transcoding

Provide in-Wasm transcoding without allocating intermediate JS objects:

// XML <-> JSON directly inside Wasm linear memory
const jsonBytes = polyxml.xmlToJsonBytes(xmlBytes);

🔬 Benchmarking & Tradeoff Analysis

To rigorously evaluate whether the Wasm boundary crossing is worth the architectural cost, a dedicated benchmark suite must be created.

Benchmark Setup & Tooling

  • Harness: reproducible timed runs across Node.js, Bun, and headless Chromium.
  • Payload Tiers:
    • Micro (1 KB - 5 KB): Small single-record payloads (REST / Webhook messages).
    • Medium (500 KB - 2 MB): Typical batch documents (e.g., SEPA pacs.008 payments).
    • Large (50 MB - 200 MB): Heavy telemetry, NeTEx public transit dumps, or defense track logs.
  • Competitors: fast-xml-parser, xml2js, native browser DOMParser.

Metrics to Measure

  1. Throughput (MB/s and ops/sec): End-to-end time to deserialize into JS objects.
  2. Memory Footprint & GC Pauses: Heap deltas (process.memoryUsage().heapUsed) and GC observations where the runtime exposes reliable measurements; report limitations.
  3. Bundle Size Impact: report Wasm and JS wrapper bytes and gzip sizes; distinguish these from an application bundle.

Hypothesized Tradeoffs & Evaluation Criteria

  • The "Small Payload" Boundary Penalty:
    • Hypothesis: For tiny XML payloads (< 5 KB), JavaScript-to-Wasm FFI boundary crossing and string copying into Wasm linear memory may be slower than a lightweight pure JS parser.
    • Decision Rule: If pure JS is faster for micro-payloads, document a hybrid threshold recommendation or provide an inline zero-overhead fast-path.
  • The "Large Payload" Evaluation:
    • Hypothesis to test: Wasm may improve throughput or reduce JS heap pressure for large documents; copies and output object creation still have costs.
    • Decision rule: Report measured throughput, end-to-end allocations, and GC pauses. Document where Wasm helps or hurts; no fixed speedup or GC threshold is required.

✅ Acceptance Criteria

Publishing the package to npm is tracked separately in #71.

Activity

  1. nth-bailey commented on Sep 23, 2026

    @nth-bailey
    CollaboratorAuthor

    Implementation is underway in PR #69. The Rust core builds for wasm32-unknown-unknown, and the first package exposes schema-free XML↔JSON conversion plus reusable inline-XSD typed conversion. The npm artifact imports in Node and Bun; Vite and Webpack production builds emit the Wasm asset.

    Two scope corrections from the prototype:

    • wasm-bindgen copies Uint8Array inputs into Wasm memory and copies output bytes back to JavaScript. The proposed API cannot honestly be called zero-copy across that boundary.
    • The current API accepts complete documents. True incremental ReadableStream parsing needs a separate core/parser design; buffering all chunks and parsing at the end would only change the input shape, not provide streaming memory behavior.

    PR #69 intentionally keeps this issue open. The remaining work is incremental streaming, representative Node/Bun/browser benchmarks at the proposed payload sizes, GC/memory measurements, and a documented decision about when Wasm helps. The 3× throughput and 60% GC thresholds in the original description remain hypotheses until those measurements exist.

  2. nth-bailey commented on Sep 23, 2026

    @nth-bailey
    CollaboratorAuthor

    PR #70 implements the record-stream API and Node/Bun/Chromium benchmark suite, with measured results in docs/benchmarks/wasm-vs-js.md. It is open for CI and review. Streaming handles repeated direct children with bounded per-record buffering; typed XSD streaming and splitting one huge record are outside this API. npm publication remains a separate release decision after the PR lands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions