Document security considerations for HTML translation #78
Description
Activity
+1, and thanks for splitting this out from #60, @reillyeon. Agreed this should be documented as a security consideration now, independently of whether or how we add official HTML-fragment support (#82, Decision 2). The two are complementary: #78 documents the hazard defensively today; #60 decides whether to remove it structurally later.
Sharpening the threat model so the text is precise:
-
The output of
translate()/ the other text APIs is model-generated and untrusted. Because these models emit natural-language text, and text can be HTML-shaped, the output can contain markup even when the input was plain text. The model is not a sanitizer, it makes no safety guarantee about its output. -
The realistic adversary is someone who controls the input being translated (e.g. user-generated content). Crafted input can induce valid markup in the output, which then executes if the developer does
element.innerHTML = await translator.translate(...). -
This is orthogonal to the input-side "don't let context act as instructions" guidance already in the algorithm sections, this is the output side: don't let model output act as markup.
Action Points
- Explainer: Proposed guidance (author-facing; I can add it in the explainer, in a Security Considerations subsection):
Treat the output of these APIs as untrusted text. Do not insert it into a document as HTML (e.g. via
innerHTML,insertAdjacentHTML(),document.write()), or otherwise interpret it as code, without sanitizing it first, useElement.setHTML()/ the Sanitizer API, or assign totextContentwhen structure isn't needed. This holds regardless of whether the input was plain text, and regardless of whether the model runs locally or in the cloud.-
Placeholder in spec: since the Translator/Language Detector spec defers its privacy & security considerations to Writing Assistance APIs §6–§7, I'd suggest the core "untrusted output" text live in the shared WA-APIs Security Considerations (it applies equally to Summarizer/Rewriter/Prompt output, which devs also
innerHTML), with a short translation-specific pointer + theinnerHTMLexample kept here in translation-api so the concrete case is visible.If What is the input **representation**, an opaque *string + format/
mimeTypetag*, or a *structured DOM-to-DOM* transform? #82 , Decision 2 (DOM-to-DOM with a structure-preservation invariant) lands, it's worth noting there as the structural mitigation: the UA reconstructs known-safe structure and only substitutes text nodes, so output can't become new active markup, but this consideration should not wait on that.I've drafted candidate text. Happy to either open that PR or hand you drafted text to land under your name since it's assigned to you, whichever you prefer: your call.
-
RESOLUTION: The group resolves to document the potential injection attack vector in the security considerations section of the Translator and Language Detector APIs specification. (issue #78)
In mozilla/standards-positions#1015 @jakearchibald raises a legitimate concern that developers will attempt to translate HTML fragments with code that amounts to,
This is dangerous because many translation models expect text and not HTML as input. Since HTML is text, treating the output of the translator as HTML risks introducing injection attacks where a maliciously crafted non-HTML input can result in valid HTML output.
While #60 discusses ways in which the API could be extended to officially support translating HTML fragments, the specification should also be amended to document this attack vector as a security consideration.