Draft proposal for discussion. The API design and scope may change based on feedback.
Web apps using on-device models should work the same across different browsers and devices. The app's behavior depends on the model and a workflow tested with one may fail with another.
For example, a web app may use on-device models to catalog and manage expense reports. The app might use one model to extract content from the scanned document and another model to give it semantic meaning. For this app to behave the same across devices, it needs to use the right model for the task and be able to use that same model in different browsers.
The Web Models API (navigator.models) lets apps request a specific open-weight model, guaranteeing predictable behavior by using the same model they were built and tested against, regardless of the device or browser. Users control access, and the browser manages execution and shares model resources across sites.
- Let apps request a known model from a registry of open-weight models.
- Manage on-device inference.
- Let users grant and revoke access for each site and model pair.
- Support text, structured output, tool calls, embeddings, image, and audio input.
- Share model weights and loaded resources across sites while keeping each site's data and permissions separate.
- Keep model setup and execution out of page code.
- Allowing sites to use models without the user's permission.
- Allowing sites to load arbitrary model files or install and remove models.
- Choosing a model for the app or substituting one for the model it requested.
- Managing chat history.
- Running tools on the app's behalf.
On-device, open-weight models allow web apps to run AI inference on the user's hardware. Storing weights on the device and sharing across apps saves bandwidth, disk space and time. Local inference can reduce latency, work offline, improve privacy, and cut costs by avoiding sending input to a cloud provider.
Developers pick a specific model by name (model ID), giving them a consistent target, so the app behaves the same regardless of which browser it's running in.
An app developer using the Web Models API has the ability to request the model that's best for their application. A small focused model may be right where speed is critical, while another task will benefit from a larger generalized model.
For example:
- Document search needs an embedding model.
- A receipt reader needs a model that handles images.
- A transcription tool needs a model that handles audio.
- A tool for authoring news articles needs a model that works well in a particular language.
Searching personal notes, messages, or documents should not require sending them to a model provider. On-device models can process them on the user's device, giving users more control over their data.
A local embedding model allows for document search without having to send anything over the network. For example, a query of a notes app such as “when is our launch date” could find a note that says “release moved to October.” The app uses the same model to index notes and search them, while the browser manages the model weights. With the model and documents available locally, this search can also work offline.
Web pages access the API through navigator.models. Each root inference call takes one model ID. A session can keep that ID bound for later calls.
| Method | Purpose | Returns |
|---|---|---|
requestAccess(model) |
Ask to use one model. | A promise for the model ID and grant status. |
generate(model, input) |
Get a complete answer. | A promise for a generation result. |
stream(model, input) |
Read output as it arrives. | An async iterable with a completed promise. |
embed(model, input) |
Turn text into vectors. | A promise for vectors, model ID, and usage. |
session(model, options) |
Reuse a model across calls. | A promise for a session. |
info(model) |
Read details after a grant. | A promise for model info. |
list() |
Read the caller's granted models. | A promise for a list of model info. |
The site asks for access when the user starts a task.
const model = "<model-id>";
writeButton.addEventListener("click", async () => {
if (!("models" in navigator)) {
showMessage("This browser does not support Web Models.");
return;
}
try {
const access = await navigator.models.requestAccess(model);
if (access.status !== "granted") {
showMessage("Model access was not granted.");
return;
}
const result = await navigator.models.generate(model, {
messages: [
{ role: "system", content: "Use short, clear sentences." },
{ role: "user", content: draftInput.value }
],
maxOutputTokens: 200
});
output.textContent = result.data ?? "";
if (result.finishReason === "length") {
showMessage("The answer reached its length limit.");
}
} catch (error) {
showMessage("The model could not finish. You can keep editing.");
}
});A string is shorthand for one user message:
const result = await navigator.models.generate(model, "What is a promise?");
console.log(result.data);The first inference call can also request access. It follows the same
user-action and consent rules as requestAccess().
The result contains the answer in data and the model ID in model. Its finishReason tells the app whether the answer finished, reached a length limit, or proposed tool calls. See the result details.
A grant covers the top-level site, the requesting frame's origin, and one model ID. It gives no access to other models.
The browser shows a permissions dialog (similar to a webcam access dialog) detailing the requested model and lets the user allow or deny access. In normal browsing, grants last until revoked or cleared with site data. Private browsing uses temporary grants which expire when the private browsing session is terminated.
The models Permissions Policy feature defaults to self. Cross-origin
frames need delegation:
<iframe src="https://app.example" allow="models"></iframe>The frame can then request access under the usual consent rules. The browser
identifies both the top-level site and the requesting frame. If policy blocks
use, navigator.models can remain present, but its methods fail with
SecurityError.
The Permissions API reports access changes without opening a prompt:
const permission = await navigator.permissions.query({
name: "models",
model
});
permission.addEventListener("change", () => {
if (permission.state !== "granted") {
showMessage("Access to the model has ended.");
}
});After a grant, the site can read the model's public details:
const info = await navigator.models.info(model);
console.log(info.displayName, info.capabilities);
const grantedModels = await navigator.models.list();list() returns only models granted to the caller, or [] if there are none. info() rejects before a grant. Permission queries must also hide whether an ungranted model exists. Checking for navigator.models reveals only whether the API is present.
Model details describe supported inputs and tasks, along with coarse resource limits. See the capability list.
stream() lets the app read output as it arrives. An AbortSignal lets the user stop generation. See the streaming example.
A session binds one model and can keep it loaded between calls. The app still passes the conversation on each call and releases the session when done. See the session example.
Apps can supply a JSON Schema to define the shape of an answer. For example, a receipt app could ask for the merchant, date, and total as separate fields. The API checks that the result matches the schema before returning it. See the structured output example.
The embed() method takes one or more strings and returns a vector for each. Apps can store and compare these vectors, using the same model to keep new inputs compatible with existing data.
const embeddingModel = "<embedding-model-id>";
const result = await navigator.models.embed(embeddingModel, [
"How do I reset my password?",
"Where can I change my account details?"
]);
console.log(result.model, result.vectors);If the model supports it, apps can request tool calls, ask for more reasoning, or send images and audio as input. The app validates any tool calls and runs the requested tools. Apps could use these calls with tools exposed through WebMCP. See the request details.
Mozilla's Prompt API review warns that sites could depend on one model's quirks, making other browsers and model updates harder to support. Web Models aims to make that dependency clear by giving models IDs separate from the browser.
Developers learn model IDs from a registry while building the app. Pages cannot scan the user's model catalog. The browser must use the requested model or fail.
An app developer who has chosen and tested a model should be able to request that same model across browsers. Each model ID refers to a fixed release, including its weight variant, tokenizer, and default settings for preparing inputs and producing outputs. If that definition changes, it gets a new ID. Browsers that support an ID must follow the same definition.
The app developer is also in the best position to decide when to move to a newer model. They can test a new release and choose when their app requests access to it. A browser may stop supporting an older release, but it must fail the request instead of choosing a replacement. The old ID cannot be reused for a different release.
A few details still need work:
- Availability across browsers. Browsers need a way to publish and recognize the same model IDs. Each browser's catalog still limits access.
- Retirement. We still need to decide how browsers announce that a release will no longer be supported and how much notice developers get.
- Fallbacks. Apps should keep unrelated features available when model access fails. Depending on a model offered by only one browser can still exclude users.
- Testing. Tests need to cover the behavior apps rely on, including how models follow instructions, use tools, handle inputs, and keep embeddings compatible under the same ID. Matching method names and valid JSON alone cannot establish compatibility.
We would like browsers to use a shared registry, so developers can request the same model releases across browsers. How that registry should work, who should maintain it, and what browsers would be expected to support remain open questions.
A public model ID should help developers find its license and terms before building with it. The API has no method for reading them, so the registry and developer documentation need to provide them for each model supplier.
Models can take gigabytes to download and substantial memory to run. With a library, each site manages its own model files and runtime. Separate sites may download and store the same model again.
With the Web Models API, a browser can share the same model weight file for use across sites, manage resource use, and let users decide which sites can access each model. These controls belong in the browser, where users can review and revoke access. A library alone cannot provide them.
Revocation stops affected work and destroys sessions. Active calls and streams fail with NotAllowedError. Navigation also ends the document's sessions. Browsers can limit resource use and pause or throttle background work to protect memory, battery, and other tabs.
Apps must handle denied access, resource limits, and unavailable sessions. Errors before a grant must not reveal whether a model exists. See the error reference.
| Approach | Benefit | Tradeoff |
|---|---|---|
| A remote provider | The app chooses a hosted model and service. | Needs network access, service integration, and a way to cover costs. |
| WebAssembly or WebGPU with a library | The app controls model files and execution. | Each site handles downloads, storage, updates, and runtime code. |
| Extend the Prompt API | Reuses an API for browser-managed generation. | Needs rules for named models, their identity, and embeddings. |
| Task-specific AI APIs | The app requests a task such as translation. | Less control for apps tied to a model or embedding space. |
| An extension or library prototype | Lets developers test ideas before standardization. | Needs extra setup and does not establish native browser support. |
The Prompt API already covers streaming and structured output, but leaves model choice to the browser. Apps that depend on a particular model still need a way to request it. Extending the Prompt API with named models and embeddings remains an option. Either approach needs clear rules for what a model ID promises.
- Fingerprinting. A user's installed models could help sites recognize them across visits. Before consent, sites cannot list installed models or check whether a particular model is present.
- Resource abuse. User action, prompt limits, quotas, revocation, and background controls limit unwanted use.
- Cross-site leaks. Implementation must keep each site's conversation data separate.
- Unsafe output. Apps should check tool names and arguments before running actions; treat model output as untrusted data.
- Media and runtime safety. Browsers must check media and isolate model execution. Errors must hide private paths, credentials, raw runtime output, and other sites' data.
Common (currently in private alpha) has an early implementation based on the spec in api-reference.md. Currently, the Common browser code and API implementation tests are not public.
We would like feedback on:
- Model IDs and retirement. Who assigns IDs? How should browsers announce that a release will no longer be supported, and how much notice should developers get?
- Prompt API. Could extending it meet the same needs? Which apps need a separate API, and why?
- User choice. Can users understand, grant, and revoke model access easily?
- Compatibility. Which differences in how browsers run a model are allowed under the same ID? How should we test behavior and embedding compatibility?
- Streams and resource limits. How should backpressure work? How should apps recover after background suspension or memory pressure?
- Workers. Should workers have access to the API, and if so, how should they inherit access?
- Catalogs and terms. How can browsers offer the same IDs? Should a universal registry be strictly enforced? How should developers find model terms?
- API notes and public references
- Mozilla's Prompt API position and discussion
- WebGPU
- Prompt API
- WebAssembly
- Permissions API
- Permissions Policy
The Web Models API explainer and API reference are licensed under the Creative Commons Attribution 4.0 International License.
See LICENSE.md for the full license text.