See how sure an AI model is about every single word it writes, and change its words while it is still writing.
Download for Windows (.exe) · Linux and macOS · Project site · Docs
Try it in your browser, no install and no API key: every-other-token.vercel.app/play. Type a prompt and watch every other token come out reversed.
every-other-token is a free LLM token stream viewer and interceptor for the command line and the browser. It sits on the live stream from OpenAI, Anthropic, Gemini, OpenRouter or a model on your own machine (Ollama, llama.cpp, vLLM, LM Studio), shows the confidence and perplexity of each token from the logprobs, and can rewrite every other token (or any fraction) as it arrives. A built-in mock provider lets you try all of it with no API key.
Who it's for: anyone curious how language models pick their words, plus people doing LLM interpretability research, red-teaming and prompt engineering.
- Intercept. It opens the model's streaming connection (SSE) itself, so it sees each chunk the moment it arrives.
- Score. Each token gets
confidence = exp(logprob)andperplexity = exp(-logprob), plus the top alternatives the model considered. - Mutate. The chosen tokens (every other one by default, or only the ones the model was unsure about) go through a transform: reverse, uppercase, noise, delete, synonym and more.
- Output. You watch it in the terminal (plain, or a full-screen view with
--tui) or the web UI, or save it as JSON lines, CSV or an HTML heatmap.
Linux (x86_64, Ubuntu 20.04+ / Debian 11+). One line, no dependencies, installs to ~/.local/bin:
mkdir -p ~/.local/bin && curl -fsSL https://gitlab.com/mattbusel/Every-Other-Token/-/releases/permalink/latest/downloads/every-other-token-linux-x86_64.tar.gz | tar xz --strip-components=1 -C ~/.local/bin --wildcards '*/every-other-token'| Other systems | |
|---|---|
| Windows | Download every-other-token-windows-x86_64.exe and run it. (Unsigned, so SmartScreen may ask: More info, then Run anyway.) |
| macOS, or from source | cargo install --locked every-other-token |
Double-click the Windows .exe and the web UI opens in your browser. Every release, with SHA-256 checksums: Releases.
All of these are real output from the mock provider (a fixed reply with fixed logprobs, run through the real pipeline), so you can reproduce them exactly without an API key.
1. Reverse every other token (the default)
$ every-other-token "Why is the sky blue?" --provider mock
The kciuq brown xof jumps revo the yzal dog. This si a kcom response rof prompt: Why si the yks blue?
24 tokens streamed, 12 transformed2. Rewrite only the words the model was unsure about (confidence at or below 0.6)
$ every-other-token "Why is the sky blue?" uppercase --provider mock --rate 1 --min-confidence 0.6
The quick BROWN fox JUMPS over the LAZY dog. THIS is a mock response for PROMPT: WHY IS THE SKY BLUE?
24 tokens streamed, 11 transformed"brown" (0.46), "jumps" (0.57), "lazy" (0.41), "This" (0.54) and "prompt" (0.59) were below the line. The echoed prompt at the end carries no logprobs in the mock, so it falls back to the plain rate.
3. Get the numbers as JSON, one line per token
$ every-other-token "Why is the sky blue?" uppercase --provider mock --json-stream
{"text":"The","original":"The","index":0,"transformed":false,"importance":0.8869204521179199,"confidence":0.88692045,"perplexity":1.1274968,"is_error":false,"arrival_ms":0}
{"text":" QUICK","original":" quick","index":1,"transformed":true,"importance":0.6376281380653381,"confidence":0.63762814,"perplexity":1.5683122,"is_error":false,"arrival_ms":0}
...4. Watch it full-screen in the terminal with every-other-token "Why is the sky blue?" --provider mock --tui. The reply streams in colored by confidence (green sure, yellow unsure, red guessing), rewritten tokens are underlined, and side panels show running stats, a confidence sparkline and the other words the model considered for the latest token. Press q to quit; the reply stays in your terminal.
5. Run it on a model on your own machine. With Ollama running, every-other-token "Why is the sky blue?" --provider ollama streams from llama3.2 with no API key. Any other server that speaks the OpenAI API works too: every-other-token "Why is the sky blue?" reverse my-model --base-url http://localhost:8080/v1.
6. Watch it in the browser with every-other-token --web --provider mock (or just double-click the .exe). Split view: the original stream on the left, the rewritten one on the right, each token underlined by its confidence.
Hosted APIs either send no token probabilities (Anthropic, Gemini) or cannot tell you which words of your prompt mattered. A model running inside the tool can do both. Build with the local feature and it downloads SmolLM2-135M-Instruct once (269 MB, to your cache folder) and runs it on the CPU with candle:
cargo install every-other-token --features local
every-other-token "Reply with only the city name. Capital of France?" --provider local --attributeReal output:
The latipac of ecnarF is siraP.
How much the reply depended on each prompt word (drop in total log-probability when the word is removed):
Reply +1.432 #########
with +1.108 #######
only +1.875 ############
the +0.357 ##
city +3.499 ######################
name. +0.932 ######
Capital +3.229 #####################
of +0.857 #####
France? +4.689 ##############################
Every token carries its exact log-probability and the model's real top-5 alternatives, in the terminal, --tui, the web UI and every export. --attribute removes each prompt word, re-scores the same reply, and reports how much its probability fell (occlusion attribution). --model takes another Llama-architecture model from the Hugging Face Hub, --seed N samples instead of greedy decoding.
Providers that return no probabilities (Anthropic, Gemini, many local servers) show no confidence rather than a made-up one.
- Get it. Use the Linux one-liner or the Windows .exe above (or
cargo install). - Try it offline. Run
every-other-token "Why is the sky blue?" --provider mock, or double-click the .exe and pick Mock (no API key) in the provider menu. - Point it at a real model. Set
OPENAI_API_KEYorANTHROPIC_API_KEY(orOPENROUTER_API_KEY,GEMINI_API_KEY), then runevery-other-token "Why is the sky blue?" --visual(terminal, colored by confidence), add--tuifor the full-screen view, or runevery-other-token --web(browser). Pick the provider with--provider openai|anthropic|ollama|openrouter|gemini. Real confidence numbers come from logprobs (OpenAI, OpenRouter, recent Ollama); providers that send none, like Anthropic and Gemini, get an estimate from how quickly each token arrived.
Run every-other-token --help for every flag, with examples at the bottom.
| Doc | What's in it |
|---|---|
| Usage guide | All transforms, --rate and --min-confidence, A/B system prompt tests, provider diff, research mode, web UI views, config file, shell completions, build from source |
| Library reference | The Rust modules behind the CLI (attribution export, drift detection, mutation lab, causal maps and more) with examples |
| Architecture | How the stream, transforms and outputs fit together |
| HTTP API and WebSocket rooms | The web server endpoints and collaborative token editing |
| Feature flags | Optional Cargo features |
| docs.rs | Full API docs for using it as a library |
| Changelog and Contributing | Release history and how to send a change |
MIT, see LICENSE.
Need this kind of engineering on your product? I take on a small number of client builds: LLM features, iOS apps and performance work, fixed price. Services and pricing · Email · LinkedIn
