Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
fc5d72b
feat: 上游 User-Agent: cli + threadId 与 x-session-id 同值
xelr233 Sep 10, 2026
b0e7cc4
feat: issue #18 上游 HTTP(S) 代理 + #20 请求树提前释放
xelr233 Sep 10, 2026
0591973
fix: 还原 CC_MAX_BODY_MB 默认值为 100MB —— #7 的证据否决了 8MB
xelr233 Sep 10, 2026
a4eaaf8
feat: 设备指纹改为按 API key 确定性派生,并对齐官方 CLI 哈希算法
xelr233 Sep 10, 2026
e167a72
docs: 修正时区建议 —— 它是客户端本机时区,不应绑定出口 IP
xelr233 Sep 10, 2026
d961276
test: fork 专属行为的回归护栏
xelr233 Sep 12, 2026
9ace422
test: 修复 fork 测试的 await 漏写,并让挂起有界
xelr233 Sep 12, 2026
18109b1
fix: 预请求也发送 User-Agent: cli(此前只有 generate 设了)
xelr233 Sep 12, 2026
e89f7d6
test: 修正 threadId 的两条契约断言
xelr233 Sep 12, 2026
e711988
test: 撤下 --test-timeout(Node 18 不支持),改用 unref 看门狗
xelr233 Sep 12, 2026
5c4f22f
Merge upstream/master (a941812..6b845b6): 协议对齐 CLI 1.53.1 + 真机修正
xelr233 Sep 14, 2026
383aed8
fix: 工具名重写对齐 CLI —— 只在 messages 里改,且只改一项
xelr233 Sep 14, 2026
3cefa6b
fix: 工具名一个都不重命名 —— 删除 toWireToolName(383aed8 的收尾)
xelr233 Sep 14, 2026
1a670f2
fix: 上游没正常走完 finish 时不再谎报成功(issue #38)
xelr233 Sep 14, 2026
7e5c261
fix: 指纹 cpuCount 改为逻辑处理器数(原表 15 项全是物理核心数)
xelr233 Sep 14, 2026
fb44078
fix: 修掉线上刷屏的日志噪音,并让上游 error 事件自带的 statusCode 生效
xelr233 Sep 15, 2026
f229a0c
fix: 流空闲超时不再 destroy 客户端连接(反代 502 / connection error 的直接成因)
xelr233 Sep 15, 2026
2ae5f13
fix: 显式设置 keepAliveTimeout,消除反代复用已关闭连接的 EPIPE
xelr233 Sep 15, 2026
31ce5e2
fix: 413 报错附带实际请求体积,便于客户端与运维定阈值
xelr233 Sep 15, 2026
003f708
test: 修掉 413 用例里的日志竞态,并把 Node 24 加进 CI 矩阵
xelr233 Sep 15, 2026
26fbae0
ci: Node 矩阵收敛到 22/24,新增 Bun 测试 job
xelr233 Oct 2, 2026
7cd2645
Merge upstream/master (6b845b6..ce5a217): #50 闪断透明重试 + #54 工具截图单发 + #…
xelr233 Oct 2, 2026
aa4a3c1
fix: 流式 /v1/responses 零输出防护失效(#56,#54 回归)——空响应不再谎报 completed
xelr233 Oct 2, 2026
639fc6e
fix: 上游闪断重试的断连监听器改挂重试循环前(#55,#50 回归)
xelr233 Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 65 additions & 0 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
name: Test

on:
push:
# 只在这些长期分支上跑 push;其余分支靠 pull_request 触发,避免同一提交跑两遍
branches: [master, release]
pull_request:
workflow_dispatch:

# 同一分支的新推送取消旧运行,避免排队浪费
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true

permissions:
contents: read

jobs:
test:
runs-on: ubuntu-latest
timeout-minutes: 10
strategy:
fail-fast: false
matrix:
# 22 是 Dockerfile 的构建版本;24 是生产部署版本(systemd 里跑的就是 v24.x)。
# 18/20 不再覆盖,engines >=18 仅作声明;另由 test-bun 覆盖 Bun 运行时。
node: ['22', '24']
name: node ${{ matrix.node }}
steps:
- name: Checkout
uses: actions/checkout@v4

- name: Set up Node.js ${{ matrix.node }}
uses: actions/setup-node@v4
with:
node-version: ${{ matrix.node }}

- name: Show Node version
run: node --version

- name: Syntax check
run: node --check proxy.mjs

- name: Run tests
run: npm test

test-bun:
# 社区用户有用 Bun 跑本代理的场景,单独用 Bun 跑一轮测试套件;
# 被测子进程经 process.execPath 派生,Bun 下整条链路都是 Bun。
# Bun 缺失的能力(node:http 不支持 CONNECT)在 test/fork.test.mjs 里显式跳过。
runs-on: ubuntu-latest
timeout-minutes: 10
name: bun
steps:
- name: Checkout
uses: actions/checkout@v4

- name: Set up Bun
uses: oven-sh/setup-bun@v2

- name: Show Bun version
run: bun --version

- name: Run tests
run: bun test
48 changes: 42 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ A reverse proxy that converts Command Code API to OpenAI / Anthropic compatible

Built by analyzing official CLI network traffic to accurately replicate the Command Code API request protocol, including device-fingerprint and lifecycle pre-requests.

**Features**: OpenAI Chat Completions / **Responses API (`/v1/responses`)** + Anthropic Messages API | Streaming & non-streaming | Tool calling (tool_use) | Multimodal image input | Reasoning effort | Dynamic model list | Cache hit metrics | Device fingerprint disguise (per-key, auto-refresh) | `x-api-key` auth (Anthropic SDK) | Client disconnect detection with upstream abort | Zero-output → 429 auto-retry | Consecutive timeout → 429 auto-retry | Privacy-aware logging
**Features**: OpenAI Chat Completions / **Responses API (`/v1/responses`)** + Anthropic Messages API | Streaming & non-streaming | Tool calling (tool_use) | Multimodal image input | Reasoning effort | Dynamic model list | Cache hit metrics | Device fingerprint disguise (per-key, auto-refresh) | `x-api-key` auth (Anthropic SDK) | Client disconnect detection with upstream abort | Zero-output guard (429 non-streaming / response.failed streaming) | Consecutive timeout → 429 auto-retry | Privacy-aware logging

**Community**: [Linux.do](https://linux.do) — a friendly Chinese tech community.

Expand Down Expand Up @@ -104,6 +104,40 @@ authority for actual retention and provider availability.

> ⚠️ **Memory amplification**: a request body exists in several copies before it reaches upstream; measured peak ≈ body size × **5.1–7.4** (7 MB → +52 MB, 20 MB → +116 MB, while a request rejected with `413` costs only ×1.05). The default `CC_MAX_BODY_MB=100` therefore implies up to ~550 MB for a **single** request, and that limit is per-request, not global. See [Memory & Deployment](#memory--deployment).

### Device fingerprint

Relevant config: `fingerprintSalt` / `CC_FINGERPRINT_SALT`, `deviceProjectDir` / `CC_DEVICE_PROJECT_DIR`.

The device fingerprint reported to `/alpha/fingerprint/record` is **derived deterministically** from the API key (`fpDigest(apiKey, field) = sha256(salt + "\\0" + apiKey + "\\0" + field)`), so one key is always one device:

| Event | Old behaviour (random) | Now (derived) |
|---|---|---|
| Process restart | Map cleared → **new machine** | same machine |
| Second instance | same key = **two machines** | same machine |
| Session expiry (12h) | `keyStateStore.delete` → **new machine every 12h** | same machine |

> The 12h case was the most visible: a real user does not replace their computer twice a day, and upstream's `device_fingerprints` table is keyed on `(userId, thumbmark)`.

**Why derived rather than "pick a device from a hash bucket"** — a fixed pool caps entropy at the pool size, so once the number of keys exceeds it, keys *must* share a fingerprint. With ~50 keys and a 1000-entry pool, ~2 keys collide; with a 100-entry pool, ~20 do. A shared `thumbmark` under two different `userId`s is direct evidence of multi-account-same-machine — exactly what you don't want to manufacture. Derivation keeps every key a distinct device (collision probability 2⁻²⁵⁶) while still being stable.

```bash
CC_FINGERPRINT_SALT=some-local-secret npm start # optional: bulk-reset every key's device identity
CC_DEVICE_PROJECT_DIR='C:\\Users\\you\\projects\\app' npm start # optional: change the fabricated project dir (slug follows)
```

The salt is optional but recommended: without it the derivation is a pure function of the API key, so anyone who knows the scheme could recompute your users' fingerprints. With it, the same key yields different devices on different deployments, at no cost.

**The signal values are fabricated too.** The official CLI reads the real machine (Windows registry MachineGuid, NIC MACs, `os.userInfo`, `git config`); this proxy derives plausible-looking substitutes from the API key — MachineGuid's `8-4-4-4-12` shape, `xx:xx:xx:xx:xx:xx` MACs, a `DESKTOP-xxxxxx` hostname, a readable git email. Those raw values never leave process memory; only their hashes go on the wire.

**Pool selection scores-and-takes-the-max rather than using modulo** — modulo would rotate *every* key's device whenever the pool grows; taking the max only affects keys where the new candidate happens to win.

The hash construction follows the official CLI (`buildMachineFingerprint` / `hashSignal` in `command-code`) — `thumbmark = sha256(IB + "\0machine\0" + [machineId, macs.join(",")].join("|"))` with `IB = "command-code:device-fingerprint:v1"`, and each component hashed as `sha256(IB + "\0" + value.toLowerCase())`. The previous implementation hashed random hex without the `IB` prefix and built the thumbmark from the component *hashes*; upstream cannot recompute either way (it never sees the raw `machineId`), so it was undetectable — but it is now aligned.

> Not addressed here: the appearance pool is still all high-end desktop CPUs and the timezone is drawn uniformly from a global pool.
>
> `timezone` is the **client machine's OS timezone** (`Intl.DateTimeFormat().resolvedOptions().timeZone`), *not* the egress IP's. So do **not** bind it to the egress IP: a user in mainland China reaching this service through a proxy normally has a machine timezone that does not match where the traffic exits, and that mismatch is the norm rather than an anomaly.
>
> The property that matters is therefore the **distribution across your own user base**, not agreement with the IP. If your users are concentrated in one region, drawing timezones uniformly from 15 global zones makes every account look like it belongs to a different continent. Set the pool to match who actually uses the deployment — this is an operator decision, and for a single-region user base it means narrowing (or weighting) `FINGERPRINT_TZS` rather than randomising it globally.
### Tool screenshot budget

Images inside tool results are sent upstream as **separate** `image` blocks (an `input_image` inside
Expand Down Expand Up @@ -345,7 +379,7 @@ Produced by the proxy itself:
| `401` | API key missing / malformed (must start with `user_`; sent via `Authorization: Bearer` or `x-api-key`) |
| `404` | Unknown path |
| `413` | Body exceeds `CC_MAX_BODY_MB` (connection kept alive and drained, not reset) |
| `429` | Zero output tokens, stream idle timeout (30s streaming / 90s non-streaming), or an upstream rate-limit mapping — all carry `Retry-After` so SDKs back off; after 3 consecutive timeouts a "reduce context" hint is returned |
| `429` | Zero output tokens (non-streaming only; streaming reports 200 + `response.failed`, see "Zero-Output Guard"), stream idle timeout (30s streaming / 90s non-streaming), or an upstream rate-limit mapping — all carry `Retry-After` so SDKs back off; after 3 consecutive timeouts a "reduce context" hint is returned |
| `502` | CC upstream error (connection-level failures such as `fetch failed` also land here) |
| `503` | `CC_MAX_INFLIGHT` is set and the in-flight cap is exceeded (`type: server_busy`) |

Expand Down Expand Up @@ -455,7 +489,7 @@ Aligned line-by-line against the official npm package source (`command-code@1.53
| Mechanism | Implementation |
|-----------|---------------|
| **Device Fingerprint** | `POST /alpha/fingerprint/record` before first request per key; signal values (Windows MachineGuid shape, real-shaped MACs, `DESKTOP-xxxxxx` hostname) are **derived deterministically from the API key** and hashed exactly like the CLI, so one key always reports the same device — across restarts, memory reclamation and multiple instances (bulk reset via `CC_FINGERPRINT_SALT`) |
| **Lifecycle Events** | `POST /alpha/lifecycle-events` (`cli_session_exists`, metadata `{sessionId, cliVersion, mode, os}`) sent in parallel with the fingerprint on key init |
| **Lifecycle Events** | `POST /alpha/lifecycle-events` (`cli_session_exists`, metadata `{sessionId, cliVersion, mode, os}`) sent in parallel with the fingerprint on key init, using the same `User-Agent: cli` as generate |
| **Per-Key Session** | One session per API key, 12h expiry + 1h random jitter |
| **Version** | `x-command-code-version` reports the **protocol version actually implemented** (currently `1.53.1`); newer npm releases only raise a drift **warning**, never a silent version bump |
| **CLI Envelope** | 9 keys: `config / memory / taste / skills / permissionMode / threadId / mode / promptCache / params` |
Expand All @@ -467,7 +501,7 @@ Aligned line-by-line against the official npm package source (`command-code@1.53
| **Key Validation** | Regex `user_[a-zA-Z0-9_-]+` on `Authorization: Bearer` or `x-api-key`, auto-cleans extra paths/prefixes, rejects `sk-xxx` format |
| **Stream Timeout** | 30s streaming / 90s non-streaming → 429 with SDK auto-retry |
| **Consecutive Timeout** | 3 consecutive timeouts before "reduce context" hint |
| **Zero-Output Guard** | outputTokens=0 → 429 `rate_limit_error` (SDK auto-retry, anti false billing) |
| **Zero-Output Guard** | outputTokens=0 with no output item: non-streaming → 429 `rate_limit_error` (SDK auto-retry, anti false billing); streaming — where `response.created` is sent eagerly per #54 — → 200 + `response.failed` (upstream_error). Empty responses are never dressed up as success ([#56](https://github.com/MAXeaglet/commandcode-proxy/issues/56)) |
| **Upstream Abort** | `AbortController` on client disconnect + all error paths |
| **Privacy Logging** | No API key fragments, no error bodies, no stack traces in logs |

Expand Down Expand Up @@ -588,6 +622,8 @@ Over the limit it returns `503` + `Retry-After: 5` + `type: server_busy` — a s

**Why it exists**: memory is `in-flight × (0.13 MB + 5.5 × body_MB)`. `CC_MAX_BODY_MB` bounds only the **per-request** term; nothing bounds the multiplier — at the default 100 MB, N concurrent requests can cost N × 550 MB.

> **Why the default body cap stays at 100 MB**: [#7](https://github.com/MAXeaglet/commandcode-proxy/issues/7) recorded a legitimate multimodal session (21 base64 images, ~10.11 MiB) hitting the old 10 MB cap, so the threshold cannot be lowered without breaking real usage — which is exactly why the concurrency side has to be bounded instead. The startup warning about implied worst-case memory is advisory; `CC_MAX_INFLIGHT` is the enforcement.

> ⚠️ Enabling this is **not** the same as being memory-safe: 32 × 550 MB still exceeds a small box. For a hard bound, lower `CC_MAX_BODY_MB` **as well**.

## Upstream Idle Timeouts
Expand Down Expand Up @@ -686,7 +722,7 @@ The body exists in several copies before being forwarded: `chunks[]` / `Buffer.c
| 20 MB | 100 MB | +116 MB (5.8×) | 200 |
| 20 MB | 8 MB | +21 MB (1.05×) | **413** |

At startup a `warn` is logged when the implied worst case is ≥ 500 MB. The limit is **per request** and the proxy does no in-flight limiting of its own — a public deployment must add both at the reverse proxy.
At startup a `warn` is logged when the implied worst case is ≥ 500 MB. The body limit is **per request** — cap the multiplier with `CC_MAX_INFLIGHT` (in-process, global only), and add per-IP / per-key limits in the reverse proxy. A public deployment should do both.

### Suggested nginx front

Expand Down Expand Up @@ -750,7 +786,7 @@ A more robust cap still belongs at the reverse proxy (`limit_conn`), since only

- **`logFile` uses `appendFileSync`** — synchronous writes on the event loop. Under public load they serialize the loop; prefer leaving it empty and collecting stdout.
- **systemd guard rails**: set `MemoryMax=` and `NODE_OPTIONS=--max-old-space-size=` so an overshoot kills the proxy, not `sshd`/`nginx`.
- **Multi-account + multiple instances**: `sessionStore` / `keyStateStore` are per-process `Map`s, so the same API key served by two instances gets two different sessions and **two different device fingerprints** — upstream sees one account on multiple machines. Scale with consistent hashing on the API key (`hash $cc_key consistent`), not round-robin.
- **Multi-account + multiple instances**: `sessionStore` is still a per-process `Map`, so the same API key served by two instances gets two different sessions. **The device fingerprint is no longer affected** — it is derived, so it is the same machine across instances and restarts (see [Device fingerprint](#device-fingerprint)). Consistent hashing on the API key (`hash $cc_key consistent`) is still recommended to keep session affinity, rather than round-robin.

## Disclaimer

Expand Down
Loading
Loading