本项目用于模拟 OpenAI 标准的模型服务接口,支持对话、Embedding 和图片生成假数据,并可通过管理接口动态模拟阿里云百炼限流。
- fastapi
- uvicorn
- pydantic
source venv/bin/activate
uvicorn main:app --reload仓库提供 config.example.json。首次启动前复制为本地配置:
cp config.example.json config.jsonconfig.json 已被 Git 忽略,可以按环境修改,不会误提交。配置在服务启动时读取,修改后需要重启服务。如果未提供 config.json,服务会使用 config.example.json 中展示的内置默认值。
chat:流式对话的首 Token 延迟、分块大小、分块间隔和总长度。embedding:向量长度、生成方式、自定义向量和响应延迟。image_generation:假图片 URL 的基础地址和响应延迟。rate_limit:启动时的限流开关、限流类型和Retry-After秒数。
对话、Embedding 和图片请求中的 config 仍可以仅覆盖当次请求的默认值。
限流的启动值由 config.json 中的 rate_limit 提供。通过 PUT /admin/rate-limit 可更新当前进程配置,更新后的模型请求立即生效;通过 GET /admin/rate-limit 查询当前配置。接口更新不会回写 config.json。
limit_type 与阿里云百炼限流错误的对应关系:
limit_type |
含义 | HTTP 状态码 | error.code |
|---|---|---|---|
rpm |
RPS/RPM 请求频率超限 | 429 | Throttling.RateQuota |
tpm |
TPS/TPM Token 配额超限 | 429 | Throttling.AllocationQuota |
burst |
请求速率短时间突增 | 429 | Throttling.BurstRate |
concurrency |
并发请求数超限 | 429 | Throttling.Concurrency |
开启 RPM 限流:
curl -X PUT http://127.0.0.1:8000/admin/rate-limit \
-H "Content-Type: application/json" \
-d '{
"enabled": true,
"limit_type": "rpm",
"retry_after_seconds": 60
}'此后请求 /v1/chat/completions、/v1/embeddings 或 /v1/images/generations 均会返回:
{
"error": {
"message": "Requests rate limit exceeded, please try again later.",
"type": "Throttling.RateQuota",
"param": null,
"code": "Throttling.RateQuota"
}
}响应状态码为 429,并带有 Retry-After 和 X-Request-Id 响应头。恢复正常请求:
curl -X PUT http://127.0.0.1:8000/admin/rate-limit \
-H "Content-Type: application/json" \
-d '{"enabled": false}'动态限流配置保存在单个服务进程的内存中,重启后恢复为
config.json中的值。如使用多个 Uvicorn worker,每个 worker 会有独立配置,此模拟功能建议使用单 worker。
-
准备项目代码
- 若已在
~/works/fake_openai目录下有 main.py、requirements.txt 等文件,可跳过此步。 - 否则可用如下命令:
git clone <your_repo_url> fake_openai cd fake_openai
- 若已在
-
创建并激活 Python 虚拟环境
python3 -m venv venv source venv/bin/activate -
安装依赖
pip install -r requirements.txt
-
创建本地配置
cp config.example.json config.json
-
启动服务
uvicorn main:app --reload
- 默认监听 http://127.0.0.1:8000
- 可用
--host 0.0.0.0 --port 8080自定义监听地址和端口
-
测试接口 可用 curl、Postman 或 httpie 测试:
curl -X POST http://127.0.0.1:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-3.5-turbo", "messages": [{"role": "user", "content": "你好"}], "stream": true, "config": { "first_token_delay": 1, "stream_chunk_size_range": [5, 10], "stream_interval_range": [0.2, 0.4], "total_length_range": [40, 80] } }'
-
(可选)生产环境部署
- 推荐去掉 --reload:
uvicorn main:app --host 0.0.0.0 --port 8000
- 可结合 supervisor、systemd、docker 等方式守护进程。
- 推荐去掉 --reload:
POST /v1/chat/completions
{
"model": "gpt-3.5-turbo",
"messages": [{"role": "user", "content": "你好"}],
"stream": true,
"config": {
"first_token_delay": 1,
"stream_chunk_size_range": [5, 10],
"stream_interval_range": [0.2, 0.4],
"total_length_range": [40, 80]
}
}