Skip to content

Repository files navigation

QAPI3

Become a sponsor

Sponsor me on afdian link

Fast, SQL-safe Api.

This program contains these features:

  • Minecraft Server validation: User can easily bind their account to api,
  • Chat Sync between several Minecraft Servers(WIP): This feature allows users to chat across servers.
  • Minecraft Server Status Query: This api will automatically show latest server status.(Need Plugin-side support)

Future RoadMaps

More codes in Kotlin

DIY player card with QCommunity WEB

Add Discord support

Guides

To install:

simply just run gradle build and that's all. Don't forget to add configurations like MySQL server and Redis!

I recommend you to just run program in docker and redirect port to nginx, etc.

Avatar images, including special card avatars, are cached under avatars/ and served by /qo/download/avatar/image. Set QAPI_AVATAR_CACHE_DIR to change the cache directory, QAPI_AVATAR_CACHE_TTL_SECONDS to change the refresh interval (default: 86400), and QAPI_PUBLIC_BASE_URL when the API response must contain an absolute cached-image URL.

To add components, sub servers, please add node object in nodes.json

[
  {
    "name": "QQ",
    "id": 0,
    "role": "SERVER",
    "token": "123456"
  },
  {
    "name": "QO",
    "id": 1,
    "role": "SERVER",
    "token": "123456"
  }
]

and then they can post messages, add new features, etc...

Database configuration

Database access is managed by Spring Data R2DBC. Connection settings continue to come from data/sql/info.json:

{
  "url": "jdbc:mysql://db.example:3306/qapi",
  "username": "qapi",
  "password": "change-me"
}

The existing jdbc:mysql: URL is converted to r2dbc:mysql: when the pool is created. The original pool defaults (initial/minimum idle 5, maximum 100, 3-second acquisition timeout, and 60-second eviction interval) are retained; do not configure a JDBC DataSource for request handling.

API performance metrics

GET /qo/metrics returns the in-memory processing-time summary for every HTTP endpoint. Each metric is keyed by HTTP method and Spring route template, for example GET /qo/download/status, and contains count, avgMs, p90Ms, and p99Ms. avgMs is calculated over all requests since startup; percentiles are calculated from the most recent 4096 requests for that endpoint. The sample limit can be changed with the qapi.metrics.max-samples-per-endpoint property.

API Endpoint

GET /qo/download/status ->

{
    "totalcount": 0,
    "mspt_3s": 2,
    "code": 0,
    "players": [],
    "game_time": 0,
    "mspt": 2.562659,
    "recent60": [],
    "onlinecount": 0,
    "timestamp": 1727272149198
}

GET /qo/msglist/download ->

{
    "messages": [
        "服务状态更新:\nService SortMC 的状态为 null\n最新的heartbeat状态为: 1 延迟 77ms",
        "[QO] 玩家Lplayer加入了服务器。",
        "[QO] 玩家Lplayer退出了服务器,本次游玩时间 4分钟",
        "[QO] 玩家Lplayer加入了服务器。",
        "服务状态更新:\nService QAPI origin 的状态为 null\n最新的heartbeat状态为: 0 延迟 0ms",
        "[QQ] <CHJWOS_|2859972822>:wow",
        "[QQ] <CHJWOS_|2859972822>:https://www.mcmod.cn/class/14114.html",
        "[QQ] <glowingstone124|1294915648>:[CQ:reply,id=986310598][CQ:at,qq=2859972822,name=CHJWOS_] fabric mod自定义gui的话",
        "[QQ] <glowingstone124|1294915648>:是不是要拿opengl嗯写",
        "[QQ] <CHJWOS_|2859972822>:高版本可能是有原版帮你写好了的",
        "[QQ] <CHJWOS_|2859972822>:不是原版风格的就自己写",
        "[QQ] <glowingstone124|1294915648>:我不想要原版",
        "[QQ] <glowingstone124|1294915648>:我想写个现代化ui",
        "[QQ] <CHJWOS_|2859972822>:矩形啥的应该是已经内置了",
        "[QQ] <CHJWOS_|2859972822>:如果有模糊啥的",
        "[QQ] <CHJWOS_|2859972822>:可能1.20.5往上也有",
        "[QQ] <glowingstone124|1294915648>:行",
        "[QQ] <glowingstone124|1294915648>:动画呢",
        "[QQ] <glowingstone124|1294915648>:是不是要自己搓",
        "[QQ] <CHJWOS_|2859972822>:是了",
        "[QQ] <東雪蓮Official|3125265713>:累了",
        "[QQ] <東雪蓮Official|3125265713>:今天操场上测了下速",
        "[QQ] <東雪蓮Official|3125265713>:23圈",
        "[QQ] <東雪蓮Official|3125265713>:一圈400",
        "[QQ] <東雪蓮Official|3125265713>:15分钟",
        "[QQ] <東雪蓮Official|3125265713>:每小时大约36km",
        "[QQ] <東雪蓮Official|3125265713>:[CQ:image,file=6EF6BBF5A4EC9B58B3754A1E7836C689.jpg,subType=1,url=https://multimedia.nt.qq.com.cn/download?appid=1407&amp;fileid=CgozMTI1MjY1NzEzEhT-HH8woFCrx7llrTrMAfDy9YbiQBiutiMg_woo9s29x5beiAMyBHByb2RQgL2jAQ&amp;spec=0&amp;rkey=CAMSKMa3OFokB_TlE5oz_MZGn_1PxOOLL_sQeAG7OFPt_2onFxvUsjDhYv0,file_size=580398]",
        "[QQ] <CHJWOS_|2859972822>:哇啊",
        "[QQ] <CHJWOS_|2859972822>:全世界的人都在教我怎么sampler2D传入图片",
        "[QQ] <CHJWOS_|2859972822>:我想要传入当前帧画面的教程啊",
        "[QQ] <glowingstone124|1294915648>:那你把这一帧变成bytemap"
    ],
    "empty": false
}

GET /qo/download/registry?name=glowingstone124 ->

{
    "qq": 1294915648,
    "code": 0,
    "frozen": false,
    "online": false,
    "economy": 0,
    "playtime": 1712,
    "last_login": 1787558400000
}

LLM Tool Calling

The OpenAI-compatible chat endpoint can execute built-in tools on Responses-backed presets in both JSON and SSE modes before returning the final assistant message.

  • get_server_status: query Minecraft server status and player counts.
  • get_current_date: return the current date and time, with an optional IANA timezone (default Asia/Shanghai).
  • get_player_rankings: query mining, placement, and cumulative playtime leaderboards.
  • query_metro_lines: search metro lines, stations, sections, and signal coordinates.
  • search_minecraft_knowledge: search the configured RAG knowledge base for Minecraft/QO information.
  • add_memory: create or update a structured per-group long-term memory.
  • search_memory: query structured memories for the current group.
  • forget_memory: delete structured memories only when explicitly requested.
  • get_member_profile: read the current user's persistent QQ-uid profile.
  • upsert_member_profile: create or update a confirmed identity, preference, summary, or group nickname under the current user's QQ uid.
  • forget_member_profile_field: delete a profile field when the current user explicitly asks to forget it.

Structured memories are stored in the automatically created MySQL llm_memories table. A memory is uniquely identified by group_id + subject + memory_key, so multiple facts about the same subject can coexist. On the first startup after upgrading, legacy data/llm/rag/<groupId>/memory.txt and data/llm/rag/groups/<groupId>/memory.txt files are imported once; completion is recorded in llm_memory_migrations. Legacy files are retained for rollback but are excluded from RAG after migration.

Member profiles are stored separately in llm_member_profiles and llm_member_profile_fields. QQ uid (qquid in LLM context) is the global unique identity shared by QQ, Kotshi Web, and linked Minecraft accounts and receives a stable generated profile_id; group_nickname remains scoped by group. Explicit /remember content fields and evidence-based observed_summary fields are injected only for the authenticated current qquid; other participants contribute identity metadata only.

Archived group messages are periodically and incrementally converted into both a multi-member group summary and per-qquid observed profiles with the provider's configured summary.model; this runs without waiting for anyone to ask the bot. For QQ group requests, the model receives the rolling summary, a recent participant list, and up to 20 recent messages with each sender's qquid, nickname, message ID, and current-sender flag. Message text is capped at 500 characters. It retrieves older or more specific messages with search_chat_history when needed. Kotshi Web receives the same authenticated qquid profile as QQ chat. Summaries are policy-versioned so older summaries are rebuilt after isolation-policy changes. Boolean environment variables accept only true and false.

Related environment variables:

  • LLM_SYSTEM_PROMPT: fixed system prompt text. When set, it takes precedence over the prompt file.
  • LLM_SYSTEM_PROMPT_FILE: system prompt file. Linux inotify events, atomic replacements, Docker bind mounts, and Kubernetes ConfigMap/Secret-style replacements are reloaded without restarting the API; invalid or blank updates keep the previous valid prompt.
  • LLM_QO_GROUP_ID: QQ group allowed to access QO server knowledge and server tools. This restriction applies to QQ/Web requests; Minecraft requests may use all registered tools regardless of group mapping. If omitted, QO RAG and server tools remain unavailable to QQ/Web requests.
  • LLM_BLOCKED_QQ_UIDS: comma- or space-separated QQ uids denied before any LLM request. If omitted, no user is blocked by this rule.
  • LLM_ULTRA_BRIEF_QQ_UIDS: comma- or space-separated QQ uids that receive one-sentence replies unless safety or factual clarification requires more.
  • LLM_STRIP_EMOJI: set to true to remove emoji from upstream answers during output sanitization. Tool-call markup and emoticons are always removed.
  • qapi.llm.web-allowed-origin-patterns: comma-separated origin patterns allowed to request Web SSE chat. Defaults to https://*.qoriginal.vip,http://localhost:*,http://127.0.0.1:*; untrusted or missing origins are rejected for stream=true requests.
  • qapi.cors.allowed-origins: comma-separated CORS origins. The default explicitly includes https://ai.qoriginal.vip and retains the existing wildcard for compatibility. Set QAPI_CORS_ALLOWED_ORIGINS in production to replace the default with an exact origin list.
  • LLM_GROUP_SUMMARY_ENABLED: enable per-group rolling fact summaries, default true. When disabled, raw history is still archived for explicit search but is not automatically injected.
  • LLM_GROUP_SUMMARY_DIR: persistent rolling-summary directory, default data/llm/summaries.
  • LLM_GROUP_SUMMARY_MAX_CHARS: maximum persisted summary characters per group, default 5000.
  • LLM_GROUP_SUMMARY_TIMEOUT_MS: maximum time spent updating a summary; on timeout, the previous safe summary is retained and raw history is not injected, default 30000.
  • LLM_PERIODIC_SUMMARY_ENABLED: enable archive-driven group and member summarization, default true.
  • LLM_PERIODIC_SUMMARY_INTERVAL_MS: delay between background summary runs, default 600000 (10 minutes).
  • LLM_PERIODIC_SUMMARY_INITIAL_DELAY_MS: startup delay before the first background run, default 120000 (2 minutes).
  • LLM_PERIODIC_SUMMARY_MAX_GROUPS: maximum archived groups checked per run, default 1000.
  • LLM_PERIODIC_SUMMARY_BATCH_MESSAGES: oldest-first message batch size per summary call, default 200.
  • LLM_PERIODIC_SUMMARY_MIN_MESSAGES: minimum pending messages before a partial batch is summarized, default 40.
  • LLM_PERIODIC_SUMMARY_MAX_WAIT_MS: maximum time the oldest message in a partial batch may wait before it is summarized even below the message threshold, default 3600000 (1 hour).
  • LLM_PERIODIC_SUMMARY_MAX_BATCHES_PER_GROUP: catch-up batches processed for one group in a run, default 1.
  • LLM_HISTORY_TTL_MS: how long an idle conversation stays cached in memory, default 1800000 (30 minutes). After eviction, its recent context and rolling summary reload from SQL on the next request.
  • LLM_IMAGE_STORE_DIR: durable directory for images referenced by archived AI turns, default data/llm/images. Keep this directory with the SQL database when moving or restoring QAPI.
  • LLM_MEMORY_CONTEXT_MAX_ITEMS: maximum relevant memories injected into a request, default 10.
  • LLM_MEMORY_CONTEXT_MAX_CHARS: maximum memory context characters, default 6000.
  • LLM_MEMBER_PROFILE_CONTEXT_MAX_ITEMS: maximum qbot high-activity member profiles injected into one request, default 50.
  • LLM_MEMBER_PROFILE_CONTEXT_MAX_FACTS: maximum self-declared facts accepted from each member profile, default 16.
  • LLM_MEMBER_PROFILE_CONTEXT_MAX_CHARS: maximum total qbot member-profile context characters, default 20000.

QQ group messages are archived in the llm_chat_history table through POST /qo/asking/v1/chat/history, using stable source IDs and INSERT IGNORE for idempotency. The periodic summarizer consumes this archive directly, so bot completion requests no longer carry the sliding raw group_context. The LLM can retrieve older, group-scoped records with the search_chat_history tool; results never cross group boundaries.

Completed AI turns from Web, QQ, and Minecraft are written to llm_conversation_turns before the request finishes. The compact recent context and rolling summary are written to llm_conversation_state, so a QAPI restart or memory-cache eviction can resume a conversation. The full turn archive remains in SQL even after older turns are summarized; explicitly deleting a Web conversation removes its archive, resumable context, and stored images. Synced messages from nodes, Web chat, and system notices are written to messages; a normal shutdown flushes pending messages. Existing short messages.message columns are expanded on first use so long messages can be archived. GET /qo/msglist/public excludes system notices (from=2).

  • LLM_TOOLS_ENABLED: enable QAPI's other local function tools, default true; Web Search remains available when disabled.
  • Web Search calls the DuckDuckGo Lite search relay at http://10.10.0.3:9123/search, replacing the SearXNG service on that port. The relay implementation and deployment steps are in docs/DUCKDUCKGO_SEARCH_SERVICE.md. QAPI sends q and gives the model up to eight results. Web conversations receive result URLs for citation. QQ and Minecraft conversations receive opaque result IDs instead of URLs and are instructed to answer without source attribution. web_fetch resolves a recent result from the same conversation, sends its URL internally to the Trafilatura service at http://10.10.0.3:9124/fetch, and returns up to 12,000 characters of page text without the URL for non-Web channels. Both tools work across Chat Completions, Responses, Anthropic, and Command Code. Chat Completions SSE waits for the tool rounds to finish before sending the final answer.
  • Trafilatura microservice interface and deployment checks: docs/WEB_FETCH_SERVICE.md.
  • LLM_PROVIDERS_FILE: provider configuration JSON path, default data/llm/providers.json. The file is watched and the configuration (including referenced token files) is periodically reloaded; invalid updates keep the last valid provider.
  • adminUids: optional top-level array in the provider configuration listing QQ UIDs allowed to grant reset cards (default empty). It is reloaded together with the provider file; see Reset cards.
  • LLM_PROVIDER: selected provider name. If omitted, the JSON defaultProvider is used and may be changed by hot-reloading the provider file. When set, this environment override remains fixed until restart.
  • LLM_TOOL_MAX_ROUNDS: maximum tool-call loops per request, default 3.
  • LLM_TOOL_METRO_MAX_RESULTS: maximum metro search results returned to the model, default 12.

LLM request exceptions are written to data/llm/llm-error.log relative to the QAPI3 working directory. Entries include the request stage, provider, requester ID, access-record ID when available, and exception stack traces. Request and response bodies and API tokens are excluded.

Tool execution failures and invalid_tool_call responses are appended to data/llm/toolcall-failure.log. New entries use timestamped multiline blocks with a parse or execute stage, readable request context, and separate, formatted diagnostic sections. Parse failures include the upstream response and assistant text, plus provider, model, and tool round; execution failures include complete arguments and results. Recognized credential fields, Bearer tokens, and Base64 data URLs are redacted. Console output contains a compact summary and the archive path. Existing log entries are preserved.

Tool-call validation inspects the first assistant choice rather than scanning response metadata. Missing, null, or empty tool_calls fields on normal answers are accepted, and DSML calls in assistant text can still be parsed in those cases. Malformed call structures, unfinished tool tags, or a tool_calls finish reason without a parseable call still produce invalid_tool_call. Failure archives preserve null fields for diagnosis.

Set enableCommandCode: true to try the endpoint-free providers.commandcode block before defaultProvider, with a fallback when the Command Code connection fails. See Command Code provider setup, including pricing, image input, and fallback behavior.

Ordinary chat requests apply the mode/source reasoning-effort policy without forcing a Fast/Thinking output-token cap. Explicit caller token limits are preserved. Chat Completions and Responses requests omit the output limit when none is supplied; Anthropic and Command Code retain their protocol adapter defaults. Builder client-tool steps and internal summaries use separate output budgets.

Provider configuration example (data/llm/providers.json; replace example URLs and model IDs with those supported by your upstream):

{
  "defaultProvider": "gateway",
  "providers": {
    "gateway": {
      "chatCompletionsUrl": "https://gateway.example/v1/chat/completions",
      "responsesUrl": "https://gateway.example/v1/responses",
      "anthropicUrl": "https://gateway.example/v1/messages",
      "tokenFile": "LLMAPITOKEN",
      "contextWindow": 524288,
      "models": {
        "fast": {
          "model": "provider-fast-model",
          "protocol": "responses"
        },
        "thinking": {
          "model": "provider-thinking-model",
          "protocol": "chat-completions"
        },
        "quality": {
          "model": "provider-claude-model",
          "protocol": "anthropic",
          "thinkingMode": "adaptive"
        }
      },
      "summary": {
        "provider": "anthropic",
        "model": "fast",
        "contextWindow": 32768
      },
      "compact": {
        "enabled": true,
        "triggerTurns": 12,
        "triggerPercent": 70,
        "keepTurns": 4,
        "maxSummaryChars": 8000
      }
    },
    "anthropic": {
      "chatCompletionsUrl": "unavaliable",
      "responsesUrl": "unavaliable",
      "anthropicUrl": "https://api.anthropic.com/v1/messages",
      "tokenFile": "data/llm/anthropic.token",
      "contextWindow": 200000,
      "models": {
        "fast": {
          "model": "your-fast-claude-model-id",
          "protocol": "anthropic"
        },
        "thinking": {
          "model": "your-thinking-claude-model-id",
          "protocol": "anthropic",
          "thinkingMode": "adaptive"
        },
        "compact": {
          "model": "your-summary-claude-model-id",
          "protocol": "anthropic"
        }
      },
      "summary": {
        "model": "compact",
        "contextWindow": 32768
      }
    }
  }
}

Every provider must declare all three endpoint fields, even when unused. Set an unsupported endpoint to the literal "unavaliable" ("unavailable" is also accepted). Endpoints are used exactly as configured; QAPI never derives one endpoint from another or falls back to a different protocol. Every models.<preset> entry must explicitly declare model and protocol, with protocol one of responses, chat-completions, or anthropic. A model selecting an unavailable endpoint makes the configuration invalid. All providers are validated, including inactive ones.

For migration, replace string model entries with these objects, remove responsesModels, and add anthropicUrl. There is no implicit protocol default. The spelling aliases antrophicUrl and protocol antrophic are accepted. balanceUrl is optional.

Anthropic requests use x-api-key and anthropic-version: 2023-06-01. Optional per-model thinkingMode is enabled (default; manual token budget), adaptive (effort-based thinking), or disabled; choose a mode supported by the upstream model. Anthropic tool/search continuation preserves original content blocks and thinking signatures. Streaming responses keep QAPI’s Chat Completion chunks, kotshi.status events, and final [DONE].

contextWindow is the main model's context-window size in tokens (default 524288). The API keeps the system prompt and newest user messages, then drops the oldest history when the estimated input would exceed that window. models accepts arbitrary preset names; request the quality preset with ?model=quality. summary.provider may reference any configured provider, while summary.model may be any preset of that provider or a model name declared in its models object. The summary request uses that model’s explicit protocol. summary.contextWindow independently limits the summary request. Summary settings default to the selected provider's fast model and the main contextWindow. Conversation autocompact uses the same summary configuration. Its provider-local compact object defaults to enabled=true, triggerTurns=12, triggerPercent=70, keepTurns=4, and maxSummaryChars=8000. It replaces older turns with one rolling summary once raw history exceeds the turn or token threshold, while keeping the newest configured turns.

Each upstream LLM request logs its source, provider name, resolved model, and API type. Provider reloads are logged as [LLM] reloaded provider old -> new; tokens and request bodies are not part of this status log. Full request-body logging remains separately controlled by LLM_DEBUG_PROMPT.

For Contributors

When integrating this project with GoCi, please notice there are some flags can be use.

[SKIP CI]: when pushing a commit which description contains this, GoCi will automatically skip build this commit.

About

new version of QAPI3. Not only for QuantumOriginal.

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages