diff --git a/AGENTS.md b/AGENTS.md index 96e6f5a..e7085ac 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -15,14 +15,14 @@ This file is the top-level orientation map. Depth lives under [docs/](docs/). | [oli_bot/api/](oli_bot/api/) | FastAPI package exposing the harness over an OpenAI-compatible REST API (`GET /v1/models`, `POST /v1/chat/completions` streaming + non-streaming, `GET /health`) plus a stateful `WS /v1/chat` WebSocket that relays every `AgentEvent` as a typed JSON envelope (`text_chunk`/`thinking`/`tool_call_executing`/`tool_call_result`/`assistant_response`/`usage`/`error`/`done`, plus `sub_agent_started`/`sub_agent_progress`/`sub_agent_completed` and `todo` frames) for real-time browser UIs. Sub-agent events are demuxed by `task_id`; `_wire_todo_relay()` pushes `builtin__todowrite` snapshots onto `mcp_manager.pending_todos`, drained by the socket loop. REST is stateless from the caller's POV; the WebSocket keeps per-connection history, backed by a server-persisted `ConversationStore` under the `"default"` server: turns are sent with an optional `session_id` (`{"content": "...", "session_id": "..."}`, `{"action": "clear", "session_id": "..."}` wipes it) and each completed turn is saved to disk, so browser and TUI share the same session files. Sessions are created/listed/loaded/renamed/deleted via `GET/POST /v1/sessions` and `GET/PUT/DELETE /v1/sessions/{id}`. A single process-private `Agent` is shared across requests and serialised with an `asyncio.Lock` held across `await` points (a `threading.RLock` would be per-thread reentrant and not serialize coroutines sharing the event loop); the WebSocket holds the lock for the duration of each run. Auto-approves permissions (no human), but offline/dry-run still apply. Also exposes `GET/POST /v1/mcp` plus `PUT/DELETE /v1/mcp/{name}` for the browser UI's MCP server configuration view (backed by `MCPClientManager` on `app.state.agent.mcp_manager`). `app.py` holds the `create_app()` factory + `init_state()`; `routers/` holds the route modules (`chat`, `ws`, `sessions`, `config`, `mcp`, `workspace`, `health`); `runner.py` implements the completion handoff, `harness.py` the agent construction + todo relay + `dispatch` handler, `session_service.py` the load/persist helpers, `errors.py` the OpenAI-style error handlers, `convert.py` the wire-message/event converters, `deps.py` the FastAPI dependencies, and `__main__.py` the `oli-server` entrypoint. [`oli_bot/api_server.py`](oli_bot/api_server.py) is a compat shim re-exporting the old single-module surface. See [docs/API_SERVER.md](docs/API_SERVER.md). | | [oli_bot/agent.py](oli_bot/agent.py) | `Agent` — mode + system prompt owner; orchestrates the tool-calling loop and streams typed events (`TextChunk`, `ThinkingChunk`, `ToolCallChunk`, `ToolCallExecuting`, `ToolCallResult`, `StreamChunk`, `UsageEvent`, `Error`, `Done`). Aggregates per-call `UsageChunk`s from each backend round into a single per-run `UsageEvent`. Also hosts `sanitize_tool_history`, `_merge_usage`, `stream_sub_agent_run`, and `AgentPool` (built from `agents.yaml` when `--use-pool` is set; located via `$OLI_AGENTS_YAML`, the package dir, the repo root, the current working directory, or `~/.config/oli`). | | [oli_bot/backends/](oli_bot/backends/) | Backend package — `ModelBackend` ABC, `OllamaBackend`, `OpenAIBackend`, `HuggingFaceBackend`, `TransformersBackend`, and the `create_model_backend()` factory. `OpenAIBackend` speaks both Chat Completions (default) and the Responses API, selected by the `OLI_OPENAI_RESPONSES_ENABLED` / `openai_responses_enabled` flag routed through `generate()` / `stream_generate()`. Also hosts the shared `_StreamingThinkParser` and per-backend message formatting (Ollama native `images`, OpenAI `image_url` or Bedrock-native blocks via `openai_vision_style`, Responses `input_text`/`input_image` items, textual placeholder for text-only backends). Every backend surfaces a trailing `UsageChunk`: exact counts from provider usage where available (OpenAI `usage`/`stream_options`, Responses `input_tokens`/`output_tokens`, Ollama `prompt_eval_count`/`eval_count`, HF `usage`), else a `~chars/4` estimate via `estimate_tokens`. See [docs/BACKENDS.md](docs/BACKENDS.md). | -| [oli_bot/screens/](oli_bot/screens/) | All `ModalScreen` subclasses: `PermissionScreen`, `ConfirmScreen`, `ModelPickerScreen`, `ServerListScreen`, `MCPSetupScreen`, `SessionListScreen`, `WorkspaceListScreen`, `SubAgentViewScreen`, `ConfigScreen`, `InputPromptScreen`, plus `taglines.py` / `todo_widget.py`. | +| [oli_bot/screens/](oli_bot/screens/) | All `ModalScreen` subclasses: `PermissionScreen`, `ConfirmScreen`, `ModelPickerScreen`, `ServerListScreen`, `MCPSetupScreen`, `SessionListScreen`, `WorkspaceListScreen`, `SubAgentViewScreen`, `ConfigScreen`, `InputPromptScreen`, `QuestionScreen` (answers the `builtin__question` tool — one modal listing the posed questions with options + per-question freeform fields and a Confirm button), plus `taglines.py` / `todo_widget.py`. | | [oli_bot/models.py](oli_bot/models.py) | Shared dataclasses: `Message`, `ToolCall`, `ModelResponse`, `HostConfig`, `MCPServerConfig`, `ProfileData`, `SubAgentRun`, `ImageAttachment`, `TodoItem` / `TodoListState`, plus `AgentEvent` variants and the `AgentRole` enum. Token accounting lives here too: `Usage` (prompt/completion/`estimated` flag), the per-call `UsageChunk` stream event, and the per-run `UsageEvent`. `Message.images` is in-memory only (dropped on session save); `ModelResponse.usage` is optionally set by backends. | | [oli_bot/config.py](oli_bot/config.py) | `AppConfig` — `pydantic_settings.BaseSettings`. Env vars prefixed `OLI_`, plus `.env` support and `OLI_TRUNCATION_SMALL` / `_LARGE` aliases via `AliasChoices`. Module-level `configs = AppConfig()` singleton. See [docs/CONFIGURE.md](docs/CONFIGURE.md). | | [oli_bot/settings.py](oli_bot/settings.py) | `SettingsManager` — load/save/merge `~/.config/oli/settings.json`; precedence `settings.json` > `OLI_*` env > SDK-standard env (`OPENAI_API_KEY`, `OPENAI_BASE_URL`, `HUGGINGFACE_API_KEY`, `HF_TOKEN`) > declared defaults. Empty API-key strings in JSON fall through to env. | | [oli_bot/profiles/](oli_bot/profiles/) | `ProfileManifest` / `PermissionsManifest` (Pydantic) in `schema.py`; `ProfilePermissionEnforcer` (layered allow/deny glob patterns, base-profile inheritance, deny-overrides-allow) in `permissions.py`; `ProfileManager` (profile CRUD, manifest loading, circular-dependency detection) in `manager.py`. Built-in profiles ship as sibling directories. | | [oli_bot/mcp_client.py](oli_bot/mcp_client.py) | `MCPClientManager` — MCP server lifecycle (stdio/http via v2 `mcp.client.Client`, `mode="auto"` handshake), tool discovery + invocation, per-server tool-list cache, offline gating. Uses v2 snake_case fields (`Tool.input_schema`, `CallToolResult.is_error`/`structured_content`). | | [oli_bot/tools/manager.py](oli_bot/tools/manager.py) | `BuiltinToolManager` — registration, profile + session permission gating, dry-run gating, offline gating, and `TruncationManager` post-processing. Awaits coroutine handlers. | -| [oli_bot/tools/](oli_bot/tools/) | Tool handlers: `files.py` (read/write/edit + `view_image` via Pillow), `directories.py` (glob/grep/list_directory/tree — filesystem work runs via `asyncio.to_thread` / `create_subprocess_exec`), `web.py` (search + fetch + specialised searches, all guarded by `_check_ssrf`), `shell.py` (allowlisted `run_command`, including read-only `git`), `parsing.py` (`compare`), `memory.py` (`think`, `todowrite`, `notebook`), `truncation.py` (per-tier char budgets), `permissions.py` (sensitive-path detection). See [docs/TOOLS.md](docs/TOOLS.md). | +| [oli_bot/tools/](oli_bot/tools/) | Tool handlers: `files.py` (read/write/edit + `view_image` via Pillow), `directories.py` (glob/grep/list_directory/tree — filesystem work runs via `asyncio.to_thread` / `create_subprocess_exec`), `web.py` (search + fetch + specialised searches, all guarded by `_check_ssrf`), `shell.py` (allowlisted `run_command`, including read-only `git`), `parsing.py` (`compare`), `memory.py` (`think`, `todowrite`, `notebook`), `question.py` (pose questions to the user via a TUI modal and block for answers; no interactive callback — e.g. the API server — yields an error result), `truncation.py` (per-tier char budgets), `permissions.py` (sensitive-path detection). See [docs/TOOLS.md](docs/TOOLS.md). | | [oli_bot/sessions.py](oli_bot/sessions.py) | `Session` (permission gating) + `ConversationStore` (per-server JSON persistence under `~/.config/oli/sessions//`) + `WorkspaceManager`. `save_session()` returns the (possibly new) id so callers can rebind after a corrupt-file rewrite. Persisted messages preserve `tool_call_id`; loads pass through `sanitize_tool_history` so poisoned histories self-heal. | | [oli_bot/backends/upstream_manager.py](oli_bot/backends/upstream_manager.py) | `UpstreamManager` — multi-server lifecycle persisted to `hosts.json`, URL validation. | | [oli_bot/voice.py](oli_bot/voice.py) | `VoiceEngine` — optional, lazy-loaded mic → STT → TTS engine for the `/voice` command (faster-whisper, Piper TTS, WebRTC VAD, pyaudio). All I/O is blocking; `chat.py` calls it via `asyncio.to_thread`. `record()` accepts a `threading.Event` so `chat.py` can interrupt an in-progress recording the instant voice mode is toggled off, instead of waiting out the silence/max-duration timeout. All seven tunables (whisper/piper models, sample rate, VAD frame duration, VAD aggressiveness, silence timeout, max record seconds) are `AppConfig` fields (`OLI_VOICE_*` env / `settings.json` `voice` section / `/config` screen); `chat.py` passes them explicitly when constructing the engine, and saving `/config` drops the engine so the next `/voice` picks up new values. | @@ -50,6 +50,7 @@ Built-in profiles: | `analyst` | Data-analyst specialist — extracts claims, triangulates sources, flags tensions | | `coder` | Software-engineer profile — read, write, and run access for end-to-end development workflows | | `reviewer` | Code-review profile — read-only analysis with test/lint execution; no file modifications | +| `editor` | Writing-editor profile — proofreading, grammar and prose health, feedback on creative writing | | `writer` | Technical writer profile — prose, documentation, READMEs, changelogs, and guides | | `planner` | Planning agent — decomposes goals into structured, saved plans; no file modifications | diff --git a/README.md b/README.md index 4c9782e..5969a04 100644 --- a/README.md +++ b/README.md @@ -38,8 +38,8 @@ Concretely, that means: ## Features - **Declarative sub-agent pooling (optional)** — with `--use-pool`, the root agent can fan tasks out concurrently to vendor-agnostic sub-agents defined in an optional [agents.yaml](agents.yaml) file via a `dispatch` tool. Each pool entry binds a model _and_ a backend, so dispatch decisions are also compute-location decisions — a frontier model can plan while sensitive work stays on a local model, or a local root can fan out to faster remote SLMs for latency-sensitive tool calls. -- **Agent profiles** — drop-in system prompts with permission manifests, base-profile inheritance, and auto-generated profiles via `/profile create`. Bundled profiles: `default`, `coder`, `reviewer`, `writer`, `planner`, `researcher`, `analyst`. -- **Rich built-in tool set** — file ops, shell access, web search/fetch, Wikipedia/GitHub/arXiv search, task tracking, reasoning scratchpad, notebook, and more. Sandbox-locked with shell allowlists, SSRF protection, and sensitive-file gating. +- **Agent profiles** — drop-in system prompts with permission manifests, base-profile inheritance, and auto-generated profiles via `/profile create`. Bundled profiles: `default`, `coder`, `reviewer`, `editor`, `writer`, `planner`, `researcher`, `analyst`. +- **Rich built-in tool set** — file ops, shell access, web search/fetch, Wikipedia/GitHub/arXiv search, task tracking, reasoning scratchpad, notebook, interactive user-question modals, and more. Sandbox-locked with shell allowlists, SSRF protection, and sensitive-file gating. - **Permission system** — write operations and sensitive reads require user approval. Session grants, workspace scoping, and profile-level allow/deny lists. - **OpenAI-compatible API server** — run the same agent harness behind `/v1/models` and `/v1/chat/completions` (streaming + non-streaming) so any workflow that speaks the OpenAI wire protocol (the `openai` Python SDK, curl, or plain REST) can drive the agent. - **Voice mode (optional, experimental)** — `/voice` toggles a hands-free mic → STT → LLM → TTS loop (faster-whisper, Piper TTS, WebRTC VAD) for the TUI. Fully local; requires the `voice` extras and a downloaded Piper model. @@ -85,6 +85,7 @@ oli --profile researcher | `default` | ✅ | ✅ | ✅ | General-purpose tasks | | `coder` | ✅ | ✅ | ✅ | Software development end-to-end | | `reviewer` | ❌ | ✅ | ❌ | Code review, quality analysis | +| `editor` | ✅ | ❌ | ❌ | Proofreading, grammar, prose, creative edits| | `writer` | ✅ | ❌ | ✅ | Docs, READMEs, changelogs, prose | | `planner` | ✅ | ❌ | ✅ | Roadmaps, task decomposition, saved plans | | `researcher` | ❌ | ❌ | ✅ | Web research with structured JSON output | diff --git a/docs/PROFILES.md b/docs/PROFILES.md index e023bf9..38e0e6f 100644 --- a/docs/PROFILES.md +++ b/docs/PROFILES.md @@ -23,6 +23,7 @@ When a profile is loaded, both `AGENTS.md` and `SKILLS.md` are combined into the | `analyst` | Specialist data-analyst agent — extracts claims, triangulates across sources, flags tensions | | `coder` | Software-engineer profile — read, write, and run access for end-to-end development workflows | | `reviewer` | Code-review profile — read-only analysis with test/lint execution; no file modifications | +| `editor` | Writing-editor profile — proofreading, grammar and prose health, creative-writing feedback | | `writer` | Technical writer profile — prose, documentation, READMEs, changelogs, and guides | | `planner` | Planning agent — decomposes goals into structured, saved plans; no file modifications | diff --git a/docs/TOOLS.md b/docs/TOOLS.md index f493988..86e7790 100644 --- a/docs/TOOLS.md +++ b/docs/TOOLS.md @@ -27,6 +27,7 @@ All built-in tools are exposed to the model as `builtin__`. | **extract_article** | Extract full text from an article URL via newspaper4k; returns title, authors, publish date, and a text preview. | | **compare** | Compare files or directories and summarize differences. | | **todowrite** | Create and maintain a structured task list for the current session; tracks progress, organizes multi-step work. | +| **question** | Pose one or more questions to the user and block until they answer. Each item takes a `question` string, optional `options` (shown as selectable choices), and an optional `recommended` (a 1-based option index or exact option text marking the model's suggestion). Rendered in the TUI as a modal widget with a "write your own answer" field per question and a final Confirm button; the answers are returned as the tool result so the model can proceed. When no interactive user is attached (e.g. the API server), the tool returns an error telling the model to proceed on its own. | | **think** | Internal reasoning scratchpad -- stores chain-of-thought in conversation history without displaying it to the user. | | **notebook** | Agent working memory -- store and retrieve Markdown notes across named pages under `~/.config/oli/notes/`. Pages named `plan-` (as used by `/mode plan`) auto-increment to `plan--2`, `-3`, ... on collision instead of overwriting. | @@ -40,7 +41,7 @@ The agent requires user approval for operations that could affect your system: - **Sensitive files** (`.env*`, `*.pem`, `*.key`, `~/.ssh/`, `~/.aws/`, files with `secret`/`credential`/`password`/`token` in the name) -- prompt on `read_file` even inside the workspace. - **Sensitive glob/grep patterns** -- requests whose pattern or `include` field references sensitive keywords also prompt. - **Outbound HTTP tools** (`fetch`, `download_file`, `upload_file`, `search_github`, `view_image` URL branch) -- no permission prompt, but every request passes through the SSRF guard (see below). -- **Unrestricted tools** -- `websearch`, `search_wikipedia`, `search_arxiv`, `search_stackoverflow`, `search_open_library`, `think`, `todowrite`, `notebook` -- no permission gating (the network ones are still blocked by offline mode). +- **Unrestricted tools** -- `websearch`, `search_wikipedia`, `search_arxiv`, `search_stackoverflow`, `search_open_library`, `think`, `todowrite`, `notebook`, `question` -- no permission gating (the network ones are still blocked by offline mode). Each permission prompt offers three choices: **Allow once**, **Allow for session**, or **Deny**. Session grants persist for the lifetime of the TUI process. diff --git a/oli_bot/api/harness.py b/oli_bot/api/harness.py index eac0369..1dce806 100644 --- a/oli_bot/api/harness.py +++ b/oli_bot/api/harness.py @@ -73,6 +73,13 @@ async def _api_confirm(description: str) -> str: return "session" +# Note: the ``builtin__question`` tool needs an interactive human. The API +# server never registers a question callback on ``BuiltinToolManager``, so a +# call to it returns an error result telling the model to proceed on its own. +# If an out-of-band ask flow is ever desired, wire ``set_question_callback`` +# here the same way ``set_todo_callback`` is wired in ``_wire_todo_relay``. + + def _wire_todo_relay(agent: Agent) -> None: """Relay ``builtin__todowrite`` updates to WebSocket clients. diff --git a/oli_bot/chat.py b/oli_bot/chat.py index 7063e40..859a4c5 100644 --- a/oli_bot/chat.py +++ b/oli_bot/chat.py @@ -76,6 +76,7 @@ WorkspaceListScreen, MCPSetupScreen, PermissionScreen, + QuestionScreen, ConfirmScreen, SessionListScreen, SubAgentViewScreen, @@ -549,6 +550,7 @@ def on_mount(self) -> None: # Wire up the todo-change callback so updates fire immediately self._builtin_tools.set_todo_callback(self._on_todos_changed) self._builtin_tools.set_sub_todo_callback(self._on_sub_todos_changed) + self._builtin_tools.set_question_callback(self._question_callback) # Set border title for the todo panel todo_panel = self.query_one("#todo-panel", TodoWidget) @@ -2629,6 +2631,14 @@ async def _permission_callback(self, description: str) -> str: async with self._permission_lock: return await self.push_screen_wait(PermissionScreen(description)) + async def _question_callback(self, questions: list[dict]) -> str | None: + """Present the agent's ``builtin__question`` batch and return the + user's answers, serialized with permission prompts so concurrent + sub-agents queue their questions instead of stacking modals. + """ + async with self._permission_lock: + return await self.push_screen_wait(QuestionScreen(questions)) + def _register_dispatch_tool(self) -> None: """Register the `dispatch` built-in tool that fans a batch of tasks out to pooled sub-agents concurrently. Schema generation is shared diff --git a/oli_bot/profiles/default/AGENTS.md b/oli_bot/profiles/default/AGENTS.md index cb773aa..cfa09a7 100644 --- a/oli_bot/profiles/default/AGENTS.md +++ b/oli_bot/profiles/default/AGENTS.md @@ -28,6 +28,7 @@ All built-in tools are called via `builtin__`. | `builtin__extract_article` | `url: string` | no permission needed (blocked by offline mode) | Extract full text from an article URL via newspaper4k; returns title, authors, publish date, and a text preview | | `builtin__think` | `thought: string` | no permission needed | Internal reasoning scratchpad (not shown to user). Use to plan multi-step work, reason about problems, or analyze before acting. | | `builtin__todowrite` | `todos: array[{content, status, priority}]` | no permission needed | Create and maintain a structured task list. Track progress, mark items complete, and verify all tasks are done. Call with the full list each time. | +| `builtin__question` | `questions: array[{question, options?: [string], recommended?: string}]` | no permission needed (interactive) | Pose questions to the user and wait for their answers before proceeding. Each question may list suggested options (mark one as `recommended` when you believe it is best — a 1-based option index or exact option text); the user can select an option or type a custom answer in the modal widget. The tool blocks until the user confirms. Use when a decision or input from the user is required. | | `builtin__compare` | `target_a: string, target_b: string, mode?: string, ignore_whitespace?: boolean` | no permission needed | Compare files or directories and summarize differences. | | `builtin__tree` | `path?: string`, `depth?: number` | same as read tools | Display directory structure as a tree. Shows recursive layout of files and subdirectories. | | `builtin__dispatch` | `batch: array[{agent, task}]` | same as read tools (root agent only); no permission needed for plain text/read-only agents | **Available only when agent pooling is enabled** (`--use-pool` / `OLI_USE_AGENT_POOL=true`) and the `agents.yaml` pool is non-empty. Fans out agent/task pairs concurrently (`asyncio.gather`) to the configured sub-agents and returns all results aggregated into a single labeled string. If pooling is not enabled or the pool is empty, this tool is **not registered** and must not be called. | @@ -37,7 +38,7 @@ All built-in tools are called via `builtin__`. - **Write tools** (`write_file`, `edit_file`, `download_file`, `upload_file`) — always require user permission - **Read tools** (`read_file`, `view_image`, `glob`, `grep`, `list_directory`, `tree`) — require permission only when targeting paths outside the session workspace - **Shell tools** (`run_command`) — require permission only when their working directory is outside the session workspace (same boundary logic as read tools) -- `websearch`, `fetch`, `search_wikipedia`, `search_github`, `search_arxiv`, `search_stackoverflow`, `search_open_library`, `extract_article`, `think`, `todowrite`, `notebook` — no permission gating +- `websearch`, `fetch`, `search_wikipedia`, `search_github`, `search_arxiv`, `search_stackoverflow`, `search_open_library`, `extract_article`, `think`, `todowrite`, `notebook`, `question` — no permission gating ## Shell command allowlist diff --git a/oli_bot/profiles/default/SKILLS.md b/oli_bot/profiles/default/SKILLS.md index d48c0d5..00c4466 100644 --- a/oli_bot/profiles/default/SKILLS.md +++ b/oli_bot/profiles/default/SKILLS.md @@ -4,6 +4,7 @@ | ------------------------- | ----------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Plan multi-step work | `builtin__think` | Internal scratchpad; not shown to user. Use before acting. | | Track task progress | `builtin__todowrite` | Create/maintain a todo list. Plan steps, update status as you go, verify completeness. Call with the full updated list each time. | +| Ask the user something | `builtin__question` | Pose questions and block until the user answers. Give concrete options when the choice is well-defined and mark one as `recommended` if clearly best; the user can also type a custom answer. Use before acting on something only the user can decide, never to stall on a decision you can make yourself. | | Delegate sub-tasks | `builtin__dispatch` | Dispatch sub-agents with stand-alone tasks relevant to the larger task at hand. **NOTE**: this tool may not always be available! It will be contingent on the user configuring agent pools. | | Explore project structure | `builtin__list_directory`, `builtin__glob`, `builtin__tree` | `list_directory` for flat listings with metadata; `tree` for recursive directory structure overview. | | Search code content | `builtin__grep` | Regex search with line numbers. | diff --git a/oli_bot/profiles/editor/AGENTS.md b/oli_bot/profiles/editor/AGENTS.md new file mode 100644 index 0000000..a6f0642 --- /dev/null +++ b/oli_bot/profiles/editor/AGENTS.md @@ -0,0 +1,95 @@ +You are a skilled writing editor. You proofread for correctness, give clear, +actionable feedback on creative and general writing, and help authors improve +the health of their prose — grammar, mechanics, and style — without drowning +out their voice. + +## Principles + +- **Read the whole piece first.** Never comment on a fragment or a single + paragraph in isolation. Understand the piece's purpose, audience, and voice + before giving feedback. +- **Respect the author's voice.** Editing is about clarity and correctness, + not homogenisation. Don't rewrite for its own sake; prefer the smallest + change that fixes the problem. +- **Show, don't just tell.** Quote the original, then give your suggested + rewrite. "This is awkward" is useless; "this is awkward — consider *X* + instead of *Y* because …" is useful. +- **Explain the why.** Tie each recommendation to a reason: a grammar rule, + an established style guide, rhythm, clarity, or reader comprehension. +- **Be honest but kind.** Name problems directly and specifically, but frame + the feedback around the work and the goal, never the author. +- **Level the concerns.** Grammar errors are not the same as stylistic + judgement calls. Report them separately. + +## What to look for + +**Proofreading (correctness)** +- Subject–verb and pronoun agreement; consistent verb tense. +- Spelling, homophones, and typos. +- Punctuation: commas, semicolons, apostrophes, quotation marks. +- Misused words ("its/it's", "affect/effect", "then/than"). +- Fragments, run-ons, and comma splices. + +**Prose health (style)** +- Sentence rhythm: chains of same-length structures; monotone pacing. +- Redundancy: phrases that say the same thing twice, filler ("in order to", + "due to the fact that"), hedging adverbs. +- Precision: vague qualifiers ("very", "really", "sort of") and clichés. +- Passive voice where the actor matters and is known. +- Word choice that clashes with the tone or register. +- Paragraph flow and transitions. + +**Creative writing (craft)** +- Pacing: scenes that drag or rush; summary vs. scene balance. +- Character voice: narration or dialogue that slips out of character. +- Show-don't-tell: told emotion where a detail or action would land harder. +- Consistency: characters, timeline, setting, tense. +- The opening and the ending: hooks and resonance. + +## Workflow + +1. **Gather the text** — read the file(s) in full with `read_file`; use + `glob`/`grep` to cross-reference names, terms, and timeline details for + consistency. +2. **Sort findings** — use `think` to rank issues by severity and decide what + is worth flagging. If the piece is large, focus: the biggest issues beat an + exhaustive list. +3. **Report** — lead with a short summary, then grouped findings in the format + below. Quote → suggestion → reason. +4. **Apply edits only when asked** — prefer `edit_file` for targeted fixes and + `write_file` for wholesale rewrites. When the author only wants feedback, + propose the changes and wait. + +## Output format + +``` +## Summary +Two or three sentences: overall assessment, strongest parts, the one thing +worth fixing first. + +## Proofreading +### [Severity] Short title (location) +Original: «as written» +Suggested: «concrete alternative» +Why: reason. + +## Prose and style +... + +## Creative feedback (when relevant) +... + +## Small wins +A bullet list of low-stakes fixes — a typo at ¶3, a missing comma at ¶7 — that +a spellchecker would flag but a reader still deserves. +``` + +**Severity levels:** +- **Critical** — a genuine error: grammar mistake, misspelling, or a factual + and consistency break that interrupts a reader. +- **Major** — impedes clarity or intent: tangled sentence, confusing + transition, pacing problem, character-voice slip. +- **Minor** — a judgement call: word choice, rhythm, smoother phrasing. + +If there is nothing to flag, say so clearly ("reads clean") rather than +inventing problems. \ No newline at end of file diff --git a/oli_bot/profiles/editor/SKILLS.md b/oli_bot/profiles/editor/SKILLS.md new file mode 100644 index 0000000..92a980e --- /dev/null +++ b/oli_bot/profiles/editor/SKILLS.md @@ -0,0 +1,56 @@ +## Tool selection + +| Task | Tool | Notes | +|------|------|-------| +| Read the piece to edit | `builtin__read_file` | Always read the full text before commenting. | +| Find the files to edit | `builtin__glob`, `builtin__list_directory`, `builtin__tree` | Locate drafts, chapters, or writing in the workspace. | +| Cross-reference for consistency | `builtin__grep` | Check names, timelines, and terminology across the piece. | +| Compare drafts and revisions | `builtin__compare` | What changed between versions; verify edits landed. | +| Apply targeted fixes | `builtin__edit_file` | **Preferred for corrections** — quote enough context for a unique match. | +| Apply wholesale rewrites | `builtin__write_file` | Full new version of a file; creates parent directories automatically. | +| Sort findings before reporting | `builtin__think` | Rank issues by severity; drop noise. | +| Track multi-file edits | `builtin__todowrite` | For longer projects spanning several chapters or documents. | +| Keep style notes | `builtin__notebook` | Save recurring errors, the author's preferences, or style-guide decisions. | + +## Feedback pattern + +For every finding, use three parts: **original → suggestion → why**. + +``` +## Proofreading +### [Major] Run-on in the second paragraph (¶2) +Original: "She ran to the door it was locked." +Suggested: "She ran to the door. It was locked." +Why: Comma splice — two independent clauses need a period or a conjunction. +``` + +Always separate three tiers: +- **Proofreading** — objective errors (grammar, spelling, punctuation, tense). +- **Prose health** — style and rhythm judgements (redundancy, pacing, precision). +- **Creative feedback** — craft-level notes (voice, show-don't-tell, structure). + +Name strengths too. A good passage left alone is informative — tell the author +what is working and why, so they can keep doing it. + +## Editing checklist + +- [ ] Read the full piece before commenting; no fragment review. +- [ ] Each finding quotes the original text verbatim. +- [ ] Every suggestion has a reason, not just a replacement. +- [ ] Severity labels are honest — a typo is not "Critical". +- [ ] Edits preserve the author's voice unless the voice is the problem. +- [ ] Consistency checked: names, pronouns, tense, timeline, tone. +- [ ] Typo-class finds bundled under "Small wins" so real issues stand out. + +## Common pitfalls + +- **Overwriting.** Rewriting a whole sentence to fix a dangling comma buries + the real issue. Prefer the smallest change that fixes the problem. +- **Voice flattening.** Leaning every sentence toward "standard correct + prose" erases a distinctive voice. Edit for clarity, not conformity. +- **Nit-picking the best passages.** If a paragraph works, leave it alone — + feedback should reward strengths, not just list flaws. +- **Big list, no priority.** Lead with the highest-impact fix; an exhaustive + catalog of nits hides the one thing worth doing. +- **Style guides as law.** Style is a tool, not a rulebook — "don't end with + a preposition" loses to a natural, readable sentence. \ No newline at end of file diff --git a/oli_bot/profiles/editor/profile.json b/oli_bot/profiles/editor/profile.json new file mode 100644 index 0000000..8cae7aa --- /dev/null +++ b/oli_bot/profiles/editor/profile.json @@ -0,0 +1,50 @@ +{ + "schema_version": 1, + "name": "editor", + "version": "1.0.0", + "description": "Writing-editor profile — proofreading, grammar and prose health, feedback on creative and general writing", + "default_model_tier": "large", + "base": null, + "required_tools": [ + "builtin__read_file", + "builtin__write_file", + "builtin__edit_file", + "builtin__glob", + "builtin__grep", + "builtin__list_directory", + "builtin__tree", + "builtin__compare", + "builtin__think", + "builtin__todowrite", + "builtin__notebook" + ], + "permissions": { + "allow_tools": [ + "builtin__read_file", + "builtin__write_file", + "builtin__edit_file", + "builtin__glob", + "builtin__grep", + "builtin__list_directory", + "builtin__tree", + "builtin__compare", + "builtin__think", + "builtin__todowrite", + "builtin__notebook", + "builtin__view_image" + ], + "deny_tools": [ + "builtin__run_command", + "builtin__download_file", + "builtin__upload_file", + "builtin__websearch", + "builtin__fetch", + "builtin__search_wikipedia", + "builtin__search_github", + "builtin__search_arxiv", + "builtin__search_stackoverflow", + "builtin__search_open_library", + "builtin__extract_article" + ] + } +} \ No newline at end of file diff --git a/oli_bot/screens/__init__.py b/oli_bot/screens/__init__.py index 4a7ea3e..45de4dd 100644 --- a/oli_bot/screens/__init__.py +++ b/oli_bot/screens/__init__.py @@ -12,6 +12,7 @@ from .mcp_setup import MCPSetupScreen from .model_picker import ModelPicker from .permission import PermissionScreen +from .question import QuestionScreen from .server_list import ServerListScreen from .session_list import SessionListScreen from .sub_agent_view import SubAgentViewScreen @@ -26,6 +27,7 @@ "MCPSetupScreen", "ModelPicker", "PermissionScreen", + "QuestionScreen", "ServerListScreen", "SessionListScreen", "SubAgentViewScreen", diff --git a/oli_bot/screens/question.py b/oli_bot/screens/question.py new file mode 100644 index 0000000..7d31050 --- /dev/null +++ b/oli_bot/screens/question.py @@ -0,0 +1,159 @@ +from __future__ import annotations + +from typing import List, Optional + +from textual.app import ComposeResult +from textual.containers import Container, Horizontal, VerticalScroll +from textual.screen import ModalScreen +from textual.widgets import Button, Input, Label, RadioButton, RadioSet + +from ..tools.question import _resolve_recommended + + +class QuestionScreen(ModalScreen[str | None]): + """Present questions posed by the agent and collect the user's answers. + + Each question is rendered as a numbered block with its suggested options + (the recommended one pre-selected when the model marked it) and a "Write + your own answer" field. A final Confirm button resolves every question's + answer (custom text wins over a selected option) and dismisses with a + text summary that becomes the ``question`` tool result. Esc cancels. + """ + + CSS = """ + #question-container { + width: 78; + height: 80%; + border: round $primary; + background: $surface; + padding: 1; + margin: 1 2; + } + #question-title { + text-style: bold; + content-align: center middle; + padding: 0 0 1 0; + } + #question-scroll { + height: 1fr; + border: round #6b7d74; + padding: 1 2; + } + .question-block { + margin: 0 0 1 0; + } + .question-text { + text-style: bold; + } + .question-meta { + color: $text-muted; + } + RadioSet { + margin: 0 0 1 0; + border: round #6b7d74; + } + RadioSet:focus { + border: round $primary; + } + .question-own-input { + margin: 1 0 1 0; + } + #question-buttons { + height: 3; + align: center middle; + } + Button { + margin: 0 1; + } + """ + + BINDINGS = [("escape", "cancel", "Cancel")] + + def __init__(self, questions: List[dict]): + super().__init__() + self._questions = list(questions) + + def compose(self) -> ComposeResult: + with Container(id="question-container"): + yield Label("Question", id="question-title") + with VerticalScroll(id="question-scroll"): + for i, q in enumerate(self._questions): + with Container(classes="question-block"): + yield Label( + f"{i + 1}. {q.get('question', '')}", + classes="question-text", + ) + if q.get("options"): + idx, freeform = _resolve_recommended( + q.get("recommended"), q.get("options") + ) + if freeform: + yield Label( + f"[Recommended action: {freeform}]", + classes="question-meta", + ) + yield RadioSet( + *[ + RadioButton( + str(opt), + id=f"q{i}-opt{j}", + ) + for j, opt in enumerate(q["options"]) + ], + id=f"q{i}-options", + ) + yield Label("Write your own answer:", classes="question-meta") + yield Input( + placeholder="Type your own answer", + id=f"q{i}-input", + classes="question-own-input", + ) + with Horizontal(id="question-buttons"): + yield Button("Confirm", variant="primary", id="question-confirm") + yield Button("Cancel", id="question-cancel") + + def on_mount(self) -> None: + for i, q in enumerate(self._questions): + if not q.get("options"): + continue + rs = self.query_one(f"#q{i}-options", RadioSet) + idx, _ = _resolve_recommended(q.get("recommended"), q.get("options")) + buttons = list(rs.query(RadioButton)) + if idx is not None and 0 <= idx < len(buttons): + buttons[idx].value = True + rs.index = idx + first_input = self.query_one("#q0-input", Input) + if first_input is not None: + first_input.focus() + + def _answer(self, i: int) -> str: + q = self._questions[i] + custom = self.query_one(f"#q{i}-input", Input).value.strip() + if custom: + return custom + if q.get("options"): + rs = self.query_one(f"#q{i}-options", RadioSet) + idx = rs.pressed_index + if idx is not None and 0 <= idx < len(rs.children): + return str(rs.children[idx].label) + return "No answer provided" + + def _format_result(self) -> str: + parts: list[str] = [] + for i, q in enumerate(self._questions): + parts.append(f"Question {i + 1}: {q.get('question', '')}") + parts.append(f"User answer: {self._answer(i)}") + parts.append("") + return "\n".join(parts).rstrip() + + def on_button_pressed(self, event: Button.Pressed) -> None: + if event.button.id == "question-confirm": + self.dismiss(self._format_result()) + elif event.button.id == "question-cancel": + self.dismiss(None) + + def action_cancel(self) -> None: + self.dismiss(None) + + +__all__ = ["QuestionScreen"] diff --git a/oli_bot/tools/manager.py b/oli_bot/tools/manager.py index 146af5a..74da761 100644 --- a/oli_bot/tools/manager.py +++ b/oli_bot/tools/manager.py @@ -51,6 +51,7 @@ "fetch", "think", "compare", + "question", "search_stackoverflow", "search_open_library", "extract_article", @@ -86,6 +87,7 @@ def __init__( self._todos: list[dict] = [] self._todo_change_callback: Optional[Callable[[list[dict]], None]] = None self._sub_todo_change_callback: Optional[Callable] = None + self._question_callback: Optional[Callable[[list[dict]], Any]] = None self._pending_attachments: list["ImageAttachment"] = [] self._pending_caption: str = "" self._register_default_tools() @@ -141,6 +143,7 @@ def _register_default_tools(self) -> None: from .shell import register_tools as reg_shell from .parsing import register_tools as reg_parsing from .memory import register_tools as reg_memory + from .question import register_tools as reg_question reg_files(self) reg_dirs(self) @@ -148,6 +151,7 @@ def _register_default_tools(self) -> None: reg_shell(self) reg_parsing(self) reg_memory(self) + reg_question(self) def register_tool( self, @@ -246,6 +250,18 @@ def set_sub_todo_callback(self, callback: Optional[Callable]) -> None: """ self._sub_todo_change_callback = callback + def set_question_callback( + self, callback: Optional[Callable[[list[dict]], Any]] + ) -> None: + """Register the ``question`` tool's interaction callback. + + The callback receives the raw question items the model posed and returns + the user's answers as a string (or ``None`` on cancel), which becomes + the tool result fed back to the model. When no callback is registered + (e.g. headless API runs), the tool returns an error instead. + """ + self._question_callback = callback + def attach_image(self, attachment: "ImageAttachment") -> None: """Push an image attachment onto the pending queue for the current tool call.""" self._pending_attachments.append(attachment) diff --git a/oli_bot/tools/question.py b/oli_bot/tools/question.py new file mode 100644 index 0000000..0059048 --- /dev/null +++ b/oli_bot/tools/question.py @@ -0,0 +1,122 @@ +from __future__ import annotations + +import inspect +from typing import TYPE_CHECKING + +if TYPE_CHECKING: + from .manager import BuiltinToolManager + + +def register_tools(manager: "BuiltinToolManager") -> None: + manager.register_tool( + name="question", + description=( + "Pose questions to the user and wait for their answers before " + "proceeding. Use this whenever you need the user to make a decision, " + "choose between options, or provide information only they know. " + "Pass each question as a separate item; add suggested options when " + "relevant and mark one as recommended if you believe it is clearly " + "best. The questions are shown to the user as a widget listing each " + "question with its options and a 'write your own answer' field; " + "the tool blocks until the user confirms their answers." + ), + parameters={ + "type": "object", + "properties": { + "questions": { + "type": "array", + "items": { + "type": "object", + "properties": { + "question": { + "type": "string", + "description": "The question to ask the user.", + }, + "options": { + "type": "array", + "items": {"type": "string"}, + "description": ( + "Optional suggested answers, shown as numbered " + "choices the user can select." + ), + }, + "recommended": { + "type": "string", + "description": ( + "Optional. The choice you recommend: a 1-based " + "index into 'options' (as text, e.g. '2') or the " + "exact text of one of the options. Omit when no " + "option is clearly best." + ), + }, + }, + "required": ["question"], + }, + "description": "The one or more questions to ask the user.", + }, + }, + "required": ["questions"], + }, + handler=lambda questions: _question_handler(questions, manager), + ) + + +def _resolve_recommended( + recommended: object, options: list | None +) -> tuple[int | None, str | None]: + """Resolve a question's ``recommended`` into (option_index, freeform_text). + + 1-based option indices (as int or digit string) and exact/partial option + text matches resolve to an option index; anything else is treated as a + freeform recommended-action string. + """ + options = options or [] + if recommended is None: + return None, None + + if isinstance(recommended, bool): + return None, None + + if isinstance(recommended, int): + idx = recommended - 1 + if 0 <= idx < len(options): + return idx, None + return None, None + + text = str(recommended).strip() + if not text: + return None, None + + if text.isdigit(): + idx = int(text) - 1 + if 0 <= idx < len(options): + return idx, None + + lowered = text.lower() + for i, opt in enumerate(options): + if str(opt).lower() == lowered or lowered in str(opt).lower(): + return i, None + + return None, text + + +async def _question_handler( + questions: list[dict], manager: "BuiltinToolManager" +) -> str: + if not questions: + return "Error: question called with no questions." + + callback = getattr(manager, "_question_callback", None) + if callback is None: + return ( + "Error: the 'question' tool requires an interactive user, but no " + "question callback is registered in this environment. Proceed using " + "your best judgment or spell out the question in your reply instead." + ) + + result = callback(questions) + if inspect.isawaitable(result): + result = await result + if result is None: + return "Error: The user declined to answer the questions." + return str(result) diff --git a/tests/integration/test_agent_tool_loop_wire.py b/tests/integration/test_agent_tool_loop_wire.py index c16f259..1cda1df 100644 --- a/tests/integration/test_agent_tool_loop_wire.py +++ b/tests/integration/test_agent_tool_loop_wire.py @@ -206,3 +206,128 @@ async def test_write_then_read_file_via_tools( assert any( e.full_text == "Wrote and read back." for e in events if isinstance(e, Done) ) + + +@pytest.mark.integration +async def test_question_tool_awaits_user_answer_and_resumes( + make_full_agent, mock_openai, auto_allow +): + agent, session = make_full_agent(workspace=None, offline_mode=False) + mocked_answers = "Question 1: Which environment?\nUser answer: Production" + + def fake_question_callback(questions): + assert questions == [ + { + "question": "Which environment?", + "options": ["Staging", "Production"], + "recommended": "2", + } + ] + return mocked_answers + + agent.mcp_manager._builtin_tools.set_question_callback(fake_question_callback) + + mock_openai.script( + ( + "stream", + [ + cc( + delta={ + "tool_calls": [ + tool_delta( + 0, + tc_id="q1", + name="builtin__question", + args=( + '{"questions": [{"question": "Which environment?", ' + '"options": ["Staging", "Production"], ' + '"recommended": "2"}]}' + ), + ) + ] + } + ), + cc( + finish="tool_calls", + usage={"prompt_tokens": 4, "completion_tokens": 2}, + ), + ], + ), + ( + "stream", + [ + cc(content="Deploying to production as you chose."), + cc(finish="stop", usage={"prompt_tokens": 6, "completion_tokens": 3}), + ], + ), + ) + + from oli_bot.models import Done, Message, ToolCallResult + + messages = [Message(role="user", content="pick the deploy target")] + events = [ev async for ev in agent.process(messages, confirm_callback=auto_allow)] + results = [e for e in events if isinstance(e, ToolCallResult)] + assert len(results) == 1 + assert results[0].name == "builtin__question" + assert results[0].result == mocked_answers + assert any( + e.full_text == "Deploying to production as you chose." + for e in events + if isinstance(e, Done) + ) + assert any( + msg.role == "tool" and "User answer: Production" in msg.content + for msg in messages + ) + # The question tool never falls through to a permission prompt, so the + # session scope list must remain empty. + assert not session._session_grants + + +@pytest.mark.integration +async def test_question_tool_without_callback_returns_error( + make_full_agent, mock_openai, auto_allow +): + agent, _ = make_full_agent(workspace=None, offline_mode=False) + mock_openai.script( + ( + "stream", + [ + cc( + delta={ + "tool_calls": [ + tool_delta( + 0, + tc_id="q1", + name="builtin__question", + args='{"questions": [{"question": "Proceed?"}]}', + ) + ] + } + ), + cc( + finish="tool_calls", + usage={"prompt_tokens": 4, "completion_tokens": 2}, + ), + ], + ), + ( + "stream", + [ + cc(content="Proceeding on my own."), + cc(finish="stop", usage={"prompt_tokens": 6, "completion_tokens": 3}), + ], + ), + ) + + from oli_bot.models import Done, Message, ToolCallResult + + messages = [Message(role="user", content="ask me something")] + events = [ev async for ev in agent.process(messages, confirm_callback=auto_allow)] + results = [e for e in events if isinstance(e, ToolCallResult)] + assert len(results) == 1 + assert results[0].result.startswith("Error:") + assert "interactive user" in results[0].result + assert any( + e.full_text == "Proceeding on my own." for e in events if isinstance(e, Done) + ) diff --git a/tests/unit/test_question_tool.py b/tests/unit/test_question_tool.py new file mode 100644 index 0000000..520bf3b --- /dev/null +++ b/tests/unit/test_question_tool.py @@ -0,0 +1,217 @@ +"""The ``builtin__question`` tool: registration, callback plumbing, and result +formatting, plus the ``QuestionScreen`` modal's confirm/cancel flow. +""" + +from __future__ import annotations + +import pytest + +from oli_bot.config import AppConfig +from oli_bot.tools.manager import BuiltinToolManager, PLAN_TOOLS, READ_ONLY_TOOLS +from oli_bot.tools.question import _resolve_recommended + + +def _manager() -> BuiltinToolManager: + return BuiltinToolManager(config=AppConfig(_env_file=None, offline_mode=False)) + + +def _spec(*questions: dict) -> dict: + return {"questions": list(questions)} + + +# ---------- registration ------------------------------------------------------ + + +def test_question_tool_registered_with_schema(): + m = _manager() + names = {t["name"] for t in m.get_tool_definitions()} + assert "builtin__question" in names + + info = next(t for t in m.get_tool_definitions() if t["name"] == "builtin__question") + props = info["parameters"]["properties"] + assert "questions" in props + assert "questions" in info["parameters"]["required"] + q = props["questions"]["items"] + assert "question" in q["required"] + + +def test_question_available_in_ask_and_plan_modes(): + m = _manager() + readonly = {t["name"] for t in m.get_readonly_tool_definitions()} + plan = {t["name"] for t in m.get_plan_tool_definitions()} + assert "builtin__question" in readonly + assert "question" in READ_ONLY_TOOLS + assert "question" in PLAN_TOOLS + assert "builtin__question" in plan + + +# ---------- handler behavior -------------------------------------------------- + + +@pytest.mark.asyncio +async def test_question_handler_returns_callback_result(): + m = _manager() + received: list | None = None + answers = "Question 1: A\nUser answer: Production" + + async def callback(questions): + nonlocal received + received = questions + return answers + + m.set_question_callback(callback) + result = await m.call_tool( + "question", + _spec({"question": "Which env?", "options": ["Staging", "Production"]}), + ) + assert result == answers + assert received == [ + {"question": "Which env?", "options": ["Staging", "Production"]} + ] + + +@pytest.mark.asyncio +async def test_question_handler_awaits_sync_callback(): + m = _manager() + m.set_question_callback(lambda questions: "User answer: yes") + result = await m.call_tool("question", _spec({"question": "Proceed?"})) + assert result == "User answer: yes" + + +@pytest.mark.asyncio +async def test_question_without_callback_returns_error(): + m = _manager() + result = await m.call_tool("question", _spec({"question": "Ping?"})) + assert result.startswith("Error:") + assert "interactive user" in result + + +@pytest.mark.asyncio +async def test_question_callback_none_means_declined(): + m = _manager() + m.set_question_callback(lambda questions: None) + result = await m.call_tool("question", _spec({"question": "Ping?"})) + assert result.startswith("Error:") + assert "declined" in result + + +@pytest.mark.asyncio +async def test_question_with_no_questions_returns_error(): + m = _manager() + m.set_question_callback(lambda questions: "n/a") + result = await m.call_tool("question", _spec()) + assert result.startswith("Error:") + assert "no questions" in result + + +# ---------- recommended resolution -------------------------------------------- + + +def test_resolve_recommended_index_int(): + assert _resolve_recommended(2, ["a", "b", "c"]) == (1, None) + + +def test_resolve_recommended_digit_string(): + assert _resolve_recommended("3", ["a", "b", "c"]) == (2, None) + + +def test_resolve_recommended_matching_option_text(): + assert _resolve_recommended("Production", ["Staging", "Production"]) == (1, None) + + +def test_resolve_recommended_partial_option_text(): + assert _resolve_recommended("prod", ["staging", "production"]) == (1, None) + + +def test_resolve_recommended_freeform(): + assert _resolve_recommended("roll back first", ["a", "b"]) == ( + None, + "roll back first", + ) + + +def test_resolve_recommended_out_of_range_falls_back(): + assert _resolve_recommended(9, ["a", "b"]) == (None, None) + assert _resolve_recommended(True, ["a", "b"]) == (None, None) + assert _resolve_recommended(None, ["a", "b"]) == (None, None) + + +# ---------- QuestionScreen modal ---------------------------------------------- + + +from textual.app import App # noqa: E402 +from textual.widgets import Input, RadioButton, RadioSet # noqa: E402 + +from oli_bot.screens import QuestionScreen # noqa: E402 + + +class _QuestionHost(App): + def __init__(self, questions): + super().__init__() + self._questions = questions + self.result = None + + def compose(self): + return [] + + def on_mount(self): + self.push_screen( + QuestionScreen(self._questions), + callback=lambda result: setattr(self, "result", result), + ) + + +async def test_question_screen_preselects_recommended(): + app = _QuestionHost( + [{"question": "Env?", "options": ["Staging", "Production"], "recommended": "2"}] + ) + async with app.run_test(): + rs = app.screen.query_one("#q0-options", RadioSet) + assert rs.index == 1 + + +async def test_question_screen_confirm_with_custom_answer(): + app = _QuestionHost([{"question": "Env?", "options": ["Staging", "Production"]}]) + async with app.run_test() as pilot: + app.screen.query_one("#q0-input", Input).value = "preprod" + await pilot.click("#question-confirm") + await pilot.pause() + assert app.result is not None + assert "Question 1: Env?" in app.result + assert "User answer: preprod" in app.result + + +def _press(rs: RadioSet, index: int) -> None: + buttons = list(rs.query(RadioButton)) + buttons[index].value = True + rs.index = index + + +async def test_question_screen_confirm_with_option_selection(): + app = _QuestionHost([{"question": "Env?", "options": ["Staging", "Production"]}]) + async with app.run_test() as pilot: + _press(app.screen.query_one("#q0-options", RadioSet), 1) + await pilot.click("#question-confirm") + await pilot.pause() + assert app.result is not None + assert "User answer: Production" in app.result + assert "No answer provided" not in app.result + + +async def test_question_screen_custom_answer_wins_over_option(): + app = _QuestionHost([{"question": "Env?", "options": ["Staging", "Production"]}]) + async with app.run_test() as pilot: + _press(app.screen.query_one("#q0-options", RadioSet), 1) + app.screen.query_one("#q0-input", Input).value = "custom-env" + await pilot.click("#question-confirm") + await pilot.pause() + assert app.result is not None + assert "User answer: custom-env" in app.result + + +async def test_question_screen_cancel_returns_none(): + app = _QuestionHost([{"question": "Proceed?"}]) + async with app.run_test() as pilot: + await pilot.press("escape") + await pilot.pause() + assert app.result is None