Skip to content

TODO

  • First-run onboarding - the assistant introduces itself, explains what it can do, and asks what to do. No system-level configuration for simple users
  • Also ask how agent focus should behave (focus_on_delegate: off/ask/on) and whether the prompt names the active role (prompt_role), both off by default
  • One-window status - /status shows sessions, running agents, and configured jobs on the current machine, with logs reachable from the same place (the journalctl-style alternative)
  • Agent health monitoring - the assistant watches endpoints (e.g. the /health of agents running as web APIs) and warns when an agent stops responding
  • Per-agent todo lists - view a delegated agent's tasks, mark items done, jump into its session, and ask for the current state (OpenCode-style)
  • Non-blocking delegation - the assistant starts sub-agents or whole teams for bigger tasks and reports their status instead of blocking the current run
  • Recurring tasks carry their own role - each job carries its own role and skills, so behavior like "make doc changes per AGENTS.md" is encoded once instead of re-prompted every time
  • Auto-improving skills (autogen/autolearn) - a role refines its skills from experience and keeps them alongside its memory, so recurring work gets better without re-prompting
  • Runtime pluginization - make models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI replaceable plugins, and move the REPL to a plugin. Five replaceable layers: access, orchestration, capability, model, storage
  • One runtime, many agents - position Polyglav as an agent harness for fleets, not a single assistant
  • ACP inward and outward - speak the Agent Client Protocol so Polyglav can host and be hosted by other harnesses, alongside MCP
  • Commands as tools - expose slash commands to the model as permission-gated tools to configure roles, teams, and skills, with command access limited per permission
  • Markdown catalogue - store roles, teams, and skills as Markdown files (front matter) referenced from JSON, and move the bundled data files (src/polyglav/*.json) into one catalogue directory near src
  • Hidden files and allowed paths - hide secrets and config from tools by default, and restrict visible paths to allowed roots
  • Context policy - keep user prompts until the task is done, cut unused references and tool results immediately, add a recall tool for role and team memory, enforce a configurable context limit, and offload large tool results to .polyglav/tmp/ referenced from the session log
  • Context UI polish - auto-compaction, context trim, highlighting, and context desaturation
  • Research cache - store web and PDF results under .polyglav/files for later reuse, with a PDF-to-text tool
  • Fleet awareness - the supervisor always knows which agents and teams run on a machine, with oversight across multiple machines
  • Open WebUI and OpenTUI connectors - drive Polyglav from external chat and terminal UIs
  • Tool-call quality enhancer - steer tool calling toward correct, cheaper calls (for example web-search parameters that find the right references) and better results
  • PlantUML plugin - an external plugin that draws configurations and workflows as node diagrams instead of ASCII
  • UI commands - /clear, /new, /providers, /providers list, /thinking [hide|show], and persist error logs
  • Feedback - capture operator feedback on a run or answer and feed it back into memory and skills
  • Inbound webhooks - accept inbound events to start a run or answer a parked ask
  • Plugin management UI - list, enable, disable, install, and update plugins from a UI surface, not only slash commands
  • Model picker - /models autocomplete and a direct numeric selection of a model
  • Centralized static strings - move hardcoded prompts and static strings into one module
  • Code cleanup pass - a dedicated pass over the core for dead code, naming, and rough edges
  • Report-back connectors - job summaries delivered out-of-band (email first, idea only) when the terminal is closed
  • Remove the legacy provider-plugin backfill migration (Config._migrate_plugins) once a stable polyglav release has shipped with the externalized providers. Existing plugins lists no longer need the automatic append, and the migration code is dead weight
  • Human-in-the-loop channels - job events outbound (proposed, will_run, failed, waiting_approval) plus inbound actions (approve/reject/run/disable) over configurable connectors (webhook first, then email, then Telegram), so a job can reach an operator who is not on the box
  • Global jobs overview across agents - polyglav jobs list --root <dir> fleet scan, a GET /jobs + POST /jobs/<name>/approve|reject|run|disable operator API on polyglav serve, then a web Control UI, so one view shows which agents run next and with which task
  • Edge / offline store-and-forward buffering - offline-capable agents with local buffering for unreliable connectivity (enterprise use case)
  • Immutable agent config - polyglav serve agents must never be able to change their own configuration, permissions, or tool list (control-plane rule from the use-case reference architecture)
  • Hash-chained / tamper-evident audit log - additive on session logs (enterprise.md recommendation), hash-chained append or WORM storage
  • Single-purpose agent fleet with "one agent per process, scoped to a folder" as the headline pattern. README + docs/fleet.md set the niche
  • Fleet and swarm orchestration as the two layers - fleet: supervisor running many scoped polyglav serve instances (port allocation, health checks, restart policy, per-agent config generation), swarm: /agent types, auditor agents, generate > check > correct
  • Community presence - decide on Discord/X channels and fill the README community link slots
  • Add multiuser capability or queue for API requests (request queue, per-token rate limits)
  • Add ReadTheDocs documentation
  • Citations / source attribution - return URL + snippet with every answer
  • Bookmarks - /bookmark add/remove/list for session pinning
  • Interactive data analysis - CSV querying, SQL execution, code eval in REPL
  • Notebook mode - persistent editable cells with run outputs
  • Hybrid web + local RAG - vector store (FAISS/Weaviate), embeddings, local document search
  • Command palette / fuzzy search - CTRL-P style history search
  • Topic-aware ranking - classifier for query intent to weight search results
  • Competitor research - validate USPs against actual peers (OpenClaw, Claude Code, opencode, agentic-infra services). Comparison notes now live in docs/compare/ and feed the feature backlog (Plan/Build modes, Web Control UI, plugin marketplace, telemetry, binary builds, sharing)
  • Self-update - polyglav update (Pi pi update --self analogue)
  • Standalone binary build - Pi-style release script producing a single executable (contentious for a zero-dep Python package)
  • Opt-in telemetry contracts - vendor-neutral event schema (OpenCode, Pi @earendil-works/pi-telemetry). Decide whether it fits the no-telemetry stance
  • Conversation sharing - web-shareable session links (OpenCode /share) or published sessions (Pi pi-share-hf), building on the planned Markdown export
  • Plugin registry / marketplace - discoverable plugin sharing (OpenClaw ClawHub analogue). PyPI entry-point source is a prerequisite

Open

  • [ ] Command audit - review the slash commands for opaque or overlapping behavior and consolidate
  • [ ] Focus keeps the run's context - re-entering a finished run uses its retained engine or resumes that run's own session, and never starts a fresh session. /focus drops the session: attach, because saved sessions belong to /load
  • [ ] Interruptible run status - the status line yields to the keyboard during a delegated or team run, with a hint line naming the keys (Enter to continue, ^O to open, ^C to cancel) and a Switch marker
  • [ ] Focus a live sub-run - /focus attaches to a running sub-run and leaves the caller reachable
  • [ ] Tester loop scope - the tester runs only the related tests during the loop and the whole suite once, as a role-prompt change
  • [ ] Thinking visibility - with show_thinking false only the + Thought N.Ns line prints, so the reasoning is invisible while its duration is shown
  • [ ] Prompt brevity budget - every role prompt states a short output budget and a fixed report shape, so a report is a few lines instead of an essay
  • [ ] Docs writer commits each added task - the docs writer commits the tasks right after the operator's prompt and then works through the single points one at a time
  • [ ] hide_confirm_input default true - the typed input is hidden on the tool confirm unless overridden
  • [ ] Batched structured asks - ask handles one question at a time, so several decisions cannot be asked in one call
  • [ ] Committer as a callable stage - a run calls the committer to land the current state as one commit with a correct message, so a long run does not accumulate uncommitted work
  • [ ] Researcher role with fresh context - a bundled researcher role with web access, started fresh for every task so research never inherits stale context
  • [ ] Per-role path scoping - a role declares the files and folders it may touch, enforced by the tool policy rather than by a prompt
  • [ ] .polyglav directory layout - decide and document the reserved subfolders and file names under .polyglav/
  • [ ] Skill definition - a skill holds tool, language, or framework instructions, not a project description
  • [ ] Skill catalog review - rework the bundled and local skills against that definition, add a python skill and a polyglav skill, and move project-description text into AGENTS.md
  • [ ] AGENTS.md as the project description - keep it the single project description and update it when the structure, conventions, or extension points change
  • [ ] Leader PM posture - the leader holds the whole picture, pushes back on a request that breaks the project, and concretizes an ambiguous prompt until the requirement is synced instead of guessing
  • [ ] memorize as a tool - memory writes happen through a tool an agent calls, triggered by the operator's prompt, not only through the slash command
  • [ ] Role instruction files - a role's long instructions live in .polyglav/roles/<name>.md, referenced from the short JSON entry and appended verbatim
  • [ ] Memory with references - a bounded summary that points at full-length Markdown and session artifacts, with a guard against a misleading reference when the context is gone
  • [ ] Root role memory - inject .polyglav/memory/roles/<role>.md in bind_root_agent through a shared prompt-composition helper, with a refresh path
  • [ ] Conclusion stage - a bundled stage with a write-scoped role that distills a finished run into role files, skills, or memory, and never commits
  • [ ] Saved-session catalog in /load - /load lists and loads saved sessions so an operator can reattach to a prior agent after a restart, while /focus stays live-runs-only
  • [ ] Ask continuation - answering a parked ask resumes and continues its origin run in place, instead of only injecting the answer into the session
  • [ ] Handoff from sub-agents - a team stage or delegated agent can hand focus to the next agent (composer > planner > developer), not only the REPL root
  • [ ] Fix the PyPI long-description screenshot - the image fails to load on the package page
  • [ ] Fix opencode permission rejection - the provider stops when a permission is rejected instead of continuing
  • [ ] Exact-args permission grants - approve the specific command (not just the tool), and let an always grant live beyond the current sub-agent run
  • [ ] REPL UI colors and prompts - a consistent color scheme and prompt behavior for the REPL:
  • [ ] User prompt - bold the >>>/... marker so it stands out in long output
  • [ ] Actions/Asks orange - activity/tool status lines and the confirm/ask prompts
  • [ ] Keep output dimmed - tool results and detail lines stay dim
  • [ ] Newline before status - a status line never prints inline after streamed text
  • [ ] [Y/n] enter-to-continue - confirm defaults to yes on an empty answer
  • [ ] Hide confirm input - hidden for the [Y/n] confirm, visible for a free-text ask
  • [ ] Errored tool calls red - the ! Error: line in red
  • [ ] Thinking/Thought blue - headers blue, reasoning body dim
  • [ ] Runs, focus, and memory redesign (see PLAN.md Runs, focus, and memory):
  • [ ] Non-blocking runs and live focus - background execution, output/input routing, cancellation
  • [ ] Remove --session-id - explicit session naming is no longer needed now that auto sessions are named ses_<ts>_<id>
  • [ ] Relocate job run sessions under .polyglav/jobs/<name>/ (kept in sessions/ for now)
  • [ ] Role-name sync - adopt assistant, composer, manager, and specialist as the canonical roles across types, prompts, and docs
  • [ ] Assistant-roles track docs - record the assistant, composer, and manager architecture and the work packages in VISION, PLAN, and TODO
  • [ ] Core dev team configuration - a bundled development team with the review loop plus the project lead/support teams and their skills
  • [ ] Manager role - a bundled role that runs one or many teams and reports, sequential first
  • [ ] Per-job report destination - a report_url (or connector list) on a job so different jobs report to different endpoints, instead of one global report.webhook
  • [ ] Per-task decide-vs-park for direction asks - a task class (or per-run switch) that lets the supervisor auto-resolve a direction ask after a timeout instead of always parking it for the operator
  • [ ] Full file_* namespace extension - if file_glob/file_grep prove better with most models, extend the prefix to list_dir/glob/grep (old names stay aliases)
  • [ ] Tool spec polish - rename grep.glob -> include (alias glob), add examples and prefer-web_fetch guidance to tool descriptions
  • [ ] Mid-run blocking job approval - an ask tool inside a running job pauses the run in place (per-tool-call waiting_approval), notifies via a connector, and resumes the same session when the operator replies. Needs resumable mid-run state, a wait loop inside the run, and the connectors/transport below (deeper than the shipped per-run --require-approval gate)
  • [ ] Job event hooks - the scheduler emits typed transitions (proposed, approved, will_run, executing, verified, failed, timeout, waiting_approval) to registered services. Channel-agnostic core, first consumers are the connectors and the operator API
  • [ ] Job connectors - bundled polyglav-core-webhook (stdlib JSON POST, zero deps, works with n8n/IFTTT/any URL) first. External email (SMTP + polling) and Telegram (urllib long-poll) plugins later, all driving the jobs operator API so operators can react in time
  • [ ] Jobs operator API - GET /jobs and POST /jobs/<name>/approve|reject|run|disable on polyglav serve, so clients (web Control UI, connectors, fleet supervisor) can see and act per agent
  • [ ] Fleet jobs overview - polyglav jobs list --root <dir> scanning agent worktrees (agent, job, status, next run, task table), then the web Control UI on top
  • [ ] Role directory scan for export/import - read .polyglav/roles/*.md (front-matter roles) to import and export roles to Markdown, paralleling the sessions Markdown export/import
  • [ ] Auto team selection - the assistant picks roles, teams, and skills from the registries for a task and delegates in sequence (team orchestration as a user-facing pattern, e.g. "compare with competitors" -> Researcher > Writer > Referencer > Editor)
  • [ ] Thinking visibility - /thinking on + reasoning config documented, per-provider reasoning_content check so reasoning shows in the REPL
  • [ ] /spawn command - launch a scoped polyglav serve agent from the REPL (home -> project path), supervise (health/list/stop) and delegate to it (docs/fleet.md)
  • [ ] Remote channels - command agents from messaging apps (OpenClaw channels parity):
  • [ ] Channel gateway - one adapter surface over the engine/serve API
  • [ ] Telegram adapter - long-polling bot, send + receive
  • [ ] WhatsApp adapter - business-API HTTP channel
  • [ ] More adapters (Discord, Signal, email)
  • [ ] Remote auth + session scoping + headless deny
  • [ ] Plugin test harness - polyglav plugins test <name> ships. Bundled plugin suites live next to the plugins (plugins/<name>/tests/, discovered by the core suite). Remaining: external plugins are expected to ship a test suite, and polyglav plugins test is the runner for them
  • [ ] Session recall - full-text search across past sessions (grep/index over .polyglav/sessions/) so an agent can answer from its own history
  • [ ] Tool dry-run mode - propose tool args/effects without executing (enterprise tool-gateway requirement)
  • [ ] Context-aware cross-plugin tool router - virtual tool names (open, search, ...) dispatch per-argument to the matching plugin handler via register_handler(name, match=...) (e.g. open https://... > polyglav-core-web, open ../... > polyglav-core-fs), with merged schemas and args-aware policy accessors
  • [ ] Swarm orchestration - agent cooperation layer (docs/swarm.md): /agent types, auditor agents, generate > check > correct, and team patterns as sub-tasks below
  • [ ] Grep text index - internal bundled plugin (stdlib) that indexes converted text files for local search, bridging toward the vector store
  • [ ] Agent folder watcher - internal bundled plugin (stdlib threading + pathlib polling) that detects new files in an agent's folder and triggers their processing (e.g. convert new PDFs on arrival), scoped capability, no deps
  • [ ] Minimal web Control UI - stdlib http.server page over the existing polyglav serve JSON API (OpenClaw Control UI analogue). Richer frameworks stay plugin-first
  • [ ] Externalize the bundled plugins (polyglav-core-web/fs/exec) into separate versioned repositories, while the bundled copies stay the shipped defaults. Global/local plugins of the same name already override them
  • [ ] PyPI plugin source - discover installed plugin packages via importlib.metadata entry points (polyglav.plugins group)
  • [ ] Shared plugin virtualenv - one venv for all plugin dependencies, injected at import
  • [ ] Per-plugin virtualenv isolation - ~/.config/polyglav/plugins/<name>/.venv. The loader injects its site-packages at import (strongest dependency separation)
  • [ ] Web scraper plugin - full page scraping beyond fetch_page's text extraction (structured content, links), shipped as an external plugin repository
  • [ ] PDF-to-text converter plugin - extract text from local/remote PDFs, shipped as an external plugin repository
  • [ ] Auditor agents - sub-agents that review/check a produced output (tests, code review, fact-check)
  • [ ] Generate > check > correct orchestration - run a main agent, an auditor, and a fix pass in a loop until passing
  • [ ] Custom system prompts per session
  • [ ] code_debug / compile - pdb/gcc/rustc wrappers (test/lint/format landed as code_test/code_lint/code_format)
  • [ ] docs_search - local grep + DuckDuckGo for documentation lookups
  • [ ] Workspace sessions - tools write into a scoped --workspace dir, optional --git sync
  • [ ] Sandboxed exec - namespace/container isolation for run_command (documented, planned for a later version)
  • [ ] /agent types - interactive type selection/run UX (type registry, sub-engine, and delegate landed. The /agent command itself remains)
  • [ ] PM/dev/tester team orchestration as a user-facing pattern (the roles + delegate primitives landed, and it needs the jobs/team-config layer to be a pattern)
  • [ ] Headless web API plugin-first - stdlib http.server fallback, richer framework (FastAPI) via the dependency plugin
  • [ ] Enterprise plugins (stdlib-first, third-party deps optional):
  • [ ] Data ingestion - read_stream / write_stream (MQTT, OPC-UA, Modbus)
  • [ ] Time-series - anomaly_detect (z-score), forecast
  • [ ] Model inference - model_infer / predict_failure (ONNX)
  • [ ] Optimization - optim_schedule (scheduling / linear programming)
  • [ ] SCADA control - scada_command (OPC-UA registers)
  • [ ] Reporting - report_gen (Markdown/PDF, email/BI push)
  • [ ] Audit logging + metrics (/metrics) for enterprise deployments
  • [ ] Onboarding wizard (polyglav wizard) for data-source / MES interface setup
  • [ ] RBAC - role-based access control for enterprise deployments
  • [ ] Queue-based scaling - many concurrent sensor/chat feeds without blocking the loop
  • [ ] Session import from Markdown/JSON

Done

  • [x] Mode-aware prompt color - cyan read, orange write, per-mode color override
  • [x] Team stage error detail - status and sub-session instead of unknown error
  • [x] Team stage self-reentry guard - a stage cannot re-run its own team
  • [x] Sub-agent denial feedback - explicit permission denied, prompt warning, loop cap
  • [x] Ollama null-content fix - tool-only assistant turns serialize empty content, not null
  • [x] Renamed the project from Replio to Polyglav
  • [x] code_test no-match detected on Python 3.11 (Ran 0 tests)
  • [x] Tool identity - git and the dev wrappers print their own glyph and verb
  • [x] Error color - a failed command echoes its output red
  • [x] Compact status params - large or multiline args render as <N chars>
  • [x] Optional output log - output_log writes the REPL transcript with colors to .replio/output/

  • [x] Renamed types to roles - Role/RoleRegistry, /roles + /role, --role, register_roles, roles.json

  • [x] Sub-run stage lines - completed delegate/team stages stay visible with duration and tokens
  • [x] Sub-run verbosity - subrun_verbosity quiet/summary/full, summary forwards writes and edits
  • [x] Reliable multi-line input - explicit close rule (delimiter at line start or blank line), backslash continuation
  • [x] ask never confirm-gated - a question is not blocked by the prompt it needs, and free-text asks stay visible
  • [x] Committer permission - a vcs category gates git_commit, so an allowed role commits unattended
  • [x] Scoped code_test - a target runs that target, empty runs are reported, timeout 300s
  • [x] Config merge fix - nested config objects merge per key across default, global, and local
  • [x] REPL UI colors and prompts - orange actions/asks, red errors, blue thinking, [Y/n] confirm
  • [x] opencode reasoning echo - send reasoning_content back on assistant messages
  • [x] Session log version - each session records the Replio version that wrote it
  • [x] Per-run output buffers - BufferUI writes each run's log, /focus log reads it
  • [x] Run status animation - Braille status spinner for delegation/team stages
  • [x] Provider session binding - provider session_id bound to the run session
  • [x] Memory scopes - shared .replio/memory/ role/team/job memory, /memorize
  • [x] Run continuation - delegate/team resume/context, warm sub_<key> removed
  • [x] Run-owned focus - run-keyed stack, run_engine resume, no agent_<role>
  • [x] Run tree navigation - /focus by run/session with ↔ Switch to <role>
  • [x] Run-to-run handoff - handoff targets runs and preserves the target session
  • [x] focus_on_delegate targets the child run (TurnResult.run_id)
  • [x] Session log fields - name -> session_name, code -> session_id
  • [x] Job and delegation names - job_<ts>_<id> and sub_<ts>_<id>
  • [x] Stable session id - six-char session_id, ses_<ts>_<id> names, and #id handles
  • [x] /print - reprint a turn (or one part) in full, cap via print_max_chars, --full, --run
  • [x] /history - numbered turn index with n/all, --thoughts [all], and a --run selector
  • [x] Session turn cutover - turn/part storage, co-located tool calls, flat messages removed
  • [x] Session turn model - turns.py turn/part shape, builders, and provider conversion
  • [x] focus_on_delegate config - off/ask/on to focus the delegated role after the run
  • [x] handoff tool - pause or finish a run and hand control to a parent/sibling/child/role/#id
  • [x] /focus command - show the run tree/log, attach by id/role/session, navigate runs
  • [x] Focus manager and role engines - REPL focus stack, stable agent_<role> sessions, prompt_role
  • [x] Run registry and call tree - in-process runs with parent/children and a flat call log
  • [x] Session role metadata - the owning role is stamped as role on every session at creation
  • [x] Assistant root role - REPL binds assistant/assistant_type, delegation-first prompt
  • [x] Composer role - bundled team composer, catalog allow, edit/team deny
  • [x] Team review loop - team loop over producer/reviewer, VERDICT pass, max_iterations
  • [x] Warm member sessions - delegate/team session_key, team warm_sessions reuse sub_
  • [x] Agent catalog tool - list/show/save/remove types, teams, skills, plus reload
  • [x] Per-invocation skills - delegate/team/stage skills layered over a role's standing skills
  • [x] End-to-end supervisor verification - replio jobs add-supervisor, governance-loop test
  • [x] Report-back - job.run.completed events, last run:/summary: lines, footer status
  • [x] replio-core-webhook - report-back connector POSTing job reports to report.webhook
  • [x] Pending-ask inbox - unattended human asks park in .replio/asks.json, answered via /asks
  • [x] Unattended mode - no stdin at any depth, confirms auto-deny, confirm_timeout
  • [x] team tool - model-invokable team pipeline, per-stage permission + ceiling/depth guards
  • [x] Delegation permission ceiling + ask permission grants - grant_permission, ask_policy
  • [x] CLI realigned to internal commands - replio models/config/plugins mirror /models, /config, /plugins
  • [x] Plural registry commands - /roles, /teams, /skills (was /type, /team, /skill)
  • [x] /session split - active session vs /sessions catalog (mirrors /model vs /models)
  • [x] Model commands - /model show/switch, /models configured, /models list [provider] live
  • [x] Loop tools - /tool delegate runs a persisted loop turn, sub-agent ask reaches the operator
  • [x] Externalize bundled providers - vendor providers to bundled plugins, base + host hints stay core
  • [x] ask tool - core: human or lead mid-run questions, sub-agent routing
  • [x] Connect any OpenAI-compatible endpoint by URL - named custom provider entries
  • [x] /connect provider rework - name/URL connect, preset defaults, re-enter key
  • [x] Model refs unfold - provider/model refs resolve, gated on approved models
  • [x] models.json - approved-model history (provider, model, timestamps), no keys
  • [x] Provider registry - providers.json, one API key per provider, engine resolves from it
  • [x] Sequential team runs - run_team stage loop, per-run briefs, team memory, /team run
  • [x] persona renamed to agent type - types.py/TypeRegistry, /type, --type, no aliases
  • [x] Bundled web plugin renamed - replio-core-websearch -> replio-core-web
  • [x] GitHub Pages website - mkdocs site + Actions workflow, docs stay in main
  • [x] OpenCode Zen + Go providers - opencode.ai endpoints, model-ref strip, URL detection
  • [x] Tool-use evaluation harness - replio eval + fixture plugin, metrics
  • [x] Project instructions file - per-worktree AGENTS.md in system prompt, capped
  • [x] run_command allowlist - tool_permission.bash_allow, heredocs/multi-line rejected
  • [x] code_test/code_lint/code_format wrappers - dev.*_cmd, bundled replio-core-dev plugin
  • [x] git tool - read-only git + gated git_commit, bundled replio-core-git plugin
  • [x] file_edit tool - search-and-replace with diff preview, bundled replio-core-edit plugin
  • [x] Default tool-result cap - tool_max_result_chars default 100k, list_dir entry cap
  • [x] Default tool names namespaced - web_fetch merges open+fetch_page, old names as aliases
  • [x] Tool schema advertises canonical names only - aliases resolve at call time (20 -> 13 defs)
  • [x] Agent-loop cancellation + tool-dialect hardening - Ctrl-C cancels the turn, unknown-tool hints
  • [x] Skills registry - SkillRegistry, local/global dirs + plugins, type skills in prompts
  • [x] Teams registry - Team/TeamRegistry, 4-layer teams.json merge, /team read/edit
  • [x] Plugin contribution hooks - register_roles/teams/skills hooks, RoleRegistry.reload()
  • [x] Fleet orchestration - replio fleet supervisor CLI: ports, health, restart, config gen
  • [x] Per-run job sessions - job_<ts>_<name> files, unified ses_/sub_/job_ naming
  • [x] Job run memory - rolling .memory.md summary injected into each run, seeded
  • [x] Job task definition as linked Markdown - --file templated task files, replio jobs edit
  • [x] Job context management - per-run job_* files, ses_/sub_ naming, 100-run cap
  • [x] Job per-run approval - --require-approval parks a run in waiting_approval
  • [x] Jobs diagnostics - replio jobs status (fired count, last error, uptime)
  • [x] Watchable job runs - replio jobs run --verbose streams the live turn
  • [x] Richer job runs - --type/--system-prompt, recurring-job prompt, tool carve
  • [x] Scheduled / durable jobs - replio jobs: registry, daemon, cron, retries, approval gate
  • [x] Per-agent permission profiles - role tool_permission drives sub-agent policy
  • [x] In-process sub-engine - Engine.run_subagent: role overrides, NullUI, sub_ session
  • [x] delegate tool - core, per-role permission resolver, delegate_echo result display
  • [x] Roles registry - global+local roles.json merge, /roles command, docs/roles.md
  • [x] Soft tool results surfaced as dimmed info lines - (empty file), (no matches) etc
  • [x] Session audit trail - permission decisions (allow/ask/deny -> granted/declined/denied)
  • [x] Stable message ids - msg_<hex> auto-assigned to every session message
  • [x] Global model registry - /connect appends models + keys, /model list/--online, picker reuse
  • [x] REPL /config --global/--local scope flags + apply() in-memory overrides
  • [x] Scoped config writes - local saves overrides only, replio config CLI
  • [x] Turn recovery - auto-continue on truncation, reasoning-only not flagged empty
  • [x] Thinking captured from reasoning deltas (ollama.com), documented
  • [x] Confirm prompt ? glyph starts at the beginning of the line, aligned with activity glyphs
  • [x] Word-level streaming buffering - buffer REPL output to word boundaries, word_streaming config
  • [x] Multi-line input - detect """/''' blocks, one composed message, framing stripped
  • [x] Config validation (test connection on change) - /connect probe, /provider warn
  • [x] Session export to Markdown - /session export and replio export CLI render logs to Markdown
  • [x] Plan/Build modes - plan (edit+bash denied) vs build, /mode cmd + --mode flag, per-message mode
  • [x] Thinking/reasoning toggle - show_thinking (display) + reasoning (request), /thinking cmd
  • [x] Thinking spinner - animated spinner while thinking when show_thinking false, cleared \r\033[K
  • [x] Activity lines + tool status ephemeral - never persisted to session files
  • [x] MCP via bundled replio-core-mcp plugin - import/expose tools over MCP, dual-era
  • [x] Alias layer - tool/param aliases (read/view, ls, bash/exec, cursor>offset) + open tool
  • [x] Glyph activity lines - dimmed status, glyph_lines config
  • [x] write_file reports resolved abs path (Created|Overwritten|Appended)
  • [x] Tab completion restored (libedit binding) + extended to paths, tool names, and subcommands
  • [x] README terminal screenshot as SVG
  • [x] deploy/ fleet templates - Dockerfile, compose, systemd + launchd units
  • [x] docs/fleet.md - single-purpose scoped-agent fleet pattern
  • [x] README fleet positioning - tagline, Features, fleet section, roadmap
  • [x] Per-turn stats on own line - footer newline when output lacked one
  • [x] One-shot retry for empty/truncated streams before surfacing error
  • [x] write_file status preview ends with dimmed summary (path, lines, chars, action)
  • [x] Thinking announce - + Thinking (or + Thought 12.3s hidden). Headless mirrors
  • [x] Human-readable tool status - [tool: arg] oneliner + dimmed detail lines
  • [x] Built-in web + machine features moved to bundled plugins
  • [x] Discovery precedence bundled < global < local, bundled can't update/uninstall
  • [x] register_services entry hook - powers the web_search: true search-then-answer mode
  • [x] plugins config list replaces plugins.enabled/deny, empty = all
  • [x] tests/test_bundled_plugins.py (10 tests)
  • [x] Plugin system - directory-based external plugins for tools, providers, and commands
  • [x] PluginManager - discovery, manifest compat ranges, single entry import
  • [x] register_tools hook - tools inherit policy, /tool, /help, logging
  • [x] register_providers hook merged into provider registry
  • [x] register_commands(commands) hook, registered after builtins at engine init
  • [x] Plugin manifest + docs - docs/plugins.md (schema, compatibility contract, security)
  • [x] Optional per-plugin deps - requires, lazy import, --deps pip-installs
  • [x] Activation via config - plugins list, enable/disable/install/uninstall maintain it
  • [x] /plugins - list/detail/enable/disable/install/update/uninstall
  • [x] replio plugins CLI - headless list/install/update/uninstall
  • [x] tests/test_plugins.py (30 tests)
  • [x] PyPI proper configuration and documentation
  • [x] Clear screen on REPL start (clear_screen config, default true)
  • [x] /config structured values - JSON parse, -a/-r, reload
  • [x] /compact summarization - append-only, compact_from boundary
  • [x] /session load compaction offer prints the summary, load records a command message
  • [x] /session preview <name> - read-only structural preview without switching sessions
  • [x] Session-name tab completion for /session load/delete
  • [x] Context-size display - dimmed (Ns, N tokens) after each response
  • [x] noise_tools config - noise results replaced by marker in sessions
  • [x] max_tokens optional (0=unset), hitting cap warns + logs error
  • [x] Provider payload from log - command filtered, dangling tool msgs skipped
  • [x] replio run - one-shot CLI (JSON/text) reusing loop + sessions
  • [x] Flags: --prompt, --provider, --model, --output=json, --verbose
  • [x] --session-id - address persistent sessions from headless mode
  • [x] Headless logging (--verbose) instead of visual activity lines
  • [x] replio serve - stdlib http.server HTTP JSON API (POST /chat) over the same agent loop
  • [x] Tests for headless entry points (mock provider, no network)
  • [x] Hardening: agent-loop turn failures always visible in the session log
  • [x] Persist streamed content when the SSE stream ends without a done event
  • [x] Record errors entry for silent failures (EOF, empty/thinking-only done)
  • [x] Non-normal finish_reason (length) logs an errors entry + prints a warning
  • [x] Catch unexpected exceptions escaping the agent loop, log them, keep the REPL alive
  • [x] Tests: token-stream-then-EOF, empty done, streamed exception
  • [x] Sessions are complete logs - every message, tool call + result, reasoning, error persisted
  • [x] Append-only - compaction and load never remove or rewrite entries
  • [x] Session.to_dict() keeps role: tool (full results, noise_tools replaced by a marker)
  • [x] Per-round thinking metadata on assistant messages, excluded from content
  • [x] Session-level errors array
  • [x] created_at / updated_at metadata (bumped on every message)
  • [x] tool_analysis (default false) - model-generated one-line analysis on each tool message
  • [x] session_tool_max_chars (default 0 = unlimited) - caps persisted tool-result content
  • [x] No index.json - session names carry timestamp + first-message slug
  • [x] Flat message list retained (maps 1:1 to provider context)
  • [x] Unified streaming agent loop
  • [x] Single SSE stream, detect tool_calls vs content from the first delta
  • [x] Eliminates double-call cost when no tools are used
  • [x] Unified dispatch
  • [x] Slash commands call the same ToolRegistry as the model
  • [x] Generic query refinement via tool metadata and collapsed _handle_message branches
  • [x] Machine tools (tools/machine.py) - read_file, list_dir, write_file, run_command
  • [x] read_file - numbered lines, offset/limit, truncation, binary/permission errors
  • [x] list_dir - sorted entries, trailing / for subdirs, file sizes
  • [x] write_file - parent-dir creation, w/a modes
  • [x] run_command - subprocess exec with timeout, stdout/stderr capture, exit code, 8k cap
  • [x] category/permission/path_arg/key_arg registration metadata
  • [x] Discovery tools - glob (recursive, noise-dir skip) + grep (regex, file:line)
  • [x] read_file always reports total line count in its header
  • [x] Tool permission model (tools/policy.py + config)
  • [x] tools.allow / tools.deny name-level policies (deny + allow-whitelist take precedence)
  • [x] tool_permission category actions - allow/ask/deny per category
  • [x] Path-scoped confirm - outside-worktree read/write/list escalate to ask
  • [x] bash: ask default - every run_command prompts y/N
  • [x] Confirm prompts - denied tools filtered, cancelled calls feed [cancelled]
  • [x] /tool command routes through the same policy
  • [x] OpenAI / Groq / Anthropic providers
  • [x] Provider auto-detection from base_url
  • [x] Base provider (OpenAI-compatible interface)
  • [x] Ollama cloud provider
  • [x] Session manager (JSON CRUD)
  • [x] /session new and /session load actually switch the active session
  • [x] Session save after file rename - JSON name field matches filename
  • [x] Auto-session naming with first user message as context hint
  • [x] Tool results and assistant tool_calls persisted in session files
  • [x] DuckDuckGo Lite search via html.parser
  • [x] Terminal + AI context formatting
  • [x] /search <query> and /web <query> commands
  • [x] Auto-search mode (web_search: true config)
  • [x] fetch_page - _TextExtractor (HTMLParser) for clean text extraction
  • [x] Search term extraction - query_refine config auto-refines short queries
  • [x] tools/registry.py - decorator-based tool registration
  • [x] tools/builtins.py - web_search and fetch_page tools
  • [x] BaseProvider.chat_nonstreaming() - non-streaming tool decision round
  • [x] Two-phase chat: non-streaming tool decision > stream final content
  • [x] _show_tool_status() - dimmed status during tool execution
  • [x] /search command integration with tool calling
  • [x] Config default: tool_calling: true
  • [x] tool_status_visible config flag (default true)
  • [x] Tool calling integration verified end-to-end
  • [x] Thinking/reasoning token detection and dimmed display
  • [x] SSE streaming survives multi-byte UTF-8 split across read chunks (byte-buffered line decoding)
  • [x] Markdown-aware streaming (disabled by default via markdown_streaming)
  • [x] Error handling improvements (network timeout, auth errors)
  • [x] Edge cases: streaming when tool_calling=true but no tools used
  • [x] Tool-call messages lost on exception - try/finally persists session
  • [x] Final assistant response missing when streaming returns empty - non-streaming fallback
  • [x] Project scaffolding (pyproject.toml, venv, dir structure)
  • [x] Config module (global + local JSON merge)
  • [x] HTTP SSE streaming utility (urllib)
  • [x] REPL loop with readline history + tab completion
  • [x] Command registry + built-in slash commands
  • [x] Streaming token display
  • [x] Documentation (README, AGENTS.md, TODO.md, CHANGELOG.md)