TODO¶
- First-run onboarding - the assistant introduces itself, explains what it can do, and asks what to do. No system-level configuration for simple users
- Also ask how agent focus should behave (
focus_on_delegate: off/ask/on) and whether the prompt names the active role (prompt_role), both off by default - One-window status -
/statusshows sessions, running agents, and configured jobs on the current machine, with logs reachable from the same place (the journalctl-style alternative) - Agent health monitoring - the assistant watches endpoints (e.g. the
/healthof agents running as web APIs) and warns when an agent stops responding - Per-agent todo lists - view a delegated agent's tasks, mark items done, jump into its session, and ask for the current state (OpenCode-style)
- Non-blocking delegation - the assistant starts sub-agents or whole teams for bigger tasks and reports their status instead of blocking the current run
- Recurring tasks carry their own role - each job carries its own role and skills, so behavior like "make doc changes per AGENTS.md" is encoded once instead of re-prompted every time
- Auto-improving skills (autogen/autolearn) - a role refines its skills from experience and keeps them alongside its memory, so recurring work gets better without re-prompting
- Runtime pluginization - make models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI replaceable plugins, and move the REPL to a plugin. Five replaceable layers: access, orchestration, capability, model, storage
- One runtime, many agents - position Polyglav as an agent harness for fleets, not a single assistant
- ACP inward and outward - speak the Agent Client Protocol so Polyglav can host and be hosted by other harnesses, alongside MCP
- Commands as tools - expose slash commands to the model as permission-gated tools to configure roles, teams, and skills, with command access limited per permission
- Markdown catalogue - store roles, teams, and skills as Markdown files (front matter) referenced from JSON, and move the bundled data files (
src/polyglav/*.json) into one catalogue directory nearsrc - Hidden files and allowed paths - hide secrets and config from tools by default, and restrict visible paths to allowed roots
- Context policy - keep user prompts until the task is done, cut unused references and tool results immediately, add a recall tool for role and team memory, enforce a configurable context limit, and offload large tool results to
.polyglav/tmp/referenced from the session log - Context UI polish - auto-compaction, context trim, highlighting, and context desaturation
- Research cache - store web and PDF results under
.polyglav/filesfor later reuse, with a PDF-to-text tool - Fleet awareness - the supervisor always knows which agents and teams run on a machine, with oversight across multiple machines
- Open WebUI and OpenTUI connectors - drive Polyglav from external chat and terminal UIs
- Tool-call quality enhancer - steer tool calling toward correct, cheaper calls (for example web-search parameters that find the right references) and better results
- PlantUML plugin - an external plugin that draws configurations and workflows as node diagrams instead of ASCII
- UI commands -
/clear,/new,/providers,/providers list,/thinking [hide|show], and persist error logs - Feedback - capture operator feedback on a run or answer and feed it back into memory and skills
- Inbound webhooks - accept inbound events to start a run or answer a parked ask
- Plugin management UI - list, enable, disable, install, and update plugins from a UI surface, not only slash commands
- Model picker -
/modelsautocomplete and a direct numeric selection of a model - Centralized static strings - move hardcoded prompts and static strings into one module
- Code cleanup pass - a dedicated pass over the core for dead code, naming, and rough edges
- Report-back connectors - job summaries delivered out-of-band (email first, idea only) when the terminal is closed
- Remove the legacy provider-plugin backfill migration (
Config._migrate_plugins) once a stable polyglav release has shipped with the externalized providers. Existingpluginslists no longer need the automatic append, and the migration code is dead weight - Human-in-the-loop channels - job events outbound (
proposed,will_run,failed,waiting_approval) plus inbound actions (approve/reject/run/disable) over configurable connectors (webhook first, then email, then Telegram), so a job can reach an operator who is not on the box - Global jobs overview across agents -
polyglav jobs list --root <dir>fleet scan, aGET /jobs+POST /jobs/<name>/approve|reject|run|disableoperator API onpolyglav serve, then a web Control UI, so one view shows which agents run next and with which task - Edge / offline store-and-forward buffering - offline-capable agents with local buffering for unreliable connectivity (enterprise use case)
- Immutable agent config -
polyglav serveagents must never be able to change their own configuration, permissions, or tool list (control-plane rule from the use-case reference architecture) - Hash-chained / tamper-evident audit log - additive on session logs (enterprise.md recommendation), hash-chained append or WORM storage
- Single-purpose agent fleet with "one agent per process, scoped to a folder" as the headline pattern. README +
docs/fleet.mdset the niche - Fleet and swarm orchestration as the two layers - fleet: supervisor running many scoped
polyglav serveinstances (port allocation, health checks, restart policy, per-agent config generation), swarm:/agenttypes, auditor agents, generate > check > correct - Community presence - decide on Discord/X channels and fill the README community link slots
- Add multiuser capability or queue for API requests (request queue, per-token rate limits)
- Add ReadTheDocs documentation
- Citations / source attribution - return URL + snippet with every answer
- Bookmarks -
/bookmark add/remove/listfor session pinning - Interactive data analysis - CSV querying, SQL execution, code eval in REPL
- Notebook mode - persistent editable cells with run outputs
- Hybrid web + local RAG - vector store (FAISS/Weaviate), embeddings, local document search
- Command palette / fuzzy search - CTRL-P style history search
- Topic-aware ranking - classifier for query intent to weight search results
- Competitor research - validate USPs against actual peers (OpenClaw, Claude Code, opencode, agentic-infra services). Comparison notes now live in
docs/compare/and feed the feature backlog (Plan/Build modes, Web Control UI, plugin marketplace, telemetry, binary builds, sharing) - Self-update -
polyglav update(Pipi update --selfanalogue) - Standalone binary build - Pi-style release script producing a single executable (contentious for a zero-dep Python package)
- Opt-in telemetry contracts - vendor-neutral event schema (OpenCode, Pi
@earendil-works/pi-telemetry). Decide whether it fits the no-telemetry stance - Conversation sharing - web-shareable session links (OpenCode
/share) or published sessions (Pipi-share-hf), building on the planned Markdown export - Plugin registry / marketplace - discoverable plugin sharing (OpenClaw ClawHub analogue). PyPI entry-point source is a prerequisite
Open¶
- [ ] Command audit - review the slash commands for opaque or overlapping behavior and consolidate
- [ ] Focus keeps the run's context - re-entering a finished run uses its retained engine or resumes that run's own session, and never starts a fresh session. /focus drops the session:
attach, because saved sessions belong to /load - [ ] Interruptible run status - the status line yields to the keyboard during a delegated or team run, with a hint line naming the keys (Enter to continue, ^O to open, ^C to cancel) and a Switch
marker - [ ] Focus a live sub-run - /focus attaches to a running sub-run and leaves the caller reachable
- [ ] Tester loop scope - the tester runs only the related tests during the loop and the whole suite once, as a role-prompt change
- [ ] Thinking visibility - with show_thinking false only the + Thought N.Ns line prints, so the reasoning is invisible while its duration is shown
- [ ] Prompt brevity budget - every role prompt states a short output budget and a fixed report shape, so a report is a few lines instead of an essay
- [ ] Docs writer commits each added task - the docs writer commits the tasks right after the operator's prompt and then works through the single points one at a time
- [ ] hide_confirm_input default true - the typed input is hidden on the tool confirm unless overridden
- [ ] Batched structured asks -
askhandles one question at a time, so several decisions cannot be asked in one call - [ ] Committer as a callable stage - a run calls the committer to land the current state as one commit with a correct message, so a long run does not accumulate uncommitted work
- [ ] Researcher role with fresh context - a bundled researcher role with web access, started fresh for every task so research never inherits stale context
- [ ] Per-role path scoping - a role declares the files and folders it may touch, enforced by the tool policy rather than by a prompt
- [ ]
.polyglavdirectory layout - decide and document the reserved subfolders and file names under.polyglav/ - [ ] Skill definition - a skill holds tool, language, or framework instructions, not a project description
- [ ] Skill catalog review - rework the bundled and local skills against that definition, add a
pythonskill and apolyglavskill, and move project-description text intoAGENTS.md - [ ]
AGENTS.mdas the project description - keep it the single project description and update it when the structure, conventions, or extension points change - [ ] Leader PM posture - the leader holds the whole picture, pushes back on a request that breaks the project, and concretizes an ambiguous prompt until the requirement is synced instead of guessing
- [ ]
memorizeas a tool - memory writes happen through a tool an agent calls, triggered by the operator's prompt, not only through the slash command - [ ] Role instruction files - a role's long instructions live in
.polyglav/roles/<name>.md, referenced from the short JSON entry and appended verbatim - [ ] Memory with references - a bounded summary that points at full-length Markdown and session artifacts, with a guard against a misleading reference when the context is gone
- [ ] Root role memory - inject
.polyglav/memory/roles/<role>.mdinbind_root_agentthrough a shared prompt-composition helper, with a refresh path - [ ] Conclusion stage - a bundled stage with a write-scoped role that distills a finished run into role files, skills, or memory, and never commits
- [ ] Saved-session catalog in
/load-/loadlists and loads saved sessions so an operator can reattach to a prior agent after a restart, while/focusstays live-runs-only - [ ] Ask continuation - answering a parked ask resumes and continues its origin run in place, instead of only injecting the answer into the session
- [ ] Handoff from sub-agents - a team stage or delegated agent can hand focus to the next agent (composer > planner > developer), not only the REPL root
- [ ] Fix the PyPI long-description screenshot - the image fails to load on the package page
- [ ] Fix opencode permission rejection - the provider stops when a permission is rejected instead of continuing
- [ ] Exact-args permission grants - approve the specific command (not just the tool), and let an
alwaysgrant live beyond the current sub-agent run - [ ] REPL UI colors and prompts - a consistent color scheme and prompt behavior for the REPL:
- [ ] User prompt - bold the
>>>/...marker so it stands out in long output - [ ] Actions/Asks orange - activity/tool status lines and the confirm/ask prompts
- [ ] Keep output dimmed - tool results and detail lines stay dim
- [ ] Newline before status - a status line never prints inline after streamed text
- [ ]
[Y/n]enter-to-continue - confirm defaults to yes on an empty answer - [ ] Hide confirm input - hidden for the
[Y/n]confirm, visible for a free-text ask - [ ] Errored tool calls red - the
! Error:line in red - [ ] Thinking/Thought blue - headers blue, reasoning body dim
- [ ] Runs, focus, and memory redesign (see PLAN.md
Runs, focus, and memory): - [ ] Non-blocking runs and live focus - background execution, output/input routing, cancellation
- [ ] Remove
--session-id- explicit session naming is no longer needed now that auto sessions are namedses_<ts>_<id> - [ ] Relocate job run sessions under
.polyglav/jobs/<name>/(kept insessions/for now) - [ ] Role-name sync - adopt assistant, composer, manager, and specialist as the canonical roles across types, prompts, and docs
- [ ] Assistant-roles track docs - record the assistant, composer, and manager architecture and the work packages in VISION, PLAN, and TODO
- [ ] Core dev team configuration - a bundled development team with the review loop plus the project lead/support teams and their skills
- [ ] Manager role - a bundled role that runs one or many teams and reports, sequential first
- [ ] Per-job report destination - a
report_url(or connector list) on a job so different jobs report to different endpoints, instead of one globalreport.webhook - [ ] Per-task decide-vs-park for
directionasks - a task class (or per-run switch) that lets the supervisor auto-resolve a direction ask after a timeout instead of always parking it for the operator - [ ] Full
file_*namespace extension - iffile_glob/file_grepprove better with most models, extend the prefix tolist_dir/glob/grep(old names stay aliases) - [ ] Tool spec polish - rename
grep.glob->include(aliasglob), add examples and prefer-web_fetchguidance to tool descriptions - [ ] Mid-run blocking job approval - an
asktool inside a running job pauses the run in place (per-tool-callwaiting_approval), notifies via a connector, and resumes the same session when the operator replies. Needs resumable mid-run state, a wait loop inside the run, and the connectors/transport below (deeper than the shipped per-run--require-approvalgate) - [ ] Job event hooks - the scheduler emits typed transitions (
proposed,approved,will_run,executing,verified,failed,timeout,waiting_approval) to registeredservices. Channel-agnostic core, first consumers are the connectors and the operator API - [ ] Job connectors - bundled
polyglav-core-webhook(stdlib JSON POST, zero deps, works with n8n/IFTTT/any URL) first. External email (SMTP + polling) and Telegram (urllib long-poll) plugins later, all driving the jobs operator API so operators can react in time - [ ] Jobs operator API -
GET /jobsandPOST /jobs/<name>/approve|reject|run|disableonpolyglav serve, so clients (web Control UI, connectors, fleet supervisor) can see and act per agent - [ ] Fleet jobs overview -
polyglav jobs list --root <dir>scanning agent worktrees (agent, job, status, next run, task table), then the web Control UI on top - [ ] Role directory scan for export/import - read
.polyglav/roles/*.md(front-matter roles) to import and export roles to Markdown, paralleling the sessions Markdown export/import - [ ] Auto team selection - the assistant picks roles, teams, and skills from the registries for a task and delegates in sequence (team orchestration as a user-facing pattern, e.g. "compare with competitors" -> Researcher > Writer > Referencer > Editor)
- [ ] Thinking visibility -
/thinking on+reasoningconfig documented, per-providerreasoning_contentcheck so reasoning shows in the REPL - [ ]
/spawncommand - launch a scopedpolyglav serveagent from the REPL (home -> project path), supervise (health/list/stop) and delegate to it (docs/fleet.md) - [ ] Remote channels - command agents from messaging apps (OpenClaw channels parity):
- [ ] Channel gateway - one adapter surface over the engine/serve API
- [ ] Telegram adapter - long-polling bot, send + receive
- [ ] WhatsApp adapter - business-API HTTP channel
- [ ] More adapters (Discord, Signal, email)
- [ ] Remote auth + session scoping + headless deny
- [ ] Plugin test harness -
polyglav plugins test <name>ships. Bundled plugin suites live next to the plugins (plugins/<name>/tests/, discovered by the core suite). Remaining: external plugins are expected to ship a test suite, andpolyglav plugins testis the runner for them - [ ] Session recall - full-text search across past sessions (grep/index over
.polyglav/sessions/) so an agent can answer from its own history - [ ] Tool dry-run mode - propose tool args/effects without executing (enterprise tool-gateway requirement)
- [ ] Context-aware cross-plugin tool router - virtual tool names (
open,search, ...) dispatch per-argument to the matching plugin handler viaregister_handler(name, match=...)(e.g.open https://...> polyglav-core-web,open ../...> polyglav-core-fs), with merged schemas and args-aware policy accessors - [ ] Swarm orchestration - agent cooperation layer (
docs/swarm.md):/agenttypes, auditor agents, generate > check > correct, and team patterns as sub-tasks below - [ ] Grep text index - internal bundled plugin (stdlib) that indexes converted text files for local search, bridging toward the vector store
- [ ] Agent folder watcher - internal bundled plugin (stdlib
threading+pathlibpolling) that detects new files in an agent's folder and triggers their processing (e.g. convert new PDFs on arrival), scoped capability, no deps - [ ] Minimal web Control UI - stdlib
http.serverpage over the existingpolyglav serveJSON API (OpenClaw Control UI analogue). Richer frameworks stay plugin-first - [ ] Externalize the bundled plugins (
polyglav-core-web/fs/exec) into separate versioned repositories, while the bundled copies stay the shipped defaults. Global/local plugins of the same name already override them - [ ] PyPI plugin source - discover installed plugin packages via
importlib.metadataentry points (polyglav.pluginsgroup) - [ ] Shared plugin virtualenv - one venv for all plugin dependencies, injected at import
- [ ] Per-plugin virtualenv isolation -
~/.config/polyglav/plugins/<name>/.venv. The loader injects its site-packages at import (strongest dependency separation) - [ ] Web scraper plugin - full page scraping beyond
fetch_page's text extraction (structured content, links), shipped as an external plugin repository - [ ] PDF-to-text converter plugin - extract text from local/remote PDFs, shipped as an external plugin repository
- [ ] Auditor agents - sub-agents that review/check a produced output (tests, code review, fact-check)
- [ ] Generate > check > correct orchestration - run a main agent, an auditor, and a fix pass in a loop until passing
- [ ] Custom system prompts per session
- [ ]
code_debug/compile- pdb/gcc/rustc wrappers (test/lint/format landed ascode_test/code_lint/code_format) - [ ]
docs_search- local grep + DuckDuckGo for documentation lookups - [ ] Workspace sessions - tools write into a scoped
--workspacedir, optional--gitsync - [ ] Sandboxed exec - namespace/container isolation for
run_command(documented, planned for a later version) - [ ]
/agenttypes - interactive type selection/run UX (type registry, sub-engine, anddelegatelanded. The/agentcommand itself remains) - [ ] PM/dev/tester team orchestration as a user-facing pattern (the roles +
delegateprimitives landed, and it needs the jobs/team-config layer to be a pattern) - [ ] Headless web API plugin-first - stdlib
http.serverfallback, richer framework (FastAPI) via the dependency plugin - [ ] Enterprise plugins (stdlib-first, third-party deps optional):
- [ ] Data ingestion -
read_stream/write_stream(MQTT, OPC-UA, Modbus) - [ ] Time-series -
anomaly_detect(z-score),forecast - [ ] Model inference -
model_infer/predict_failure(ONNX) - [ ] Optimization -
optim_schedule(scheduling / linear programming) - [ ] SCADA control -
scada_command(OPC-UA registers) - [ ] Reporting -
report_gen(Markdown/PDF, email/BI push) - [ ] Audit logging + metrics (
/metrics) for enterprise deployments - [ ] Onboarding wizard (
polyglav wizard) for data-source / MES interface setup - [ ] RBAC - role-based access control for enterprise deployments
- [ ] Queue-based scaling - many concurrent sensor/chat feeds without blocking the loop
- [ ] Session import from Markdown/JSON
Done¶
- [x] Mode-aware prompt color - cyan read, orange write, per-mode
coloroverride - [x] Team stage error detail - status and sub-session instead of
unknown error - [x] Team stage self-reentry guard - a stage cannot re-run its own team
- [x] Sub-agent denial feedback - explicit
permission denied, prompt warning, loop cap - [x] Ollama null-content fix - tool-only assistant turns serialize empty content, not null
- [x] Renamed the project from Replio to Polyglav
- [x]
code_testno-match detected on Python 3.11 (Ran 0 tests) - [x] Tool identity - git and the dev wrappers print their own glyph and verb
- [x] Error color - a failed command echoes its output red
- [x] Compact status params - large or multiline args render as
<N chars> -
[x] Optional output log -
output_logwrites the REPL transcript with colors to.replio/output/ -
[x] Renamed types to roles - Role/RoleRegistry, /roles + /role, --role, register_roles, roles.json
- [x] Sub-run stage lines - completed delegate/team stages stay visible with duration and tokens
- [x] Sub-run verbosity -
subrun_verbosityquiet/summary/full, summary forwards writes and edits - [x] Reliable multi-line input - explicit close rule (delimiter at line start or blank line), backslash continuation
- [x]
asknever confirm-gated - a question is not blocked by the prompt it needs, and free-text asks stay visible - [x] Committer permission - a vcs category gates git_commit, so an allowed role commits unattended
- [x] Scoped
code_test- a target runs that target, empty runs are reported, timeout 300s - [x] Config merge fix - nested config objects merge per key across default, global, and local
- [x] REPL UI colors and prompts - orange actions/asks, red errors, blue thinking,
[Y/n]confirm - [x] opencode reasoning echo - send
reasoning_contentback on assistant messages - [x] Session log version - each session records the Replio version that wrote it
- [x] Per-run output buffers -
BufferUIwrites each run's log,/focus logreads it - [x] Run status animation - Braille status spinner for delegation/team stages
- [x] Provider session binding - provider
session_idbound to the run session - [x] Memory scopes - shared
.replio/memory/role/team/job memory,/memorize - [x] Run continuation -
delegate/teamresume/context, warmsub_<key>removed - [x] Run-owned focus - run-keyed stack,
run_engineresume, noagent_<role> - [x] Run tree navigation -
/focusby run/session with↔ Switch to <role> - [x] Run-to-run handoff -
handofftargets runs and preserves the target session - [x]
focus_on_delegatetargets the child run (TurnResult.run_id) - [x] Session log fields -
name->session_name,code->session_id - [x] Job and delegation names -
job_<ts>_<id>andsub_<ts>_<id> - [x] Stable session id - six-char
session_id,ses_<ts>_<id>names, and#idhandles - [x]
/print- reprint a turn (or one part) in full, cap viaprint_max_chars,--full,--run - [x]
/history- numbered turn index withn/all,--thoughts [all], and a--runselector - [x] Session turn cutover - turn/part storage, co-located tool calls, flat messages removed
- [x] Session turn model -
turns.pyturn/part shape, builders, and provider conversion - [x]
focus_on_delegateconfig - off/ask/on to focus the delegated role after the run - [x]
handofftool - pause or finish a run and hand control to a parent/sibling/child/role/#id - [x]
/focuscommand - show the run tree/log, attach by id/role/session, navigate runs - [x] Focus manager and role engines - REPL focus stack, stable
agent_<role>sessions,prompt_role - [x] Run registry and call tree - in-process runs with parent/children and a flat call log
- [x] Session role metadata - the owning role is stamped as
roleon every session at creation - [x] Assistant root role - REPL binds assistant/assistant_type, delegation-first prompt
- [x] Composer role - bundled team composer, catalog allow, edit/team deny
- [x] Team review loop - team loop over producer/reviewer, VERDICT pass, max_iterations
- [x] Warm member sessions - delegate/team session_key, team warm_sessions reuse sub_
- [x] Agent catalog tool - list/show/save/remove types, teams, skills, plus reload
- [x] Per-invocation skills - delegate/team/stage skills layered over a role's standing skills
- [x] End-to-end supervisor verification -
replio jobs add-supervisor, governance-loop test - [x] Report-back -
job.run.completedevents,last run:/summary:lines, footer status - [x]
replio-core-webhook- report-back connector POSTing job reports toreport.webhook - [x] Pending-ask inbox - unattended human asks park in
.replio/asks.json, answered via/asks - [x] Unattended mode - no stdin at any depth, confirms auto-deny,
confirm_timeout - [x]
teamtool - model-invokable team pipeline, per-stage permission + ceiling/depth guards - [x] Delegation permission ceiling +
askpermission grants -grant_permission,ask_policy - [x] CLI realigned to internal commands - replio models/config/plugins mirror /models, /config, /plugins
- [x] Plural registry commands - /roles, /teams, /skills (was /type, /team, /skill)
- [x]
/sessionsplit - active session vs/sessionscatalog (mirrors /model vs /models) - [x] Model commands - /model show/switch, /models configured, /models list [provider] live
- [x] Loop tools -
/tool delegateruns a persisted loop turn, sub-agent ask reaches the operator - [x] Externalize bundled providers - vendor providers to bundled plugins, base + host hints stay core
- [x]
asktool - core: human or lead mid-run questions, sub-agent routing - [x] Connect any OpenAI-compatible endpoint by URL - named custom provider entries
- [x] /connect provider rework - name/URL connect, preset defaults, re-enter key
- [x] Model refs unfold - provider/model refs resolve, gated on approved models
- [x] models.json - approved-model history (provider, model, timestamps), no keys
- [x] Provider registry - providers.json, one API key per provider, engine resolves from it
- [x] Sequential team runs -
run_teamstage loop, per-run briefs, team memory,/team run - [x]
personarenamed toagent type- types.py/TypeRegistry, /type, --type, no aliases - [x] Bundled web plugin renamed - replio-core-websearch -> replio-core-web
- [x] GitHub Pages website - mkdocs site + Actions workflow, docs stay in main
- [x] OpenCode Zen + Go providers - opencode.ai endpoints, model-ref strip, URL detection
- [x] Tool-use evaluation harness - replio eval + fixture plugin, metrics
- [x] Project instructions file - per-worktree
AGENTS.mdin system prompt, capped - [x]
run_commandallowlist -tool_permission.bash_allow, heredocs/multi-line rejected - [x]
code_test/code_lint/code_formatwrappers - dev.*_cmd, bundled replio-core-dev plugin - [x]
gittool - read-only git + gatedgit_commit, bundled replio-core-git plugin - [x]
file_edittool - search-and-replace with diff preview, bundled replio-core-edit plugin - [x] Default tool-result cap - tool_max_result_chars default 100k, list_dir entry cap
- [x] Default tool names namespaced -
web_fetchmergesopen+fetch_page, old names as aliases - [x] Tool schema advertises canonical names only - aliases resolve at call time (20 -> 13 defs)
- [x] Agent-loop cancellation + tool-dialect hardening - Ctrl-C cancels the turn, unknown-tool hints
- [x] Skills registry -
SkillRegistry, local/global dirs + plugins, typeskillsin prompts - [x] Teams registry -
Team/TeamRegistry, 4-layerteams.jsonmerge,/teamread/edit - [x] Plugin contribution hooks - register_roles/teams/skills hooks,
RoleRegistry.reload() - [x] Fleet orchestration -
replio fleetsupervisor CLI: ports, health, restart, config gen - [x] Per-run job sessions -
job_<ts>_<name>files, unifiedses_/sub_/job_naming - [x] Job run memory - rolling
.memory.mdsummary injected into each run, seeded - [x] Job task definition as linked Markdown -
--filetemplated task files,replio jobs edit - [x] Job context management - per-run
job_*files,ses_/sub_naming, 100-run cap - [x] Job per-run approval -
--require-approvalparks a run inwaiting_approval - [x] Jobs diagnostics -
replio jobs status(fired count, last error, uptime) - [x] Watchable job runs -
replio jobs run --verbosestreams the live turn - [x] Richer job runs -
--type/--system-prompt, recurring-job prompt, tool carve - [x] Scheduled / durable jobs -
replio jobs: registry, daemon, cron, retries, approval gate - [x] Per-agent permission profiles - role tool_permission drives sub-agent policy
- [x] In-process sub-engine - Engine.run_subagent: role overrides, NullUI, sub_ session
- [x]
delegatetool - core, per-role permission resolver, delegate_echo result display - [x] Roles registry - global+local roles.json merge, /roles command, docs/roles.md
- [x] Soft tool results surfaced as dimmed info lines -
(empty file),(no matches)etc - [x] Session audit trail - permission decisions (allow/ask/deny -> granted/declined/denied)
- [x] Stable message ids -
msg_<hex>auto-assigned to every session message - [x] Global model registry - /connect appends models + keys, /model list/--online, picker reuse
- [x] REPL
/config --global/--localscope flags + apply() in-memory overrides - [x] Scoped config writes - local saves overrides only, replio config CLI
- [x] Turn recovery - auto-continue on truncation, reasoning-only not flagged empty
- [x] Thinking captured from
reasoningdeltas (ollama.com), documented - [x] Confirm prompt
?glyph starts at the beginning of the line, aligned with activity glyphs - [x] Word-level streaming buffering - buffer REPL output to word boundaries, word_streaming config
- [x] Multi-line input - detect
"""/'''blocks, one composed message, framing stripped - [x] Config validation (test connection on change) - /connect probe, /provider warn
- [x] Session export to Markdown - /session export and replio export CLI render logs to Markdown
- [x] Plan/Build modes - plan (edit+bash denied) vs build, /mode cmd + --mode flag, per-message mode
- [x] Thinking/reasoning toggle - show_thinking (display) + reasoning (request), /thinking cmd
- [x] Thinking spinner - animated spinner while thinking when show_thinking false, cleared \r\033[K
- [x] Activity lines + tool status ephemeral - never persisted to session files
- [x] MCP via bundled replio-core-mcp plugin - import/expose tools over MCP, dual-era
- [x] Alias layer - tool/param aliases (read/view, ls, bash/exec, cursor>offset) + open tool
- [x] Glyph activity lines - dimmed
status, glyph_lines config - [x] write_file reports resolved abs path (Created|Overwritten|Appended)
- [x] Tab completion restored (libedit binding) + extended to paths, tool names, and subcommands
- [x] README terminal screenshot as SVG
- [x] deploy/ fleet templates - Dockerfile, compose, systemd + launchd units
- [x] docs/fleet.md - single-purpose scoped-agent fleet pattern
- [x] README fleet positioning - tagline, Features, fleet section, roadmap
- [x] Per-turn stats on own line - footer newline when output lacked one
- [x] One-shot retry for empty/truncated streams before surfacing error
- [x] write_file status preview ends with dimmed summary (path, lines, chars, action)
- [x] Thinking announce - + Thinking (or + Thought 12.3s hidden). Headless mirrors
- [x] Human-readable tool status - [tool: arg] oneliner + dimmed detail lines
- [x] Built-in web + machine features moved to bundled plugins
- [x] Discovery precedence bundled < global < local, bundled can't update/uninstall
- [x]
register_servicesentry hook - powers theweb_search: truesearch-then-answer mode - [x] plugins config list replaces plugins.enabled/deny, empty = all
- [x]
tests/test_bundled_plugins.py(10 tests) - [x] Plugin system - directory-based external plugins for tools, providers, and commands
- [x] PluginManager - discovery, manifest compat ranges, single entry import
- [x] register_tools hook - tools inherit policy, /tool, /help, logging
- [x] register_providers hook merged into provider registry
- [x]
register_commands(commands)hook, registered after builtins at engine init - [x] Plugin manifest + docs -
docs/plugins.md(schema, compatibility contract, security) - [x] Optional per-plugin deps - requires, lazy import, --deps pip-installs
- [x] Activation via config -
pluginslist, enable/disable/install/uninstall maintain it - [x] /plugins - list/detail/enable/disable/install/update/uninstall
- [x]
replio pluginsCLI - headless list/install/update/uninstall - [x]
tests/test_plugins.py(30 tests) - [x] PyPI proper configuration and documentation
- [x] Clear screen on REPL start (
clear_screenconfig, defaulttrue) - [x] /config structured values - JSON parse, -a/-r, reload
- [x] /compact summarization - append-only, compact_from boundary
- [x]
/session loadcompaction offer prints the summary, load records acommandmessage - [x]
/session preview <name>- read-only structural preview without switching sessions - [x] Session-name tab completion for
/session load/delete - [x] Context-size display - dimmed (Ns, N tokens) after each response
- [x] noise_tools config - noise results replaced by marker in sessions
- [x] max_tokens optional (0=unset), hitting cap warns + logs error
- [x] Provider payload from log - command filtered, dangling tool msgs skipped
- [x] replio run - one-shot CLI (JSON/text) reusing loop + sessions
- [x] Flags:
--prompt,--provider,--model,--output=json,--verbose - [x]
--session-id- address persistent sessions from headless mode - [x] Headless logging (
--verbose) instead of visual activity lines - [x]
replio serve- stdlibhttp.serverHTTP JSON API (POST /chat) over the same agent loop - [x] Tests for headless entry points (mock provider, no network)
- [x] Hardening: agent-loop turn failures always visible in the session log
- [x] Persist streamed content when the SSE stream ends without a
doneevent - [x] Record errors entry for silent failures (EOF, empty/thinking-only done)
- [x] Non-normal
finish_reason(length) logs anerrorsentry + prints a warning - [x] Catch unexpected exceptions escaping the agent loop, log them, keep the REPL alive
- [x] Tests: token-stream-then-EOF, empty
done, streamed exception - [x] Sessions are complete logs - every message, tool call + result, reasoning, error persisted
- [x] Append-only - compaction and load never remove or rewrite entries
- [x]
Session.to_dict()keepsrole: tool(full results,noise_toolsreplaced by a marker) - [x] Per-round
thinkingmetadata on assistant messages, excluded fromcontent - [x] Session-level
errorsarray - [x]
created_at/updated_atmetadata (bumped on every message) - [x]
tool_analysis(defaultfalse) - model-generated one-line analysis on each tool message - [x]
session_tool_max_chars(default 0 = unlimited) - caps persisted tool-result content - [x] No
index.json- session names carry timestamp + first-message slug - [x] Flat message list retained (maps 1:1 to provider context)
- [x] Unified streaming agent loop
- [x] Single SSE stream, detect
tool_callsvs content from the first delta - [x] Eliminates double-call cost when no tools are used
- [x] Unified dispatch
- [x] Slash commands call the same
ToolRegistryas the model - [x] Generic query refinement via tool metadata and collapsed
_handle_messagebranches - [x] Machine tools (
tools/machine.py) -read_file,list_dir,write_file,run_command - [x]
read_file- numbered lines,offset/limit, truncation, binary/permission errors - [x]
list_dir- sorted entries, trailing/for subdirs, file sizes - [x]
write_file- parent-dir creation,w/amodes - [x]
run_command- subprocess exec with timeout, stdout/stderr capture, exit code, 8k cap - [x]
category/permission/path_arg/key_argregistration metadata - [x] Discovery tools - glob (recursive, noise-dir skip) + grep (regex, file:line)
- [x]
read_filealways reports total line count in its header - [x] Tool permission model (
tools/policy.py+ config) - [x]
tools.allow/tools.denyname-level policies (deny + allow-whitelist take precedence) - [x] tool_permission category actions - allow/ask/deny per category
- [x] Path-scoped confirm - outside-worktree read/write/list escalate to
ask - [x]
bash: askdefault - everyrun_commandprompts y/N - [x] Confirm prompts - denied tools filtered, cancelled calls feed [cancelled]
- [x]
/toolcommand routes through the same policy - [x] OpenAI / Groq / Anthropic providers
- [x] Provider auto-detection from base_url
- [x] Base provider (OpenAI-compatible interface)
- [x] Ollama cloud provider
- [x] Session manager (JSON CRUD)
- [x]
/session newand/session loadactually switch the active session - [x] Session save after file rename - JSON
namefield matches filename - [x] Auto-session naming with first user message as context hint
- [x] Tool results and assistant
tool_callspersisted in session files - [x] DuckDuckGo Lite search via
html.parser - [x] Terminal + AI context formatting
- [x]
/search <query>and/web <query>commands - [x] Auto-search mode (
web_search: trueconfig) - [x]
fetch_page-_TextExtractor(HTMLParser) for clean text extraction - [x] Search term extraction -
query_refineconfig auto-refines short queries - [x]
tools/registry.py- decorator-based tool registration - [x]
tools/builtins.py-web_searchandfetch_pagetools - [x]
BaseProvider.chat_nonstreaming()- non-streaming tool decision round - [x] Two-phase chat: non-streaming tool decision > stream final content
- [x]
_show_tool_status()- dimmed status during tool execution - [x]
/searchcommand integration with tool calling - [x] Config default:
tool_calling: true - [x]
tool_status_visibleconfig flag (defaulttrue) - [x] Tool calling integration verified end-to-end
- [x] Thinking/reasoning token detection and dimmed display
- [x] SSE streaming survives multi-byte UTF-8 split across read chunks (byte-buffered line decoding)
- [x] Markdown-aware streaming (disabled by default via
markdown_streaming) - [x] Error handling improvements (network timeout, auth errors)
- [x] Edge cases: streaming when
tool_calling=truebut no tools used - [x] Tool-call messages lost on exception -
try/finallypersists session - [x] Final assistant response missing when streaming returns empty - non-streaming fallback
- [x] Project scaffolding (pyproject.toml, venv, dir structure)
- [x] Config module (global + local JSON merge)
- [x] HTTP SSE streaming utility (urllib)
- [x] REPL loop with readline history + tab completion
- [x] Command registry + built-in slash commands
- [x] Streaming token display
- [x] Documentation (README, AGENTS.md, TODO.md, CHANGELOG.md)