Skip to content

Testing

Tests live in tests/ and use the stdlib unittest framework, with no external runner, no network, and no API key. Mock tests patch provider responses so the agent loop, engine, CLI, and server are exercised without a real model.

Running tests

Run all tests:

python -m unittest discover tests

The canonical setup installs the package (pip install -e .), which is also what CI does. On a source checkout without installing, use PYTHONPATH=$PWD/src (absolute, because the detached-fleet daemon changes directory, so a relative src would not resolve). PluginManager then falls back to the repo-root plugins/ directory for the bundled plugins.

Run a single file:

python -m unittest tests.test_tool_calling

Run tests before committing changes to verify core logic isn't broken.

Test coverage

File Covers
test_agent_loop.py Agent-loop behavior: single round trip, thinking persistence, graceful error bail, empty/truncated-stream multi-attempt retry, recovery hint after failed tool-call rounds, truncation error messages (configured cap vs provider default), auto-continue on truncation (stitch + cap + continue instruction), reasoning-only turn not flagged empty, empty-done retried, KeyboardInterrupt cancels the turn (mid-stream and after a tool call) with partial output persisted, unknown-tool results list available tools
test_ask.py ask tool (core): schema/registration (category ask, permission ask, key_arg question, tool_permission.ask default allow), human answer / empty-cancelled / context+options rendering, no-channel error (headless), target='lead' lead-model consultation (question + delegated task in the prompt) + silent-lead fallback to human + root-fallback, sub-engine _lead/_ask_ui inheritance, full-loop ask -> answer tool part -> continue
test_asks.py Pending-ask inbox: AskStore add/find/list/status-filter/answer round-trip + persistence + incrementing ids + unknown/empty answer, inject_answer appends the user part to the origin session (+ missing-session skip), parking (unattended root/sub-agent human ask -> [parked] + record with no stdin, target='lead' lead-less root parks, attended still prompts, permission-ask human route parks with kind=permission while auto route still grants once to the lead), resume-in-context after reload, /asks command (list/show/answer/usage/unknown-id, answer reloads a current origin session), serve GET /asks + POST /asks/<id>/answer (200/404/400)
test_autonomous_supervisor.py End-to-end governance loop (deterministic, mock provider): a leader-typed supervisor job runs a team, a stage's ask parks under unattended mode, the run verifies and the report payload carries the parked ask. Answering injects into the origin session and a resumed turn sees it. The grant_permission ceiling lets the leader delegate bash it denies itself
test_bundled_plugins.py Bundled plugin discovery, tool registration, search/mcp_server/report service, bundled update/uninstall blocking, default-config plugin set, provider registration via the hook
test_cli.py polyglav run: JSON/text output, session-id persistence, exit codes, one-shot overrides applied but never persisted, _engine_from_args approval wiring (explicit --model auto-approves, --approve-model grants, default does not). polyglav export: default/custom/stdout targets, unknown session. polyglav models: bare configured listing (grouped, active marker, key marker, empty) and list [provider] probe (default/<provider>, error/empty), main dispatch for both. polyglav plugins: list/install/uninstall, enable/disable plugins-list toggle + unknown name, main dispatch
test_commands.py Slash-command registration and /help output (aliases, subcommands, tools listed under /tool, mode-filtered listings), /connect provider flow (interactive picker list/number/URL/bad number/stored-custom reconnect, named connect preset defaults + probe args, stored-key keep/re-enter, unknown-name error, URL known-host + plugin default-URL match + custom name derivation + [name] override + bare hostname, probe-before-commit decline/accept, connect_check off, config model untouched), /model show/switch-touch (provider/model ref unfold + approval prompt), /config scope flags (--global/--local, api_key as a normal key incl. global writes, -a/-r scope), /models configured listing and list probe (key marker sourced from providers.json, provider filter, error/empty, configured-model note), /provider warn
test_completion.py Readline tab completion: commands, session names, plugin names, tool names
test_config.py Config scopes: local-only saves, --global writes, apply() in-memory overrides (never written), unset fallback/origin, global>local merge, empty-local-does-not-shadow-global, polyglav config CLI (get/set/unset/reload, JSON values, show-origin). api_key is an ordinary key, with no global forcing, 0600, or migration
test_delegate.py delegate tool: role allow default (no prompt) / ask confirm grant-decline / unknown-role deny, delegate_echo on/off display + sub footer, /tool delegate single print, empty-content log-summary fallback, sub-agent session persistence + resolver actions
test_delegation_permissions.py Permission ceiling + grants: clamp_action/resolve_permissions/resolve_grant_ceiling, ToolPolicy one-shot/always grants + consume + category/tool matching, no-escalation (a role carve cannot widen above the caller's grant_permission), supervisor grants a category it cannot use itself, one-shot grants do not propagate to children, ask(kind='permission') routing (auto/human/deny, ceiling denial, lead grants once only, operator grants once/always), granted tool appears then is consumed from the schema, audit entries
test_engine.py Engine.chat turn result, thinking/content separation, load-or-create sessions, auto ses_<ts>_<id> naming (stable across turns), plan-mode schema filtering, instruction injection, per-turn mode, glyph param suffix gating, ! error-line rendering and show_errors gating, soft-result note-line rendering and show_notes gating, check_connection/list_models probe resolution and overrides without state mutation, _reinit_provider provider-registry API key resolution (no config fallback, registry custom base_url fallback when config empty), model-ref unfold + approval gate (unfolded provider/base_url/model, headless deny, approve_models grant, chat short-circuit, subrole-model unfold + gate, team-run pre-check deny)
test_eval.py Eval harness: fixture model + loading, declarative verifier (exact/must_include/avoid/max_calls/min_calls/args), metric computation (accuracy, redundant, errors, tokens), fixture discovery + precedence (plugin/global/local), cwd isolation and restore, suite aggregation
test_fleet.py Fleet supervisor: port allocation (preferred/bind-probe fallback/in-use skip/exhaustion), /health probe ok + failure, manifest/state round-trip + corrupt tolerance, spawn > health > crash > restart > down via sys.executable -c mock servers, max_restarts give-up, unhealthy-threshold restart, disabled gate, log files, env seams (POLYGLAV_FLEET_PORT), polyglav fleet CLI (init/add/remove/status/restart/config incl. role inline + unknown-role error), detached-daemon end-to-end (up --detach, status, down)
test_focus.py FocusManager: root start, run-keyed attach to a run's own session and role (agent_<role> removed), re-enter identity, back/reset stack walking, root-run reset, unknown-run error, engines listing, run_engine resume of a saved run, focus_session by session name. REPL routing to the active engine for turns and slash commands, role prompt (prompt_role) gating. /focus command: show, #id/#run/#<session_id>/session:<name>/root attach, back/child/parent/sibling/next/prev, role no longer a target, unknown target, absent manager
test_focus_on_delegate.py focus_on_delegate config: default off, off leaves focus alone, on records a pending focus on the child run, ask prompts and focuses on yes / leaves focus on no, unattended ask skips the prompt, an invalid mode falls back to off, a sub-agent without a focus manager is a no-op. The pending focus does not stop the agent loop (the caller's answer still streams) and ChatLoop applies it after the turn
test_handoff.py handoff tool (core): registration + handoff permission, run target sets the pending handoff and pauses the run, done finishes it, unknown target errors without pausing, parent from a child run resolves the root run, child/sibling/#id/#run/session name/root resolution, sub-agent without a focus manager errors. Agent loop stops after the handoff batch (no second provider round), ChatLoop applies the pending focus by run id and resets on the root run
test_history.py /history command: default last-10 / n / all listings by absolute turn index, header with total turns and role, line fields (status, duration, tool count, prompt), command-only turns, prompt first-line clip, --thoughts excerpt vs --thoughts all, --run target resolution (session name, bare name, live #run, saved-run #run, live #id, saved #id, live role run, saved session fallback) without switching focus or current session, unknown target and usage errors, empty-session and no-focus-manager cases
test_http.py SSE streaming: data parsing, done marker, multi-byte split across chunks, HTTP errors, POST-preserving redirects (loopback server)
test_jobs.py Job model + registry (round-trip incl. require_approval/task_file/approve_model/created_at, runnable + ready-to-run gates, corrupt file tolerance), cron parser (steps/ranges/lists/dom/dow/the restrictive day rule, leap day), next_run/compute_next_run/parse_dt, scheduler run/tick under a mocked engine (verified/failed, retries with backoff, per-attempt history, unknown role, one-shot at, approval gates, per-run require_approval park/re-arm, run memory write + injection, per-run session naming + collision dedupe + --session stability, run content capture), _build_engine role-skill injection (present/missing), task file template/linkage/missing-file failure, status/list/show rendering (incl. the last run: summary line), polyglav jobs CLI (add/approve/list, --file, edit, status output, stop, auto-approval, bad cron, duplicates, run exit codes + content printing), report-back dispatch (verified/failed/build-failure events, payload fields, injected service, no-service no-op, service-exception tolerance)
test_memory.py Shared memory module: path under .polyglav/memory/<scope>/, read/write round-trip, missing-empty, legacy job/team fallback reads, memory_enabled global + per-scope, cap_memory truncation. Engine role memory: injection into the sub-agent system prompt, scope-disabled skip, write after run_subagent, disabled skip, Engine.memorize + /memorize command, team memory on the shared path
test_models.py ModelRegistry (approved-model history in global models.json): path under GLOBAL_DIR, put/find by (provider, model), dedupe + last_used, distinct models keep separate entries, touch, remove, grouped, reload, corrupt-file tolerance, old per-model-key shape dropped without migration, GLOBAL_DIR default
test_modes.py Mode resolution and policy merging: built-ins (build/plan), custom modes, unknown fallback, instruction composition
test_ollama_provider.py Streaming provider: fragmented tool-call reassembly, thinking events (reasoning_content and reasoning keys), payload construction. Lives in the polyglav-core-ollama plugin suite (see the bundled-plugin suites note below)
test_plugins.py Plugin manager: manifest compat ranges, discovery precedence, registration hooks (tools/providers/commands/services/roles/teams/skills + hook-failure status), _bundled_dir fallback (source layout + forced import failure), install/update/uninstall, polyglav plugins test (+ load_plugin_test_suite)
test_print.py /print command: whole-turn reprint with the full metadata block (empty fields as -, including a command turn), part-only <n>.<m>, out-of-range part, unknown turn, print_max_chars cap vs --full, --run targets (session name, bare name, live role run, saved session fallback, live #run, live #id) without switching focus, unknown target and usage/invalid-spec errors
test_provider_session.py Provider session binding: BaseProvider/OpenAICompatibleProvider store session_id (default empty), an owned provider is rebound to the run's session, a shared provider keeps its binding, a provider without session_id is skipped, and _current_session_id handles a missing session
test_providers.py Core provider substrate: OpenAICompatibleProvider defaults/headers/endpoints, POST-preserving redirects, check_connection probe (success/empty/model note/HTTP/network), list_models silent-on-error, base reasoning payload, host-pattern detect_provider (known hosts, longest-pattern precedence, empty/unknown fallback), core PROVIDERS registry (generic fallback only) and merged_providers (bundled providers included). Vendor provider defaults/endpoints/reasoning/headers live in the per-plugin suites (polyglav-core-ollama, -openai, -groq, -anthropic, -opencode)
test_providers_registry.py ProviderRegistry (global providers.json): path under GLOBAL_DIR, put/find/dedupe per provider, empty key keeps existing, base_url stored only when given, key/base_url lookup, touch, 0600 when keyed, reload, corrupt-file tolerance, remove, GLOBAL_DIR default. resolve_model_ref: known/unknown provider, bare model, no-default provider, empty parts, plugin providers
test_repl_input.py REPL input: multi-line """/''' block detection, framing strip (pure, lead-in, indentation preserved), EOF exit during an open block, slash commands single-line
test_run_buffer.py Per-run output buffers: Run.append_buffer/buffer_text/clear_buffer (leading blank skip, cap trims oldest), BufferUI plain-text line formatting (tokens split on newlines, partial flush on footer/info, tool/activity/error/result lines, max_lines sets the run cap), confirm denies and ask returns None, /focus log [n] prints the focused run's buffer (tail limit, empty)
test_run_resume.py Run resume and context control: fresh sub_ sessions by default, resume by run id / session name / session id reuses the session and reactivates the run, unknown target raises, context='new' ignores resume, context='compact' summarizes first, team resume seeds the first stage, delegate/team tools forward resume/context, team field round-trip without warm_sessions/session_key
test_runs.py RunRegistry: sequential ids, parent/children tree and creation-order call log, finish status set-once, task recording, unknown-run tolerance. Engine runs: root run registration, bind_root_agent updates the run role, sub-engine shares the registry and links the parent, run_subagent finishes the child run, load_or_create_session updates the run session and session id
test_server.py polyglav serve HTTP API: /chat, /sessions, /health, /version
test_session_id.py Session id: session_id_hash determinism/length/base36, coded_session_name/session_id_from_name parsing (ses_/job_/sub_, explicit names ignored), Session.session_id stamping and round-trip, missing id loads empty, auto ses_<ts>_<id> naming and uniqueness, explicit names have no id, SessionManager.find_by_session_id hit/normalize/miss, RunRegistry.start(session_id=) + find_by_session_id, engine run/session id mirroring, load_or_create_session sync, id stable across turns, saved-session lookup
test_session_log.py Session model: turn/part storage and the explicit turn/part API, turn index/status/metadata, append-only serialization, tool_max_chars/noise_tools transforms on tool-part output, errors/permissions round-trip, session_id field round-trip, files without the current format do not load (and stay listed), compaction boundary and provider reconstruction, agent-loop turn persistence (thinking, tool parts, analysis)
test_session_render.py Session Markdown export: turn/part renderer (user, assistant, thinking, tool call arguments + result + analysis, command + compaction, system), error section, /sessions export dispatch and file/stdout targets. Also the terminal turn_summary (line fields, tool count, prompt clip, command turn, dim thinking excerpt vs full all) and render_turn (metadata block incl. empty-as--, all part types, tool error + analysis, compaction, part selection, cap vs --full)
test_skills.py SkillRegistry: local/global dir scans, plugin/global/local merge and precedence, origins, put/remove round-trip, reload (disk re-read + plugin-manager re-apply), skills_section, /skills command (list/show/new override/remove, plugin remove rejected)
test_subagent.py In-process sub-engine: provider/plugin/worktree inheritance, role prompt/mode/tool_permission application, role-skill system-prompt injection (present/missing/empty/no-skills), model override, BufferUI, unknown role, full run_subagent flow + persisted sub_* session with parent_id, ask-gated tool cancellation, parent sub_sessions linkage
test_team_run.py Engine.run_team: brief builder (task + prior results + handoff + memory + task hint, prior-result truncation), sequential stage execution with per-stage sub_* sessions + parent linkage + exact brief persistence, stage mode override + caller-mode inheritance, stop-on-failure, unknown-stage-role stop, zero stages, rolling team-memory write (summarized + prior-seeded + fallback), /teams run command output (stages + final result, unknown team, usage)
test_team_tool.py team tool: schema/registration (category delegate, loop), pipeline run + final result + sub-sessions, unknown team / no stages / unknown stage role errors, resolver allow/deny/ask per stage role, clamped-stage note, run_team cycle + depth guards, _team_depth/_team_stack propagation, full agent-loop team call
test_teams.py TeamRegistry: bundled/plugin/global/local merge and precedence, origins, stage round-trip (dict + short-string forms), put/remove, reload (disk re-read + plugin-manager re-apply), /teams command (list/show/new override/remove, list <tag> filter, bundled remove rejected)
test_tool_calling.py Tool-calling flow: single and multiple calls, unknown tools, query refinement
test_tool_policy.py ToolPolicy: allow/ask/deny, worktree escalation, deny/allowlist precedence, per-invocation resolver (refines non-deny base, skipped without args, cannot override deny list)
test_tool_registry.py Tool registration metadata, schema, refine flags, note-result predicates, _config pass-through, activity params strings, fs tool glyphs (* List / * Grep), permission_fn storage + resolver_for
test_turns.py Turn model (src/polyglav/sessions/turns.py): turn/part shape and builders for the six part types, finish_turn/finish_tool, tool co-location, provider reconstruction (a turn's text plus consecutive tool parts grouped into one assistant tool_calls message followed by its tool results, thinking attached to the answer or tool round, reasoning-only turns, command parts dropped, system passthrough, deterministic tool-call ids, JSON-serialized arguments), and compaction/provider_messages (summary plus turn-index boundary)
test_unattended.py Unattended mode: config defaults, root _ask_ui dropped (config, --unattended, /unattended) and restored by set_unattended, sub-engine inherits unattended + no _ask_ui, confirm auto-deny (no ui.confirm call), root human ask non-blocking / sub-agent human ask routes to the lead (no input), full-loop run_command confirm auto-deny, model approval not prompted unless approve_models, /unattended command status + toggle, confirm_timeout deny-on-timeout / answer-when-ready / zero-timeout passthrough for confirm and ask
test_roles.py RoleRegistry: bundled/plugin/global/local merge and precedence, origins (bundled/plugin/local/global origin), tags roundtrip + merge, put/remove/reload (disk re-read + plugin-manager re-apply), /roles command (list/show/new override/remove, list <tag> filter, bundled remove rejected)
test_ui.py UI sinks: glyph activity lines, status oneliner fallback, headless verbose rendering, ! tool-error lines, word-streaming buffering (boundary flush, tail flush, off-mode immediate writes, markdown across boundaries, flush before status/confirm), thinking spinner, delegation status spinner (status_begin/status_end label, status_spinner off, Null no-op, run_subagent wraps it), confirm ? glyph at line start, confirm re-raises KeyboardInterrupt / returns False on EOF

tests/helpers.py provides make_chat(config_data), a ChatLoop with a mocked provider, used by most tests to drive the engine without a model.

Plugin test suites

Each bundled plugin ships its unit tests in its own directory (plugins/<name>/tests/), covering the plugin's tools and helpers without touching the core. The core suite discovers them through tests/test_plugin_suites.py (registered via load_tests), so python -m unittest discover tests runs everything. A single plugin's suite runs standalone with python plugins/<name>/tests/<file>.py, and headless via polyglav plugins test <name> (or polyglav plugins test for every plugin with a suite).

Live testing

Manual live tests against a real provider API are done ad-hoc, not automated. Use a local model (Ollama) or a disposable API key, and verify a turn end-to-end: streaming output, a tool call round trip, and session persistence.

The overnight autonomous run is verified live with a real model: polyglav jobs add-supervisor night --interval 86400 --task "...", run it (polyglav jobs run night or the daemon), watch /asks for a parked question, answer it, and confirm the run completes and (with report.webhook set) the report arrives. See the Supervisor overnight run section in docs/jobs.md.