# AbstractRuntime — full documentation

> Durable workflow runtime for AbstractFramework: workflows run as a persisted
> state machine (interrupt, checkpoint, resume) with explicit waits and an
> append-only execution ledger. The AbstractCore integration adds LLM calls,
> tools, live token streaming, model residency and media generation.

This file concatenates the documentation pages listed below, in full. It
contains no source code; the repository's code and tests are the source of
truth. Relative links inside each page are relative to that page's own path.
See llms.txt for the linked index.

## Document index

- README.md
- docs/getting-started.md
- docs/architecture.md
- docs/api.md
- docs/faq.md
- docs/troubleshooting.md
- docs/integrations/abstractcore.md
- docs/proposal.md
- docs/limits.md
- docs/artifacts.md
- docs/automations.md
- docs/tool-approval.md
- docs/tools-comms.md
- docs/entity-runtime.md
- docs/mcp-worker.md
- docs/evidence.md
- docs/snapshots.md
- docs/provenance.md
- docs/workflow-bundles.md
- docs/manual_testing.md
- docs/adr/README.md
- examples/README.md
- SECURITY.md
- CONTRIBUTING.md
- CODE_OF_CONDUCT.md
- ACKNOWLEDGMENTS.md
- ROADMAP.md
- docs/README.md


==============================================================================
# FILE: README.md
==============================================================================

# AbstractRuntime

**AbstractRuntime** is a durable workflow runtime (interrupt → checkpoint → resume) with an append-only execution ledger.

It is designed for long-running workflows that must survive restarts and explicitly model blocking (human input, timers, external events, subworkflows) without keeping Python stacks alive.

**Version:** 0.7.0 • **Python:** 3.10+

**Status:** pre-1.0 (API may evolve). For production use, pin versions and follow `CHANGELOG.md`.

## AbstractFramework ecosystem

AbstractRuntime is one component of the wider [AbstractFramework](https://github.com/lpalbou/AbstractFramework) ecosystem:
- **AbstractRuntime** (this repo) — durable workflow kernel (`src/abstractruntime/core/*`)
- **AbstractCore** — LLM + tools integration (wired via `src/abstractruntime/integrations/abstractcore/*`)  
  Repo: [lpalbou/abstractcore](https://github.com/lpalbou/abstractcore)

At a high level, hosts define workflow graphs (`WorkflowSpec`) and AbstractRuntime executes them durably. When nodes request LLM/tool work (`EffectType.LLM_CALL`, `EffectType.TOOL_CALLS`), those effects are typically handled via AbstractCore.

```mermaid
flowchart LR
  Host["Host app / orchestrator"] -->|"WorkflowSpec"| RT["AbstractRuntime"]
  RT -->|"LLM_CALL / TOOL_CALLS"| AC["AbstractCore"]
  AC -->|"results / waits"| RT
```

## Install

Remote-light runtime:

```bash
pip install abstractruntime
```

The base install includes AbstractCore 2.18.0 or newer with remote provider,
tool, vision, voice, audio, and music integration, plus the
`abstractruntime-mcp-worker` entry point. It keeps inference remote/light by
default: local engines such as MLX, vLLM, HuggingFace/Torch, Diffusers, and
sentence-transformer embeddings are not selected unless you choose a hardware
profile or another package-specific local extra.

VisualFlow document nodes use permissive dependencies in Runtime's base install:
`Read PDF` extracts text and metadata with `pypdf`, `Write PDF` renders text or
Markdown-style report content to real PDF bytes with `reportlab`, and
`Write DOCX` renders Markdown-style report content to Word-compatible `.docx`
bytes with the Python standard library.

Native Python hardware profiles add local inferencer stacks:

```bash
pip install "abstractruntime[apple]"
pip install "abstractruntime[gpu]"
```

`abstractruntime[apple]` delegates to AbstractCore's native Apple aggregate;
`abstractruntime[gpu]` delegates to AbstractCore's GPU aggregate.

## Quick start (pause + resume)

```python
from abstractruntime import Effect, EffectType, Runtime, StepPlan, WorkflowSpec
from abstractruntime.storage import InMemoryLedgerStore, InMemoryRunStore


def ask(run, ctx):
    return StepPlan(
        node_id="ask",
        effect=Effect(
            type=EffectType.ASK_USER,
            payload={"prompt": "Continue?"},
            result_key="user_answer",
        ),
        next_node="done",
    )


def done(run, ctx):
    answer = run.vars.get("user_answer") or {}
    text = answer.get("text") if isinstance(answer, dict) else None
    return StepPlan(node_id="done", complete_output={"answer": text})


wf = WorkflowSpec(workflow_id="demo", entry_node="ask", nodes={"ask": ask, "done": done})
rt = Runtime(run_store=InMemoryRunStore(), ledger_store=InMemoryLedgerStore())

run_id = rt.start(workflow=wf)
state = rt.tick(workflow=wf, run_id=run_id)
assert state.status.value == "waiting"

state = rt.resume(
    workflow=wf,
    run_id=run_id,
    wait_key=state.waiting.wait_key,
    payload={"text": "yes"},
)
assert state.status.value == "completed"
```

## What’s included (v0.7.0)

Kernel (import-light):
- workflow graphs: `WorkflowSpec` (`src/abstractruntime/core/spec.py`)
- durable execution: `Runtime.start/tick/resume` (`src/abstractruntime/core/runtime.py`)
- durable waits/events: `WAIT_EVENT`, `WAIT_UNTIL`, `ASK_USER`, `EMIT_EVENT`
- append-only ledger (`StepRecord`) + node traces (`vars["_runtime"]["node_traces"]`)
- retries/idempotency hooks: `src/abstractruntime/core/policy.py`
- runtime-aware limits (`_limits`) with a default iteration budget of 20 (`docs/limits.md`)
- Stop that reaches the running effect: `Runtime.cancel_run(...)` signals the in-flight model or tool call, which ends as a `cancelled` ledger record (never retried); a model unload stops the calls using that model first (`docs/api.md`)
- host pause at step boundaries: `Runtime.tick(..., step_gate=...)`
- run-tree tool ceiling: an explicit `allowed_tools` list can only narrow across child runs, and approval policy never widens it
- automations: run a workflow on a schedule (`schedule@1`, fixed UTC intervals) or on request (`manual@1`) as a durable controller run whose occurrences are deterministic child runs and session turns, with commands, independent or growing context, discussions forked at any occurrence (own workspace, automation workspace mounted read-only), tool approval, typed waits and quiet-by-default notifications (`docs/automations.md`)
- explicit run ids: `Runtime.start(..., run_id=...)` creates a run only if the id is free (`RunStore.create_if_absent`), and `run_mutation_lock(run_id)` serializes a run's writers in a process

Durability + storage:
- stores: in-memory, JSON/JSONL, SQLite (`src/abstractruntime/storage/*`)
- durable command inbox primitives (idempotent, append-only): `CommandStore`, `CommandCursorStore` (`src/abstractruntime/storage/commands.py`, `src/abstractruntime/storage/sqlite.py`)
- artifacts + offloading (store large payloads by reference)
- snapshots/bookmarks (`docs/snapshots.md`)
- tamper-evident hash-chained ledger (`docs/provenance.md`)

Drivers + distribution:
- scheduler: `create_scheduled_runtime()` (`src/abstractruntime/scheduler/*`)
- VisualFlow compiler + WorkflowBundles (`src/abstractruntime/visualflow_compiler/*`, `src/abstractruntime/workflow_bundle/*`)
- VisualFlow multi-entry execution lowering for fan-in routes and per-entry input overrides (`docs/workflow-bundles.md`)
- VisualFlow LLM Call and Agent nodes propagate Core generation params such as
  `thinking` and `speculation` through Runtime effects; `_runtime.speculation`
  sets a run-wide MTP preference that nested workflows and Agent loops inherit. Provider Models nodes can apply Core
  `capability_route` filters so run-time model discovery matches Gateway/Flow
  authoring.
- VisualFlow image/video nodes and Runtime media helpers preserve task-specific
  Core media controls, including `count`/`n`, `seeds`, ordered
  `lora_adapters`, and video `flow_shift`, while keeping provider/model/task
  truth in AbstractCore and AbstractVision.
- `history_bundle` exports replay-safe `resolved_actions` summaries derived
  from Core request/output route resolution, so thin clients can reconstruct
  what capability ran without reverse-engineering prompt prose.
- VisualFlow structured LLM/Agent results preserve `response` as text and expose
  the schema-conformant object on `data`, so Break Object and Switch can consume
  fields without reparsing the response string.
- run history export: `export_run_history_bundle(...)` (`src/abstractruntime/history_bundle.py`)

Runtime-owned integrations:
- AbstractCore (LLM + tools, `MODEL_RESIDENCY` with residency locks and context estimates, public discovery/host/run facades, cached sessions with per-session prompt-cache listing/clearing, host memory snapshots, local-only prompt-cache export/import admin, durable bloc prompt-cache controls, bindings, lifecycle operations, generated image/video/voice/music outputs with progress events, host email helpers, Telegram host wrappers, and tool approval waits): `docs/integrations/abstractcore.md`
- AbstractCore models and engines for hosts: `config_facade` passes through the host profile, local-engine status and installs, the model catalog with fit verdicts, installed models, deletes, host jobs (with who-cancelled attribution) and the embeddable console screens: `docs/integrations/abstractcore.md#models-engines-and-host-jobs-config-facade`
- For outbound comms, use the durable run facade when the send belongs to a run: `get_abstractcore_run_facade(...).send_email(...)` / `send_telegram_message(...)`. If that child run pauses for approval or passthrough execution, resume it through `resume_tool_calls(...)`. Direct host-facade send helpers and the standalone email comms facade remain host-local and nondurable.
- AbstractMemory TripleStore integration for `MEMORY_KG_*` effects. Runtime
  depends on the light AbstractMemory contract; hosts choose storage backends
  such as LanceDB, SQLite, or in-memory stores.
- comms toolset gating (email/WhatsApp/Telegram): `docs/tools-comms.md`
- live token streaming: `Runtime.set_live_delta_sink(sink)` plus `_runtime.stream: true` delivers `llm.delta` / `llm.delta_end` events while an answer is generated, outside the ledger (`docs/integrations/abstractcore.md#live-token-streaming`)
- model switching in a shared process: changing the default model unloads the previous in-process model unless another owner in the process still uses it (`docs/integrations/abstractcore.md#unloading-and-switching-models`)
- workspace-scoped file and shell tools driven by `workspace_*` run vars, including host-only protected folders that are enforced without being shown to the model (`docs/integrations/abstractcore.md#workspace-scoped-tools`)

## Built-in scheduler (zero-config)

```python
from abstractruntime import create_scheduled_runtime

sr = create_scheduled_runtime()
run_id, state = sr.run(my_workflow)

if state.status.value == "waiting":
    state = sr.respond(run_id, {"text": "yes"})

sr.stop()
```

For persistent storage:

```python
from abstractruntime import create_scheduled_runtime, JsonFileRunStore, JsonlLedgerStore

sr = create_scheduled_runtime(
    run_store=JsonFileRunStore("./data"),
    ledger_store=JsonlLedgerStore("./data"),
)
```

## Documentation

| Document | Description |
|----------|-------------|
| [Getting Started](docs/getting-started.md) | Install + first durable workflow |
| [API Reference](docs/api.md) | Public API surface (imports + pointers) |
| [Docs Index](docs/README.md) | Full docs map (guides + reference) |
| [FAQ](docs/faq.md) | Common questions and gotchas |
| [Troubleshooting](docs/troubleshooting.md) | Symptom-oriented setup, runtime, and integration fixes |
| [Architecture](docs/architecture.md) | Component map + diagrams |
| [Overview](docs/proposal.md) | Design goals, core concepts, and scope |
| [AbstractCore Integration](docs/integrations/abstractcore.md) | `LLM_CALL` / `TOOL_CALLS`, host facades, models/engines/host jobs |
| [Automations](docs/automations.md) | Scheduled and on-request workflows: controller, triggers, commands, context, discussions, storage guarantees |
| [Artifacts](docs/artifacts.md) | Artifact identity, descriptors, catalog search and access stats |
| [Tool Approval](docs/tool-approval.md) | Tool risk tiers, run-policy ceiling, per-call refiners |
| [Comms Toolset](docs/tools-comms.md) | Opt-in email/WhatsApp/Telegram tools |
| [Entity Runtime](docs/entity-runtime.md) | Per-entity runtimes, homes, leases, visit waits |
| [Snapshots](docs/snapshots.md) | Named checkpoints for run state |
| [Provenance](docs/provenance.md) | Tamper-evident ledger documentation |
| [Evidence](docs/evidence.md) | Artifact-backed evidence capture for web/command tools |
| [Limits](docs/limits.md) | `_limits` namespace and RuntimeConfig |
| [WorkflowBundles](docs/workflow-bundles.md) | `.flow` bundle format (VisualFlow distribution) |
| [MCP Worker](docs/mcp-worker.md) | `abstractruntime-mcp-worker` CLI |
| [Changelog](CHANGELOG.md) | Release notes |
| [Contributing](CONTRIBUTING.md) | How to build/test and submit changes |
| [Code of Conduct](CODE_OF_CONDUCT.md) | Contributor conduct expectations |
| [Security](SECURITY.md) | Responsible vulnerability reporting |
| [Acknowledgments](ACKNOWLEDGMENTS.md) | Credits |
| [ROADMAP](ROADMAP.md) | Current status and longer-term direction |

## Development

```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
python -m pip install -e ".[test,docs]"
python -m pytest -q
```

See `CONTRIBUTING.md` for contribution guidelines and doc conventions.


==============================================================================
# FILE: docs/getting-started.md
==============================================================================

# Getting started

This guide gets you from install to a durable **pause → resume** workflow quickly.

If you only read one doc after `README.md`, read this one.

## Install

Remote-light runtime:

```bash
pip install abstractruntime
```

This installs AbstractCore 2.16.0 or newer with remote provider, vision,
voice, audio, music, tool, and MCP-worker support. The base install is
remote-light: it can route multimodal workflows to hosted or OpenAI-compatible
endpoints, but it does not select local inferencer stacks such as MLX, vLLM,
HuggingFace/Torch, Diffusers, or sentence-transformer embeddings.

Use hardware profiles only when this Runtime host should execute local engines:

```bash
pip install "abstractruntime[apple]"
pip install "abstractruntime[gpu]"
```

## Mental model (source of truth)

- Workflows are in-memory graphs: `WorkflowSpec` (`src/abstractruntime/core/spec.py`)
- Runs are durable checkpoints: `RunState` (`src/abstractruntime/core/models.py`)
- Nodes return “what to do next”: `StepPlan` (`src/abstractruntime/core/models.py`)
- Side effects are requested, not executed directly: `Effect` / `EffectType` (`src/abstractruntime/core/models.py`)
- Blocking is explicit and durable: `WaitState` (`src/abstractruntime/core/models.py`)
- Every step is append-only in the ledger: `StepRecord` (`src/abstractruntime/core/models.py`)

The execution loop is implemented in `Runtime.start/tick/resume` (`src/abstractruntime/core/runtime.py`).

## Quick start: pause + resume

```python
from abstractruntime import Effect, EffectType, Runtime, StepPlan, WorkflowSpec
from abstractruntime.storage import InMemoryLedgerStore, InMemoryRunStore


def ask(run, ctx):
    return StepPlan(
        node_id="ask",
        effect=Effect(type=EffectType.ASK_USER, payload={"prompt": "Continue?"}, result_key="answer"),
        next_node="done",
    )


def done(run, ctx):
    answer = run.vars.get("answer") or {}
    text = answer.get("text") if isinstance(answer, dict) else None
    return StepPlan(node_id="done", complete_output={"answer": text})


wf = WorkflowSpec(workflow_id="demo", entry_node="ask", nodes={"ask": ask, "done": done})
rt = Runtime(run_store=InMemoryRunStore(), ledger_store=InMemoryLedgerStore())

run_id = rt.start(workflow=wf)
state = rt.tick(workflow=wf, run_id=run_id)
assert state.status.value == "waiting"

state = rt.resume(workflow=wf, run_id=run_id, wait_key=state.waiting.wait_key, payload={"text": "yes"})
assert state.status.value == "completed"
print(state.output)
```

## Recommended: use the built-in scheduler wrapper

For most apps, use `create_scheduled_runtime()` which bundles:
- `Runtime`
- `Scheduler` (in-process polling driver)
- `WorkflowRegistry` (maps `workflow_id` → `WorkflowSpec`)

Implementation: `src/abstractruntime/scheduler/*`.

```python
from datetime import datetime, timedelta, timezone
import time

from abstractruntime import create_scheduled_runtime, Effect, EffectType, StepPlan, WorkflowSpec, RunStatus

def wait(run, ctx):
    until = (datetime.now(timezone.utc) + timedelta(seconds=2)).isoformat()
    return StepPlan(
        node_id="wait",
        effect=Effect(type=EffectType.WAIT_UNTIL, payload={"until": until}),
        next_node="done",
    )

def done(run, ctx):
    return StepPlan(node_id="done", complete_output={"ok": True})

wf = WorkflowSpec(workflow_id="demo_wait_until", entry_node="wait", nodes={"wait": wait, "done": done})

sr = create_scheduled_runtime(poll_interval_s=0.2)  # in-memory stores, scheduler auto-starts
run_id, state = sr.run(wf)
assert state.status == RunStatus.WAITING

time.sleep(3)
state = sr.get_state(run_id)
print(state.status.value, state.output)

sr.stop()
```

## Persist runs + ledgers (survive restarts)

Use file-backed stores:

```python
from abstractruntime import create_scheduled_runtime, JsonFileRunStore, JsonlLedgerStore

sr = create_scheduled_runtime(
    run_store=JsonFileRunStore("./data"),
    ledger_store=JsonlLedgerStore("./data"),
)
```

Notes:
- `JsonFileRunStore` stores `run_<run_id>.json`
- `JsonlLedgerStore` stores `ledger_<run_id>.jsonl`
- Durable state must be JSON-serializable; for large payloads use `ArtifactStore`/offloading (see `architecture.md`)

## Optional: LLM + tools (AbstractCore)

AbstractRuntime’s kernel is dependency-light; LLM/tool execution is wired via the AbstractCore integration:
- docs: `integrations/abstractcore.md`
- code: `src/abstractruntime/integrations/abstractcore/*`

Typical local mode:

```python
from abstractruntime.integrations.abstractcore import create_local_runtime

rt = create_local_runtime(provider="ollama", model="qwen3:4b")
```

## Next reading

- `faq.md` — common questions and gotchas
- `api.md` — public API surface (imports + pointers)
- `architecture.md` — component map + durability invariants (with diagrams)
- `manual_testing.md` — smoke tests and how to run `pytest`
- `../examples/README.md` — runnable scripts
- `integrations/abstractcore.md` — `LLM_CALL` / `TOOL_CALLS`, cached sessions, durable bloc prompt-cache, media inputs, generated media


==============================================================================
# FILE: docs/architecture.md
==============================================================================

# AbstractRuntime — Architecture

> Updated: 2026-09-28
> Version: 0.7.0
> Scope: this describes **what is implemented in this repository**.

AbstractRuntime is a **durable workflow runtime**: it executes workflow graphs as a persisted state machine with explicit waits (user, time, events, jobs, subworkflows). A run can pause for hours/days and resume **without** keeping Python stacks/coroutines alive.

## Ecosystem (AbstractFramework)

AbstractRuntime is the durable execution kernel inside the wider AbstractFramework ecosystem:
- AbstractFramework umbrella: [lpalbou/AbstractFramework](https://github.com/lpalbou/AbstractFramework)
- AbstractCore (LLM + tools): [lpalbou/abstractcore](https://github.com/lpalbou/abstractcore)

The runtime stays dependency-light and delegates LLM/tool execution to integrations (notably AbstractCore): `src/abstractruntime/integrations/abstractcore/*`. Runtime also depends on the light AbstractMemory contract so `MEMORY_KG_*` effects always have the TripleStore model types available; hosts still choose the concrete memory backend.

```mermaid
flowchart LR
  Host["Host app / AbstractFlow / service"] -->|"WorkflowSpec"| RT["AbstractRuntime"]
  RT -->|"LLM_CALL / TOOL_CALLS"| AC["AbstractCore"]
  AC -->|"results / waits"| RT
```

Key invariants (enforced by code, not convention):
- **Durable state is JSON-safe**: `RunState.vars` must remain JSON-serializable (`src/abstractruntime/core/models.py`). Large payloads should be stored as artifacts and referenced (`src/abstractruntime/storage/artifacts.py`, `src/abstractruntime/storage/offloading.py`).
- **Append-only observability**: every step is recorded as a `StepRecord` in a `LedgerStore` (`src/abstractruntime/core/models.py`, `src/abstractruntime/storage/base.py`).
- **Side effects are mediated**: nodes request work via `Effect`/`EffectType`; execution happens via effect handlers (`src/abstractruntime/core/runtime.py`).

## AbstractCore capability boundary

AbstractRuntime's job is persistence and orchestration. AbstractCore owns model/provider capability execution: chat, structured output, cached sessions/prompt cache, media input analysis, image generation, video generation, voice/audio generation, music generation, and transcription.

The boundary is intentionally narrow:
- Workflow nodes request model work with `EffectType.LLM_CALL`; the runtime persists the request/result and delegates execution to the configured AbstractCore client.
- `media` inputs remain JSON-safe in the effect payload. Artifact refs are materialized into temporary provider-ready files for the call, then cleaned up.
- Generated binary outputs are written to `ArtifactStore` and returned as `artifact_id` / `artifact_ref`, keeping `RunState.vars` and ledger records bounded and JSON-safe.
- Remote chat media is sent to AbstractCore Server as provider-ready content arrays, but persisted provider-request metadata redacts data URLs so checkpoints and ledgers do not embed media bytes.
- Provider sessions and prompt-cache objects are not runtime state. Runtime may carry stable cache keys, while AbstractCore clients/servers manage warm caches. Runtime stamps session attribution onto the session-derived cache keys it injects so hosts can list and clear caches per session through the host facade; caller-owned keys stay unstamped, and completing a run does not clear session caches.
- Hosts should use Runtime-owned AbstractCore facades for discovery snapshots, prompt-cache/model-residency control operations, and durable run-scoped media/comms child runs instead of reaching through private runtime attachments or importing Core internals directly.
- Local execution can use richer AbstractCore capability plugins. Remote and hybrid execution map the common media cases to AbstractCore Server endpoints and OpenAI-compatible content arrays, while hybrid keeps tool execution local.
- Gateway and other hosts compose Runtime with the desired memory and local-inference profile. Runtime's base package includes the AbstractMemory contract, AbstractCore remote/tool capability integration, Runtime-owned permissive PDF read/write and standard-library DOCX write support, and the MCP worker entry point, but not backend extras such as LanceDB, Core media document stacks, or local inferencer stacks. Hosts choose storage, embeddings, readiness policy, and whether to add `abstractruntime[apple]` or `abstractruntime[gpu]`.
- Remote and hybrid clients use explicit Core server URLs and auth headers supplied by the host. Runtime does not read Gateway auth environment variables for provider/model/auth decisions or treat Gateway bearer tokens as Core server/provider credentials.

### Host-facing facades

Hosts such as AbstractGateway talk to AbstractCore through Runtime-owned facades, so they never import AbstractCore themselves. Durable model work goes through the runtime (effects, ledger, artifacts); host operations (discovery, residency, prompt caches, configuration, models, engines and host jobs) go through the facades, which call AbstractCore in-process or its server.

```mermaid
flowchart LR
  subgraph Host["Host (AbstractGateway, apps)"]
    HostAPI["HTTP routes / UI"]
  end

  subgraph RT["AbstractRuntime"]
    Runtime["Runtime
start / tick / resume"]
    Handlers["effect handlers
LLM_CALL / TOOL_CALLS / MODEL_RESIDENCY"]
    RunFacade["run_facade
durable media + comms child runs"]
    HostFacade["host_facade / discovery_facade
residency, prompt caches, discovery"]
    ConfigFacade["config_facade
capability defaults, models, engines, host jobs"]
    Stores["RunStore / LedgerStore / ArtifactStore"]
  end

  subgraph AC["AbstractCore"]
    Providers["providers + tools
(local or AbstractCore Server)"]
    CoreConfig["config, model catalog,
engine installer, host job registry"]
  end

  HostAPI -->|"WorkflowSpec, resume, cancel"| Runtime
  HostAPI --> RunFacade
  HostAPI --> HostFacade
  HostAPI --> ConfigFacade
  RunFacade --> Runtime
  Runtime --> Handlers
  Runtime --> Stores
  Handlers -->|"generate / tools"| Providers
  HostFacade --> Providers
  ConfigFacade --> CoreConfig
```

This keeps the runtime usable by `../abstractgateway` and application layers such as `../abstractflow`, `../abstractassistant`, `../abstractobserver`, and `../abstractcode` without embedding provider-specific model logic in the durable kernel.

### Live token streaming (outside the ledger)

Live text travels on a side channel next to the durable path. A run opts in with `_runtime.stream: true`, the host registers one sink with `Runtime.set_live_delta_sink(sink)`, and each `LLM_CALL` hands a per-call emitter (`src/abstractruntime/core/live_deltas.py`) to the AbstractCore client. The emitter batches fragments (about 40 ms), splits reasoning from content, holds back tool-call markup, and always closes the call with one `llm.delta_end` after the durable record is appended. Nothing on the live path is persisted, so replay and the ledger are the same with streaming on or off. See [Live token streaming](integrations/abstractcore.md#live-token-streaming).

```mermaid
flowchart LR
  Provider["AbstractCore provider
(stream=True)"] -->|"chunks"| Client["LLM client
think / harmony split,
tool-markup hold-back"]
  Client -->|"content / reasoning fragments"| Emitter["LiveDeltaEmitter
(per call, ~40 ms batches)"]
  Emitter -->|"llm.delta / llm.delta_end"| Sink["host sink
set_live_delta_sink"]
  Client -->|"final result
(content, usage, raw_response, prompt_cache)"| Handler["LLM_CALL handler"]
  Handler -->|"StepRecord"| Ledger["LedgerStore"]
  Handler -.->|"after the record: delta_end"| Emitter
```

### Model residency in a shared process

In-process models (MLX, HuggingFace, embeddings) belong to the process, not to one runtime. Every multi-local client registers what it still needs (pooled, per-override, locked, being built, default) with AbstractCore's process residency registry. An unload after a default switch or a failed load goes through `eject_unclaimed`, which checks every owner's claims and unloads under one lock, so one service never unloads a model another service, user, entity runtime or the AbstractCore server still uses. A model that is still generating is unloaded when its call ends. See [Unloading and switching models](integrations/abstractcore.md#unloading-and-switching-models).

```mermaid
flowchart TB
  subgraph Process["One host process"]
    ClientA["multi-local client
(service A)"]
    ClientB["multi-local client
(service B / entity runtime)"]
    Server["AbstractCore server
managed runtimes"]
    Registry["process residency registry
claims: pool, override, lock,
building, default"]
    Weights["in-process weights
MLX / HuggingFace / embeddings"]
  end

  ClientA -->|"register claims"| Registry
  ClientB -->|"register claims"| Registry
  Server -->|"register claims"| Registry
  ClientA -->|"default switch:
eject_unclaimed(old model)"| Registry
  Registry -->|"unload only if no owner claims it"| Weights
```

## Component map

```mermaid
flowchart TB
  subgraph Core["core/ (execution kernel)"]
    Models["models.py\nRunState / StepPlan / Effect / WaitState / StepRecord"]
    Runtime["runtime.py\nRuntime.start / tick / resume"]
    Spec["spec.py\nWorkflowSpec"]
    Policy["policy.py\nEffectPolicy + idempotency"]
    Vars["vars.py\nnamespaces + _limits + node_traces"]
    Config["config.py\nRuntimeConfig"]
  end

  subgraph Storage["storage/ (durability)"]
    RunStore["RunStore\nin_memory / json_files / sqlite"]
    LedgerStore["LedgerStore\n(+observable + hash_chain)"]
    Commands["commands.py\nCommandStore / CommandCursorStore"]
    Artifacts["ArtifactStore\nin_memory / file"]
    Offload["Offloading* wrappers\nstore large values by ref"]
    Snapshots["SnapshotStore\nin_memory / json"]
  end

  subgraph Scheduler["scheduler/ (drivers)"]
    Registry["WorkflowRegistry"]
    SchedulerMod["Scheduler\npoll + resume"]
    SR["ScheduledRuntime\nconvenience"]
  end

  subgraph Distribution["workflow_bundle/ + visualflow_compiler/"]
    Bundles["WorkflowBundles (.flow)\nmanifest + flows/*.json"]
    Compiler["VisualFlow compiler\nVisualFlow JSON -> WorkflowSpec"]
  end

  subgraph Integrations["integrations/ (optional wiring)"]
    AC["abstractcore/\nLLM_CALL, TOOL_CALLS, MCP worker,\nhost / run / discovery / config facades"]
    AM["abstractmemory/\nMEMORY_KG_* handlers"]
  end

  Runtime --> RunStore
  Runtime --> LedgerStore
  Runtime --> Artifacts
  Runtime --> Policy
  Runtime --> Vars
  Runtime --> Config

  SR --> SchedulerMod
  SR --> Runtime
  SchedulerMod --> RunStore
  SchedulerMod --> Registry

  Bundles --> Compiler
  AC --> Runtime
  AM --> Runtime
```

## Durable execution model

### WorkflowSpec and node handlers
- A workflow is a `WorkflowSpec` (`src/abstractruntime/core/spec.py`): `workflow_id`, `entry_node`, `nodes: dict[node_id, handler]`.
- A node handler returns a `StepPlan` (`src/abstractruntime/core/models.py`):
  - `effect`: optional side-effect request (`Effect`)
  - `next_node`: move the execution cursor
  - `complete_output`: finish the run

### RunState and the ledger
- `RunState` is the durable checkpoint stored by a `RunStore` (`src/abstractruntime/core/models.py`, `src/abstractruntime/storage/base.py`).
- Each executed step appends a `StepRecord` to a `LedgerStore` (`src/abstractruntime/core/models.py`, `src/abstractruntime/storage/base.py`).

**Invariant:** values stored in `RunState.vars` must be JSON-serializable. Use artifact references for large values (`src/abstractruntime/storage/artifacts.py`) or wrap stores with `OffloadingRunStore` / `OffloadingLedgerStore` (`src/abstractruntime/storage/offloading.py`).

### Crash-ordering invariant (reliability)

The run store is truth; the ledger is evidence. Every state transition obeys
ONE ordering law ([backlog 0045](backlog/planned/runtime_systemic_reliability/0045_crash_replay_harness_and_ordering_invariant.md); every durable feature is a consumer):

1. **State save is the commit point.** A transition is true when — and only
   when — `RunStore.save(run)` returns. Everything before the save must be
   safe to repeat; everything after must be pure observation.
2. **Ledger records for a transition append BEFORE the save that makes the
   transition true.** A crash between append and save leaves the ledger
   *ahead* of the state — the recoverable direction: replay re-executes from
   the last saved state and idempotency keys (`_runtime.effect_seq`-scoped)
   deduplicate both effect results and terminal records. The reverse order
   (save-then-append) leaves a truth the evidence never recorded — e.g. a
   COMPLETED run whose ledger never terminates, which stalls ledger-following
   clients forever.
3. **Terminal-path append failures are never silent.** A failed append on a
   completion/resume path increments `RuntimeHealth` (`record_error`) and
   logs a warning. Best-effort observability records (progress events,
   status events) may stay quiet; records that replay depends on may not.

Kill-and-replay coverage for this law lives in the crash harness
(`tests/harness.py`, backlog 0045): for every persistence point, kill →
restart → assert terminal-output equivalence, ledger convergence (terminal
record exactly once), and no doubled side effects.

## Runtime loop (start / tick / resume)

Implemented in `src/abstractruntime/core/runtime.py`:

- `Runtime.start(...)` creates a new `RunState` and initializes `_limits` from `RuntimeConfig` (`src/abstractruntime/core/config.py`).
- `Runtime.tick(...)` executes nodes until the run becomes `WAITING`, `COMPLETED`, `FAILED`, or is `CANCELLED`.
- `Runtime.resume(...)` validates the `wait_key`, writes the payload to `WaitState.result_key`, and continues execution **from** `WaitState.resume_to_node`.

```mermaid
sequenceDiagram
  participant Host
  participant RT as Runtime
  participant Node as Node handler
  participant EH as Effect handler
  participant RS as RunStore
  participant LS as LedgerStore

  Host->>RT: start(workflow, vars)
  RT->>RS: save(RunState RUNNING)

  Host->>RT: tick(run_id)
  RT->>Node: handler(run, ctx)
  Node-->>RT: StepPlan(effect?, next_node?, complete_output?)

  alt StepPlan.effect
    RT->>LS: append(StepRecord STARTED)
    RT->>EH: handle(effect)
    EH-->>RT: outcome (completed|waiting|failed)
    RT->>LS: append(StepRecord COMPLETED/WAITING/FAILED)
  end

  alt outcome=waiting
    RT->>RS: save(RunState WAITING + WaitState)
    RT-->>Host: RunState(waiting)
  else outcome=completed
    RT->>RS: save(RunState RUNNING/COMPLETED)
  end

  Host->>RT: resume(run_id, wait_key, payload)
  RT->>RS: save(RunState RUNNING)
  RT->>RT: tick(...)
```

## Effects: built-in vs wired by hosts

### Built-in (kernel-owned) effects
Registered in `Runtime._register_builtin_handlers()` (`src/abstractruntime/core/runtime.py`):
- waits: `WAIT_EVENT`, `WAIT_UNTIL`, `ASK_USER`, `ANSWER_USER`
- durable events: `EMIT_EVENT` (resumes matching `WAIT_EVENT` runs; requires `QueryableRunStore` and a workflow registry when listeners exist)
- subworkflows: `START_SUBWORKFLOW` (requires `runtime.workflow_registry`; see `src/abstractruntime/scheduler/registry.py`)
- memory primitives (JSON-safe): `MEMORY_NOTE`, `MEMORY_QUERY`, `MEMORY_TAG`, `MEMORY_COMPACT`, `MEMORY_REHYDRATE`
  - `MEMORY_COMPACT` requires an `ArtifactStore`. It uses an injected `chat_summarizer` when available; otherwise it runs an internal `LLM_CALL` subworkflow and therefore requires an `LLM_CALL` handler to be wired.
- inspection: `VARS_QUERY` (read-only access to `RunState.vars` paths; parsing helpers in `src/abstractruntime/core/vars.py`)

### Host-wired effects
The kernel defines the protocol; concrete integrations provide handlers:
- `LLM_CALL`, `TOOL_CALLS`, `MODEL_RESIDENCY`: provided by AbstractCore integration (`src/abstractruntime/integrations/abstractcore/effect_handlers.py`). The integration supports local/remote/hybrid execution, cached sessions/prompt-cache control with per-session attribution and clearing, host memory snapshots, discovery/catalog snapshots, model residency including host-wide provider-server rows in listings, residency locks with `force`-gated unload and context-estimate relays, durable run-scoped media child runs, media inputs, generated media outputs, provider progress callbacks as ledger events, provider-key header routing for remote servers, passthrough tools, approval-gated local tool execution, and workspace-scoped file and shell tools (run vars `workspace_*`, inherited by child runs; the host's built-in deny prefixes are enforced without being rendered into the model's prompt).
- `MEMORY_KG_*`: provided by the AbstractMemory bridge (`src/abstractruntime/integrations/abstractmemory/effect_handlers.py`)

### Reliability: retries + idempotency
- Policies live in `src/abstractruntime/core/policy.py` (e.g., `RetryPolicy`, `NoRetryPolicy`, `compute_idempotency_key()`).
- The runtime records `idempotency_key` and `attempt` on ledger records (`StepRecord`) and can reuse prior results after restarts (`src/abstractruntime/core/runtime.py`).

## Storage layer

Interfaces: `src/abstractruntime/storage/base.py`.

Included backends:
- in-memory (tests/dev): `src/abstractruntime/storage/in_memory.py`
- filesystem JSON/JSONL: `src/abstractruntime/storage/json_files.py`
- SQLite: `src/abstractruntime/storage/sqlite.py`

Decorators/helpers:
- `ObservableLedgerStore` for in-process subscriptions (`src/abstractruntime/storage/observable.py`, exposed via `Runtime.subscribe_ledger()`)
- `HashChainedLedgerStore` for tamper-evidence (`src/abstractruntime/storage/ledger_chain.py`)
- `ArtifactStore` + helpers (`src/abstractruntime/storage/artifacts.py`)
- `OffloadingRunStore` / `OffloadingLedgerStore` to keep checkpoints bounded (`src/abstractruntime/storage/offloading.py`)
- snapshots/bookmarks (`src/abstractruntime/storage/snapshots.py`)

## Drivers: scheduler

The scheduler is an in-process driver loop that resumes due waits and can deliver external events:
- `Scheduler` (`src/abstractruntime/scheduler/scheduler.py`) polls `QueryableRunStore.list_due_wait_until(...)`
- `ScheduledRuntime` + `create_scheduled_runtime()` (`src/abstractruntime/scheduler/convenience.py`) is the "zero-config" wrapper used in `examples/`

## Automations

An automation runs a target workflow on a trigger, and is itself a durable run. The controller is a packaged
VisualFlow bundle (`abstractframework.automation-controller@1.0.0`) whose nodes are `automation` adapters; it waits on
one `WAIT_EVENT` with a deadline, starts each occurrence as an asynchronous `START_SUBWORKFLOW` child with a
deterministic id, and records every transition as an `automation.*` ledger record. No new effect type or driver is
involved: a host drives a controller like any other run. See [automations.md](automations.md).

```mermaid
flowchart LR
  subgraph Host["Host (AbstractGateway run loop, or drive_automation)"]
    Loop["tick due runs /<br/>resume finished subworkflow parents"]
    API["commands + reads"]
  end
  subgraph Automations["automations/ + triggers/"]
    Commands["apply_automation_command"]
    Controller["controller run<br/>(automation-controller@1.0.0)"]
    Triggers["trigger registry<br/>schedule@1, manual@1,<br/>entry-point sources"]
    Queries["get_automation / list_occurrences /<br/>list_attention / pending_waits /<br/>list_automations"]
  end
  subgraph Runs["runs"]
    Occ["occurrence run<br/>(session turn)"]
    Desc["descendant runs"]
    Disc["discussion run<br/>(own workspace + read-only mount)"]
  end
  subgraph Stores["storage/"]
    RS["RunStore<br/>create_if_absent + run index"]
    LS["LedgerStore<br/>automation.* records"]
  end

  API --> Commands
  API --> Queries
  Commands -->|"run_mutation_lock,<br/>decision, wake"| Controller
  Loop --> Controller
  Loop --> Occ
  Controller -->|"prepare / admit / rearm"| Triggers
  Controller -->|"START_SUBWORKFLOW run_id"| Occ
  Occ --> Desc
  Controller --> RS
  Controller --> LS
  Occ --> RS
  Disc --> RS
  Queries --> RS
  Queries --> LS
```

Invariants the design rests on:

- **Create-if-absent.** Controllers, occurrences and discussions are created under explicit ids through
  `RunStore.create_if_absent`, so replaying a creation after a crash finds the same run and never starts a second
  one (`Runtime.start(..., run_id=...)`, `START_SUBWORKFLOW` `payload.run_id`).
- **One decision protocol.** Every controller step and command reconciles, looks its decision key up exactly, saves
  an intent, appends the record and applies it, so a crash at any point yields exactly one record and one
  application.
- **One lock per run.** `run_mutation_lock(run_id)` serializes `tick`, `resume` commits and automation commands in a
  process. v1 supports one writer process per store.
- **Turns are defined once.** `select_session_turns` decides which runs are a session's turns (parent-less runs
  except controllers, plus occurrences). History bundles, session replay, growing context and discussion seeds all
  use it, and `list_run_index(root_only=True)` returns the same turn roots.
- **Attribution is indexed.** Each run index row carries `automation_id`, `role`, `occurrence_index` and
  `session_kind`, derived from inline `vars._meta` that the offloading store never moves. `Runtime.start` uses
  `session_attribution` to anchor every root run in a discussion session to the discussion's validated root and its
  workspace setup (its own writable workspace plus the automation's workspace as a read-only mount), and refuses the
  start when the attribution cannot be read.

## Observability: what you can export

- Ledger (source of truth): `Runtime.get_ledger(run_id)` (`src/abstractruntime/core/runtime.py`)
- Runtime-owned node traces (bounded): stored at `vars["_runtime"]["node_traces"]` (`src/abstractruntime/core/runtime.py`, helpers in `src/abstractruntime/core/vars.py`)
- Evidence capture for external-boundary tools (`web_search`, `fetch_url`, `execute_command`):
  - recorder: `src/abstractruntime/evidence/recorder.py`
  - API: `Runtime.list_evidence(...)` / `Runtime.load_evidence(...)` (`src/abstractruntime/core/runtime.py`)
- Run history bundle export (portable replay artifact):
  - `export_run_history_bundle(...)` (`src/abstractruntime/history_bundle.py`)
  - `history_bundle["resolved_actions"]` is the portable capability/action summary layer built from
    Runtime/Core execution metadata

## VisualFlow + WorkflowBundles

AbstractRuntime includes a compiler and a portable bundle format:
- VisualFlow compiler: `src/abstractruntime/visualflow_compiler/*` (VisualFlow JSON -> `WorkflowSpec`)
- Multi-entry VisualFlow authoring routes are lowered into internal `join_exec` and `path_mux` nodes when a target has multiple incoming `exec-in` edges plus per-route input overrides. This keeps authoring JSON clean while making runtime behavior explicit and restart-safe. See `workflow-bundles.md` for the concrete metadata shape.
- VisualFlow document nodes include `read_pdf`, `write_pdf`, and `write_docx`.
  They use workspace-scoped paths, keep document bytes out of checkpoints, and
  store only JSON-safe extracted text, metadata, hashes, content types, and file
  paths in node outputs.
- WorkflowBundles (`.flow`): `src/abstractruntime/workflow_bundle/*` (manifest + flows + assets)
  - pack/unpack helpers: `pack_workflow_bundle(...)`, `open_workflow_bundle(...)`

## See also
- `../README.md` — install + quick start
- `getting-started.md` — first steps
- `api.md` — public API surface (imports + pointers)
- `automations.md` — automations: controller, triggers, commands, context, discussions, storage guarantees
- `limits.md` — `_limits` and RuntimeConfig
- `snapshots.md` — snapshot/bookmark stores
- `provenance.md` — hash chain and verification
- `evidence.md` — artifact-backed evidence capture
- `workflow-bundles.md` — `.flow` bundles + VisualFlow distribution
- `mcp-worker.md` — MCP worker entrypoint (`abstractruntime-mcp-worker`)
- `integrations/abstractcore.md` — AbstractCore wiring
- `manual_testing.md` — end-to-end smoke tests
- `adr/README.md` — rationale (why)


==============================================================================
# FILE: docs/api.md
==============================================================================

# API reference

This document summarizes the **public Python API** of AbstractRuntime and points to the **source of truth in code**.

Public exports live in `src/abstractruntime/__init__.py`. If you are unsure what is supported for external use, start there.

Stability guideline:
- Prefer imports from `abstractruntime` (package root) and `abstractruntime.storage`.
- Deep imports from `abstractruntime.core.*` / `abstractruntime.storage.*` are fine for advanced use, but treat them as lower-stability unless they are explicitly documented/re-exported.

## Recommended imports

Core kernel:

```python
from abstractruntime import Effect, EffectType, Runtime, StepPlan, WorkflowSpec
```

Storage helpers (common stores):

```python
from abstractruntime.storage import (
    InMemoryLedgerStore,
    InMemoryRunStore,
    JsonFileRunStore,
    JsonlLedgerStore,
)
```

Scheduler convenience wrapper:

```python
from abstractruntime import create_scheduled_runtime
```

AbstractCore integration (included in the base `abstractruntime` install):

```python
from abstractruntime.integrations.abstractcore import (
    ApprovalToolExecutor,
    MappingToolExecutor,
    ToolApprovalPolicy,
    create_local_runtime,
)
```

See also: `getting-started.md` (end-to-end runnable examples).

## Core types (durable workflow semantics)

Implementation: `src/abstractruntime/core/models.py`, `src/abstractruntime/core/spec.py`.

- `WorkflowSpec`: in-memory workflow graph (`workflow_id`, `entry_node`, `nodes`).
- `StepPlan`: node return value (what happens next): `effect`, `next_node`, or `complete_output`.
- `Effect` / `EffectType`: durable side-effect request protocol (the runtime mediates execution).
- `RunState` / `RunStatus`: durable checkpoint for a run, persisted by a `RunStore`.
- `WaitState` / `WaitReason`: durable pause metadata for `WAIT_*` / `ASK_USER` / passthrough tool waits.

Durability invariant: `RunState.vars` must remain JSON-serializable (`src/abstractruntime/core/models.py`). For large payloads use artifacts/offloading (`src/abstractruntime/storage/artifacts.py`, `src/abstractruntime/storage/offloading.py`).

## Runtime (start / tick / resume)

Implementation: `src/abstractruntime/core/runtime.py`.

- `Runtime.start(workflow, vars=..., actor_id=..., session_id=..., parent_run_id=..., run_id=...) -> run_id`
  - creates and persists a new `RunState`
  - `run_id` starts the run under an id you choose (`[A-Za-z0-9_-]+`). The run is created only if that id is free; starting it again with the same identity (workflow, session, parent, `vars._meta.occurrence`, `vars._meta.creation_digest`) returns the existing run untouched, and a different identity raises `RunIdentityConflict` (a `ValueError`, `reason_code = "identity_conflict"`). The store must support `create_if_absent` (otherwise `NotImplementedError`). `START_SUBWORKFLOW` accepts the same explicit id as `payload.run_id`
  - a root run that names a session is checked against the session's attribution: in a discussion session it gets the discussion's workspace setup (its own writable workspace and the read-only mount) and provenance, and when the attribution cannot be read (an invalid discussion root, or a store without a run index) the start raises `SessionAttributionError` (see [automations.md](automations.md#discussions))
- `Runtime.tick(workflow, run_id, max_steps=..., step_gate=None) -> RunState`
  - executes node handlers and effects until the run becomes `WAITING`, `COMPLETED`, `FAILED`, or `CANCELLED`
  - optional `step_gate()` is consulted at every step boundary; when it returns `False` the tick returns the persisted state and the run stays `RUNNING` (a later tick continues where it stopped)
- `Runtime.resume(workflow, run_id, wait_key, payload, max_steps=...) -> RunState`
  - validates the `wait_key`, writes `payload` to `WaitState.result_key` (if set), and continues from `WaitState.resume_to_node`
  - a wait is resumed at most once: a resume of a run that is no longer waiting, or that waits on another key, raises `StaleResumeError` (a `ValueError`), for example when another caller resumed it first
  - the check and the commit run under a per-run lock, and tools approved with `{"approved": true}` execute while that lock is held, so a long tool delays a competing resume of the same run, which is then refused
- `run_mutation_lock(run_id)` (package root): the per-run, per-process, re-entrant lock that `tick` holds for the whole tick and `resume` for its commit; take it around your own read-modify-save of a run so a tick cannot overwrite your change
- `Runtime.get_state(run_id) -> RunState` and `Runtime.get_ledger(run_id) -> list[dict]`
  - host-facing read APIs for checkpoints and the append-only ledger
- `Runtime.cancel_run(run_id, reason=None, cancelled_by="api") -> RunState`
  - persists `CANCELLED`, then signals every effect of that run (and its in-flight descendants) executing in this process; the running attempt is recorded as `cancelled` (`StepStatus.CANCELLED`, never retried) with `cancelled_by` and `reason`, and no further effect starts
- `Runtime.set_default_provider_model(provider=..., model=...)`
  - re-points the default provider/model seeded into new runs; pair it with the pooled client's `set_default_provider_model(...)` so both agree
- `Runtime.set_live_delta_sink(sink)` where `sink(event: dict) -> None`, or `None` to remove it
  - receives live token deltas (`llm.delta`, `llm.delta_end`) for LLM calls of runs started with `_runtime.stream: true` (a boolean; `start` refuses any other value); nothing is written to the ledger. See [Live token streaming](integrations/abstractcore.md#live-token-streaming)

Effect cancellation helpers (`src/abstractruntime/core/effect_cancellation.py`):
- `inflight_effects(run_ids=None)` lists executing effects (run, step, provider, model, elapsed time)
- `request_model_effects_cancel(provider, model, cancelled_by="model_eject")` stops the effects using a model before it is unloaded; local `unload_model_residency` calls it for you
- `kill_inflight_effect(step_id, killed_by=...)` is an in-process hard stop for a call that ignores its cancel event (it cannot interrupt a thread blocked inside one native call)

Tool scope: an explicit `allowed_tools` list in `run.vars["_runtime"]`, in a child run's `_runtime`, or in a `TOOL_CALLS` payload is a ceiling intersected across the run tree (`src/abstractruntime/core/tool_scope.py`). Approval policy can remove a prompt but can never grant a tool outside it; a missing key means unrestricted and an empty list denies every tool.

For the execution model (ledger records, effect outcomes, waits), see `architecture.md`.

## Scheduler convenience API

Implementation: `src/abstractruntime/scheduler/*`.

Use `create_scheduled_runtime()` for a zero-config wrapper that bundles `Runtime` + an in-process polling `Scheduler`:
- `ScheduledRuntime.run(workflow, vars=..., actor_id=..., max_steps=...) -> (run_id, state)` (`src/abstractruntime/scheduler/convenience.py`)
- `ScheduledRuntime.respond(run_id, payload) -> RunState` (resumes a waiting run using its stored `wait_key`)
- `ScheduledRuntime.stop()` (stops the scheduler thread/loop)

For time-based waits, the scheduler polls due runs via `QueryableRunStore.list_due_wait_until(...)` (`src/abstractruntime/storage/base.py`, `src/abstractruntime/scheduler/scheduler.py`).

## Storage layer (durability backends)

Interfaces: `RunStore`, `LedgerStore`, and `QueryableRunStore` are defined in `src/abstractruntime/storage/base.py`.

Included backends:
- In-memory (tests/dev): `InMemoryRunStore`, `InMemoryLedgerStore` (`src/abstractruntime/storage/in_memory.py`)
- Filesystem:
  - checkpoints: `JsonFileRunStore` (`src/abstractruntime/storage/json_files.py`)
  - append-only ledger: `JsonlLedgerStore` (`src/abstractruntime/storage/json_files.py`)
- SQLite:
  - `SqliteRunStore`, `SqliteLedgerStore` (`src/abstractruntime/storage/sqlite.py`)

Notes:
- `abstractruntime.storage` intentionally exports only the most common store types. SQLite types are available via:
  - `from abstractruntime import SqliteRunStore, SqliteLedgerStore`, or
  - `from abstractruntime.storage.sqlite import SqliteRunStore, SqliteLedgerStore`

Run index and creation (all built-in run stores, including the offloading wrapper):
- `create_if_absent(run) -> (run, created)`: creates a run only if its id is free and never overwrites an existing one; `store_supports_create_if_absent(store)` / `require_create_if_absent(store)` check a store first. Process-crash safe; power-loss durability is not claimed
- `list_run_index(status=, workflow_id=, session_id=, root_only=, limit=, oldest_first=, automation_id=, role=, session_kind=)`: lightweight rows carrying `automation_id`, `role`, `occurrence_index`, `session_kind` and `workspace_root` (the run's top-level `vars["workspace_root"]` as stored, stripped; `None` when absent); the three attribution filters take a value, a comma-separated string or a list; `root_only=True` returns turn roots (parent-less runs except automation controllers, plus automation occurrences)
- `session_kinds(session_id) -> frozenset` and `latest_occurrence_row(automation_id)`: indexed lookups used by session attribution and automation listings
- `JsonFileRunStore.warm_session_index()` builds the session and children indexes at host startup. Store objects and processes sharing one JSON run folder see each other's created and deleted runs through the creation journal `.runs_created.log`; v1 supports one writer process per store

Common decorators:
- `ObservableLedgerStore` for subscriptions (`src/abstractruntime/storage/observable.py`)
- `HashChainedLedgerStore` + `verify_ledger_chain(...)` for tamper-evidence (`src/abstractruntime/storage/ledger_chain.py`)
- `OffloadingRunStore` / `OffloadingLedgerStore` to store large values by artifact reference (`src/abstractruntime/storage/offloading.py`)

## Commands (durable control-plane inbox)

AbstractRuntime ships append-only, idempotent **command inbox** primitives designed for gateways/workers that must accept retries safely:
- models + interfaces: `CommandRecord`, `CommandStore`, `CommandCursorStore` (`src/abstractruntime/storage/commands.py`)
- backends: in-memory + JSONL (`src/abstractruntime/storage/commands.py`), SQLite (`src/abstractruntime/storage/sqlite.py`)

These APIs are exported at the package root (see `src/abstractruntime/__init__.py`).

## Artifacts (store by reference)

Implementation: `src/abstractruntime/storage/artifacts.py`.
Deep dive: `artifacts.md`.

Key types:
- `ArtifactStore` (interface), `InMemoryArtifactStore`, `FileArtifactStore`
- helpers: `artifact_ref(...)`, `resolve_artifact(...)`, `is_artifact_ref(...)`

The store keeps payload bytes out of run state and persists structured metadata:
- `ArtifactDescriptor` is the Runtime-owned descriptor used by Gateway and Observer. It separates `semantic_kind` such as `voice`, `music`, `sound`, or `image` from `render_kind` such as `audio`, `markdown`, `html`, or `json`, and can carry workflow/node/turn links, media facts, generation/provenance data, source refs, security, and action links.
- `ArtifactAccessStats` records explicit metadata/content/preview/download/export actions when HTTP or UI layers call `record_access(...)`. Plain `load(...)` and `get_metadata(...)` remain side-effect free.
- `search(...)`, `count(...)`, `facet_counts(...)`, and `stats(...)` provide metadata queries for host control planes. `FileArtifactStore` serves these from a repairable SQLite catalog when possible, including exact `total`, `total_bytes`, and requested facet counts without forcing Gateway/Observer to load every matching artifact.

Artifacts are used by:
- offloading wrappers (`src/abstractruntime/storage/offloading.py`)
- evidence capture (`docs/evidence.md`, `src/abstractruntime/evidence/recorder.py`)
- AbstractCore media integration: input artifact refs can be materialized for LLM calls, and generated image/video/voice/music/audio outputs are stored as artifact refs

## Snapshots / bookmarks

Implementation: `src/abstractruntime/storage/snapshots.py`.

- `SnapshotStore` interface + `InMemorySnapshotStore`, `JsonSnapshotStore`
- `Snapshot` model (a named bookmark of run state)

Docs: `snapshots.md`.

## Effect policies (retries + idempotency)

Implementation: `src/abstractruntime/core/policy.py`.

- `EffectPolicy` protocol and implementations: `DefaultEffectPolicy`, `RetryPolicy`, `NoRetryPolicy`
- `compute_idempotency_key(...)` helper

Docs: `architecture.md` (reliability section).

## WorkflowBundles (`.flow`) and VisualFlow distribution

Implementation:
- bundles: `src/abstractruntime/workflow_bundle/*`
- compiler: `src/abstractruntime/visualflow_compiler/*`

VisualFlow compiler helpers are available from `abstractruntime.visualflow_compiler`:
- `load_visualflow_json(...)` normalizes VisualFlow JSON into the stdlib model.
- `visual_to_flow(...)` lowers VisualFlow into the internal Flow IR.
- `compile_visualflow(...)` and `compile_visualflow_tree(...)` compile VisualFlow JSON into executable `WorkflowSpec` objects.

VisualFlow authoring note (media and document nodes):
- Runtime recognizes first-class VisualFlow media nodes such as `generate_image`, `edit_image`, `image_to_image`, `upscale_image`, `image_upscale`, `generate_video`, `text_to_video`, `image_to_video`, `generate_voice`, `generate_music`, `transcribe_audio`, and `listen_voice`.
- Generated-media and transcription nodes lower to a durable `EffectType.LLM_CALL` with an `output` selector (for example `{"modality":"music","task":"music_generation"}`), while `listen_voice` lowers to `WAIT_EVENT`. Hosts should persist the authoring node type rather than pre-lowering to `llm_call`.
- Runtime also recognizes file/document nodes. `read_file` and `write_file`
  handle UTF-8 text/JSON workspace paths. In Gateway-hosted runs, those paths
  follow the shared canonical contract: `rel/path` for the main workspace root
  and `mount_alias/rel/path` for approved mounts. `read_pdf` extracts text and
  metadata from PDF paths with `pypdf`; `write_pdf` renders text or
  Markdown-style content to real PDF bytes with `reportlab`; `write_docx`
  renders Markdown-style content to real `.docx` bytes with the standard
  library; `list_folder_files` enumerates workspace-scoped folders with
  family/extension filters;
  `import_workspace_file` snapshots a workspace file into a durable artifact;
  `read_artifact` projects saved file content back out as text/JSON/bounded
  binary metadata; and `export_artifact` writes a durable artifact back to a
  workspace path. PDF/DOCX bytes are written to the workspace path and only
  JSON-safe metadata/path values are stored in run state. In local Runtime-only
  runs with no workspace scope, relative file-node paths still fall back to the
  process working directory.

Public bundle APIs are exported from `src/abstractruntime/workflow_bundle/__init__.py` and re-exported in `src/abstractruntime/__init__.py`:
- open: `open_workflow_bundle(...)`
- registry: `WorkflowBundleRegistry`
- pack/unpack: `pack_workflow_bundle(...)`, `unpack_workflow_bundle(...)`

Docs: `workflow-bundles.md`.

## Automations

Implementation: `src/abstractruntime/automations/*`, `src/abstractruntime/triggers/*`, `src/abstractruntime/automation_queries.py`.
Deep dive: [automations.md](automations.md).

```python
from abstractruntime.automations import (
    apply_automation_command,   # pause / resume / run_now / revise / stop_current / archive
    create_automation,          # -> (automation_id, revision)
    drive_automation,           # standalone run loop
    get_automation,
    list_attention,
    list_occurrences,
    pending_waits,
    register_controller_bundle,
    start_discussion,
)
from abstractruntime.automation_queries import latest_occurrence, list_automations
from abstractruntime.triggers import get_trigger_adapter, trigger_sources
```

- `create_automation(runtime, request, *, now=None, actor_id=None)`: creates the controller root run (the automation) with a deterministic id; raises `AutomationError` (`invalid_definition`, `unsupported_feature`, `unknown_trigger_source`, `identity_conflict`)
- `apply_automation_command(runtime, *, automation_id, command_id, type, payload=None, actor=None, expected_revision=None)`: the only writer of automation state; returns `{status: "applied" | "rejected", error?, duplicate}`; idempotent per `command_id`
- `record_automation_command_result(...)`: records a host-side failure of a command as a rejected result
- reads: `get_automation(run_store, id)`, `list_occurrences(runtime, id, cursor=, limit=)`, `list_attention(ledger_store, id, after_seq=, cursor=, limit=)`, `pending_waits(run_store, id, limit=)` with typed waits (`ask_user`, `tool_approval`, `event`) and `ANSWER_PAYLOADS`; `list_automations(run_store, status=, cursor=, limit=)`, `automation_summary(run)`, `latest_occurrence(run_store, id)`
- `start_discussion(runtime, *, automation_id, occurrence_index, request_id, prompt, workspace_root, actor_id=None)`: a separate conversation forked at any occurrence N, seeded with the automation's whole timeline 1..N, working in its own writable `workspace_root` with the automation's workspace mounted read-only
- controller bundle: `register_controller_bundle(registry)`, `controller_workflow_spec()`, `controller_bundle_path()`; `CONTROLLER_WORKFLOW_ID` is `abstractframework.automation-controller@1.0.0:controller`
- trigger sources: `trigger_sources()`, `get_trigger_adapter(id, version)`, built-ins `schedule@1` and `manual@1`, third-party sources through the `abstractruntime.trigger_sources` entry-point group
- `adopt_legacy_schedule_projection(run)`: read-only summary of a legacy gateway `scheduled:*` root
- read-only mounts: `_runtime.workspace_read_only_paths` (absolute folders; file write tools and VisualFlow writers refused inside, reads and command/code tools allowed); helpers `read_only_paths(vars)`, `path_is_read_only(vars, path)` and `READ_ONLY_PATHS_KEY` in `abstractruntime.utils.workspace_paths`
- read-only workspaces: run vars `workspace_read_only: true` (or `_runtime.workspace_read_only: true`); `abstractruntime.integrations.abstractcore.tool_effects.TOOL_EFFECT_CLASSES` classifies every exposable tool as `read`, `write`, `exec`, `delegate`, `comms` or `memory-write`

## Sessions and history

Implementation: `src/abstractruntime/session_turns.py`, `src/abstractruntime/session_history.py`, `src/abstractruntime/core/run_attribution.py`.

- `select_session_turns(run_store, session_id, *, include_occurrences=True, until_ms=None, automation_id=None, through_occurrence=None, include_drafts=False, limit=50)` (package root): a session's turns, oldest first. Turns are parent-less runs except automation controllers, plus automation occurrences (a retried occurrence counts once, as its newest attempt); child runs, runtime-internal runs, legacy scheduled wrappers and draft-test runs are left out
- `session_chat_messages(run_store=, ledger_store=, artifact_store=, session_id=, max_tokens=HISTORY_REPLAY_MAX_TOKENS, until_ms=None, exclude_run_ids=None, automation_id=None, through_occurrence=None, strict=False)` (package root): the session's completed turns as user/assistant message pairs; in a discussion session the discussion's seed comes first. Returns a `ReplayedHistory` (a list of messages) whose `.report` says what was replayed and dropped. `strict=True` raises `SessionHistoryError` (`reason_code = "history_unavailable"`) instead of returning a partial history. `max_messages`, `max_total_chars` and `max_chars_per_message` (keyword-only, default `None`) are the caps retired in 0.7.0: they are accepted so hosts built against 0.6 (AbstractGateway 0.6.0) still get their history, and ignored. Passing any of them logs one warning and names them in `report["ignored_inputs"]`
- The history window: `HISTORY_REPLAY_MAX_TOKENS = 50_000` (package root). Replay keeps the most recent turns that fit 50,000 estimated tokens, newest first, as whole messages. There is no message-count cap and no character cap, and no message is ever cut. The fold stops at the first older turn that does not fit, so the window has no gaps. A total of exactly 50,000 tokens fits. One exception: when the newest turn alone is larger than the window, it is kept whole (`oversize_turn_kept: true`), because cutting it would lose content and dropping it would replay nothing. The model can use the rest of its context window; if a turn is too large for the model, the provider reports the error. Tokens are counted with `abstractruntime.memory.token_budget.estimate_message_tokens`. `max_tokens` must be a positive int
- `ReplayedHistory.report`: `{policy: "most_recent_whole_turns", max_tokens, token_estimator, replayed_messages, replayed_tokens, dropped_messages, dropped_tokens, dropped_counts_complete, oversize_turn_kept}`. Replay stops reading once the window is full, so in a long session `dropped_counts_complete` is `false` and the dropped counts cover only the turns it read. When turns were dropped, the oldest replayed message starts with a `[#TRUNCATION: N earlier message(s) ... were dropped from replay by the history window ...]` line (ADR-0026). Hosts record the report in the run as `vars._runtime.session_history`
- `fold_history_window(pairs, *, max_tokens=HISTORY_REPLAY_MAX_TOKENS)` (package root): the window itself, applied to chronological turns (lists of whole messages); returns `(kept, report)`
- `announce_dropped(messages, report, *, stamp_metadata=True)` (package root): writes the window's `[#TRUNCATION: ...]` notice at the start of the oldest kept message, in place, when `report` says turns were dropped. Use it after your own `fold_history_window` call (for example on client-sent history) so the notice matches session replay. `stamp_metadata=True` also adds `metadata.replay_truncated` and `metadata.history_window` to that message; pass `False` for plain role/content messages. `_announce_dropped` is kept as an alias of the pre-0.7.0 private name
- `window_transcript(messages, *, max_tokens=HISTORY_REPLAY_MAX_TOKENS, current_turn_start=None)` (package root): the same window over a transcript you already hold (a turn is a user message and the messages after it). `current_turn_start` is the index of the message that opened the turn in progress: from there on everything is one turn, the newest, always kept whole (`oversize_turn_kept` when it alone exceeds the window), whatever user-role messages a loop adds inside it; an index out of range raises `ValueError`. Returns a `ReplayedHistory` of copies with the notice written by `announce_dropped(..., stamp_metadata=False)`. The entity chat driver and the entity visit workflow use it for their prompts. The visit records the report in `vars._runtime.session_history`; the chat driver puts it on `TurnReport.history_window`. On the react visit arm, BRIDGE sets `_runtime.history_window_tokens` (50,000) and `_runtime.history_window_turn_start` (the visitor's message) and the AbstractAgent react loop (0.3.17 or newer) sends `window_transcript(context.messages)` on each call and records the report; the stored transcript stays whole. The report carries `window_applied: true`; with an older AbstractAgent that records no window the turn still completes, a warning is logged and the report is `{"window_applied": false, "reason": "agent_too_old", ...}` (the whole transcript was sent). When the oldest kept message carries a `<runtime_metadata>` envelope, the notice is written after it
- `session_attribution(run_store, session_id)` (`abstractruntime.core.run_attribution`): `None` or `{"kind": "chat" | "automation" | "occurrence" | "discussion", ...}`; a discussion adds its validated root, automation, occurrence and workspace. Raises `SessionAttributionError` when the lookup cannot be completed
- `is_draft_lifecycle(run_lifecycle)` (package root): whether a run's lifecycle value marks a draft test run

History bundles (`export_run_history_bundle`) and session replay use the same turn selection, so automation occurrences appear as turns (kind `occurrence`, with `automation_id` and `occurrence_index`) in both.

## Run history bundle export (portable replay artifact)

Implementation: `src/abstractruntime/history_bundle.py`.

- `export_run_history_bundle(...)`
- `persist_workflow_snapshot(...)`

This produces a portable record of a run’s state + ledger + artifacts suitable for debugging/review.
When available, it also includes `resolved_actions`: bounded capability/action summaries derived
from Core route resolution and persisted from `LLM_CALL` results.

## Runtime-owned integrations

### AbstractCore (LLM + tools)

Requires: `pip install abstractruntime` (AbstractCore 2.18.0 or newer is part of the base install).

Implementation: `src/abstractruntime/integrations/abstractcore/*`.

Entry points:
- `create_local_runtime(...)`, `create_remote_runtime(...)`, `create_hybrid_runtime(...)` (`src/abstractruntime/integrations/abstractcore/factory.py`)
- public discovery facade: `AbstractCoreDiscoveryFacade`, `get_abstractcore_discovery_facade(...)` (`src/abstractruntime/integrations/abstractcore/discovery_facade.py`)
- public host facade: `AbstractCoreHostFacade`, `get_abstractcore_host_facade(...)` (`src/abstractruntime/integrations/abstractcore/host_facade.py`)
- public email comms wrappers: `list_email_accounts(...)`, `list_emails(...)`, `read_email(...)`, `send_email(...)` (`src/abstractruntime/integrations/abstractcore/comms_facade.py`)
- public Telegram host wrappers: `TelegramTdlibNotAvailable`, `bootstrap_telegram_auth_from_env(...)`, `get_global_telegram_client(...)`, `stop_global_telegram_client()`, `send_telegram_message(...)` (`src/abstractruntime/integrations/abstractcore/telegram_facade.py`)
- public durable run facade: `AbstractCoreRunFacade`, `get_abstractcore_run_facade(...)` (`src/abstractruntime/integrations/abstractcore/run_facade.py`)
- effect handler wiring: `build_effect_handlers(...)` (`src/abstractruntime/integrations/abstractcore/effect_handlers.py`)
- tool executors: `MappingToolExecutor`, `AbstractCoreToolExecutor`, `PassthroughToolExecutor`, `ApprovalToolExecutor`, `ToolApprovalPolicy` (`src/abstractruntime/integrations/abstractcore/tool_executor.py`)
- discovery-facade delegation is implemented by the configured AbstractCore LLM clients in `src/abstractruntime/integrations/abstractcore/llm_client.py` (`list_providers`, `list_provider_models`, `get_voice_catalog`, `list_tts_models`, `list_stt_models`, `list_music_providers`, `list_music_models`, `list_vision_provider_models`, `list_cached_vision_models`, `list_vision_adapters`)
- the host's OpenAI credential for AbstractVoice's `openai` engines, `voice_openai_api_key`, travels two ways. For discovery, pass it to `get_voice_catalog`, `list_tts_models` or `list_stt_models`; the local clients hand it to the voice plugin's config, and a remote client never sends it (a remote AbstractCore server uses its own voice configuration). For execution, put it in `create_local_runtime(llm_kwargs={"voice_openai_api_key": ...})`. Constructor kwargs become the provider's `config`, which is also the capability-plugin config. Providers read only the keys they name, so a `voice_*` key changes no text request (tests `tests/test_voice_openai_api_key_path.py`)
- host-facade client delegation is implemented by the configured AbstractCore LLM clients in `src/abstractruntime/integrations/abstractcore/llm_client.py` (`get_prompt_cache_capabilities`, `get_prompt_cache_stats`, `prompt_cache_set`, `prompt_cache_update`, `prompt_cache_fork`, `prompt_cache_clear`, `prompt_cache_prepare_modules`, `upsert_text_bloc`, `get_bloc_record`, `list_blocs`, `get_bloc_kv_manifest`, `ensure_bloc_kv_artifact`, `load_bloc_kv_artifact`, `list_bloc_kv_artifacts`, `delete_bloc_kv_artifact`, `prune_bloc_kv_artifacts`, `delete_bloc`, `get_model_residency_capabilities`, `list_model_residency`, `load_model_residency`, `unload_model_residency`, `get_memory_snapshot`, `list_session_prompt_caches`, `clear_session_prompt_caches`, `lock_model_residency`, `unlock_model_residency`, `get_context_estimate`); the last six are optional in the client contract, and the facade returns `{"ok": false, "supported": false, ...}` when the configured client does not implement one
- switching the default model (`set_default_provider_model` on the pooled client) unloads the previous in-process model unless an owner in the process still uses it; `list_model_residency` diagnostics report `pending_ejects` and `last_switch_ejects`, and `unload_model_residency(runtime_id=...)` accepts any listed `local:<task>:<provider>:<model>` id, including `local:embedding:<provider>:<model>` rows (see `integrations/abstractcore.md#unloading-and-switching-models`)
- `lock_model_residency`, `unlock_model_residency`, and `get_context_estimate` accept an optional payload mapping and/or keyword arguments (keyword arguments win on conflicts); a lock requires provider-verified residency — locking a non-resident pair refuses with a structured `model_not_resident` payload (load with `lock: true` instead; remotely the Core server's HTTP 409 refusal is converted to the same payload), while unlock never requires residency so a locked-but-evicted pair is always releasable; a locked model refuses `unload_model_residency` with a structured `model_locked` payload — locally as a client-side per-`(provider, model)` lock, remotely by converting the Core server's HTTP 409 refusal — and unloads only with `force=true`, releasing the lock only after the unload succeeds
- the `MODEL_RESIDENCY` effect supports `list_loaded`, `load`, `unload`, `lock`, and `unlock` operations with the same soft-fail semantics; on `unload`, `force` is forwarded only when authored in the effect payload
- local `list_model_residency` merges AbstractCore's host-wide loaded-model sweep into text-generation listings (sweep-only rows carry `source: "provider_server"` and no `task` label), and residency rows normalize provider size extras to `size_bytes` / `size_vram_bytes` while keeping the originals; local text rows also carry registry-declared `modalities` (omitted on a registry miss), host identity (`host_id` / `host_name`), and runtime-owned lock truth (`locked` / `lockable`, with `pinned` a truthful alias of `locked` — never the default-identity flag, which `default` alone carries), while remote listings keep the Core server's row identity
- host-local prompt-cache export/import admin also lives on the host facade and client delegation layer (`list_prompt_cache_exports`, `prompt_cache_export`, `prompt_cache_import`) and is intentionally local-only
- host-facade email helpers delegate to Runtime's host-local comms facade/export layer (`list_email_accounts`, `list_emails`, `read_email`, `send_email`)
- run-facade helpers create and resume durable child runs for existing runs (`execute_llm_call`, `execute_tool_calls`, `resume_tool_calls`, `generate_image`, `edit_image`, `upscale_image`, `generate_video`, `image_to_video`, `generate_voice`, `generate_music`, `transcribe_audio`, `send_email`, `send_telegram_message`)
- task-specific image/video helpers preserve batch and adapter controls such as
  `count`/`n`, `seeds`, ordered `lora_adapters`, and video `flow_shift`; local
  subprocess isolation stays within the same public contract.

Execution controls and cancellation:
- `params.thinking` (bool or reasoning level) and `params.speculation` (`False`, `True`, or a Core speculation object such as `{"mode": "native_mtp", "num_draft_tokens": 2}`) are forwarded locally and remotely; `run.vars["_runtime"]["speculation"]` sets a run-wide preference inherited by subworkflows, Agent loops and delegated calls, and `False` stays Off across every boundary (see `integrations/abstractcore.md#execution-controls-and-local-concurrency`)
- `get_abstractcore_discovery_facade(...).get_execution_capabilities(model_name=None, provider=None)` asks the actual execution host (local or remote) what it supports, without loading a model
- the `LLM_CALL` handler hands the effect's cancel event to AbstractCore (`generate(..., cancel_event=...)`); remote clients close the request to the AbstractCore server on Stop, which the server treats as a cancel
- every `LLM_CALL` is offered a progress callback; providers that report text phases (prefill / generate / complete) produce `abstract.progress` ledger events with `kind: "llm"`
- `abstractruntime.turn_grounding.stamp_user_turn_grounding(messages, grounding=...)` writes the grounding envelope once into the stored user turn, so the prompt sent to the model stays a byte prefix of the next turn's prompt

`LLM_CALL` payloads are JSON-safe effect payloads. Common fields:
- `prompt`, `messages`, `system_prompt`, and convenience `text`
- `media`: a media path, artifact ref (`{"$artifact": "..."}` or `{"artifact_id": "..."}`), media dict, or list of those
- `output`: AbstractCore output selector; top-level `outputs` is accepted as a runtime alias
- `params`: provider/model routing, generation controls, prompt-cache keys or `prompt_cache_binding`, structured-output schema options, and tracing metadata

Multimodal support:
- common remote-light AbstractCore media, vision, voice, audio, and music dependencies are part of the base Runtime install
- local clients call AbstractCore's unified `generate(..., media=..., output=...)`
- remote and hybrid clients support AbstractCore Server chat media content arrays plus image generation, image edits, image upscaling, text-to-video, image-to-video, speech, music generation, and transcription endpoints; pass an output-specific `model` for remote media provider routing, otherwise the server endpoint can use its configured capability default
- remote transcription requires one audio media item that resolves to a local file path or artifact-backed temporary file
- generated image/video/voice/music/audio bytes require a runtime `ArtifactStore`; the result contains `artifact_id` / `artifact_ref` instead of inline bytes
- media-only normalized results expose `runtime_provider` / `runtime_model` separately from `media_provider` / `media_model`
- optional local media residency failures complete with `status_hint="warning"` and `degraded=true`; unsupported local media warmup for `image_generation`, `image_upscale`, `video_generation`, `text_to_video`, `image_to_video`, `tts`, `stt`, and `music_generation` reports `requires_long_lived_server=true`, and generated image/video tasks also report `execution_mode="local_one_shot_subprocess"`
- Gateway/hosts remain responsible for explicit Core server URLs, Core server auth headers, provider/model defaults, selected local-inference profiles, and translation of Gateway-owned env/config into explicit Runtime inputs; Runtime persists only JSON-safe routing metadata and artifact refs

Prompt cache / cached sessions:
- LLM clients expose cache control methods listed above for host-side preparation and inspection
- `LLM_CALL.params.prompt_cache_key` selects a cache key for a call; runtime can also derive a session-scoped key from `run.vars["_runtime"]["prompt_cache"]` or the Runtime-owned `ABSTRACTRUNTIME_PROMPT_CACHE` process default
- `LLM_CALL.params.prompt_cache_binding` is the durable exact-reuse input for bloc-backed prompt caching; if a binding includes `key`, Runtime adopts it as the effective prompt-cache key and refuses mismatches before provider execution
- Runtime only auto-derives session prompt-cache keys for text/chat calls; non-text output selectors such as image, voice, music, and transcription keep explicit `prompt_cache_binding` support but do not receive an inferred cache key
- when Runtime injects the derived session key, the client stamps session attribution (`session_id`, `run_id`, `workflow_id`, `node_id`, `namespace`) into the cache entry's metadata after each generate; caller-supplied `prompt_cache_key`s and binding keys are never stamped, so session-scoped clearing cannot destroy caches shared across sessions
- `get_abstractcore_host_facade(...)` exposes `get_memory_snapshot()`, `list_session_prompt_caches(session_id=None)`, and `clear_session_prompt_caches(session_id)`; session caches persist until a host clears them, the owning model is unloaded, or the provider evicts them — completing a run does not clear them
- `get_abstractcore_host_facade(...)` also exposes durable bloc helpers (`upsert_text_bloc`, `get_bloc_record`, `list_blocs`, `get_bloc_kv_manifest`, `ensure_bloc_kv_artifact`, `load_bloc_kv_artifact`, `list_bloc_kv_artifacts`, `delete_bloc_kv_artifact`, `prune_bloc_kv_artifacts`, `delete_bloc`)
- local Runtime owns the bloc root policy: `~/.abstractruntime/blocs` by default, `<base_dir>/blocs` for `create_local_file_runtime(...)`, and explicit `bloc_root_dir=...` overrides when needed
- provider cache/session handles are not durable runtime state and should not be stored in `RunState.vars`

Workspace scope:
- run vars `workspace_root`, `workspace_access_mode`, `workspace_allowed_paths`, `workspace_ignored_paths`, and the host-only `workspace_builtin_deny_prefixes` / `workspace_builtin_allow` scope file and shell tools; child runs inherit them (see `integrations/abstractcore.md#workspace-scoped-tools`)

Attachment registration limits:
- `TOOL_CALLS.payload.max_attachment_bytes`, `run.vars["_runtime"]["max_attachment_bytes"]`, or `ABSTRACTRUNTIME_MAX_ATTACHMENT_BYTES` bound the bytes Runtime stores when local `read_file` outputs are captured as session attachments

Docs: `integrations/abstractcore.md`.

### AbstractMemory bridge (KG effects)

Implementation: `src/abstractruntime/integrations/abstractmemory/effect_handlers.py`.

This provides handlers for `MEMORY_KG_*` effects (opt-in wiring layer).

## Utilities (host UX)

- Rendering helpers: `abstractruntime.rendering.stringify_json(...)` and `abstractruntime.rendering.render_agent_trace_markdown(...)` (`src/abstractruntime/rendering/*`)
- Active-context helpers (what is sent to the LLM): `ActiveContextPolicy`, `TimeRange` (`src/abstractruntime/memory/active_context.py`, exports in `src/abstractruntime/memory/__init__.py`)

## See also

- `../README.md` — install + quick start
- `getting-started.md` — first durable workflow
- `architecture.md` — component map + durability invariants
- `faq.md` — common questions and gotchas
- `integrations/abstractcore.md` — `LLM_CALL` / `TOOL_CALLS` wiring


==============================================================================
# FILE: docs/faq.md
==============================================================================

# FAQ

## What is AbstractRuntime (in one sentence)?

AbstractRuntime is a **durable workflow runtime**: it runs workflow graphs as a persisted state machine with explicit waits (pause → resume) and an append-only execution ledger.  
Code: `src/abstractruntime/core/runtime.py`, `src/abstractruntime/core/models.py`.

## Is AbstractRuntime an agent framework?

No. AbstractRuntime is the **execution substrate**. Agent logic (ReAct/CodeAct loops, prompt policies, etc.) is built *on top* of it.  
Docs: `proposal.md`. Code: `src/abstractruntime/core/*`.

## How does AbstractRuntime relate to AbstractCore / AbstractFramework?

AbstractRuntime is the **durable execution kernel**. In the AbstractFramework ecosystem, it is commonly paired with:
- **AbstractCore** for LLM + tool execution (`EffectType.LLM_CALL`, `EffectType.TOOL_CALLS`)  
  Code: `src/abstractruntime/integrations/abstractcore/*`. Repo: [lpalbou/abstractcore](https://github.com/lpalbou/abstractcore)

AbstractFramework umbrella: [lpalbou/AbstractFramework](https://github.com/lpalbou/AbstractFramework)

## Where is the public API documented?

- API guide: `api.md`
- Canonical export list: `src/abstractruntime/__init__.py`

## How do pause/resume work?

- A node returns a `StepPlan` with an `Effect` (e.g. `ASK_USER`, `WAIT_UNTIL`, `WAIT_EVENT`).
- The runtime persists a `WaitState` into `RunState.waiting` and returns `status=waiting`.
- You resume by calling `Runtime.resume(...)` (or `ScheduledRuntime.respond(...)`) with the matching `wait_key`.

Docs: `getting-started.md`, `architecture.md`. Code: `src/abstractruntime/core/runtime.py` (`tick`, `resume`) and `src/abstractruntime/core/models.py` (`WaitState`).

## Does time-based waiting (`WAIT_UNTIL`) progress automatically?

Only if **something drives the runtime**:
- `Runtime.tick(...)` will auto-unblock a due `WAIT_UNTIL` run *when called*.
- The built-in `Scheduler` provides a driver loop that polls due waits and ticks runs.

Docs: `getting-started.md`, `architecture.md`. Code: `src/abstractruntime/core/runtime.py` (`tick`), `src/abstractruntime/scheduler/scheduler.py`.

## How do I resume a waiting run?

- If you have the `WorkflowSpec`: call `Runtime.resume(workflow=..., run_id=..., wait_key=..., payload=...)`.
- If you use `create_scheduled_runtime()`: call `sr.respond(run_id, payload)` (it uses `state.waiting.wait_key`).

Docs: `getting-started.md`. Code: `src/abstractruntime/core/runtime.py`, `src/abstractruntime/scheduler/convenience.py`.

## How do I run a workflow on a schedule or on request?

Create an automation. An automation is a durable controller run that starts the target workflow as a child run (an
occurrence) on each trigger: `schedule@1` (fixed UTC intervals such as `5m` or `24h`) or `manual@1` (run now only).
`create_automation(...)` creates it, `apply_automation_command(...)` pauses, resumes, runs now, edits, stops the
current occurrence or archives it, and the host drives it like any run (`drive_automation` does this in-process).

The `Scheduler` is different: it is a driver loop that resumes due waits of existing runs. It does not create runs.

Docs: `automations.md`. Code: `src/abstractruntime/automations/*`, `src/abstractruntime/triggers/*`.

## What happens to scheduled ticks while the host is down?

When the controller next runs, the missed ticks are coalesced: one occurrence runs, for the latest due tick, and its
event payload reports the skipped range (`coalesced: {first_tick, last_tick, missed_count}`). Resuming a paused
automation never catches up: it re-arms at the first tick after now. Occurrences are created under deterministic ids
through create-if-absent, so a crash never loses one or starts one twice.

Docs: `automations.md#schedule1`, `automations.md#crash-safety`.

## Why does an automation appear as a chat in my session list?

Session views are built from turn roots: runs without a parent, except automation controllers, plus automation
occurrences. In growing mode every occurrence joins the automation's session (`automation:<id>`), so that session
reads as a chat whose turns are the occurrences. In independent mode (the default) each occurrence has its own
session. Filter the run index with `session_kind="chat,discussion"` to list only interactive sessions.

Docs: `automations.md#runs-sessions-and-history`, `api.md#sessions-and-history`.

## Do an automation's tools ask for approval?

Not by default. With `policy.tool_approval: "auto"`, creating the automation is the consent: each occurrence gets a
frozen grant for the target's tools (its `allowed_tools`, or every tool the runtime classifies). Tools the runtime
does not classify, such as third-party MCP tools, still ask, and `ask_user` questions still wait for a person. Use
`"ask"` to approve each tool batch; pending approvals appear in `pending_waits` as `tool_approval` waits.

Docs: `automations.md#tool-approval`, `automations.md#waits-on-a-person`.

## Can a discussion of an automation run change my files?

Not through its file tools. A discussion works in its own writable workspace and sees the automation's workspace as
a read-only mount (`_runtime.workspace_read_only_paths`): file tools and VisualFlow writers that target a path inside
the mount are refused, reads work. Commands and code (`execute_command`, `execute_python`) are allowed, because the
shell cannot be sandboxed; a command can therefore still change files if it chooses to, so keep tool approval on for
commands where that matters. Every later turn in the discussion session keeps the same setup, whoever starts it, and
the discussion never writes back into the automation's session, state or ledger.

Docs: `automations.md#discussions`, `automations.md#read-only-mounts`.

## Can several processes share one run store?

v1 supports one writer process per store: one process ticks, resumes and commands runs. `run_mutation_lock` serializes
writers inside that process only. Several store objects or read-only processes on one JSON run folder stay
consistent through the creation journal (`.runs_created.log`).

Docs: `automations.md#storage-guarantees`.

## Why is my `ASK_USER` answer a dict?

`Runtime.resume(..., payload=...)` always takes a **dict** payload. If the wait has a `result_key`, the runtime stores that dict into `RunState.vars` at `result_key`.  
Code: `src/abstractruntime/core/runtime.py` (`Runtime.resume`) and `src/abstractruntime/core/models.py` (`WaitState.result_key`).

Common pattern:
- resume with `{"text": "..."}` (host-side)
- read `run.vars["my_result_key"]["text"]` (node-side)

## What storage backends are included?

AbstractRuntime includes:
- in-memory: `InMemoryRunStore`, `InMemoryLedgerStore`
- filesystem: `JsonFileRunStore` (checkpoints), `JsonlLedgerStore` (append-only JSONL ledger)
- SQLite: `SqliteRunStore`, `SqliteLedgerStore`

Docs: `architecture.md`. Code: `src/abstractruntime/storage/*`.

## What must be JSON-serializable (and why)?

Everything stored in `RunState.vars` must be JSON-serializable because it is persisted as durable state.  
Code: `src/abstractruntime/core/models.py` (`RunState`) and store implementations under `src/abstractruntime/storage/`.

For large values, use:
- `ArtifactStore` references (`src/abstractruntime/storage/artifacts.py`)
- offloading wrappers (`OffloadingRunStore`, `OffloadingLedgerStore`) (`src/abstractruntime/storage/offloading.py`)

Docs: `architecture.md`.

## How do I run LLM calls and tools?

LLM and tool execution are wired via the **AbstractCore integration**:
- `EffectType.LLM_CALL`
- `EffectType.TOOL_CALLS`

Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/*`.

## Can `LLM_CALL` analyze images, audio, or files?

Yes, when the configured AbstractCore provider/model supports the media. Pass `payload.media` as a path, a media dict, an artifact ref such as `{"$artifact": "..."}`, or a list of those. The runtime keeps the effect payload JSON-safe and materializes artifact refs into temporary provider-ready files for the call.

Common remote-light media/vision/audio/music dependencies are included in the base `abstractruntime` install. Use `abstractruntime[apple]` or `abstractruntime[gpu]` only when this host should execute local inferencer stacks.
Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/effect_handlers.py`, `src/abstractruntime/integrations/abstractcore/llm_client.py`.

## How do I generate images, video, voice/audio, or music?

Use `LLM_CALL` with AbstractCore's `output` selector:

```python
{"text": "A red cube on a white table", "output": {"modality": "image", "format": "png"}}
{"text": "A logo reveal", "output": {"modality": "video", "task": "text_to_video", "provider": "mlx-gen", "model": "Wan-AI/Wan2.2-TI2V-5B-Diffusers", "format": "mp4"}}
{"text": "Hello from Runtime", "output": {"modality": "voice", "voice": "alloy", "format": "wav"}}
{"text": "Warm lo-fi piano with brushed drums", "output": {"modality": "music", "provider": "acemusic", "model": "ace-step", "format": "wav"}}
```

Generated bytes require a runtime `ArtifactStore`. The durable result contains `artifact_id` / `artifact_ref`, not inline binary data. Remote and hybrid runtimes support common AbstractCore Server endpoints for image generation, image edits, text-to-video, image-to-video, speech, music generation, transcription, and chat media. Local runtimes can use richer AbstractCore capability plugins for voice cloning, reference-guided generation, local text-to-music, and local video generation when those AbstractCore capabilities are installed.

## Does AbstractRuntime implement image, voice, music, or video engines?

No. AbstractRuntime provides the durable graph runner, checkpoint/ledger model, waits, and artifact boundary. AbstractCore provides the LLM/media generation and analysis capabilities. Image, video, voice, transcription, and music all flow through the same JSON-safe `output` selector plus artifact-backed result shape; Runtime does not implement provider engines itself.

## Where should cached session or prompt-cache state live?

Store stable cache selectors or cache configuration in runtime-visible JSON. There are two main tracks:

- best-effort session reuse: `payload.params.prompt_cache_key`, `run.vars["_runtime"]["prompt_cache"]`, or the Runtime-owned `ABSTRACTRUNTIME_PROMPT_CACHE`
- durable exact reuse: `payload.params.prompt_cache_binding` from a previously loaded bloc/KV artifact

If a binding includes `key`, Runtime uses it as the effective prompt-cache key and does not derive a competing session key. Do not store provider session objects, cache handles, clients, or warm-cache state in `RunState.vars`. AbstractCore clients/servers own those objects, and runtime correctness should still hold when a cache is cold.

Gateway-specific prompt-cache environment variables should be consumed by Gateway and passed to Runtime explicitly; Runtime does not read the Gateway env namespace directly.

Hosts can inspect, prepare, and clean up caches through `abstractruntime.integrations.abstractcore.get_abstractcore_host_facade(runtime)`, which exposes the normal prompt-cache/model-residency controls plus durable bloc helpers such as `upsert_text_bloc(...)`, `ensure_bloc_kv_artifact(...)`, `load_bloc_kv_artifact(...)`, `list_bloc_kv_artifacts(...)`, `delete_bloc_kv_artifact(...)`, and `delete_bloc(...)` without depending on the private runtime attachment directly.
Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/host_facade.py`, `src/abstractruntime/integrations/abstractcore/llm_client.py`.

## Can a host still export or import local provider prompt caches?

Yes, but treat that as **host-local operator tooling**, not the main durable
workflow memory model.

Use the Runtime host facade:
- `list_prompt_cache_exports(...)`
- `prompt_cache_export(...)`
- `prompt_cache_import(...)`

Important limits:
- this surface is **local-only**; remote and hybrid runtimes return
  `prompt_cache_local_only`
- Runtime owns the export root policy:
  - `~/.abstractruntime/prompt_cache_exports` by default
  - `<base_dir>/prompt_cache_exports` for `create_local_file_runtime(...)`
- exports are partitioned per provider/model, so the same logical export name
  can coexist cleanly across different local backends

For durable replay-safe workflow reuse, prefer `prompt_cache_binding` from
durable bloc/KV artifacts instead of host-local provider cache exports.
Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/host_facade.py`, `src/abstractruntime/integrations/abstractcore/llm_client.py`.

## Does Runtime duplicate durable bloc text? How do per-model caches relate to it?

For local runtimes, Runtime owns the bloc root and stores one durable **text snapshot** per SHA256 within that root. That bloc is the source of truth. The provider/model cache is a **derived artifact** under that bloc, not a second independent memory model.

So the intended shape is:
- one text/file bloc per content hash inside one Runtime bloc root
- zero or more derived cache artifacts, one per provider/model pair

That means the same bloc text can back several model-specific caches, but those caches are intentionally separate because provider/model-native KV formats are not portable.
Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/llm_client.py`, `../abstractcore/abstractcore/core/file_blocs.py`.

## Can I delete a specific durable bloc or prune old bloc caches?

Yes.

Use the Runtime host facade:
- `list_blocs(...)`
- `list_bloc_kv_artifacts(...)`
- `delete_bloc_kv_artifact(...)`
- `prune_bloc_kv_artifacts(...)`
- `delete_bloc(...)`

The important safety flags are:
- `dry_run=True` to preview the affected artifact or bloc set
- `clear_loaded=True` to clear matching live prompt-cache keys before deletion when Runtime can see that live state
- `force=True` only when you intentionally want to bypass the live-binding safety check

The important scope distinction is:
- `delete_bloc_kv_artifact(...)`: delete one provider/model artifact, keep the durable text bloc
- `delete_bloc(...)`: delete the durable text bloc itself and, by default, all derived KV artifacts under it

## Where should a host get provider / voice / music / vision catalogs from?

From Runtime. Use `abstractruntime.integrations.abstractcore.get_abstractcore_discovery_facade(runtime)` for
provider discovery, provider models, model capability lookup, voice/TTS/STT catalogs, music provider/model catalogs,
vision provider catalogs, and cached vision model snapshots.

These are snapshot/query reads, not durable `LLM_CALL` effects, so replay should use the recorded snapshot rather than
re-querying the current machine or server and pretending the answer is unchanged.
Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/discovery_facade.py`, `src/abstractruntime/integrations/abstractcore/discovery_queries.py`, `src/abstractruntime/integrations/abstractcore/llm_client.py`.

## Should Gateway or another host import AbstractCore comms or Telegram helpers directly?

No. For the remaining host/operator paths, use Runtime's public wrappers instead:

- `get_abstractcore_host_facade(runtime).list_email_accounts(...)`
- `...list_emails(...)`
- `...read_email(...)`
- `...send_email(...)`
- `abstractruntime.integrations.abstractcore.list_email_accounts(...)`
- `...list_emails(...)`
- `...read_email(...)`
- `...send_email(...)`
- `abstractruntime.integrations.abstractcore.telegram_facade.bootstrap_telegram_auth_from_env(...)`
- `...get_global_telegram_client(...)`
- `...stop_global_telegram_client()`
- `...send_telegram_message(...)`

Important nuance: the read/bootstrap wrappers are still **host-local**. They do not proxy through a remote Core
server, and they do not write durable Runtime history on their own. They exist so hosts can depend
on Runtime as the package boundary instead of importing `abstractcore.tools.comms_tools`,
`abstractcore.tools.telegram_tdlib`, or `abstractcore.tools.telegram_tools` directly.

For outbound sends that belong to a run, use the durable run facade instead:

- `get_abstractcore_run_facade(runtime).send_email(...)`
- `get_abstractcore_run_facade(runtime).send_telegram_message(...)`
- `get_abstractcore_run_facade(runtime).resume_tool_calls(...)` when an approval-gated or passthrough tool child run needs to continue

Those create child runs, record the send request and outcome in the ledger, and replay should show
the recorded result rather than resending the external message.

## Should a host execute image / TTS / music / STT directly for an existing run?

No. If the work is run-scoped and should become part of durable run history, the host should ask Runtime to execute it. Use `abstractruntime.integrations.abstractcore.get_abstractcore_run_facade(runtime)` and create a child run with `generate_image(...)`, `edit_image(...)`, `upscale_image(...)`, `generate_voice(...)`, `stream_voice(...)`, `generate_music(...)`, `transcribe_audio(...)`, or the lower-level `execute_llm_call(...)`.

That keeps the ledger, artifacts, and replay surface Runtime-authored instead of synthesizing history after host-side work already happened.
Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/run_facade.py`.

## Why can local media residency return `ok:false` without failing the run?

Because local media warmup is not always a meaningful reusable state. In particular, local image generation may execute through a one-shot subprocess isolation boundary, so a prior warmup cannot be reused by the next request. Runtime therefore reports unsupported local media residency explicitly instead of pretending success.

For optional residency (`required=false`), the effect still completes durably but includes `status_hint="warning"` and `degraded=true`. Unsupported local media responses also report `requires_long_lived_server=true` and a `config_hint` that points at `ABSTRACTCORE_SERVER_BASE_URL`; image generation additionally reports `execution_mode="local_one_shot_subprocess"`.
Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/effect_handlers.py`, `src/abstractruntime/integrations/abstractcore/llm_client.py`.

## What are “local / remote / hybrid” execution modes?

They refer to where LLM and tools execute:
- **Local**: in-process LLM + local tool execution
- **Remote**: HTTP to an AbstractCore server + tools typically passthrough
- **Hybrid**: remote LLM + local tools

`create_local_runtime(...)` currently uses `MultiLocalAbstractCoreLLMClient` under the hood. That client is still
local-only: it can keep multiple in-process `(provider, model)` local clients warm and route between them per request,
but it does not switch between local and remote AbstractCore backends. If you want remote model execution, use
`create_remote_runtime(...)` or `create_hybrid_runtime(...)`.

Docs: `integrations/abstractcore.md`, `../docs/adr/0002_execution_modes_local_remote_hybrid.md`. Code: `src/abstractruntime/integrations/abstractcore/factory.py`.

## What does passthrough tool mode mean?

In passthrough mode, tool calls are **not executed** in-process:
- the `TOOL_CALLS` handler returns `WAITING` with tool call details
- an external worker/operator executes the tools
- the host resumes the run with the tool results

Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/tool_executor.py` (`PassthroughToolExecutor`).

## How do I require approval before tools run?

Use `ApprovalToolExecutor` around a trusted local executor. Safe read-only/default bridge tools can execute immediately; write, command, email/WhatsApp, and unknown tools produce a durable approval wait. Resume with `{"approved": true}` to run the pending calls or `{"approved": false, "reason": "..."}` to return structured tool errors.

Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/tool_executor.py`.

## How should provider API keys be passed to a remote AbstractCore server?

Use `Authorization: Bearer <server-key>` for AbstractCore server authentication. If a request needs a per-request upstream provider key, pass `params.provider_api_key` (or legacy `params.api_key`) in the runtime payload; Runtime converts it to the `X-AbstractCore-Provider-API-Key` header. Current AbstractCore servers reject provider keys in query strings or JSON bodies for security.

Docs: `integrations/abstractcore.md`. Code: `src/abstractruntime/integrations/abstractcore/llm_client.py`.

## Does AbstractRuntime retry effects (LLM/tools)? Is it idempotent?

Retry and idempotency are controlled via `EffectPolicy`:
- idempotency keys are used to reuse prior completed results after restarts
- retry behavior is configurable (e.g. `RetryPolicy`)

Docs: `architecture.md`. Code: `src/abstractruntime/core/policy.py`, `src/abstractruntime/core/runtime.py` (effect execution + reuse).

## Is the ledger tamper-proof?

No. The built-in provenance feature is **tamper-evident** (hash chain), not signature-backed non-forgeability.

Docs: `provenance.md`. Code: `src/abstractruntime/storage/ledger_chain.py`.

## How do I stream progress updates?

If your `LedgerStore` supports subscriptions (or is wrapped with `ObservableLedgerStore`), you can subscribe in-process:
- `Runtime.subscribe_ledger(callback, run_id=...)`

Long-running generated media uses the same ledger stream. Runtime converts provider progress callbacks into `EMIT_EVENT` ledger records named `abstract.progress` with JSON-safe payloads such as `phase`, `step`, `total_steps`, `frame`, `total_frames`, and `progress`.

For the text of an answer while it is generated, register a live sink with `Runtime.set_live_delta_sink(sink)` and start the run with `_runtime.stream: True`. Live deltas never reach the ledger; see `integrations/abstractcore.md#live-token-streaming`.

Docs: `architecture.md`. Code: `src/abstractruntime/core/runtime.py` (`subscribe_ledger`), `src/abstractruntime/storage/observable.py`.

## What is “evidence capture”?

Evidence capture records durable, artifact-backed evidence for selected external-boundary tools:
- `web_search`, `fetch_url`, `execute_command`

It runs best-effort after successful `TOOL_CALLS` and requires an `ArtifactStore`.  
Docs: `evidence.md`. Code: `src/abstractruntime/evidence/recorder.py`, `src/abstractruntime/core/runtime.py` (`_maybe_record_tool_evidence`, `list_evidence`, `load_evidence`).

## What are snapshots and are they safe to restore?

Snapshots are named bookmarks of run state. Restoring a snapshot is a host-level operation (load + write back into your RunStore).  
Safety depends on whether workflow code/spec has changed since the snapshot was taken.

Docs: `snapshots.md`. Code: `src/abstractruntime/storage/snapshots.py`.

## How do WorkflowBundles (`.flow`) relate to `WorkflowSpec`?

`WorkflowSpec` is an in-memory graph of Python callables (not portable). WorkflowBundles (`.flow`) distribute **VisualFlow JSON** plus a manifest; hosts compile VisualFlow JSON into `WorkflowSpec` using the VisualFlow compiler.

Docs: `workflow-bundles.md`, `architecture.md`. Code: `src/abstractruntime/workflow_bundle/*`, `src/abstractruntime/visualflow_compiler/*`.

## How do I run the MCP worker?

Use the `abstractruntime-mcp-worker` CLI from the base Runtime install and select toolsets explicitly.

Docs: `mcp-worker.md`. Code: `src/abstractruntime/integrations/abstractcore/mcp_worker.py`.

## Where should I look for runnable examples?

- `../examples/README.md` (runnable scripts)
- `manual_testing.md` (smoke tests)


==============================================================================
# FILE: docs/troubleshooting.md
==============================================================================

# Troubleshooting

This page is for symptom-oriented fixes. For concepts and limits, see `faq.md`; for setup and examples, see
`getting-started.md`.

## Importing `abstractruntime.integrations.abstractcore` fails

Symptom:
- Importing the AbstractCore integration raises an `ImportError` about the required AbstractCore version or missing
  optional dependencies.

Checks:

```bash
python -m pip show AbstractRuntime abstractcore
python -m pip install -U abstractruntime
```

Fix:
- Install or upgrade the base Runtime package. LLM/tools integration, common
  remote-light multimodal dependencies, and the MCP worker entry point are part
  of the base install.
- The current AbstractCore integration expects `abstractcore>=2.18.0`.

Verify:

```bash
python -c "import abstractruntime.integrations.abstractcore as ac; print(ac.__all__[:5])"
```

## A run is waiting and does not continue

Symptom:
- `Runtime.tick(...)` returns `status=waiting`, and the run stays paused.

Likely causes:
- The run is waiting for `ASK_USER`, `WAIT_EVENT`, passthrough tools, or tool approval.
- `WAIT_UNTIL` needs a driver loop, or a host must call `tick(...)` again after the due time.

Fix:
- For user/event/tool waits, resume with the exact `wait_key` from `state.waiting.wait_key`.
- For time-based waits, use `create_scheduled_runtime()` or another host driver that periodically ticks due runs.

Verify:

```python
state = rt.get_state(run_id)
print(state.status, state.waiting.wait_key if state.waiting else None)
```

## Generated media is missing or too large for run state

Symptom:
- Image, voice, music, or transcription outputs fail, or binary content is not present in `RunState.vars`.

Likely causes:
- Generated binary outputs require a runtime `ArtifactStore`.
- Runtime stores generated bytes by artifact reference instead of embedding raw bytes in checkpoints or ledger records.

Fix:
- Construct the runtime with an artifact store such as `InMemoryArtifactStore` or `FileArtifactStore`.
- Read `artifact_id` / `artifact_ref` from the durable result and load the artifact from the configured store.

Docs:
- `api.md#artifacts-store-by-reference`
- `integrations/abstractcore.md#multimodal-generation`

## Remote media calls use the wrong model or fail with input-media errors

Symptom:
- Remote image/TTS/STT calls do not use the expected media model.
- Remote image generation fails when media is supplied.
- Remote transcription rejects media that is not a local file or artifact-backed temporary file.

Fix:
- Put endpoint-specific routing in the `output` selector, not in the chat model unless you intend to route chat.
- Use `output.task="image_edit"` for image edits with one source image and optional mask.
- Use exactly one audio media item for remote STT, and make sure it resolves to a local path or artifact-backed file.

Docs:
- `integrations/abstractcore.md#multimodal-generation`

## Local media residency returns `model_residency_unsupported`

Symptom:
- A local `MODEL_RESIDENCY` load for `image_generation`, `image_upscale`, `video_generation`, `text_to_video`, `image_to_video`, `tts`, `stt`, or `music_generation` returns `ok=false` with
  `code="model_residency_unsupported"`.

Meaning:
- Runtime is reporting that the current local execution topology cannot truthfully keep that media backend resident for
  later reuse.

Fix:
- Use a configured long-lived AbstractCore server for media residency, then let Runtime relay `/acore/models/*`.
- Keep local media warmup optional (`required=false`) if it is only an optimization.
- Do not infer loaded state from model defaults, catalogs, downloaded weights, or provider names.

Docs:
- `integrations/abstractcore.md#prompt-cache-control-plane-and-durable-blocs`
- `faq.md#why-can-local-media-residency-return-okfalse-without-failing-the-run`

## Prompt-cache behavior is not reused for generated media

Symptom:
- A workflow-level prompt-cache flag creates session cache keys for text/chat calls, but not for image, voice, music, or
  transcription output selectors.

Meaning:
- Runtime only auto-derives session prompt-cache keys for text/chat calls. Non-text output selectors may still carry an
  explicit `prompt_cache_binding`, but Runtime does not invent one.

Fix:
- Use `params.prompt_cache_binding` for durable exact text/chat prefix reuse.
- Treat generated media calls as separate capability executions unless the selected AbstractCore backend documents a
  task-specific cache contract.

Docs:
- `integrations/abstractcore.md#prompt-cache-control-plane-and-durable-blocs`
- `faq.md#where-should-cached-session-or-prompt-cache-state-live`

## Live replies do not stream

Symptom:
- A host registered a sink with `Runtime.set_live_delta_sink(...)`, but an answer arrives only when the call ends, or the
  sink receives only an `llm.delta_end`.

Checks:
- The run must set `_runtime.stream` to the boolean `True`; `Runtime.start` refuses any other value with a `ValueError`.
- Read the call's `llm.delta_end`: `reason: "unavailable"` carries a `detail`, and the `LLM_CALL` ledger record carries
  the same value as `metadata._runtime_observability.stream_unavailable`.

Fix, by `detail`:
- `remote_core`: remote mode does not stream; use a local or multi-local runtime for live text.
- `node_stream_off`: the node sets `params.stream: False`; remove it to stream that node.
- `structured_output`: structured and media-output calls never stream; this is expected.
- `usage_unavailable`: the provider's server cannot report token usage in streams; configure the server to accept
  `stream_options`, or accept non-streamed answers for that model.
- `tool_envelope_holdback`: the whole answer was a tool call or a hidden channel; this is expected.
- `sink_error`: your sink raised. Keep the sink fast and non-raising (put the event on a queue and return).

Verify:
- Start a run with `vars={"_runtime": {"stream": True}}` and check that the sink receives `llm.delta` events before the
  `llm.delta_end` with `reason: "completed"`.

Docs:
- `integrations/abstractcore.md#live-token-streaming`

## The previous model stays in memory after a default switch

Symptom:
- After `set_default_provider_model(...)`, the previous MLX, HuggingFace or embedding model still appears in
  `list_model_residency` and memory does not drop.

Checks:
- Read `list_model_residency` diagnostics. `pending_ejects` lists a model waiting for its running call to end.
  `last_switch_ejects` shows, per model, whether it was unloaded, kept because another owner still uses or locked it
  (`skipped` with a `reason`), or failed (`ok: false` with the remaining holders).

Fix:
- A pending model is unloaded when its call ends; no action is needed.
- A model kept for another owner stays until that owner releases it. Unlock or unload it there. An explicit
  `unload_model_residency(runtime_id=...)` frees the model from every holder in the process, including owners that
  still use it, so use it only when you intend that.
- A `skipped` reason that names the claim registry means the installed AbstractCore predates it; upgrade AbstractCore.

Docs:
- `integrations/abstractcore.md#unloading-and-switching-models`

## An automation does not fire

Symptom:
- `get_automation(...)` shows the automation, but no new occurrence appears in `list_occurrences(...)`.

Likely causes:
- Nothing drives the controller. Without a host run loop, `drive_automation` returns when the controller parks; a
  later tick needs the controller to be ticked again after its deadline.
- The automation is `paused`, `archived`, `completed` (its schedule is exhausted: `count` reached, `until` passed, or
  a one-shot schedule already fired) or `failed`.
- An occurrence is still running or waiting on a person, or is in retry backoff. Occurrences run one at a time.
- The trigger is `manual@1`, which only fires on `automation.run_now`.

Checks:

```python
from abstractruntime.automations import get_automation, list_occurrences, pending_waits

info = get_automation(runtime.run_store, automation_id)
print(info["status"], info["next_fire_at"], info["state"]["pending_occurrence"])
print(pending_waits(runtime.run_store, automation_id))
```

Fix:
- Standalone hosts: tick the controller after `next_fire_at`, then drive it:
  `runtime.tick(workflow=controller_workflow_spec(), run_id=automation_id)` and `drive_automation(runtime, automation_id)`.
- Answer a pending wait with `Runtime.resume(...)` and the payload of its `kind`, or send
  `automation.stop_current`.
- Send `automation.resume` for a paused automation. It re-arms at the next tick; `automation.run_now` fires once
  immediately.

Verify:
- `list_occurrences(...)` shows a new item, and `next_fire_at` moves to the following tick.

Docs:
- `automations.md#the-controller`, `automations.md#commands`

## A start fails with `identity_conflict`

Symptom:
- `Runtime.start(..., run_id=...)` raises `RunIdentityConflict`, or `create_automation`, `start_discussion` or
  `apply_automation_command` report `identity_conflict`.

Likely causes:
- The id (or `request_id`, or `command_id`) was already used for a different request. Ids are idempotency keys: the
  same request returns the existing run or result, a different one is refused.

Fix:
- Use a new `request_id` / `command_id` for a new request, or resend exactly the original request to get the
  existing result.

Docs:
- `api.md#runtime-start--tick--resume`, `automations.md#creating-an-automation`, `automations.md#commands`

## A start fails with `SessionAttributionError`

Symptom:
- `Runtime.start(..., session_id=...)` raises `SessionAttributionError` (`reason_code = "session_attribution_failed"`).

Likely causes:
- The run store has no run index (neither `session_kinds` nor `list_run_index`), so the runtime cannot tell whether
  the session is a discussion.
- The session is a discussion whose root cannot be validated: its runs name different roots, or the root is missing,
  has a parent, lives in another session, carries no seed, or has no `workspace_root`.

Fix:
- Use one of the built-in run stores (SQLite, JSON files, in-memory, optionally wrapped by the offloading store), or
  add `list_run_index` / `session_kinds` to your store.
- For a damaged discussion session, start a new discussion with `start_discussion(...)` instead of adding turns to it.

Docs:
- `automations.md#discussions`, `automations.md#storage-guarantees`

## Growing context or a discussion fails with `history_unavailable`

Symptom:
- An automation's admission fails, or `start_discussion` raises `SessionHistoryError` (`reason_code =
  "history_unavailable"`).

Likely causes:
- The run store has no run index; strict history needs one.
- A discussion's seed is missing, malformed or offloaded to an artifact the artifact store cannot load.
- `through_occurrence` names an occurrence the session does not hold.

Fix:
- Use a store with a run index and pass the runtime's artifact store to history reads.
- Check the occurrence number against `list_occurrences(...)`.

Docs:
- `automations.md#context-independent-or-growing`, `automations.md#runs-sessions-and-history`

## A tool is refused because the workspace is read-only

Symptom:
- A tool or VisualFlow node fails with "is inside a read-only mount" or "this workspace is read-only", or a run
  fails because a read-only `workspace_root` does not exist.

Likely causes:
- "inside a read-only mount": the run is a discussion (or lists the folder in `_runtime.workspace_read_only_paths`)
  and the tool tried to write into the automation's workspace. Reads and commands are allowed; file writes into the
  mount are refused.
- "this workspace is read-only": the run was started with `workspace_read_only: true`. Tools classified `write` or
  `exec`, tools the runtime does not classify, and file-writing nodes are refused. A read-only workspace folder is
  never created.

Fix:
- In a discussion, write into its own workspace (`workspace_root`) instead of the mount.
- Otherwise continue the work in the automation itself or in an ordinary chat session, where the workspace is
  writable.
- Check a tool's class with `tool_effect_class(name)` from
  `abstractruntime.integrations.abstractcore.tool_effects`.

Docs:
- `automations.md#read-only-mounts`, `automations.md#read-only-workspaces`

## MCP worker command is not found

Symptom:
- `abstractruntime-mcp-worker` is not available on the command line.

Fix:

```bash
python -m pip install -U abstractruntime
abstractruntime-mcp-worker --help
```

Docs:
- `mcp-worker.md`


==============================================================================
# FILE: docs/integrations/abstractcore.md
==============================================================================

# AbstractCore integration

This integration wires AbstractRuntime effects to AbstractCore so workflows can execute:
- `EffectType.LLM_CALL`
- `EffectType.TOOL_CALLS`

Implementation pointers (this repo):
- factories: `src/abstractruntime/integrations/abstractcore/factory.py`
- effect handlers: `src/abstractruntime/integrations/abstractcore/effect_handlers.py`
- tool executors: `src/abstractruntime/integrations/abstractcore/tool_executor.py`
- default toolsets (incl. comms gating): `src/abstractruntime/integrations/abstractcore/default_tools.py`

## Install

```bash
pip install abstractruntime
```

The base install includes AbstractCore 2.16.0 or newer. That is the supported baseline for the current server auth split (`Authorization` for server auth, `X-AbstractCore-Provider-API-Key` for provider overrides), generated-media contracts, image upscaling, capability catalog, prompt-cache control-plane endpoints (including session attribution via `/acore/prompt_cache/key_meta`), host memory snapshots, the host-wide loaded-model sweep used by residency listings, the process-wide MLX residency report and eject used by model unload, durable bloc prompt-cache helpers, bindings and lifecycle operations, task-aware model residency for text/image/video/TTS/STT, current tool catalog, AbstractCore's public output-selector contract, async/sync text-generation output-selector parity, video generation endpoints, the public local vision-cache catalog helper used by Runtime discovery, vision adapter discovery plus batch/LoRA media controls, and the released shared workspace/file-filter utility surface used by Runtime packaging and integration checks.

The base install also includes the remote-light media/capability plugins needed
for AbstractCore's multimodal `generate(..., output=...)` path. Local
image/video/voice/music generation still depends on configured AbstractCore
capability backends and hardware profiles:

```bash
pip install "abstractruntime[apple]"
pip install "abstractruntime[gpu]"
```

With `abstractmusic>=0.1.12`, the base music integration includes the lightweight remote ACE Music backend without local model-runtime extras. The MCP worker entrypoint is included in the base Runtime install.

## Execution modes

The factories implement three execution modes (ADR-0002):
- **Local**: in-process AbstractCore providers + local tool execution
- **Remote**: HTTP to an AbstractCore server (`/v1/chat/completions`) + tool passthrough
- **Hybrid**: remote LLM + local tool execution

Local mode currently uses `MultiLocalAbstractCoreLLMClient` as the built-in LLM router. Despite the name, it is not a
local+remote combo client: it routes among multiple in-process local `(provider, model)` clients and keeps them warm in
the current process. Remote model execution is a separate topology exposed through `create_remote_runtime(...)` and
`create_hybrid_runtime(...)`.

Factory functions (exported from `abstractruntime.integrations.abstractcore`):
- `create_local_runtime(...)`
- `create_remote_runtime(...)`
- `create_hybrid_runtime(...)`

Runtime stays explicit at the boundary: Gateway/hosts construct these clients with the Core server URL, Core server auth headers, provider/model defaults, retry policy, tool executor, and artifact store they intend to use. Runtime does not read `ABSTRACTGATEWAY_*` environment variables directly and does not reinterpret Gateway bearer tokens as Core server tokens or provider keys. Gateway-owned config should be consumed by Gateway, then passed to Runtime through explicit run state, effect payloads, constructor arguments, or Runtime-owned environment variables.

## Minimal LLM workflow

```python
from abstractruntime import Effect, EffectType, StepPlan, WorkflowSpec
from abstractruntime.integrations.abstractcore import create_local_runtime


def ask_model(run, ctx):
    return StepPlan(
        node_id="ask_model",
        effect=Effect(
            type=EffectType.LLM_CALL,
            payload={
                "prompt": "Answer in one sentence: what is durable workflow state?",
                "params": {"temperature": 0.0, "max_tokens": 128},
            },
            result_key="llm",
        ),
        next_node="done",
    )


def done(run, ctx):
    llm = run.vars.get("llm") or {}
    return StepPlan(node_id="done", complete_output={"answer": llm.get("content")})


workflow = WorkflowSpec(
    workflow_id="abstractcore_llm_demo",
    entry_node="ask_model",
    nodes={"ask_model": ask_model, "done": done},
)

rt = create_local_runtime(provider="ollama", model="qwen3:4b")
run_id = rt.start(workflow=workflow)
state = rt.tick(workflow=workflow, run_id=run_id)
print(state.output)
```

## `LLM_CALL` payload (recommended shape)

`Effect(type=EffectType.LLM_CALL, payload=...)`

```json
{
  "prompt": "...",
  "request": {
    "text": "...",
    "messages": [{"role": "user", "content": "..."}],
    "media": ["path/or/artifact-ref"]
  },
  "text": "optional text alias, useful for TTS",
  "messages": [{"role": "user", "content": "..."}],
  "system_prompt": "...",
  "media": ["path/or/artifact-ref"],
  "output": {"modality": "text|image|video|voice|sound|music", "task": "optional"},
  "tools": [{"name": "...", "description": "...", "parameters": {...}}],
  "params": {
    "temperature": 0.0,
    "max_tokens": 256,
    "base_url": null
  }
}
```

Notes:
- `request` is the lower-level Core semantic request shape. Legacy top-level `prompt`, `text`,
  `messages`, and `media` fields still work and normalize to the same Core request contract.
- Remote mode supports per-request dynamic routing by forwarding `params.base_url` to the AbstractCore server request body (`src/abstractruntime/integrations/abstractcore/llm_client.py`).
- Remote mode sends per-request provider key overrides from `params.api_key` / `params.provider_api_key` as `X-AbstractCore-Provider-API-Key` headers. Server/master auth should be supplied separately through the client's configured headers, usually `Authorization: Bearer <ABSTRACTCORE_SERVER_API_KEY>`.
- Local mode treats `base_url` and provider API keys as provider-construction concerns. `MultiLocalAbstractCoreLLMClient` can construct a per-call client when a host injects `params.base_url` plus `params.api_key` or `params.provider_api_key` (for example from a Gateway provider endpoint profile), then strips those fields before calling the provider.
- `media` accepts one item or a list. Durable artifact refs such as `{"$artifact": "...", "filename": "speech.wav"}` are materialized to temporary files for AbstractCore and never stored as raw bytes in `RunState`.
- `output` may be top-level or inside `params`; top-level `outputs` is accepted as a runtime alias for AbstractCore's `output`.
- `output.tags`, when present, are merged into the generated artifact metadata. Runtime metadata such as `run_id` and `tags` is used by AbstractRuntime's ArtifactStore boundary and is not forwarded as provider-specific generation kwargs.
- Host-supplied run defaults such as `run.vars["_runtime"]["provider"]` and `run.vars["_runtime"]["model"]` are persisted as JSON-safe routing metadata; provider clients, auth objects, downloaded model handles, and server sessions are not durable runtime state.

## Execution controls and local concurrency

For text/chat calls, pass `params.thinking` as a boolean or reasoning level and
`params.speculation` as `False`, `True`, or a Core speculation object, for example
`{"mode": "native_mtp", "num_draft_tokens": 3, "require_acceleration": true}`.
Explicit `False` overrides the loaded provider's default; omission leaves that
default available. Core owns support validation and actual inference behavior.
For native MLX, configure the head and `mlx_batching=True` when constructing/loading
the provider; a per-call parameter alone does not enable an unloaded head/scheduler.

Local Runtime clients allow concurrent submission only when the loaded MLX instance
explicitly advertises safe scheduling. Ordinary instances retain serialization
through completion, including streamed iteration. Scheduler admission is not a
promise that every request batches: MTP uses fixed cohorts, and separate prefix-cache
managers remain separate scheduling groups. Token callbacks on one client can
interleave; use request results for attribution.

Remote text/chat clients forward these controls and expose Core's returned
`execution`, `speculation`, `performance`, and `prompt_cache` fields in result
`metadata`, without replacing Runtime-owned request/trace provenance. Remote chat
still returns an aggregated response. These controls require a compatible Core
backend; they do not implement HF/GGUF native MTP by themselves.

Set `_runtime.speculation` for a run-wide preference. Subworkflows, Agent loops,
delegated children and structured-output follow-up calls inherit it unless an
explicit call/node/child setting overrides it. Missing or `None` inherits; `False`
means Off and survives every boundary. Dictionaries are copied between scopes.

Core owns the configured default at `input.text.options.speculation` in capability
routes (`output.text` is the same route). Fresh Core configurations use depth 2
where native MTP is supported; existing configurations are preserved. Applications
should leave the control unset to follow that default. Scoped Core configuration
is supplied before local provider construction, so cached head preparation uses
the correct policy. Remote discovery queries the actual Core execution host and
does not substitute local model-registry assumptions.

## Runtime grounding

AbstractRuntime records per-call grounding as structured response metadata under `metadata.runtime_grounding`. The current fields include local datetime, timezone when detectable, country, source, whether prompt injection occurred, and an optional user identity when supplied by trace metadata or local environment.

For text/chat LLM calls only, the same grounding is rendered into the current user turn as a tagged runtime envelope:

```text
<runtime_metadata>{"country":"FR","local_datetime":"2026-05-13T18:00:00+02:00"}</runtime_metadata>
hello
```

This makes time/location/user context visible to the LLM without mutating the durable human message into a natural-language prefix. If a model echoes the runtime-owned envelope, AbstractRuntime removes that envelope from user-facing response text while preserving `metadata.runtime_grounding` for audit.

Direct media requests, including image generation, TTS, and transcription, do not receive prompt-injected grounding. They still receive trace headers/tags for observability and artifact ownership, but TTS `input` and image prompts remain the literal text supplied by the workflow.

## Multimodal generation

AbstractRuntime forwards AbstractCore's unified `generate(..., output=...)` selector and normalizes multimodal responses into JSON-safe, artifact-backed results.

Generate an image:

```python
Effect(
    type=EffectType.LLM_CALL,
    payload={
        "prompt": "A red ceramic mug on a white table.",
        "output": {"modality": "image", "format": "png", "width": 1024, "height": 1024},
    },
    result_key="image_result",
)
```

Generate speech:

```python
Effect(
    type=EffectType.LLM_CALL,
    payload={
        "text": "Hello from AbstractRuntime.",
        "output": {"modality": "voice", "voice": "coral", "format": "wav"},
    },
    result_key="speech_result",
)
```

Generate music:

```python
Effect(
    type=EffectType.LLM_CALL,
    payload={
        "text": "Warm lo-fi piano with brushed drums.",
        "output": {"modality": "music", "provider": "acemusic", "model": "ace-step", "format": "wav"},
    },
    result_key="music_result",
)
```

Generate video:

```python
Effect(
    type=EffectType.LLM_CALL,
    payload={
        "prompt": "Glowing data streams converge into a geometric logo.",
        "output": {
            "modality": "video",
            "task": "text_to_video",
            "provider": "mlx-gen",
            "model": "Wan-AI/Wan2.2-TI2V-5B-Diffusers",
            "format": "mp4",
            "num_frames": 41,
            "fps": 24,
            "steps": 10,
        },
    },
    result_key="video_result",
)
```

Image to video:

```python
Effect(
    type=EffectType.LLM_CALL,
    payload={
        "prompt": "Add a slow camera orbit.",
        "media": {"$artifact": "source_image_artifact_id", "type": "image", "role": "source"},
        "output": {
            "modality": "video",
            "task": "image_to_video",
            "provider": "mlx-gen",
            "model": "Wan-AI/Wan2.2-TI2V-5B-Diffusers",
            "format": "mp4",
        },
    },
    result_key="video_result",
)
```

Transcribe/analyze audio:

```python
Effect(
    type=EffectType.LLM_CALL,
    payload={
        "media": {"$artifact": "audio_artifact_id", "filename": "speech.wav"},
        "output": "text",
    },
    result_key="transcript",
)
```

Generated binary media requires a runtime `ArtifactStore` and is stored there. The persisted result contains artifact references:

```json
{
  "outputs": {
    "image": [
      {
        "modality": "image",
        "task": "image_generation",
        "artifact_id": "...",
        "artifact_ref": {"$artifact": "...", "content_type": "image/png"}
      }
    ]
  }
}
```

LLM-call results that flow through the Core request/output path also expose a replay-safe
`metadata._runtime_resolved_action` summary. `history_bundle` collects those into top-level
`resolved_actions`, so another client can replay which capability family, task, normalized
request/output summary, and effective route actually ran.

Media-only normalized results distinguish orchestration identity from the actual media backend:

- `runtime_provider` / `runtime_model`: the runtime-side orchestration identity, when relevant
- `media_provider` / `media_model`: the actual image/video/voice/music backend identity surfaced from the generated output

For local one-shot subprocess image generation, runtime metadata also records `execution_mode="local_one_shot_subprocess"`.

Long-running generated media may expose provider progress callbacks. Runtime offers a transient `on_progress` callback during `LLM_CALL` execution and persists each callback as an `EMIT_EVENT` ledger record named `abstract.progress`. The callback itself is never stored in the effect payload or run vars.

**It is not carried in the payload either.** `Effect.payload` is JSON to every consumer downstream (ledger, SSE stream, host handlers that `json.dumps` it), so the callback travels BESIDE the effect: `Runtime._execute_effect_with_retry` installs it on the `contextvars.ContextVar` in `abstractruntime/core/progress_channel.py` for the duration of the handler call, and `make_llm_call_handler` reads it there into its own params dict — which is what becomes provider kwargs. An explicit `params["on_progress"]` from the caller always wins; an effect with no progress channel is offered `None`, never the previous effect's callback. A handler that hands work to a bare thread must carry the callback object (it does — `params` holds it), not re-read the ContextVar.

### Text phase feedback (prefill vs generation)

The callback is injected for **every** `LLM_CALL`, not only generated-media ones. AbstractCore decides per PROVIDER whether it has an honest signal (`BaseProvider.supports_text_progress_events()`, default `False`); a provider that has none never calls back, and that run's ledger gains no progress records at all. Providers that do (today: every MLX lane) report the prefill→generation boundary, so a client can replace "Thinking…" with what the call is actually doing.

Text phase payloads carry `kind: "llm"` — the discriminator that separates them from the media shapes on the same event name — plus `phase` (`prefill` | `generate` | `complete`), `prompt_tokens`, `cached_tokens`, `fed_tokens`, `generated_tokens`, `ttft_s`, `tokens_per_second` and `elapsed_s`, on top of the runtime identity (`run_id`, `workflow_id`, `node_id`, `step_id`, `idempotency_key`, `attempt`) every progress payload gets. Keys with no measured value are absent, never zero.

One real record, from a `basic-agent` chat turn on `mlx/Qwen3.5-4B-4bit`:

```json
{"run_id": "6ef69750-…", "node_id": "reason", "status": "completed",
 "effect": {"type": "emit_event", "payload": {"name": "abstract.progress", "scope": "run",
   "payload": {"run_id": "6ef69750-…", "workflow_id": "visual_react_agent_basic-agent_0_0_4_81795ea9_node-2",
               "node_id": "reason", "step_id": "1cdb2c55-…", "attempt": 1,
               "kind": "llm", "phase": "prefill", "event_index": 0, "elapsed_s": 0.045,
               "provider": "mlx", "model": "mlx-community/Qwen3.5-4B-4bit",
               "prompt_tokens": 808, "cached_tokens": 690, "fed_tokens": 118,
               "generated_tokens": 0}}},
 "idempotency_key": "system:progress:1cdb2c55-…:c9c572d067aa"}
```

**Cost.** AbstractCore, not Runtime, bounds the volume: `prefill`, the first-token event and the terminal `complete` always fire, cadence events no closer than 0.5 s apart, for the whole call, never capped by count (ADR-0026). A 3.7 s answer produced 9 records. Runtime appends each already-completed with a uuid suffix, so no progress tick re-parses the ledger.

**Scope.** These records are `scope: "run"` and land in the ledger of the run that made the call. In a chat bundle the `llm_call` runs inside the Agent node's SUBWORKFLOW, so a client sees them only once it follows child ledgers — the same requirement that already applies to the `abstract.status` "Thinking…" event (`result.wait.details.sub_run_id` on the parent's waiting record).

Remote runtimes support chat media by sending OpenAI-compatible data URL content arrays to AbstractCore Server. They also support image generation (`/v1/images/generations`), image edits (`/v1/images/edits` or `/{provider}/v1/images/edits`), image upscaling (`/v1/images/upscale` or `/{provider}/v1/images/upscale`), text-to-video (`/v1/videos/generations`), image-to-video (`/v1/videos/edits` or `/{provider}/v1/videos/edits`), TTS (`/v1/audio/speech`), music generation (`/v1/audio/music`), and STT (`/v1/audio/transcriptions`) with the same artifact-backed result shape. The Runtime/Core request surface forwards task-specific media controls including `count`/`n`, `seeds`, ordered `lora_adapters`, and video `flow_shift`. Remote media endpoint calls do not inherit the chat model by default; pass an output-specific `model` only when you want a remote provider/model instead of the server's configured capability default. Remote STT requires exactly one audio media item that resolves to a local file path or artifact-backed temporary file. Remote image edits, image upscaling, and image-to-video require one source image media item resolving to a local path or artifact-backed temporary file. For voice clone/register or reference-guided TTS, use local execution so AbstractCore can use its in-process capability dispatcher. Runtime does not import `abstractmusic` directly; local music support comes through the configured AbstractCore capability stack.

Remote multimodal generation currently supports one `output` selector per `LLM_CALL`. Hybrid runtimes use the same remote LLM/media path as remote mode while executing tools locally. Local runtimes can use AbstractCore's in-process multimodal dispatcher for richer capability plugin behavior.

Local media residency is intentionally explicit when unsupported. `MODEL_RESIDENCY` results for local `image_generation`, `image_upscale`, `video_generation`, `text_to_video`, `image_to_video`, `tts`, `stt`, and `music_generation` return:

- `code="model_residency_unsupported"`
- `requires_long_lived_server=true`
- `config_hint` pointing to `ABSTRACTCORE_SERVER_BASE_URL`

Image/video-generation residency responses also include `execution_mode="local_one_shot_subprocess"` because local generated media can be isolated into one-shot workers unless a long-lived Core server owns the media backend.

When the workflow marks residency as optional (`required=false`), the effect still completes durably but includes `status_hint="warning"` and `degraded=true` so hosts can render the no-op honestly.

Remote auth example:

```python
from abstractruntime.integrations.abstractcore import create_remote_runtime

rt = create_remote_runtime(
    server_base_url="http://127.0.0.1:8000",
    model="openai/gpt-4o-mini",
    headers={"Authorization": "Bearer server-master-key"},
)
```

Then pass a per-request upstream provider key through `params.provider_api_key` only when the AbstractCore server is acting as a provider proxy for that request:

```python
payload = {
    "prompt": "Summarize this in one sentence.",
    "params": {
        "provider_api_key": "sk-provider-key",
        "base_url": "http://127.0.0.1:1234/v1",
    },
}
```

## Live token streaming

A host can show an answer while it is being generated. Streaming is opt-in per run and needs two things:

1. The run asks for it: `run.vars["_runtime"]["stream"] = True`. The value must be a boolean (`False` is an explicit off); `Runtime.start` refuses anything else, such as the string `"true"`, with a `ValueError`.
2. The host registers a sink on the runtime: `runtime.set_live_delta_sink(sink)`.

When both hold, every `LLM_CALL` of the run streams from the provider and the runtime calls `sink(event)` with plain dicts:

```python
{"kind": "llm.delta", "run_id": "…", "parent_run_id": None, "node_id": "reason",
 "call_id": "<step_id>", "seq": 0, "text": "Hello", "channel": "content"}
{"kind": "llm.delta_end", "run_id": "…", "parent_run_id": None, "node_id": "reason",
 "call_id": "<step_id>", "seq": 5, "reason": "completed"}
```

- `parent_run_id` is the emitting run's parent (`None` for a root run), so a host can show a child run's text in its root conversation.
- `call_id` is the `step_id` of the `LLM_CALL` ledger record that later holds the final answer, so a client can replace its live text with the durable result.
- `seq` counts the events of one call from 0, `llm.delta_end` included.
- `channel` is `content` (the answer) or `reasoning` (the model's thinking). Reasoning arrives from the provider's reasoning stream, and inline `<think>…</think>` markup is split out of the content live, so clients can hide or fold it. The markup itself is never sent.
- `reason` is `completed`, `failed`, `cancelled` or `unavailable`. Every live fragment of a call is sent before the call's durable record is written (the last batch is flushed first, and later fragments are dropped); every call then ends with exactly one `llm.delta_end`, sent after that record is in the ledger. Each retry attempt is its own call with its own `call_id`; so is the rare re-run of an attempt interrupted by a stray kill (`<step_id>:reinvoke`), whose interrupted first run ends `cancelled` with `detail: "reinvoked"` (the run goes on; a client can say "reply restarted" rather than "stopped").
- Tool-call markup a local model writes as text (`<tool_call>`, `<|tool_call|>`, `<function_call>`, …) never reaches the live text. AbstractCore already removes it from streamed content; the runtime also stops the live content of a call at the first such marker. The final record carries the tool calls.
- Harmony transcripts (gpt-oss) are split live: the `final` channel streams as content, `analysis` as reasoning, and a tool call (`commentary to=functions.…`) is held back. The framing tokens are never sent.
- Text is coalesced: the first fragment is sent immediately, later fragments are grouped over about 40 ms.

```python
events = []
runtime.set_live_delta_sink(events.append)
run_id = runtime.start(workflow=wf, vars={"_runtime": {"stream": True}})
runtime.tick(workflow=wf, run_id=run_id)
```

**The ledger is unchanged.** Deltas are never persisted. A streamed run records the same `LLM_CALL` result as a non-streamed one (`content`, `reasoning`, `tool_calls`, `usage`, `raw_response`), apart from the `stream` flag in the recorded provider request and the timings the stream measures (`gen_time`, `ttft_ms`). In streaming mode `raw_response` is the provider's terminal chunk.

**Child runs inherit the switch.** Subworkflows, and the delegated children of AbstractAgent loops, receive the parent's `_runtime.stream` unless they set their own value.

**When a call does not stream, it says why.** The call runs non-streamed with the same complete answer, its `llm.delta_end` has `reason: "unavailable"` and a `detail`, and the same detail is recorded on the `LLM_CALL` record as `metadata._runtime_observability.stream_unavailable`:

| `detail` | Meaning |
| --- | --- |
| `usage_unavailable` | The provider cannot report token usage when streaming (for example an OpenAI-compatible server that rejects `stream_options`). The call that lacked usage is reported on its `LLM_CALL` record; if its text had already streamed, its `llm.delta_end` still says `completed` (the reply did stream). Such a record carries `metadata.usage_estimated: true`: without usage, the aborted-generation check judges a streamed answer from its `finish_reason` and text (no terminal reason, or a short reply ending with ":" and no tool call). Later calls on that model run non-streamed only when the provider also reports that its server rejects usage in streams; they stream again as soon as a streamed answer brings usage back (one call in ten streams to re-check). A server that rejects the option but sends usage anyway (LM Studio) keeps streaming. |
| `prompt_cache_unavailable` | Reserved for a provider whose streamed answers do not carry `metadata.prompt_cache` while its non-streamed answers do. No provider is in this case with a current AbstractCore. |
| `structured_output` | Structured or media-output calls are never streamed. |
| `provider_cannot_stream` | The provider answered in one piece. |
| `remote_core` | Remote mode: the AbstractCore server call is not streamed. |
| `node_stream_off` | The `LLM_CALL` payload sets `params.stream: False`. |
| `sink_error` | The host sink raised; the rest of the call was not delivered live. |
| `reinvoked` (with `reason: "cancelled"`) | A stray kill interrupted the call and the runtime re-runs it under a new `call_id`; the run continues and answers. |
| `tool_envelope_holdback` | The whole answer was a tool call (or a channel that is not shown), so no answer text was streamed. |

A streamed call records the same `usage`, `raw_response` and `metadata.prompt_cache` as a non-streamed one; where it cannot, the call does not stream.

**AbstractCore dependency for MLX.** MLX calls with a prompt-cache key stream with a complete record only on an AbstractCore release that puts the prompt-cache record, usage and finish reason on the last streamed chunk. With an older AbstractCore these calls still stream, but their record has no `metadata.prompt_cache`; upgrade AbstractCore to keep streamed and non-streamed records identical.

**Harmony output on raw lanes.** On a provider lane where AbstractCore itself buffers harmony (gpt-oss) output until the end of the answer, the text reaches the sink in one piece when the call ends. Lanes that separate reasoning on the server side (for example LM Studio) stream normally.

**Writing a sink.** The sink is called from the provider's thread, in `seq` order for a given call. Keep it fast: put the event on a queue and return. A sink that raises is switched off for the rest of that call and the call itself continues normally.

## `TOOL_CALLS` payload

```json
{
  "tool_calls": [
    {
      "name": "tool_name",
      "arguments": {"x": 1},
      "call_id": "optional (provider id)",
      "runtime_call_id": "optional (stable; runtime-generated)"
    }
  ],
  "allowed_tools": ["optional allowlist (order-insensitive)"]
}
```

Notes:
- `runtime_call_id` is generated/normalized by the runtime for durability (`src/abstractruntime/core/runtime.py`).
- In remote/passthrough mode, a host/worker boundary can use `runtime_call_id` as an idempotency key.

## Tool execution modes

Tool execution is controlled by the configured `ToolExecutor` (`src/abstractruntime/integrations/abstractcore/tool_executor.py`):

- **Executed (trusted local)**: use `MappingToolExecutor` (recommended) or `AbstractCoreToolExecutor`.
- **Passthrough (untrusted/server/edge)**: use `PassthroughToolExecutor`.
  - The `TOOL_CALLS` handler returns a durable `WAITING` run state.
  - The host executes the tool calls externally and resumes the run with results (`Runtime.resume(...)` / `Scheduler.resume_event(...)`).
- **Approval-gated local execution**: wrap a trusted executor with `ApprovalToolExecutor`.
  - Safe read-only/default bridge tools can run immediately.
  - Riskier or unknown tools return a durable `approval_required` wait.
  - A thin client can resume with `{"approved": true}` to execute the approved calls in-runtime, or `{"approved": false, "reason": "..."}` to return structured tool errors.

Approval example:

```python
from abstractruntime.integrations.abstractcore import (
    ApprovalToolExecutor,
    MappingToolExecutor,
    ToolApprovalPolicy,
    create_local_runtime,
)


def write_file(*, path: str, content: str):
    with open(path, "w", encoding="utf-8") as f:
        f.write(content)
    return {"path": path, "bytes": len(content.encode("utf-8"))}


tools = ApprovalToolExecutor(
    delegate=MappingToolExecutor({"write_file": write_file}),
    policy=ToolApprovalPolicy(),
)
rt = create_local_runtime(provider="ollama", model="qwen3:4b", tool_executor=tools)
```

## Workspace-scoped tools

When a run sets `workspace_root` in its vars, file tools (`read_file`, `write_file`, `edit_file`, `list_files`, `search_files`, `skim_files`, `skim_folders`, …) resolve paths against that folder, and the shell tools (`execute_command`, `shell_exec`, `local_helper_start`) start there. The policy is read from run vars (`src/abstractruntime/integrations/abstractcore/workspace_scoped_tools.py`):

| Run var | Meaning |
| --- | --- |
| `workspace_root` | Base folder for relative paths and the shell's starting folder. Relative roots resolve against `ABSTRACT_WORKSPACE_BASE_DIR` when set. |
| `workspace_access_mode` | `workspace_only` (default: paths stay under the root), `workspace_or_allowed` (also under `workspace_allowed_paths`), or `all_except_ignored` (anywhere except `workspace_ignored_paths`). |
| `workspace_allowed_paths` | Extra folders a `workspace_or_allowed` run may use. |
| `workspace_ignored_paths` | The operator's exclusions; always refused. |
| `workspace_builtin_deny_prefixes` | The host's own protected folders (for example a gateway's data folder and credential folders), as path prefixes. Anything under one is refused. |
| `workspace_builtin_allow` | Exceptions inside the host's protected folders (for example the run's own folder inside the data folder). The operator's `workspace_ignored_paths` still win over them. |

Lists accept a JSON array, a JSON array pasted as a string, or newline-separated entries.

- When an `LLM_CALL` offers tools and the run has a workspace scope, the runtime appends a short description of the scope to the system prompt: the root, the access mode, the authorized roots and the operator's `workspace_ignored_paths`. The host's built-in deny prefixes and allow entries are enforced but never described to the model, so the system prompt stays identical from turn to turn while the host's folders change, and the prompt cache stays warm. A refused path is reported as `Path is not accessible (protected by the host)` without listing the host's rules.
- Child runs (subworkflows and agent delegation) inherit all six vars unless they set their own; VisualFlow file and document nodes (`read_file`, `write_file`, `read_pdf`, `write_pdf`, `write_docx`, `write_chart`) apply the same scope.
- Shell tools are pinned to their starting folder only. A command can still reach other paths (`cd ..`, absolute paths); the workspace scope is a file-tool policy, not a sandbox.

## Prompt-cache control plane and durable blocs

AbstractRuntime's AbstractCore integration exposes a public host-control facade for prompt-cache, durable bloc/KV prompt-cache operations, and model-residency operations:

- `get_abstractcore_host_facade(runtime)`
- `AbstractCoreHostFacade`
- `get_prompt_cache_capabilities(...)`
- `get_prompt_cache_stats(...)`
- `prompt_cache_set(...)`
- `prompt_cache_update(...)`
- `prompt_cache_fork(...)`
- `prompt_cache_clear(...)`
- `prompt_cache_prepare_modules(...)`
- `list_prompt_cache_exports(...)`
- `prompt_cache_export(...)`
- `prompt_cache_import(...)`
- `upsert_text_bloc(...)`
- `get_bloc_record(...)`
- `list_blocs(...)`
- `get_bloc_kv_manifest(...)`
- `ensure_bloc_kv_artifact(...)`
- `load_bloc_kv_artifact(...)`
- `list_bloc_kv_artifacts(...)`
- `delete_bloc_kv_artifact(...)`
- `prune_bloc_kv_artifacts(...)`
- `delete_bloc(...)`
- `get_model_residency_capabilities(...)`
- `list_model_residency(...)`
- `load_model_residency(...)`
- `unload_model_residency(...)`
- `lock_model_residency(...)`
- `unlock_model_residency(...)`
- `get_context_estimate(...)`
- `get_memory_snapshot(...)`
- `list_session_prompt_caches(...)`
- `clear_session_prompt_caches(...)`

Behavior by execution mode:

- **Local** (`MultiLocalAbstractCoreLLMClient` / `LocalAbstractCoreLLMClient`): delegates to the in-process AbstractCore provider and normalizes responses into the same JSON-safe shape used by the endpoint.
- **Remote / Hybrid** (`RemoteAbstractCoreLLMClient`): proxies `/acore/prompt_cache/*`, `/acore/models/*`, and `/acore/memory` on the configured AbstractCore server.
  - When the remote target is the multi-provider AbstractCore server proxy rather than a direct AbstractEndpoint, callers can forward upstream `base_url` through these prompt-cache methods. Per-request provider key overrides supplied as `api_key` / `provider_api_key` are converted to `X-AbstractCore-Provider-API-Key` headers, not request bodies or query strings.
  - For durable bloc/KV methods, `base_url` takes precedence over local loaded-runtime selectors. Runtime omits `provider`, `model`, and `runtime_id` when `base_url` is supplied so Core takes the upstream endpoint branch cleanly.

Contract notes:

- Capability discovery is explicit: callers can branch on `capabilities.mode` (`none`, `keyed`, `local_control_plane`) and `supports_*` flags.
- Unsupported operations return structured payloads with `supported=false`, `operation`, `code`, and `capabilities`.
- When a provider reports `mode=local_control_plane` (for example MLX, or GGUF models whose llama.cpp chat format has an exact cached renderer), the runtime can maintain a compartmentalized `system | tools | history` cache path automatically.
- When a provider reports `mode=keyed`, the runtime still forwards stable `prompt_cache_key`s but skips module preparation/fork/update orchestration.
- This surface is intentionally host-oriented; the runtime effect handlers still only use prompt caching during LLM execution, and gateway/CLI hosts manage prompt caches and durable bloc/KV artifacts through the public facade instead of reaching through to provider internals.
- Automatic per-session prompt-cache keys are enabled by `run.vars["_runtime"]["prompt_cache"]`, `LLM_CALL.params.prompt_cache_key`, or the Runtime-owned `ABSTRACTRUNTIME_PROMPT_CACHE` process default. Gateway-specific prompt-cache env vars should be translated by Gateway into `_runtime.prompt_cache`.
- Durable exact reuse uses `LLM_CALL.params.prompt_cache_binding`. If a binding includes `key`, Runtime adopts it as the effective cache key, rejects mismatches before provider execution, and skips auto-derived session-key injection for that call.
- Automatic prompt-cache key derivation is text/chat-only. Non-text output selectors such as image, voice, music, and transcription may carry an explicit `prompt_cache_binding`, but Runtime does not derive a session cache key for them.
- Local Runtime owns the bloc store root policy:
  - default local root: `~/.abstractruntime/blocs`
  - default file-runtime root: `<base_dir>/blocs`
  - explicit `bloc_root_dir=...` overrides are allowed when hosts need a different root
- The three prompt-cache tracks are distinct:
  - session prompt cache: best-effort volatile reuse
  - durable bloc prompt cache: exact reuse through bloc/KV/binding
  - host-local prompt-cache export/import admin: optional operator tooling
    around live local provider cache state, separate from durable workflow
    memory
- `get_memory_snapshot(...)`, `list_session_prompt_caches(...)`, `clear_session_prompt_caches(...)`, `lock_model_residency(...)`, `unlock_model_residency(...)`, and `get_context_estimate(...)` are optional in the LLM-client contract. A configured client that does not implement one still binds to the facade; the facade answers `{"ok": false, "supported": false, "operation": ...}` for that call instead of failing at construction time.
- `lock_model_residency(...)`, `unlock_model_residency(...)`, and `get_context_estimate(...)` accept an optional payload mapping and/or keyword arguments. Keyword arguments merge over a copy of the payload and win on key conflicts; the first positional argument is only ever the payload mapping, and a non-mapping positional raises `TypeError`.

Host-side prompt-cache example:

```python
from abstractruntime.integrations.abstractcore import (
    create_local_runtime,
    get_abstractcore_host_facade,
)

rt = create_local_runtime(provider="mlx", model="mlx-community/Qwen3-4B-4bit")
facade = get_abstractcore_host_facade(rt)

caps = facade.get_prompt_cache_capabilities()
if caps.get("capabilities", {}).get("supports_prepare_modules"):
    facade.prompt_cache_prepare_modules(
        namespace="assistant",
        modules=[
            {"module_id": "system", "system_prompt": "You are concise."},
            {"module_id": "tools", "tools": [{"name": "read_file", "parameters": {"type": "object"}}]},
        ],
    )
```

Host-side durable bloc example:

```python
from abstractruntime.integrations.abstractcore import (
    create_local_file_runtime,
    get_abstractcore_host_facade,
)

rt = create_local_file_runtime(
    base_dir="./runtime-data",
    provider="mlx",
    model="mlx-community/Qwen3-4B-4bit",
)
facade = get_abstractcore_host_facade(rt)

record = facade.upsert_text_bloc(
    path="assistant/system.txt",
    content="Long-lived system prompt or memory text",
)
artifact = facade.ensure_bloc_kv_artifact(
    provider="mlx",
    model="mlx-community/Qwen3-4B-4bit",
    sha256=record["sha256"],
)
loaded = facade.load_bloc_kv_artifact(
    provider="mlx",
    model="mlx-community/Qwen3-4B-4bit",
    sha256=record["sha256"],
)

binding = loaded["artifact"]["prompt_cache_binding"]
```

Host-local prompt-cache export/import example:

```python
saved = facade.prompt_cache_export(
    name="orbit-cache",
    key="sess:orbit",
    q8=True,
)
listed = facade.list_prompt_cache_exports()
loaded_cache = facade.prompt_cache_import(
    name="orbit-cache",
    key="loaded:orbit",
    clear_existing=True,
)
```

Host-local export/import contract:

- This surface is **local-only**. Remote and hybrid runtimes return structured
  `prompt_cache_local_only` payloads instead of proxying host filesystem state
  through Core Server.
- Runtime owns the export root policy:
  - default local root: `~/.abstractruntime/prompt_cache_exports`
  - default file-runtime root: `<base_dir>/prompt_cache_exports`
  - explicit `prompt_cache_export_root_dir=...` overrides are allowed when a
    host needs a different local catalog root
- Exports stay partitioned by provider/model under that root, so the same
  logical export name can coexist safely across different local backends.
- This is a **secondary operator/admin feature**, not the primary durable app
  contract. For replay-safe exact reuse inside workflows, prefer
  `prompt_cache_binding` from durable bloc/KV artifacts instead.

Host-side durable bloc lifecycle example:

```python
records = facade.list_blocs()
artifacts = facade.list_bloc_kv_artifacts(bloc_id=record["record"]["bloc_id"])

# Preview a safe delete first.
preview = facade.delete_bloc_kv_artifact(
    bloc_id=record["record"]["bloc_id"],
    artifact_path=artifacts["artifacts"][0]["artifact_path"],
    dry_run=True,
)

# Remove one derived KV artifact but keep the durable text bloc.
facade.delete_bloc_kv_artifact(
    bloc_id=record["record"]["bloc_id"],
    artifact_path=artifacts["artifacts"][0]["artifact_path"],
    clear_loaded=True,
)

# Remove the whole bloc and all derived artifacts under it.
facade.delete_bloc(
    bloc_id=record["record"]["bloc_id"],
    clear_loaded=True,
)
```

Then use the binding in a normal runtime `LLM_CALL`:

```python
Effect(
    type=EffectType.LLM_CALL,
    payload={
        "prompt": "Use the durable cached prefix.",
        "params": {"prompt_cache_binding": binding},
    },
    result_key="llm",
)
```

### Storage semantics

- For **local** and **local-file** runtimes, `upsert_text_bloc(...)` persists one durable text snapshot under the Runtime-owned bloc root. Runtime chooses the root (`~/.abstractruntime/blocs` by default, or `<base_dir>/blocs` for `create_local_file_runtime(...)`), while AbstractCore's `FileBlocStore` defines the on-disk layout under that root.
- Within one bloc root, the durable source of truth is **content-addressed by SHA256**. Re-upserting the same text/file hash reuses or updates the same bloc record; it does not intentionally create several independent bloc copies under that same root.
- Deduplication is therefore **per bloc root**, not global across every Runtime instance. If several runtimes should share one durable bloc store, point them at the same `bloc_root_dir`. Separate roots intentionally isolate storage and can hold separate copies of the same text.
- The durable text bloc and the provider/model cache are different layers:
  - one bloc: durable extracted text plus metadata
  - zero or more derived KV artifacts: one per `(provider, model)` pair, stored under that bloc's `kv/` area
- Derived KV artifacts are **not portable** across providers or models. The same text bloc can legitimately have several provider/model-native artifacts, but each artifact remains tied to one provider/backend/model rendering path.
- `prompt_cache_binding` is a request-time proof that a specific runtime cache key still points at the exact loaded bloc artifact. It is not the durable text itself.
- For **remote** and **hybrid** runtimes using `base_url`, Runtime does not create its own local bloc copy; it proxies the bloc/KV operation to the configured AbstractCore server or upstream endpoint, and that remote side owns the store.

### Lifecycle operations

- Runtime exposes public host methods for:
  - listing durable bloc records
  - listing provider/model KV artifacts under those blocs
  - deleting one derived KV artifact while keeping the bloc text
  - pruning matching KV artifacts by filter
  - deleting one durable bloc and, by default, its derived KV artifacts
- Safety behavior mirrors the public AbstractCore contract:
  - `dry_run=True` previews the delete/prune result without mutating storage
  - `clear_loaded=True` clears matching live prompt-cache keys before deletion when the relevant provider/model is resident in the current runtime or the remote Core server
  - `force=True` bypasses that safety check and should be treated as an explicit operator choice
- `delete_bloc_kv_artifact(...)` deletes exactly one artifact. If the selector matches several provider/model artifacts, Runtime returns a structured error rather than guessing.
- `delete_bloc(...)` removes the durable text bloc itself. By default it also removes derived KV artifacts under that bloc; pass `delete_kv=False` only if you intentionally want to leave those artifacts behind.

### Session prompt-cache attribution and lifecycle

When Runtime derives the session-scoped prompt-cache key for a text/chat `LLM_CALL`, the LLM client stamps session attribution — `session_id`, `run_id`, `workflow_id`, `node_id`, and `namespace` — into the cache entry's metadata after each generate that used the derived key. Local clients stamp through the provider's public key-meta contract; remote clients post `POST /acore/prompt_cache/key_meta` to the configured AbstractCore server, and skip the stamp when the server does not expose that route. Stamping is best-effort and never affects the LLM call result.

Caller-owned keys are deliberately never stamped: an explicit `LLM_CALL.params.prompt_cache_key`, a `prompt_cache_binding` key, or a `_runtime.prompt_cache.key` override may be shared across sessions, so session-scoped clearing never touches them.

Hosts inspect and manage session caches through the facade:

- `list_session_prompt_caches(session_id=None)` returns `{"ok": true, "caches": [...]}` with one row per live cache key: `key`, `provider`, `model`, `runtime_id`, `session_id`, `token_count`, `bytes`, `created_at_s`, `last_used_at_s`, and the raw `meta`. Fields the backend does not report are `null` (`last_used_at_s` currently is). Without a filter, all live keys are listed; keys without attribution carry `session_id: null`.
- `clear_session_prompt_caches(session_id)` enumerates the session's keys and clears them one by one. It returns `{"ok": true, "cleared": [...], "count": <successes>}` with a per-row `cleared` flag; per-key failures are reported in the rows rather than raised. A `session_id` is required.
- `get_memory_snapshot()` returns the Core-owned host memory snapshot: `ram` (total/available/used), `process` (`rss_bytes`), and `device` (backend, allocated/total/free bytes; `null` where the backend has no query). To confirm that an unload released memory, compare `device.allocated_bytes`; process RSS is not a reliable release signal because freed device buffers can return to the process heap.

Session caches outlive runs by design: completing a run does not clear the session's caches, so later runs in the same session keep their warm prefixes. Caches go away when a host clears them explicitly, when the owning model is unloaded (`unload_model_residency` clears the provider's cache stores and Runtime's client-side mirrors), or when the provider evicts them.

```python
snapshot = facade.get_memory_snapshot()
caches = facade.list_session_prompt_caches(session_id="support-session-1")
result = facade.clear_session_prompt_caches("support-session-1")
```

## Models, engines and host jobs (config facade)

AbstractCore implements the model browser, the local-engine installer and a
host job registry. These passthroughs need AbstractCore 2.16.0 (the
base-install floor; 2.15.1 is the first release that records who cancelled a job).
`abstractruntime.integrations.abstractcore.config_facade`
passes them through unchanged, so a host such as AbstractGateway serves the same
payloads as `abstractcore models|engines … --json` and the `/acore/*` routes
without importing AbstractCore itself.

| Function | Returns |
|---|---|
| `host_profile(refresh=False)` | `host_profile_v1`: OS, accelerator, memory ceiling, free disk per model store |
| `engine_inventory(probe=False)` | `engines_status_v1`: Ollama, LM Studio, MLX, llama.cpp, vLLM, Hugging Face; installed, running, install plan |
| `engine_status(engine_id, probe=False)` | one engine row |
| `engine_install_plan(engine_id)` | the fixed install command (`argv`), method, download URL and notes for this host |
| `engine_download_url(engine_id)` | the vendor download page |
| `engine_install(engine_id, dry_run=False, force=False, allow=None)` | a `host_job_v1` job of kind `engine_install` |
| `model_catalog(q=None, engine=None, fits_only=False, hub=False)` | `model_catalog_v1`: downloadable models with presence and a fit verdict |
| `list_installed_models(provider=None)` | `models_installed_v1`: installed models per engine, with sizes |
| `model_delete_blockers(provider, artifact)` | what would stop a delete (reads only) |
| `delete_model_artifact(provider, artifact, dry_run=False, force=False)` | a `host_job_v1` job of kind `delete` |
| `start_model_download_job(provider, artifact, dry_run=False)` | a `host_job_v1` job of kind `download` (single-flight per artifact) |
| `host_jobs_list(kind=None, status=None)` | `{"schema": "host_jobs_v1", "jobs": [...]}`, newest first, including jobs started by other processes on the host |
| `host_job(job_id)` / `host_job_cancel(job_id, *, by="api", user=None)` | one job, or `None` when the id is unknown; a cancel records `by` (`api`, `console`) and the signed-in `user` on the job (`cancelled_by`, `cancelled_by_user`, `ended_reason`; AbstractCore 2.15.1+) |
| `console_fragment(kind)` | AbstractCore's embeddable web screen (`models` or `engines`): `{"html", "js", "css"}` |
| `models_engines_support()` | `{"available", "abstractcore_version", "required", "missing"}`; never raises |

Job-starting calls run in the background and return the queued job; a dry run
(and `run_inline=True`) finishes on the calling thread and returns the finished
job.

Errors a host can map to HTTP statuses:

- `AbstractCoreTooOld` (a `NotImplementedError`): the installed AbstractCore
  predates these modules. The message names the installed version and the
  upgrade command. Map it to 501.
- `RuntimeError`: AbstractCore is not installed. Map it to 503.
- `HostActionRefused`: a refusal with `status_code` and `payload()`. Examples:
  403 `not_allowed` (engine installs disabled by the host's policy), 404
  `not_found` / `unknown_engine`, 409 `busy` (an engine install is already
  running), 409 `refused` with `delete_blockers` (the model is loaded or shares
  a cache; pass `force=True` to override).

```python
from abstractruntime.integrations.abstractcore import config_facade as core

core.engine_inventory(probe=True)["engines"][0]["install"]["argv"]  # e.g. ["brew", "install", "ollama"]
job = core.engine_install("ollama", dry_run=True)                    # finished job, nothing runs
try:
    core.delete_model_artifact("ollama", "qwen3:8b", dry_run=True)
except core.HostActionRefused as refused:
    print(refused.status_code, refused.payload())
```

## Model residency listings

`list_model_residency(...)` relays Core-owned residency truth (ADR-0007) and, for text-generation listings, includes the whole host:

- Local and multi-local runtimes merge AbstractCore's host-wide loaded-model sweep into the listing, so models resident on host-local provider servers (for example Ollama or LM Studio) appear even when they were not loaded through this runtime. Sweep-only rows carry `source: "provider_server"` and no `task` label — Runtime relays what the sweep reports and does not invent task assignments. Remote runtimes receive the same host-wide view from the Core server.
- Rows for models this runtime loaded win deduplication against sweep rows and absorb the sweep's `size_bytes` / `size_vram_bytes` when they lack their own.
- Provider-reported size extras are normalized across local and remote listings: `size` is surfaced as `size_bytes` and `size_vram` as `size_vram_bytes`, with the original fields kept.
- Local text rows carry AbstractCore's registry-declared `modalities` (capability route keys such as `input.text`, `input.image`, `output.text`). When the registry has no entry for the model, the field is omitted rather than guessed. When the provider instance reports its vision lane unusable, `input.image` is removed and `modalities_note: "vision_unusable"` records the divergence.
- Local rows also carry the serving host's identity (`host_id`, `host_name`) from AbstractCore's host-info utility; rows already carrying another host's attribution — for example sweep rows — are not overwritten. Remote listings relay the Core server's rows verbatim, including the server's host identity; Runtime never re-stamps them with the client's identity.
- Rows report lock truth: rows for models this runtime manages carry `locked` and `lockable: true`, while sweep-only provider-server rows carry `lockable: false` because this runtime cannot enforce a lock on them. Lock state is runtime-owned — provider claims cannot supply or override `locked`, `lockable`, or `locked_at`. `pinned` is a truthful alias of `locked` (same value, Core parity), never the default-identity flag: `default` alone marks the client's default pair, so a configured capability default is never presented as pinned or loaded.
- Local and multi-local listings also show every model the process itself still holds, even when this runtime's pool no longer references it: MLX and HuggingFace text models (`provider_state: "resident_via_other_holders"`, with `process_holders` and `held_bytes`; HuggingFace holders are full copies of the weights), and in-process embedding models as `task: "embedding"` rows with `runtime_id` `local:embedding:huggingface:<model>`, `backend: "embeddings"`, `process_holders` and `held_bytes`.
- Every row the listing serves can be unloaded by its `runtime_id` alone. `unload_model_residency(runtime_id=...)` reads the task, provider and model from any `local:<task>:<provider>:<model>` id (text, TTS, STT, image, music or embedding) and finds a backend's own id (for example an image backend's `diffusers/<model>`) in the current listing.

## Unloading and switching models

- Unloading an in-process model (MLX, HuggingFace, embeddings) frees it from every holder in the process, not only from this runtime's instance, and the result carries a `process_eject` report. The result is `ok: false` when weights remain in memory. An embedding model loads again on the next embedding request.
- Changing the default text model (`set_default_provider_model`, which the gateway calls when the console default changes) unloads the previous in-process model, unless the pool still uses it or it is locked. The previous model is freed before the new default is loaded, so the two are never in memory together. If the previous model is still generating, the switch does not cancel that call; the model is unloaded when the call ends.
- A model is unloaded only when nothing else in the process still uses it: other services and users, entity runtimes and the AbstractCore server's runtimes register the models they pool, override, lock, are loading or use as their default, and the unload skips a model any of them uses. Multi-local clients (the client `create_local_runtime(...)` builds) register these claims; a standalone `LocalAbstractCoreLLMClient` does not, so a host that shares a process between several runtimes should build them with `create_local_runtime(...)`.
- The switch unload and the failed-load cleanup need an AbstractCore release with the process residency claim registry (`abstractcore.providers.process_residency.eject_unclaimed`). With an older AbstractCore the previous model stays loaded and `last_switch_ejects` reports `skipped` with the reason; unload it explicitly or upgrade AbstractCore. `list_model_residency` diagnostics list `pending_ejects` (waiting for a running call to end) and `last_switch_ejects` (unloaded, kept because it is still in use, or failed, with the reason).
- Chat compaction summarizes a run with the model the run uses. Only a run with no model of its own is summarized with the current default model; the summarizer does not keep the model that was the default at startup.
- A load with `ttl_s` or `keep_alive` for an in-process model, or for a model that is already loaded, lists them under `unsupported_options` with a warning. A load that fails part-way unloads what it loaded.
- Changing a capability default (image, voice, music) unloads the models the old capability routes had loaded before the new routes take effect.
- MLX has no idle or time-based unload. A load with `ttl_s` or `keep_alive` on MLX reports them under `unsupported_options`, with a warning, and the model stays loaded until it is unloaded.

## Model residency locks and context estimates

The host facade and all execution modes expose model-residency lock controls and a context-fit estimate relay:

- `lock_model_residency(payload=None, **kwargs)`
- `unlock_model_residency(payload=None, **kwargs)`
- `get_context_estimate(payload=None, **kwargs)`

Each accepts an optional payload mapping and/or keyword arguments; keyword arguments win on conflicts. Locks apply to text-generation runtimes only — a request naming another task returns a structured `model_residency_unsupported` payload instead of being relayed onto a text runtime.

**Lock rule (all execution modes): lock requires provider-verified residency.** A lock is a promise the model stays in memory, so it only applies to a model the provider verifies as resident (`provider_resident: true`). A warm client or configured default alone is configuration, not memory — locking a non-resident pair refuses with `{"ok": false, "error": "model_not_resident", ...}`; load the model first (load with `lock: true`) to lock it at load time. Unlock never requires residency, so a locked-but-since-evicted pair can always be released.

Lock behavior by execution mode:

- **Local / multi-local**: the lock is enforced client-side per `(provider, model)` pair. `unload_model_residency` refuses a locked pair with a soft `{"ok": false, "error": "model_locked", ...}` payload — never an exception. Passing `force=true` unloads the pair, and the lock is released only after the provider unload succeeds; a failed forced unload leaves the pair resident and locked. For Ollama, lock and unlock also apply the server-side `keep_alive` knob best-effort (`-1` on lock, the `"5m"` default on unlock) and report the outcome under `provider_side` (`supported` / `applied` / optional `detail`); the client-side flag remains the enforcement truth for every provider, so a knob failure does not fail the lock. Because the knob rides Ollama's native load request, unlocking a pair whose model was since evicted skips the keep-alive restore (`applied: false` with a detail) — unlock never loads a model back as a side effect, and the residency-required lock rule prevents the lock-side equivalent.
- **Remote / hybrid**: `lock_model_residency` and `unlock_model_residency` relay `POST /acore/models/lock` and `POST /acore/models/unlock` on the configured AbstractCore server, which owns the lock (and enforces the same residency-required rule). When the server refuses to unload a locked runtime with HTTP 409, `unload_model_residency` converts that refusal into the same structured `{"ok": false, "error": "model_locked", "status_code": 409, ...}` payload instead of raising; a lock refused for a non-resident model (HTTP 409) is likewise converted to `{"ok": false, "error": "model_not_resident", "status_code": 409, ...}`. `force` rides the unload request body only when true. Transport or server errors from these residency operations report `supported: true`; `supported: false` is reserved for a client that does not implement the operation at all.

Request selectors:

- Address the runtime with explicit `provider` / `model` fields or with a `runtime_id`. Local clients accept `local:text_generation:<provider>:<model>` runtime ids; a runtime id that does not address a local text runtime returns a not-found payload rather than acting on a runtime the caller did not name. A request with no selector at all applies to the client's own default identity.
- Locking requires the pair to be warm in the local client or pool AND provider-verified resident (the lock rule above). Unlocking additionally reaches a locked pair whose pooled client is no longer warm, so a lock can always be released.

When a multi-local pool re-points its default provider/model, locked pairs are exempt from the pool eviction so their in-process weights stay resident behind the lock. The one exception is a locked pair that becomes the new default identity itself: its pooled client is rebuilt with the new connection settings and its lock is cleared (logged as a warning); re-lock to pin the rebuilt client.

`get_context_estimate(...)` relays AbstractCore's analytical context-fit estimator. Local and multi-local clients call the estimator in-process, defaulting `provider` / `model` to the client's identity when omitted; remote clients relay `GET /acore/models/context_estimate` with `provider`, `model`, and an optional integer-coerced `context_length`. The estimator response is relayed verbatim — including the tri-state `fits_weights` / `fits_requested_context` split, `predicted_max_context` (the context that fits beside the weights), and `budget_bytes` (real-ceiling budget; basis and reserve stated in `notes`); it is advisory only — no Runtime load path gates on it. When the local estimator utility is unavailable, the call degrades to `{"ok": false, "supported": false, "operation": "context_estimate", ...}`.

Example:

```python
facade = get_abstractcore_host_facade(rt)

locked = facade.lock_model_residency(provider="ollama", model="qwen3:4b")
refused = facade.unload_model_residency(provider="ollama", model="qwen3:4b")  # model_locked payload
unloaded = facade.unload_model_residency(provider="ollama", model="qwen3:4b", force=True)
fit = facade.get_context_estimate(provider="ollama", model="qwen3:4b", context_length=32768)
```

### `MODEL_RESIDENCY` effect operations

The `MODEL_RESIDENCY` effect supports the operations `list_loaded`, `load`, `unload`, `lock`, and `unlock`, so workflows can author residency control durably:

```json
{"operation": "lock", "provider": "ollama", "model": "qwen3:4b"}
{"operation": "unlock", "runtime_id": "local:text_generation:ollama:qwen3:4b"}
{"operation": "unload", "provider": "ollama", "model": "qwen3:4b", "force": true, "required": false}
```

`lock` / `unlock` accept the same selector fields as the client methods (`task`, `provider`, `model`, `runtime_id`, plus `base_url` / `timeout_s` / provider-key overrides for remote relays). On `unload`, `force` is forwarded only when it is authored in the effect payload. All residency operations keep soft-fail semantics: unless the payload sets `required: true`, a refusal such as `model_locked` completes the step with `status_hint: "warning"` and `degraded: true` instead of failing the run.

## Host-local comms and Telegram wrappers

Runtime also exposes the remaining Gateway-facing host/operator wrappers for
email and Telegram:

- `get_abstractcore_host_facade(runtime)` includes:
  - `list_email_accounts(...)`
  - `list_emails(...)`
  - `read_email(...)`
  - `send_email(...)`
- `abstractruntime.integrations.abstractcore.comms_facade` also exposes:
  - `list_email_accounts(...)`
  - `list_emails(...)`
  - `read_email(...)`
  - `send_email(...)`
- `abstractruntime.integrations.abstractcore.telegram_facade` exposes:
  - `TelegramTdlibNotAvailable`
  - `bootstrap_telegram_auth_from_env(...)`
  - `get_global_telegram_client(start=False)`
  - `stop_global_telegram_client()`
  - `send_telegram_message(...)`

Contract notes:

- These are **host-local** wrappers over current public AbstractCore tool
  modules. They do not proxy through the remote AbstractCore server.
- The host facade email methods and the standalone `comms_facade` functions use
  the same Runtime-owned email wrapper layer; choose whichever is more natural
  for the host surface you are building.
- They are intentionally **nondurable**. They do not write Runtime run history
  on their own.
- Direct `send_email(...)` on the host facade and direct
  `telegram_facade.send_telegram_message(...)` are for operator-owned
  host-local flows only. If the outbound send belongs to a workflow/run, prefer
  the durable run facade methods shown below.
- Even for **remote** and **hybrid** runtimes, they still use the current host
  process env/config, local TDLib installation, and the host's own outbound
  network access.
- The Telegram global client is process-wide, not runtime-instance scoped.

Host-side operator example:

```python
from abstractruntime.integrations.abstractcore import (
    create_local_runtime,
    get_abstractcore_host_facade,
)
from abstractruntime.integrations.abstractcore.telegram_facade import (
    TelegramTdlibNotAvailable,
    bootstrap_telegram_auth_from_env,
    send_telegram_message,
)

rt = create_local_runtime(provider="ollama", model="qwen3:4b")
facade = get_abstractcore_host_facade(rt)

accounts = facade.list_email_accounts()
sent = facade.send_email(
    ["ops@example.com"],
    "Runtime status",
    body_text="All green.",
)

try:
    bootstrap = bootstrap_telegram_auth_from_env(timeout_s=30)
except TelegramTdlibNotAvailable:
    bootstrap = {"success": False, "error": "TDLib is not installed on this host."}

notify = send_telegram_message(chat_id=123456, text="Runtime check complete.")
```

## Discovery snapshots

AbstractRuntime's AbstractCore integration also exposes a public host discovery facade for snapshot/query reads:

- `get_abstractcore_discovery_facade(runtime)`
- `AbstractCoreDiscoveryFacade`
- `list_providers(...)`
- `list_provider_models(...)`
- `get_model_capabilities(...)`
- `get_voice_catalog(...)`
- `list_tts_models(...)`
- `list_stt_models(...)`
- `list_music_providers(...)`
- `list_music_models(...)`
- `list_vision_provider_models(...)`
- `list_cached_vision_models(...)`

Behavior by execution mode:

- **Local** (`MultiLocalAbstractCoreLLMClient` / `LocalAbstractCoreLLMClient`): uses public AbstractCore registries,
  capability facades, and local vision cache inspection to return JSON-safe snapshot payloads.
- **Remote / Hybrid** (`RemoteAbstractCoreLLMClient`): proxies `/providers`, `/v1/models`, `/v1/audio/*`, and
  `/v1/vision/*` on the configured AbstractCore server. Per-request provider key overrides supplied as `api_key` /
  `provider_api_key` become `X-AbstractCore-Provider-API-Key` headers.

`list_provider_models(provider, ...)` accepts the legacy `input_type` and
`output_type` filters plus Core's precise `capability_route` filter. Local mode
normalizes route filters before calling AbstractCore's provider registry; remote
mode forwards them to `/v1/models?capability_route=...`:

```python
models = facade.list_provider_models(
    "lmstudio",
    capability_route=["input.image", "output.text"],
)
embeddings = facade.list_provider_models("lmstudio", capability_route="embedding.text")
```

Contract notes:

- This surface is query-oriented. It does not create durable Runtime history on its own.
- Hosts should still ask Runtime for these reads instead of rebuilding Core catalog logic or importing Core server
  helpers directly.
- Model capability lookup is static metadata, not a live server probe. Replay should treat it as a recorded snapshot,
  not as a query to re-run.
- `list_cached_vision_models(...)` may still depend on the current local machine state. It is a Runtime-owned snapshot
  query, not durable run truth.
- Remote discovery methods accept `timeout_s=...` through facade kwargs. Local discovery remains synchronous helper
  code; async hosts should offload it to a worker thread if they do not want to block their event loop.

Host-side discovery example:

```python
from abstractruntime.integrations.abstractcore import (
    create_remote_runtime,
    get_abstractcore_discovery_facade,
)

rt = create_remote_runtime(
    server_base_url="http://127.0.0.1:8000",
    model="openai/gpt-4o-mini",
    headers={"Authorization": "Bearer server-master-key"},
)
facade = get_abstractcore_discovery_facade(rt)

providers = facade.list_providers(include_models=False)
voices = facade.get_voice_catalog(provider="openai", providers_only=True)
music = facade.list_music_providers(task="text_to_music")
vision = facade.list_vision_provider_models(task="text_to_image", providers_only=True)
upscalers = facade.list_vision_provider_models(task="image_upscale")
adapters = facade.list_vision_adapters(
    task="text_to_video",
    model="AbstractFramework/wan2.2-t2v-a14b-diffusers-8bit",
)
```

## Durable run-scoped media and comms execution

Hosts sometimes need to trigger image/TTS/music/STT work or outbound comms sends for an existing run. That work should still execute through Runtime so the child run ledger, artifact ownership, and replay surface remain Runtime-authored.

Public durable entry points:

- `get_abstractcore_run_facade(runtime)`
- `AbstractCoreRunFacade`
- `execute_llm_call(...)`
- `execute_tool_calls(...)`
- `resume_tool_calls(...)`
- `generate_image(...)`
- `edit_image(...)`
- `upscale_image(...)`
- `generate_video(...)`
- `image_to_video(...)`
- `generate_voice(...)`
- `stream_voice(...)`
- `generate_music(...)`
- `transcribe_audio(...)`
- `send_email(...)`
- `send_telegram_message(...)`

These helpers create child runs under an existing parent run and execute the real
`LLM_CALL`, `TOOL_CALLS`, or stream finalization through Runtime rather than
doing external work in host/controller code. `stream_voice(...)` yields TTS
stream events for progressive playback and completes the child run with the
final audio artifact when the stream succeeds.

Example:

```python
from abstractruntime.integrations.abstractcore import (
    create_local_runtime,
    get_abstractcore_run_facade,
)

rt = create_local_runtime(provider="mlx", model="qwen-chat")
facade = get_abstractcore_run_facade(rt)

child = facade.generate_image(
    "existing-parent-run-id",
    prompt="A red mug on a white table.",
    output={
        "provider": "mlx-gen",
        "model": "AbstractFramework/qwen-image-2512-8bit",
        "format": "png",
        "count": 2,
        "seeds": [101, 102],
        "lora_adapters": [
            {"id": "pixel-art", "scale": 0.7},
            {"id": "cool-grade", "scale": 0.2},
        ],
    },
)

assert child.status.value == "completed"
result = child.output["result"]
print(child.run_id, result["media_model"], result["outputs"]["image"][0]["artifact_id"])
```

Image upscaling uses the same durable child-run boundary and Core-owned `image_upscale` selector:

```python
child = facade.upscale_image(
    "existing-parent-run-id",
    media={"$artifact": "source-image-artifact-id", "type": "image"},
    output={
        "provider": "mlx-gen",
        "format": "png",
        "scale": 2,
        "resolution": 1024,
    },
)

print(child.run_id, child.output["result"]["outputs"]["image"][0]["artifact_id"])
```

For video, use the same child-run boundary:

```python
child = facade.generate_video(
    "existing-parent-run-id",
    prompt="Glowing data streams converge into a geometric logo.",
    output={
        "provider": "mlx-gen",
        "model": "AbstractFramework/wan2.2-t2v-a14b-diffusers-8bit",
        "format": "mp4",
        "num_frames": 41,
        "count": 2,
        "seeds": [401, 402],
        "flow_shift": 3.0,
        "lora_adapters": [{"id": "documentary-motion", "scale": 0.6}],
    },
)

print(child.run_id, child.output["result"]["outputs"]["video"][0]["artifact_id"])
```

Outbound comms sends that belong to a run should use the same durable child-run surface:

```python
email_child = facade.send_email(
    "existing-parent-run-id",
    to=["ops@example.com"],
    subject="Workflow alert",
    body_text="The workflow completed.",
)

telegram_child = facade.send_telegram_message(
    "existing-parent-run-id",
    chat_id=123456,
    text="Workflow completed.",
)
```

Contract notes for durable comms sends:

- Runtime records the send request and the send outcome in the child run ledger.
- Replay should show the recorded result; it should **not** resend the external email or Telegram message.
- Local and hybrid runtimes usually execute those sends immediately when the configured tool executor can run them.
- Remote runtimes may still enter a durable tool wait if the configured tool executor is passthrough/delegated or approval-gated. That wait/resume path is still Runtime-authored truth.
- To resume a waiting durable comms/tool child run through the same public boundary, use `get_abstractcore_run_facade(runtime).resume_tool_calls(child_run_id, payload=...)`.

## Attachment registration limits

When local `read_file` tool outputs are captured as session attachments, Runtime bounds the file bytes it stores. The limit is resolved in this order:

- `TOOL_CALLS.payload.max_attachment_bytes`
- `run.vars["_runtime"]["max_attachment_bytes"]`
- `ABSTRACTRUNTIME_MAX_ATTACHMENT_BYTES`
- the default of 25 MiB

Gateway-specific attachment env vars should be translated by Gateway into one of the explicit Runtime inputs above.

## Default toolsets (incl. comms)

`default_tools.get_default_toolsets()` provides a host-side convenience catalog of common tools:
- file/web/system tools
- optional comms tools behind env-var gating (`docs/tools-comms.md`)

This is useful when building a `MappingToolExecutor` quickly.

## See also

- `../architecture.md` — effect handler boundaries and durability invariants
- `../tools-comms.md` — enabling email/WhatsApp/Telegram tools
- `../adr/0002_execution_modes_local_remote_hybrid.md` — rationale for local/remote/hybrid


==============================================================================
# FILE: docs/proposal.md
==============================================================================

# AbstractRuntime — Overview (v0.4.x)

**AbstractRuntime** is a low-level *durable workflow runtime*:
- execute workflow graphs (state machines)
- support **interrupt → checkpoint → resume** without keeping Python stacks alive
- record an append-only **execution journal** (“ledger”) for observability/audit/debug

**Scope boundary:** AbstractRuntime is not a UI builder and not an agent framework. It is the execution substrate that higher-level orchestration (e.g., visual authoring hosts) and agent loops can build on.

## Ecosystem

AbstractRuntime is part of the wider AbstractFramework ecosystem:
- AbstractFramework umbrella: [lpalbou/AbstractFramework](https://github.com/lpalbou/AbstractFramework)
- AbstractCore (LLM + tools): [lpalbou/abstractcore](https://github.com/lpalbou/abstractcore)

In this repo, AbstractCore wiring lives under `src/abstractruntime/integrations/abstractcore/*` and is described in `integrations/abstractcore.md`.

## What problem this solves

Once a workflow can:
- ask a user and wait hours/days
- wait until a scheduled time
- wait for an external job/event

…you need durable semantics: persisted checkpoints + a journal. Keeping a Python stack alive is not reliable across restarts.

## Core concepts (durable model)

All core durable types are stdlib-only and live in `src/abstractruntime/core/models.py`:

- `WorkflowSpec`: in-memory workflow graph (node handlers keyed by id) (`src/abstractruntime/core/spec.py`)
- `RunState`: durable run checkpoint (`run_id`, `status`, `current_node`, `vars`, `waiting`, `output`, `error`, provenance fields)
- `StepPlan`: what a node returns (`effect?`, `next_node?`, `complete_output?`)
- `Effect` + `EffectType`: request for side effects (LLM, tools, waits, memory ops, etc.)
- `WaitState`: durable blocking state (`reason`, `wait_key`/`until`, `resume_to_node`, `result_key`)
- `StepRecord`: append-only ledger entry

**Non-negotiable constraint:** values stored in `RunState.vars` must be JSON-serializable. For large payloads, use `ArtifactStore` references (`src/abstractruntime/storage/artifacts.py`) or offloading wrappers (`src/abstractruntime/storage/offloading.py`).

## Minimal runtime API

Implemented in `src/abstractruntime/core/runtime.py`:
- `start(workflow, vars, actor_id, session_id) -> run_id`
- `tick(workflow, run_id) -> RunState` (progress until waiting/completed/failed)
- `resume(workflow, run_id, wait_key, payload) -> RunState`
- `get_state(run_id) -> RunState`
- `get_ledger(run_id) -> list[dict]`

Resume semantics (important):
- when a run blocks, the runtime stores `WaitState.resume_to_node`
- on resume, execution continues **from that node** (it does not re-run the waiting node)

## Persistence

Interfaces live in `src/abstractruntime/storage/base.py`.

Included backends:
- `InMemory*` (tests/dev): `src/abstractruntime/storage/in_memory.py`
- file-based JSON/JSONL: `src/abstractruntime/storage/json_files.py`
- SQLite: `src/abstractruntime/storage/sqlite.py`

Related features:
- snapshots/bookmarks: `src/abstractruntime/storage/snapshots.py` (`docs/snapshots.md`)
- tamper-evident ledger: `src/abstractruntime/storage/ledger_chain.py` (`docs/provenance.md`)
- in-process ledger subscriptions: `src/abstractruntime/storage/observable.py`

## Scheduling (driver loop)

AbstractRuntime ships a simple in-process scheduler:
- `Scheduler`, `ScheduledRuntime`, `create_scheduled_runtime()` (`src/abstractruntime/scheduler/*`)

This is a driver loop (polls due waits, resumes runs). It is not a distributed orchestrator.

## Integrations (optional)

AbstractRuntime stays dependency-light at the kernel level; concrete integrations are opt-in:
- AbstractCore (LLM + tools): `src/abstractruntime/integrations/abstractcore/*` (`integrations/abstractcore.md`)
- AbstractMemory bridge (KG assertions/queries): `src/abstractruntime/integrations/abstractmemory/*`

## Status (implemented in this repository)

As of v0.4.9:
- durable kernel: `RunState`, `WaitState`, `Runtime.start/tick/resume`
- built-in waits + events: `WAIT_EVENT`, `WAIT_UNTIL`, `ASK_USER`, `EMIT_EVENT`
- persistence backends: in-memory, JSON/JSONL, SQLite
- artifacts/offloading: store large payloads by reference
- retries/idempotency policy hooks: `src/abstractruntime/core/policy.py`
- snapshots, tamper-evident ledger chain, ledger subscriptions
- VisualFlow compiler + WorkflowBundles (`src/abstractruntime/visualflow_compiler/*`, `src/abstractruntime/workflow_bundle/*`)
- VisualFlow multi-entry lowering for fan-in execution routes (`join_exec` / `path_mux`)
- AbstractCore integration with local/remote/hybrid LLM execution, cached sessions/prompt-cache control, media inputs, generated media outputs, tool approval waits, and provider-key header routing (`integrations/abstractcore.md`)
- evidence capture helpers (`src/abstractruntime/evidence/recorder.py`, `Runtime.list_evidence/load_evidence`)
- run history bundle export (`src/abstractruntime/history_bundle.py`)

## See also

- `../README.md` — install + quick start
- `getting-started.md` — first steps
- `architecture.md` — full component map and diagrams
- `integrations/abstractcore.md` — AbstractCore wiring
- `limits.md` — `_limits` / RuntimeConfig


==============================================================================
# FILE: docs/limits.md
==============================================================================

# Runtime limits (`_limits`)

AbstractRuntime stores runtime-facing limits in a canonical `RunState.vars["_limits"]` dict. This is used for **durable configuration** (persisted in checkpoints) and for **host/agent introspection**.

Implementation pointers:
- `_limits` helpers/namespace constants: `src/abstractruntime/core/vars.py`
- limit config source of truth: `src/abstractruntime/core/config.py` (`RuntimeConfig`)
- limit APIs: `src/abstractruntime/core/runtime.py` (`get_limit_status`, `check_limits`, `update_limits`)

## What `_limits` contains

`Runtime.start(...)` initializes `_limits` from `RuntimeConfig.to_limits_dict()` when missing. (`src/abstractruntime/core/runtime.py`, `src/abstractruntime/core/config.py`)

Shape (keys are stable; values may be `None` when unknown):

```python
run.vars["_limits"] = {
    "max_iterations": 20,
    "current_iteration": 0,
    "max_tokens": 32768,          # context window (fallback when unknown)
    "max_output_tokens": None,    # provider/model dependent
    "max_input_tokens": None,     # optional budget cap for inputs
    "estimated_tokens_used": 0,   # best-effort (see below)
    "max_history_messages": -1,   # -1 = unlimited
    "warn_iterations_pct": 80,
    "warn_tokens_pct": 80,
}
```

Notes (as implemented today):
- `current_iteration` is **not** automatically incremented by the runtime; higher-level loops (agents/workflows) should update it if they want iteration budgeting.
- `estimated_tokens_used` is a **best-effort, last-known** value. The runtime updates it from `LLM_CALL` usage metadata when available (`src/abstractruntime/core/runtime.py`). It is not guaranteed to be tokenizer-accurate and is not accumulated across calls.

## Configuring limits

You can pass a `RuntimeConfig` when constructing a `Runtime`:

```python
from abstractruntime.core import Runtime, RuntimeConfig
from abstractruntime.storage import InMemoryLedgerStore, InMemoryRunStore

rt = Runtime(
    run_store=InMemoryRunStore(),
    ledger_store=InMemoryLedgerStore(),
    config=RuntimeConfig(
        max_iterations=50,
        max_tokens=65536,
        warn_iterations_pct=75,
    ),
)
```

If you use the AbstractCore convenience factories, they also accept `config=` and may populate model capabilities (`src/abstractruntime/integrations/abstractcore/factory.py`).

## Introspection and updates

### `Runtime.get_limit_status(run_id)`

Returns a structured dict for UI/status display. (`src/abstractruntime/core/runtime.py`)

### `Runtime.check_limits(run_state)`

Returns a list of `LimitWarning` objects for limits approaching/exceeded. (`src/abstractruntime/core/models.py`, `src/abstractruntime/core/runtime.py`)

As of v0.4.9, warnings are computed for:
- `iterations` (`current_iteration` vs `max_iterations`)
- `tokens` (`estimated_tokens_used` vs `max_tokens`)

### `Runtime.update_limits(run_id, updates)`

Updates selected keys in `_limits` durably (saved via the configured `RunStore`). Unknown keys are ignored. (`src/abstractruntime/core/runtime.py`)

Example:

```python
rt.update_limits(run_id, {"max_tokens": 131072, "warn_tokens_pct": 85})
```

## See also

- `architecture.md` — where `_limits` fits in the runtime
- `integrations/abstractcore.md` — where token usage metadata typically comes from (`LLM_CALL`)
- `manual_testing.md` — quick manual checks + running tests


==============================================================================
# FILE: docs/artifacts.md
==============================================================================

# Runtime artifacts

AbstractRuntime stores large payloads as artifacts so run state, ledger records,
and workflow outputs remain JSON-safe. Artifacts are the durable record for
files, generated media, tool evidence, exported history bundles, and other
payloads that should be referenced by id rather than embedded inline.

The implementation lives in `src/abstractruntime/storage/artifacts.py`.

## File-like vocabulary boundary

Runtime artifacts are not the same thing as live filesystem paths:

- `Artifact`: a Runtime-owned durable payload safe to persist, search, reuse,
  and pass between runs by reference.
- `Workspace File` / `Workspace Folder`: a server-side path capability under
  Gateway/runtime workspace policy. These are path values, not durable payloads.
- `Local File` / `Local Folder`: a client-side intake source. In hosted/browser
  mode they should be uploaded and normalized into artifacts before durable
  execution.

Gateway/Flow product copy may say `Server File` / `Server Folder` when users
choose an origin, but Runtime stays anchored on artifact refs versus
workspace-scoped paths.

## Artifact identity

An artifact has a stable `artifact_id`, a `blob_id` for content deduplication
when the backend supports it, and a `run_id` scope. JSON state and Gateway APIs
pass artifacts with refs such as:

```json
{
  "$artifact": "a7050ebc5c8330...",
  "artifact_id": "a7050ebc5c8330...",
  "run_id": "9e19bd6a-ba07-4c2e-86c6-94ec7ca0a373",
  "content_type": "audio/wav",
  "size_bytes": 5293012
}
```

The ref is a pointer, not authorization. Hosts such as AbstractGateway are
responsible for deciding which principal may list or read an artifact.

## Stored metadata

Runtime stores two metadata layers:

- `tags`: string fields for compatibility and simple lookups.
- `metadata`: structured JSON for producer-specific details.
- `descriptor`: a Runtime-owned `ArtifactDescriptor` that normalizes the fields
  Gateway and Observer should use.

The descriptor separates display format from semantic meaning:

- `render_kind`: how the artifact should be rendered, such as `image`, `audio`,
  `video`, `markdown`, `html`, `json`, `text`, or `document`.
- `semantic_kind`: what the artifact represents, such as `voice`, `music`,
  `sound`, `transcript`, `evidence`, `workflow_snapshot`, or `image`.
- `classification_source`: whether the classification came from the producer,
  runtime tags, MIME inference, or legacy fallback.

The descriptor can also carry `session_id`, `workflow_id`, `node_id`, `turn_id`,
`ledger_cursor`, `producer`, `generation`, `media`, `source_refs`, `links`,
`security`, and action metadata. Producer metadata may be sparse; consumers
should show missing fields as unavailable rather than guessing.

## Generated media provenance

When generated media is produced through the Runtime AbstractCore integration
with output selectors, Runtime stores descriptor and metadata alongside the
bytes. Current generated outputs include image, video, voice/TTS, music, sound
or audio outputs, and transcription-style text outputs where the host route
stores a transcript artifact.

Generated-media descriptors record available producer facts:

- package/capability route, provider, model, backend, and runtime provider/model
  when they differ;
- prompt or TTS text, requested format, output index, negative prompt, and
  redacted generation parameters;
- source artifact refs for edit, image-to-video, cloned/reference voice, or
  other source-media flows when provided;
- measured media facts such as duration, sample rate, dimensions, channels, or
  frame counts when the store can inspect the bytes.

Runtime redacts obvious secret fields and bounds large metadata values. It does
not store raw provider requests as indexed descriptor fields.

Gateway or package producers that store artifacts outside the main Runtime
generated-media path should use `build_artifact_descriptor_payload(...)` from
`abstractruntime.storage.artifacts`. That helper applies the Runtime descriptor
schema, bounded secret-key redaction, and prompt/text sensitivity labels without
making Gateway invent a parallel descriptor contract.

## Tool-output offload

Large tool outputs are stored as session artifacts instead of being carried
inline in prompts and ledger records. The agent keeps a bounded preview plus an
`open_attachment` handle, so full content stays retrievable on demand while the
durable record stays lean.

Offload applies to:

- `read_file` content above the inline threshold;
- `execute_command` `stdout` and `stderr`, regardless of exit code (verbose
  failures are offloaded like verbose successes);
- any other host tool whose string output exceeds the inline threshold.

Thresholds (environment-configurable):

- `ABSTRACTRUNTIME_MAX_INLINE_BYTES` (default 256 KiB): outputs at or below
  this size stay inline.
- `ABSTRACTRUNTIME_MAX_ATTACHMENT_BYTES` (default 50 MB): the retention cap for
  offloaded outputs.

Output above the retention cap is not stored and is never kept inline. The tool
result carries an explicit notice stating the size, the cap, and how to proceed
(narrow the command, for example with `head` or `grep`, or redirect to a file
and read a bounded range). The decision on how to proceed belongs to the agent
or user; the runtime never drops output silently.

## Catalog and search

`InMemoryArtifactStore` and `FileArtifactStore` support:

- `store(...)`, `load(...)`, and `get_metadata(...)`;
- `update_metadata(...)` for descriptor or structured metadata enrichment;
- `search(...)` for bounded pages;
- `count(...)`, `facet_counts(...)`, and `stats(...)` for exact totals and
  filter chips;
- `record_access(...)` for explicit metadata/content/preview/download/export
  counters.

`FileArtifactStore` maintains a repairable SQLite catalog for descriptor fields,
time filters, exact counts, byte totals, and facets. The catalog is an index of
Runtime-owned artifact metadata; the payload bytes and metadata files remain the
source of truth.

Plain `load(...)` and `get_metadata(...)` are side-effect free. UI and HTTP
layers that want access statistics must call `record_access(...)` or use Gateway
content routes that label the action.

## Gateway and Observer

AbstractGateway exposes Runtime artifacts through bounded HTTP APIs. Search
responses include `artifact_envelope_v1`, which projects Runtime descriptors,
media facts, access stats, and action links while preserving legacy row fields.
Gateway only forwards descriptor action links that are relative Gateway/UI
links; arbitrary external provider URLs should be represented as trace
availability or Gateway-owned trace records instead.

AbstractObserver renders Gateway envelopes. It should not infer canonical
artifact meaning from filenames or raw content except as a visible legacy
fallback. Use the Runtime tab for artifact inventory and the Observe tab for the
workflow/ledger narrative.

## Retrieval boundaries

Artifact search answers questions about stored files and media by metadata,
scope, type, time, producer, and links back to runs. It is not a semantic memory
search system.

- Use the ledger for "what happened in this workflow?"
- Use artifact search for "what files/media exist and how were they produced?"
- Use AbstractMemory/KG retrieval for "what knowledge or relationships were
  learned?"
- Use Gateway audit/provider links for system-level request traces when the
  envelope reports them.

## Limits

Legacy artifacts may only have MIME type and tags. Runtime projects them with
fallback descriptor fields so clients can still list and preview them, but
producer-level prompt/model/source provenance is only available for artifacts
created through descriptor-aware paths.


==============================================================================
# FILE: docs/automations.md
==============================================================================

# Automations

An automation runs a workflow again and again on a trigger: "check the price of ACME every 5 minutes", "report this
machine's memory use every 2 minutes", "run this report when I ask". AbstractRuntime runs automations itself, as
ordinary durable runs, so a runtime and a run store are enough to run one. A host such as AbstractGateway exposes them
over HTTP and drives them in its run loop.

This page covers the model, the definition, triggers, the controller, commands, context, discussions, tool approval,
notifications and the storage guarantees automations rely on. For the surrounding runtime concepts see
[architecture.md](architecture.md); for the import surface see [api.md](api.md#automations).

Code: `abstractruntime.automations` (controller, commands, service API), `abstractruntime.triggers` (trigger
sources), `abstractruntime.automation_queries` (listing), `abstractruntime.session_turns` and
`abstractruntime.session_history` (what counts as a turn and how history is replayed).

## Mental model

- **An automation is a run.** It is a durable root run, the *controller*, that executes the packaged flow
  `abstractframework.automation-controller@1.0.0`. The controller's run id is the automation id.
- **Each firing is an occurrence.** An occurrence is a child run of the controller with a deterministic id. At most
  one occurrence runs at a time. Its child runs (tools, subworkflows) are its *descendants*.
- **Occurrences are turns.** In the session views and history the runtime builds, an occurrence counts as one
  conversation turn: its prompt is the user turn and its answer is the assistant turn.
- **Context is independent or growing.** By default each occurrence starts fresh in its own session. In growing mode
  every occurrence joins the automation's session and sees the previous ones as history.
- **Automations are quiet.** An occurrence asks for attention only when its output says `notify`, when it still
  fails after its last retry, or while it waits on a person.
- **Creating an automation is consent for its tools.** With the default `tool_approval: "auto"` an occurrence's tools
  run without asking; set `"ask"` to approve them one batch at a time.
- **Everything is recorded.** Every admission, dispatch, retry, completion and command is an `automation.*` record in
  the controller's ledger. A crash at any point never loses an occurrence and never starts one twice.

```mermaid
flowchart LR
  Trigger["Trigger source<br/>schedule@1 / manual@1"] -->|"due tick or run now"| Controller["Controller run<br/>(automation id)<br/>vars._meta.automation<br/>vars._runtime.automation"]
  Commands["apply_automation_command<br/>pause / resume / run_now /<br/>revise / stop_current / archive"] -->|"decision + wake"| Controller
  Controller -->|"START_SUBWORKFLOW<br/>deterministic run_id"| Occurrence["Occurrence run<br/>role: occurrence<br/>(a session turn)"]
  Occurrence --> Descendants["Descendant runs<br/>role: descendant"]
  Occurrence -->|"output: answer, notify"| Controller
  Controller -->|"automation.* records"| Ledger["Controller ledger<br/>occurrences, attention, commands"]
  Occurrence -.->|"fork: own workspace + read-only mount"| Discussion["Discussion run<br/>own session"]
```

## Quick start

```python
import os

from abstractruntime import Runtime
from abstractruntime.automations import (
    controller_workflow_spec,
    create_automation,
    drive_automation,
    register_controller_bundle,
)
from abstractruntime.scheduler.registry import WorkflowRegistry
from abstractruntime.storage.sqlite import SqliteDatabase, SqliteLedgerStore, SqliteRunStore

db = SqliteDatabase("automations.sqlite")
registry = WorkflowRegistry()
registry.register(memory_check)                   # the target: any WorkflowSpec
register_controller_bundle(registry)              # the packaged controller flow
runtime = Runtime(run_store=SqliteRunStore(db), ledger_store=SqliteLedgerStore(db), workflow_registry=registry)

automation_id, revision = create_automation(runtime, {
    "request_id": "memory-watch-1",               # the same request finds the same automation
    "title": "Memory watch",
    "target": {"workflow_id": memory_check.workflow_id, "bundle_ref": "local@0.0.0",
               "flow_id": "memory_check", "input_data": {"prompt": "Report memory use; notify me above 90%."}},
    "trigger": {"source_id": "schedule", "source_version": 1, "config": {"every": "2m"}},
    "workspace_root": os.path.abspath("automations/memory-watch"),
})
drive_automation(runtime, automation_id)          # runs occurrence 1 now, then parks until the next tick
```

Without `start_at`, the schedule starts at creation time, so the first occurrence runs on the controller's first
step. `drive_automation` ticks the controller and its current occurrence, resumes the controller when the occurrence
ends, and returns when the controller parks (on its wake wait, or on an occurrence that waits on a person).

To fire later ticks without a host loop, tick the controller once its deadline has passed (a due wake wait is
released by `Runtime.tick`) and drive it again:

```python
runtime.tick(workflow=controller_workflow_spec(), run_id=automation_id)
drive_automation(runtime, automation_id)
```

A host with its own run loop does not need `drive_automation`: the controller is an ordinary run that waits on an
event with a deadline (listed by `list_due_wait_until`), and waits on its occurrence like any asynchronous
subworkflow parent.

## The definition

The definition is stored in `vars._meta.automation` of the controller run. Every change creates a new revision, and
unknown fields are rejected at every level.

| Field | Meaning |
|---|---|
| `schema_version` | `1` |
| `revision` | starts at 1; each applied `automation.revise` adds 1 |
| `title` | non-empty, at most 120 characters |
| `controller` | `{bundle_ref: "abstractframework.automation-controller@1.0.0", flow_id: "controller"}` |
| `target` | `{workflow_id, bundle_ref, flow_id, input_data}`: a concrete workflow; a host resolves `@default` before creating (`flow_id: "@default"` is refused) |
| `trigger` | `{binding_id, source_id, source_version, config}`; the runtime creates `binding_id` |
| `context` | `{mode: "independent" \| "growing", growing: {}}` |
| `policy` | `{serial: true, misfire: "coalesce", failure: "continue", retry, tool_approval}` |
| `session_id` | `automation:<automation_id>` |
| `workspace_root` | absolute path given to every occurrence |
| `created_at`, `archived_at` | UTC timestamps; `archived_at` is `null` until the automation is archived |

`policy.retry` is `{max_attempts, backoff: {initial, factor, max}}`, by default
`{max_attempts: 3, backoff: {initial: "30s", factor: 2, max: "10m"}}`. `max_attempts` is 1 to 10, `factor` is 1 to 10,
and `initial` and `max` are durations (see [`schedule@1`](#schedule1)). `serial`, `misfire` and `failure` accept only
the values shown; anything else is refused with `unsupported_feature`. `policy.tool_approval` is `"auto"` (default) or
`"ask"`, see [Tool approval](#tool-approval).

### Creating an automation

`create_automation(runtime, request, *, now=None, actor_id=None) -> (automation_id, revision)`.

The request is `{request_id, title, target, trigger: {source_id, source_version, config}, context?, policy?,
workspace_root, tenant?, user?}`.

- The automation id is `uuid5(AUTOMATION_NAMESPACE, "<tenant>:<user>:<request_id>")`; `tenant` and `user` default to
  `local`. Sending the same request again returns the same automation, unchanged. Reusing a `request_id` for a
  different request raises `AutomationError` with `reason_code = "identity_conflict"`.
- `actor_id` (the owner) is stamped on the controller run in the same create-if-absent step.
- An `automation.created` record is written once.
- Validation failures raise `AutomationError`. Its `reason_code` is `invalid_definition`, `unsupported_feature` or
  `unknown_trigger_source`, and its `field` names the offending field (`trigger.config.every`, for example).
- `context.growing.summary` is refused with `unsupported_feature`: automatic summaries of growing history are not part
  of v1.

## Triggers

A trigger source is an adapter with six methods (`validate`, `initial_state`, `prepare`, `admit`, `rearm`,
`normalize`; see `abstractruntime.triggers.protocol`). It reads persisted state and the time it is given, and never
does I/O, so the controller, the command applier and tests always get the same answer for the same inputs.

`trigger_sources()` lists every discovered source as `{descriptor, available, unavailable_reason?, name?}`. Each
descriptor carries `id`, `version`, `label`, `config_schema`, `event_schema` and `capabilities.kind` (`time`,
`manual` or `event`). `get_trigger_adapter(id, version)` returns an available source or raises
`UnknownTriggerSource`.

### `schedule@1`

Config: `{start_at?, every?, until?, count?, anchor?}`.

- Timestamps are RFC 3339 with an explicit offset (`2026-10-01T08:00:00Z`); a timestamp without an offset is refused.
- `every` is a whole number followed by `s`, `m`, `h` or `d` (`^[1-9][0-9]*[smhd]$`), at most `366d`. Units have
  fixed lengths in UTC, so there are no months and no daylight-saving shifts: `24h` is exactly 24 hours. Write weeks
  as `7d`.
- Ticks sit on a fixed grid, `T_k = anchor + k·every` (k = 0, 1, ...). They do not drift: an occurrence that starts
  7 seconds late does not move the next tick.
- `start_at` defaults to the creation time, and the first tick is `start_at` itself. `anchor` must equal `start_at`
  in v1.
- `until` is exclusive and must be after `start_at`. `count` (1 to 1,000,000) counts scheduled admissions; manual
  runs and retries do not count. `count` greater than 1 requires `every`.
- Without `every`, the automation fires once, at `start_at`, and is then exhausted.
- Missed ticks are **coalesced**. When several ticks are due at once (downtime, or a long occurrence), one occurrence
  runs, for the latest due tick. Its event payload is `{tick, scheduled_at, coalesced: {first_tick, last_tick,
  missed_count}}`, and an `automation.coalesced` record is written.
- The event id of a scheduled tick is `schedule@1:<binding_id>:<tick>`. A revision that changes the trigger gets a
  new `binding_id`, so the new schedule never reuses an old event id.

### `manual@1`

Config: `{}`. The automation runs only when asked with `automation.run_now`. A manual run of any automation, whatever
its trigger, carries a `manual` envelope with event id `manual:<command_id>`.

### Adding a source

Declare the adapter in the entry-point group `abstractruntime.trigger_sources`:

```toml
[project.entry-points."abstractruntime.trigger_sources"]
my_source = "my_package.triggers:MySourceAdapter"
```

- The entry-point name must equal `descriptor.id`, `descriptor.version` is an integer of at least 1, and
  `capabilities.kind` is `time`, `manual` or `event`.
- A third-party source that fails to load, has a bad descriptor, or duplicates another `id@version` is listed with
  `available: false` and a reason. It cannot be selected, and the other sources keep working.
- The built-in sources are required: if one is missing or broken, discovery raises `TriggerRegistryError`.
- `reset_trigger_registry()` forgets cached discovery (after installing a source package, or in tests).
- The v1 controller waits on deadlines (`prepare` returning `until`), on commands only (`idle`), or ends
  (`exhausted`). A source whose `prepare` returns an `event` wait fails the controller step.

## The controller

The controller is the packaged VisualFlow bundle `abstractframework.automation-controller@1.0.0` (package data under
`abstractruntime/automations/bundles/`). Every controller run keeps the workflow id
`abstractframework.automation-controller@1.0.0:controller`, so a restarted host resolves exactly this version.
`register_controller_bundle(registry)` registers the compiled flow, `controller_workflow_spec()` returns it, and
`controller_bundle_path()` returns the bundle directory for hosts that load bundles from disk.

```mermaid
flowchart TD
  start([start]) --> read_definition
  read_definition -- "continue" --> wait
  read_definition -- "end: archived or exhausted,<br/>nothing running" --> done([end])
  wait -- "go: due tick, run now,<br/>retry due, or occurrence pending" --> admit
  wait -- "park: WAIT_EVENT automation:id:wake<br/>until next tick / retry, or no deadline" --> wait
  wait -- "end" --> done
  admit -- "prepare: new occurrence admitted" --> prepare_context
  admit -- "dispatch: occurrence already pending" --> dispatch
  admit -- "rearm: nothing due (stale wake),<br/>backoff not due, or stopped" --> read_definition
  prepare_context --> dispatch
  dispatch -- "START_SUBWORKFLOW async + wait,<br/>deterministic run_id" --> record_outcome
  record_outcome -- "dispatch: child not finished" --> dispatch
  record_outcome -- "next: completed, failed,<br/>cancelled or retry scheduled" --> next_step[next]
  next_step --> read_definition
```

- **read_definition** activates the latest revision. When the automation is archived or its trigger is exhausted, and
  no occurrence is pending, the controller ends.
- **wait** decides from persisted state. With nothing to do it parks on one `WAIT_EVENT` with key
  `automation:<automation_id>:wake`, whose deadline is the next tick or the retry time; there is no deadline while
  paused or for a manual trigger. Commands wake this wait. A wake only makes the controller re-read its persisted
  state, so a stale or repeated wake never starts an extra occurrence.
- **admit** admits an occurrence only for a scheduled tick that is due while the automation is neither paused nor
  archived, or for a pending manual run. It freezes the occurrence's inputs (`prepared`: workflow, session,
  workspace, input data), the trigger envelope, the retry policy and the session kind. Every attempt uses these, so a
  later revision never changes an occurrence that has already been admitted.
- **prepare_context** checks that the frozen inputs resolve (values offloaded to the artifact store must load).
- **dispatch** records the attempt and starts the child with `START_SUBWORKFLOW {async: true, wait: true, run_id}`.
  The child runs outside the controller's tick, driven by the host. Its id is deterministic and it is created only if
  absent, so a dispatch replayed after a crash re-attaches to the same child.
- **record_outcome** reloads the child and reads its output; it never relies on the resume payload. Success completes
  the occurrence. A failure schedules a retry while attempts remain, and otherwise completes the occurrence as
  `failed`. A cancelled child completes as `cancelled`.

A controller that cannot find a seam it requires (the run and ledger stores on the tick, a child created under the
deterministic id) fails its step with `ControllerSeamError` instead of continuing.

### Occurrences

Each occurrence run carries `vars._meta.occurrence = {automation_id, occurrence_index, attempt, event_id, revision,
role: "occurrence", session_kind, fired_at, trigger_envelope}` and `workspace_root`. Its children carry the same
object with `role: "descendant"`; the runtime sets it on every hop and a child cannot clear or change it.

When the target's `input_data.prompt` is a string, it is prefixed on its own line with
`[Trigger schedule@1 · occurrence 3 · fired 2026-10-01T08:06:00+00:00]`, and the result is the occurrence's user
turn.

Run ids are deterministic:

| Occurrence | Run id |
|---|---|
| scheduled, attempt 1 | `uuid5(automation_id, "<revision>:<index>")` |
| manual (run now), attempt 1 | `uuid5(automation_id, "manual:<command_id>")` |
| attempt n ≥ 2 | the attempt-1 name with `":a<n>"` appended |

`occurrence_run_id(automation_id, revision=, index=, attempt=, command_id=)` computes them.

## Commands

`apply_automation_command(runtime, *, automation_id, command_id, type, payload=None, actor=None,
expected_revision=None, now=None)` is the only way to change an automation. It returns `{status: "applied" |
"rejected", error?, duplicate}` and returns rejections instead of raising them.

| Type | Effect | Rejected with |
|---|---|---|
| `automation.pause` | Stops **scheduled** admissions. The current occurrence and its retries finish, and `run_now` still works. This is separate from the runtime's own pause. | `invalid_state` when archived or finished |
| `automation.resume` | Re-arms the schedule at the first tick after now. It never fires on resume and never catches up the paused time. Resuming an automation that is not paused is an applied no-op. | `invalid_state` when archived or finished |
| `automation.run_now` | Runs one occurrence at the controller's next step, even while paused (the automation stays paused). There is no queue. | `automation_busy` while an occurrence or a manual run is pending; `invalid_state` when archived, finished or exhausted |
| `automation.revise` | `payload.changes = {title?, target?, trigger?, context?, policy?}` (at least one). `title`, `target`, `trigger` and `context` are replaced whole; `policy` fields are merged, so a field you do not send keeps its value. The next revision applies from the controller's next step. A changed trigger gets a new binding and is re-armed after now, so no past tick fires; an unchanged trigger keeps its binding. | `invalid_state` when archived or finished; the definition's own reason codes for invalid changes |
| `automation.stop_current` | Cancels the running occurrence tree, or its pending retry. The occurrence completes as `cancelled`, quietly. | `invalid_state` when nothing is running |
| `automation.archive` | Stops further admissions; the current occurrence finishes and the controller then ends. History is kept and the automation stays listed. Archiving twice is an applied no-op. | — |

Every command can also be rejected with:

- `automation_not_found`: no automation has that id;
- `invalid_request`: missing `command_id` or unknown `type`;
- `revision_conflict` (field `expected_revision`): `expected_revision` was given and differs from the current
  revision when the command is applied;
- `identity_conflict` (field `command_id`): the `command_id` was already used for a different command.

Commands are idempotent per `command_id`. A `command_id` is tied to its whole command (type, payload and
`expected_revision`). The `automation.command_result` record, keyed
`automation:command_result:<automation_id>:<command_id>`, is the decision. Sending the same command again returns the
recorded result with `duplicate: true` and repeats only the follow-ups that are safe to repeat: the observation
record (`automation.paused`, `automation.resumed`, `automation.revised`, `automation.archived`), waking the
controller, and cancelling the stopped child. A host that fails to carry out a command itself records the failure
with `record_automation_command_result(...)`, passing the same `payload` and `expected_revision` so a later replay is
recognized.

Commands run under the controller's `run_mutation_lock`, so they never interleave with a controller step. The
controller is woken with `max_steps=0`; the host (or `drive_automation`) then ticks it like any resumed run.

## Context: independent or growing

| | Independent (default) | Growing |
|---|---|---|
| Session | a new session per occurrence, named after its attempt-1 run id | the automation's session, `automation:<automation_id>` |
| History given to the occurrence | none | the automation's previous turns as `input_data.context.messages`: the most recent 50,000 tokens of whole turns (the session history window) |
| `use_context` given to the target | `false` | `true` |
| `_meta.occurrence.session_kind` | `occurrence` | `automation` |

The context mode alone decides whether the target reads history: at admission the runtime sets the target's
`use_context` input (and `include_context` when the target has it) to `true` for growing and `false` for
independent, whatever the definition's `target.input_data` says. A discussion always gets `use_context: true`.
The run records the decision as `_runtime.automation_context = {mode, use_context, target_use_context}`, where
`target_use_context` is the value the definition carried. Automations created before this rule, with
`use_context: false` in their target, now replay their history without being revised.

Growing history is read at admission and frozen with the occurrence's inputs, so every retry sees the same history.
It is read strictly: when it cannot be read (a store without a run index, for example), the admission fails instead
of running the occurrence without its context. History keeps whole turns, newest first, and says so in the oldest
kept message when older turns were dropped. The occurrence run records what was replayed in
`vars._runtime.session_history` (`replayed_messages`, `replayed_tokens`, `dropped_messages`, `dropped_tokens`,
`max_tokens`, ...), and the `automation.admitted` record carries the same values in its frozen inputs.

Only completed turns with both a prompt and an answer are replayed. A retried occurrence counts once, as its last
attempt.

## Discussions

`start_discussion(runtime, *, automation_id, occurrence_index, request_id, prompt, workspace_root,
actor_id=None)` starts a separate conversation about the automation as it stood at occurrence N, and returns
`{session_id, run_id, session_kind: "discussion"}`. Forking at a different point in time means choosing another N.

- It creates a new **root** run, `uuid5(automation_id, "discuss:<request_id>")`, in its own session
  `discussion-session:<run id>`. Request ids are therefore scoped to the automation. The same request returns the same
  discussion; a different request under the same `request_id` on that automation raises `identity_conflict`.
- The run uses the occurrence's workflow and frozen inputs, with `prompt` as the new user turn. The workflow must be
  registered on the runtime.
- It is seeded **once** with the automation's whole conversation through occurrence N, whatever the context mode:
  one user/assistant pair per finished occurrence 1..N (its last attempt), the occurrence's trigger/task turn and its
  answer, oldest first (`automation_timeline_messages(...)`). A failed or stopped occurrence stays in the timeline,
  its answer saying so. The pairs go through the session history window (the most recent 50,000 tokens of whole
  turns, no message ever cut); the oldest are dropped first. The first user message starts with a summary line: `[Automation "<title>": <n>
  occurrence(s) through occurrence N, showing the last K. The automation's files are mounted READ-ONLY at <path>;
  your own workspace <path> is writable.]`. The seed is stored in `_meta.discussion.seed_messages` of the root run.
- It works in its **own writable workspace**, `workspace_root`, which the host allocates (it must differ from the
  automation's). The automation's workspace is **mounted read-only** alongside it: it is reachable
  (`workspace_access_mode: "workspace_or_allowed"`, listed in `workspace_allowed_paths`) and protected by
  `_runtime.workspace_read_only_paths`, so reads work and writes, edits and moves into it are refused. Commands and
  tools run normally in the discussion's own workspace. `_meta.discussion.mounted_workspace` names the mount (see
  [Read-only mounts](#read-only-mounts)).
  When the occurrence's inputs carry the host's built-in protection (`workspace_builtin_deny_prefixes`, e.g. the
  gateway's data dir), the discussion keeps those deny prefixes unchanged and sets `workspace_builtin_allow` to
  exactly its own workspace and the mount, so both roots are usable and nothing else in the protected folders is.
- The automation's tool grant is removed: tools in a discussion ask for approval as in any chat.
- Nothing is ever written back into the automation's session, state, ledger or workspace.

Errors: `occurrence_not_found` (the automation has no occurrence N), `invalid_request` (empty `prompt` or
`request_id`), `automation_not_found`, `identity_conflict`, and `SessionHistoryError` when the seed cannot be read.

**Later turns stay anchored.** Every later root run started in a discussion session, by any caller, gets the
discussion's provenance (`_meta.discussion` without the seed) and the root's own workspace setup, whatever the caller
passed: `workspace_root`, `workspace_access_mode`, `workspace_allowed_paths`, the root's exact
`workspace_builtin_allow` when it has one (so the mount stays readable under the host's built-in protection and a
caller cannot widen the list; deny prefixes are left as they are) and `_runtime.workspace_read_only_paths` (a caller
may add mounts, never remove one). The whole-workspace `workspace_read_only` flag is applied only when the
root carries it. The discussion's root is validated first: every discussion run of the session must name
the same root, and that root must be a parent-less run of this session that carries the seed. If this check fails,
`Runtime.start` raises `SessionAttributionError` and the run is not created. Children of discussion runs carry the
discussion provenance too.

In a discussion session, `session_chat_messages` replays the seed first, as the oldest history, and drops it first
under the history window.

### Read-only mounts

`_runtime.workspace_read_only_paths` lists absolute folders that a run may read but not change, while its own
`workspace_root` stays writable. Discussions use it for the automation's workspace. Entries are resolved like
`pwd -P` (symlinks followed); the same key at the top level of the run vars is honoured too and can only add mounts.

- File tools classified `write` (`write_file`, `edit_file`, ...) are refused when their target path lies inside a
  mount; the message comes from `read_only_refusal(name, path=...)`.
- Reading inside a mount works (`read_file`, `list_files`, `search_files`, ...).
- Commands and code (`execute_command`, `shell_exec`, `execute_python`, ...) are **allowed**: the shell cannot be
  sandboxed, so a mount protects the file tools and VisualFlow writers, not what a command does.
- VisualFlow nodes that write files (`write_file`, `write_pdf`, `write_docx`, `write_chart`, `export_artifact`) are
  refused for a path inside a mount.
- Child runs and VisualFlow nodes inherit the mounts; they can add more but never clear or shrink them.
- A run with mounts and no `workspace_root` is refused.

Helpers in `abstractruntime.utils.workspace_paths`: `read_only_paths(vars)` (the resolved mounts) and
`path_is_read_only(vars, path)`; `READ_ONLY_PATHS_KEY` is the key name inside `_runtime`.

### Read-only workspaces

A run is read-only when its vars carry `workspace_read_only: true` or the trusted runtime key
`_runtime.workspace_read_only: true`. (Discussions use read-only MOUNTS instead: see above.) Under a read-only
workspace:

- tools classified `write` or `exec` in `TOOL_EFFECT_CLASSES` are refused (`write_file`, `edit_file`,
  `execute_command`, `shell_exec`, `local_helper_start`, `execute_python`, `self_improve`, ...), and so is every tool
  the table does not classify;
- tools classified `read`, `comms`, `delegate` and `memory-write` still run (`delegate` children inherit the
  read-only setting);
- VisualFlow nodes that write files (`write_file`, `write_pdf`, `write_docx`, `write_chart`, `export_artifact`) are
  refused;
- the workspace folder is never created, and a read-only run without a `workspace_root` is refused;
- child runs and VisualFlow nodes inherit the setting and cannot turn it off.

`abstractruntime.integrations.abstractcore.tool_effects` holds the table: `TOOL_EFFECT_CLASSES` maps each tool the
runtime can expose to `read`, `write`, `exec`, `delegate`, `comms` or `memory-write`, and `read_only_refusal(name)`
returns the refusal message for a tool, or `None` when it is allowed.

## Tool approval

`policy.tool_approval` decides whether the target's tools may run without asking.

- **`"auto"` (default).** An automation runs unattended and cannot ask a person before every tool call, so creating
  the automation is the consent; client forms state that its tools run without asking and list them. At admission,
  the runtime freezes the run's tool grant into the occurrence's inputs:
  `_runtime.tool_policy = {auto_approve_tools: [...], source: "automation-policy"}`. The tools named are the target's
  explicit `_runtime.allowed_tools` when it has that list, and otherwise every tool in `TOOL_EFFECT_CLASSES`. Child
  runs inherit the grant.
  - A name outside the run's tool ceiling (`allowed_tools`) grants nothing: approval never widens the ceiling.
  - A tool outside `TOOL_EFFECT_CLASSES` (a third-party MCP tool, for example) is not in the grant and still asks.
  - A `tool_policy` that the target's own `input_data._runtime` already carries is left as it is.
- **`"ask"`.** No grant. A tool call that needs approval waits on a `tool_approval` wait, as in a chat.

The grant is frozen with the rest of the occurrence's inputs: a revision of `tool_approval` applies from the next
occurrence. Questions a flow asks a person (`ask_user`) still wait in both modes. Discussions never inherit the grant.
See [tool-approval.md](tool-approval.md) for how `auto_approve_tools` combines with risk tiers.

## Waits on a person

`pending_waits(run_store, automation_id, *, limit=20)` lists the runs of the automation's occurrence trees that are
waiting on a person. Each item is typed from the wait record's structure, never from its text:

```text
{run_id, wait_key, kind: "ask_user" | "tool_approval" | "event", reason, index, prompt?, choices?, details?}
```

| `kind` | When | `details` | Answer with `Runtime.resume(run_id=..., wait_key=..., payload=...)` |
|---|---|---|---|
| `ask_user` | a `USER` wait that is not a pause (a question from the flow) | none | `{"response": "..."}` |
| `tool_approval` | a tool batch waiting for approval (`details.mode == "approval_required"`) | the calls approving will run, `[{name, arguments, call_id?}]` | `{"approved": true}` or `{"approved": false}` |
| `event` | an `EVENT` wait that carries a prompt or choices | `{scope, name}` when the wait uses the runtime's event key | `{"payload": ...}` |

`index` is the occurrence number. Paused runs and the controller's own wake wait are never listed.
`ANSWER_PAYLOADS`, `wait_kind(run)`, `typed_wait(run)` and `is_interactive_wait(run)` expose the same rules to hosts.
Waits on a person are live facts read from run state; they are not attention items.

## Notifications, attention and retries

### What asks for attention

Automations are **quiet by default**. An occurrence creates an attention item only when:

- its output carries `notify: true`: the item's title is the automation title and its body is the answer, cut to 280
  characters;
- its output carries `notify: {title, body}`: title at most 120 characters (the automation title when empty), body at
  most 2,000 characters. A missing, `false` or empty `notify` stays quiet;
- it **failed after its last retry**: the title is "<automation title> failed" and the body is the error. A failure
  that a retry fixed is quiet, and so is a cancelled occurrence.

The output is read by structure. An agent-style target ends with `response` / `success` / `meta`; a plain flow ends
with its end-node values or `{success, result}`. An output with `success: false` counts as a failure. A workflow that
must flag a result adds `notify` next to its answer.

Each notify or final failure creates exactly one attention item per occurrence. The item is carried by the
occurrence's `automation.completed` record with a per-automation sequence number.

`list_attention(ledger_store, automation_id, *, after_seq=0, cursor=None, limit=50)` returns
`{items, next_cursor}`, oldest first. Each item is `{kind: "notify" | "failure", automation_id, run_id, index, at,
title, body?, seq, cursor}`, and cursors look like `att1:<seq>`. A client should acknowledge only the cursor of the
last item it displayed, so items it never showed stay unseen.

### Retries

`policy.retry` allows 3 attempts in total by default. The delay before attempt `n + 1` is
`min(initial · factor^(n−1), max)`: by default 30 s, then 60 s, capped at 10 minutes. Each attempt is its own child
run (`...:a2`, `...:a3`) with the same frozen inputs and session. An `automation.retry_scheduled` record marks each
backoff, and `automation.completed` carries `attempts`. While an occurrence waits for its retry, scheduled ticks that
fall due are coalesced into one admission after it completes.

Retries do not make external effects exactly-once: a target that sends an email may send it again when it is retried.

## Reading automations

| Function | Returns |
|---|---|
| `get_automation(run_store, automation_id)` | `{automation_id, definition, active_revision, state, status, next_fire_at, current_occurrence}` |
| `list_occurrences(runtime, automation_id, *, cursor=None, limit=50)` | `{items, next_cursor}`, newest first, built from the ledger. Items: `{index, run_id, run_ids, attempts, revision, event_id, fired_at, trigger: {source_id, source_version}, user_turn, status, finished_at, notify, attention}`; `status` is `admitted`, `running`, `backoff`, `completed`, `failed` or `cancelled`; cursors look like `occ1:<index>` |
| `list_attention(...)`, `pending_waits(...)` | see above |
| `automation_queries.list_automations(run_store, *, status=None, cursor=None, limit=50)` | a `Page(items, next_cursor)` of automation summaries, newest first, with a cursor that survives restarts; archived automations stay listed; `status` filters on a value, a comma-separated string or a list |
| `automation_queries.automation_summary(controller_run)` | one summary: `{automation_id, title, status, revision, trigger, context_mode, target, session_id, workspace_root, next_fire_at, retry_at, current_occurrence, occurrence_count, pending_occurrence, last_outcome, archived_at, created_at, updated_at}` |
| `automation_queries.latest_occurrence(run_store, automation_id)` | the index row of the highest-numbered occurrence (its newest attempt), or `None` |
| `adopt_legacy_schedule_projection(run)` | a read-only summary of a legacy `scheduled:*` gateway root, marked `legacy: true` with capabilities `pause`, `resume` and `cancel`; legacy roots are never migrated |

Status is, in order of precedence: `archived` (the definition has `archived_at`, whatever state the controller
ended in), `failed` (the controller run failed, or was cancelled without being archived: it can never run again),
`completed` (the controller ended), `paused`, `active`. `get_automation` and `list_automations` share this one rule
(`automations.models.automation_status`).

`next_fire_at` is when the next occurrence will be admitted, decided exactly as the controller will decide it, so
a client never computes schedules itself. When the controller is parked it is the deadline of its wake wait: the
next tick, or the next attempt while an occurrence is in backoff (the summary then also sets `retry_at`). While an
occurrence runs it is the trigger adapter's answer on the persisted cursor: the next grid tick if it is still ahead,
otherwise the tick a coalesced admission will fire as soon as the running occurrence ends. It is `None` for a
manual trigger, when exhausted, and when the automation is not active (paused, archived, completed, failed).
`current_occurrence` is the occurrence in flight, `{index, run_id, attempt, status: "admitted" | "running" |
"backoff"}`, or `None`; clients read it instead of inferring "running" from the last outcome.
`get_automation` and the list summaries share both projections (`automations.controller.next_fire_at`,
`current_occurrence`).

`list_automations(changed_since=...)` is refused with `ChangedSinceUnsupported` (`unsupported_feature`): clients
poll complete pages.

Reading history is a pure read: it makes no provider or tool calls and writes nothing.

## Runs, sessions and history

Automations add structure to the run index that every built-in store keeps (SQLite, JSON files, in-memory, and the
offloading wrapper).

### Run attribution

Each run index row carries four fields derived from the run's `vars._meta` when it is saved:

| Run | `role` | `session_kind` | `automation_id` | `occurrence_index` |
|---|---|---|---|---|
| controller (`_meta.automation`) | `controller` | `automation` | its own id | — |
| occurrence (`_meta.occurrence`) | `occurrence` | `automation` (growing) or `occurrence` (independent) | the automation | N |
| child of an occurrence | `descendant` | as its occurrence | the automation | N |
| discussion (`_meta.discussion`) and its children | `discussion` | `discussion` | the automation | N |
| legacy gateway scheduled root (`_meta.schedule.kind == "scheduled_run"`) | `legacy_schedule` | `automation` | — | — |
| any other run | empty | `chat` | — | — |

`list_run_index(..., automation_id=, role=, session_kind=)` filters on these fields; each filter takes a value, a
comma-separated string (`session_kind="chat,discussion"`) or a list. Existing SQLite stores fill the columns once
when opened; the JSON store re-reads each run file once. Identity metadata (`vars._meta.automation`, `.occurrence`,
`.discussion`, `.creation_digest`) always stays inline: the offloading store never moves it to the artifact store.

Every row also carries `workspace_root`: the folder the run executes in, read from its top-level
`vars["workspace_root"]` as stored (host overrides and discussion folders included), only stripped of surrounding
whitespace and never resolved; `None` when the run has none. Existing SQLite stores fill it once when opened; the
JSON store re-reads each run file once. The offloading store never moves the key.

`session_attribution(run_store, session_id)` (`abstractruntime.core.run_attribution`) returns `None` for a session
with no runs, or `{"kind": ...}` with the session's most specific kind (`discussion`, then `automation`, then
`occurrence`, then `chat`). A discussion adds `discussion_root_run_id`, `automation_id`, `occurrence_index`,
`revision`, `workspace_root`, `discussion` (the root's `_meta.discussion` without the seed) and `workspace_policy`
(the root's workspace keys that later turns inherit). Every store exposes
`session_kinds(session_id)`, which answers from an index without scanning runs.

### Turn roots

A session's **turns** are its turn roots: runs without a parent, except automation controllers, plus automation
occurrences. `list_run_index(root_only=True)` returns turn roots, so an app that folds root runs into sessions shows
an automation's session as a chat whose turns are its occurrences.

`select_session_turns(run_store, session_id, *, include_occurrences=True, until_ms=None, automation_id=None,
through_occurrence=None, include_drafts=False, limit=50)` is the one definition of a session's turns, used by history
bundles and session replay. It returns turns oldest first and never includes child runs, automation controllers,
runtime-internal runs, legacy scheduled wrappers or (unless asked) draft-test runs. A retried occurrence is one turn,
its newest attempt. `automation_id` keeps only that automation's occurrences (other turns stay);
`through_occurrence=N` returns the history as it stood when occurrence N ran, however old N is, and raises
`OccurrenceNotInSession` when the session has no occurrence N, or `SessionHistoryError` when the store has no run
index to look it up in.

`session_chat_messages(..., automation_id=None, through_occurrence=None, strict=False)` replays those turns as
user/assistant message pairs under the session history window (the most recent `HISTORY_REPLAY_MAX_TOKENS` = 50,000
tokens of whole turns; see [API](api.md#sessions-and-history)). With `strict=True` it raises
`SessionHistoryError` (`reason_code = "history_unavailable"`) instead of returning a partial history: a store without
a run index, a missing occurrence, or a discussion whose seed is missing or cannot be read. Automation admission and
discussion seeding always read strictly.

## Storage guarantees

Automations rely on four store guarantees. Hosts that bring their own store must provide them.

- **Create-if-absent.** `RunStore.create_if_absent(run) -> (run, created)` creates a run only if its id is free. It is
  implemented by the SQLite, JSON-file, in-memory and offloading stores. The JSON store publishes a fully written
  temp file with an atomic hard link, so an existing run file is never replaced (filesystems without hard links raise
  instead). `Runtime.start(..., run_id=...)` and `START_SUBWORKFLOW` with `payload.run_id` use it: a start with the
  same identity (workflow, session, parent, `vars._meta.occurrence`, `vars._meta.creation_digest`) returns the
  existing run untouched, and a different identity raises `RunIdentityConflict` (`reason_code =
  "identity_conflict"`). A store without the primitive raises `NotImplementedError` on an explicit-id start;
  `store_supports_create_if_absent(store)` checks a store first. Recovery after a process crash is covered;
  power-loss durability is not claimed.
- **Per-run tick lock.** `run_mutation_lock(run_id)` is held by `Runtime.tick` for the whole tick, by
  `Runtime.resume` for its commit, and by every automation command. A host may take it around its own
  read-modify-save of a run; it must not modify a controller's `vars._meta.automation` or `vars._runtime.automation`
  directly, because `apply_automation_command` is the only supported writer of automation state. The lock is
  re-entrant and per process.
- **One writer process per store.** v1 supports one process that ticks, resumes and commands runs on a given store.
  Several store objects or read-only processes on one JSON run folder stay consistent: every creation and deletion is
  appended to a small creation journal (`.runs_created.log`), and each store applies new journal lines before a
  session or children lookup, so a discussion created through one store object is enforced read-only by another.
  `JsonFileRunStore.warm_session_index()` builds the session and children indexes at startup instead of on the first
  chat. The JSON store also removes run temp files older than 10 minutes, left behind by a crash, when it opens.
- **A run index for sessions.** `Runtime.start` checks the attribution of every root run that names a session. A
  store without a run index (no `session_kinds` and no `list_run_index`) cannot answer, and the start is refused with
  `SessionAttributionError` rather than risk starting an unanchored run in a discussion session. Strict history reads
  and `latest_occurrence` (which uses the store's `latest_occurrence_row`) need the index too.

### Crash safety

Each transition, whether a command or a controller step, is a *decision*:

1. Reconcile any unfinished decision.
2. Look up the decision's key exactly (`find_by_idempotency_key`). If it exists, the decision was already made and
   nothing new is appended.
3. Decide from the persisted state.
4. Save the intent (`_runtime.automation.intent`).
5. Append the record, which carries the full state change and its `state_version`.
6. Apply the change and save.

After a crash, the next step either applies the recorded change or drops the intent, so a crash at any point leaves
exactly one record and one application. A parent that crashes after starting a child with an explicit id, but before
saving its wait, finds the same child on replay and waits on it again; if the child already finished, the parent
receives its result directly.

## Limits in v1

- Schedules are fixed UTC intervals (`s`, `m`, `h`, `d`, at most `366d`). There are no cron expressions, calendar
  months, time zones or daylight-saving rules, and `anchor` must equal `start_at`.
- The shipped trigger sources are `schedule@1` and `manual@1`. There are no external or event triggers; the v1
  controller does not wait on `event` sources.
- Occurrences run one at a time (`serial`), missed ticks coalesce, and a failed occurrence never stops the
  automation (`failure: "continue"`). These policies cannot be changed.
- Growing history is the most recent 50,000 tokens of whole turns and is not summarized automatically; older turns
  drop out of the replay (they stay in the store).
- Tools outside `TOOL_EFFECT_CLASSES`, such as third-party MCP tools, still ask for approval under `auto`.
- Retries repeat external effects.
- One writer process per store; the mutation lock does not coordinate separate processes.
- `list_automations` has no change cursor (`changed_since` is refused); poll complete pages.
- Legacy gateway `scheduled:*` roots are listed read-only and are never migrated.

## See also

- [api.md](api.md#automations): the import surface
- [architecture.md](architecture.md#automations): where automations sit in the runtime
- [tool-approval.md](tool-approval.md): tool risk tiers and the run policy
- [faq.md](faq.md): common questions
- [troubleshooting.md](troubleshooting.md#an-automation-does-not-fire): symptom-oriented fixes


==============================================================================
# FILE: docs/tool-approval.md
==============================================================================

# Tool approval: risk tiers, the run-policy ceiling, and per-call refiners

AbstractRuntime decides, per tool call, whether to **auto-run** a tool or
**ask** the operator first. This is the runtime enforcement half of the
framework-wide *tool tiers* concept (operator ruling, tool-tiers wave 2026-07-23):
the gateway serves defaults and apps override them, but the decision that
actually gates execution runs here, at the effect boundary.

Implementation pointers (this repo):
- approval sets + policy: `src/abstractruntime/integrations/abstractcore/tool_executor.py`
- run-policy consumer + refiner dispatch: `src/abstractruntime/integrations/abstractcore/effect_handlers.py` (`_execute_with_run_policy`)
- risk facts + the served row shape: `src/abstractruntime/identity/tools.py` (`walled_tool_rows`), `src/abstractruntime/integrations/abstractcore/tool_inventory_facade.py` (`annotate_tool_rows`, `derive_risk_assessment`)
- the fact→tier mapping is **hosted by AbstractCore** (`abstractcore/tools/risk_facts.py`); runtime imports it (import-never-copy) and degrades to a byte-identical seed only under version skew.

## Two gates, never one

A tool call passes two independent gates:

1. **Availability** — is the tool present in the run's granted set at all?
   A tool above the run's grant is *absent* (not registered), not merely
   asked. Life-plane entity tools (memory/diary/reflection) are available by
   channel structure and are **not on the consent ladder** (`grantable:false`):
   stripping them is an identity risk the consent surface must be unable to
   express.
2. **Approval** — for an available, mutating-or-risky tool, auto-run or ask?

The gateway's grant model expresses both; this page documents the approval
gate, which is what `_execute_with_run_policy` computes.

## Risk facts and the derived tier

Every served tool row carries declared **facts** (never policy):
`mutating`, `remote_write_capable`, `comms_send`, `captures_environment`,
`standing_effect`, `destructive_capable`, plus the band-neutral
`model_controlled_destination`. A single versioned mapping derives the
operator-facing risk from those facts (max-wins):

| band (`risk_tier`) | `risk_rank` | facts |
| --- | --- | --- |
| `observe` | 1 | every declared fact false (read-only) |
| `act` | 2 | `mutating` or `remote_write_capable` |
| `outreach` | 3 | `comms_send` / `captures_environment` / `standing_effect` |
| `destroy` | 4 | `destructive_capable` (e.g. a shell reaching `rm`/`git reset`) |

Wire shape on every row: `risk_tier` is the **band word** (stable identity),
`risk_rank` is the **integer** (display/compare ordinal), `risk_presentation`
is the render word. A **factless** row (no declared facts — e.g. an
undeclared MCP tool) derives `risk_rank` 4 but `risk_presentation`
`"unvetted"`: gated at the top, never *rendered* "destroy" (deny-safe, but
honest that it is unvetted rather than proven-destructive).

## The run-policy ceiling

A run carries an optional policy under the model-unwritable
`_runtime.tool_policy` key (the gateway injects it; the model cannot set it):

- `auto_approve_tools` / `require_approval_tools` — explicit name lists, the
  finest grain. **`require` always wins** over any tier ceiling.
- `auto_approve_max_risk_rank` — a ceiling: a call whose tool derives
  `risk_rank <= N` auto-runs; above it, asks.

`model_controlled_destination` tools (the model chooses where output goes —
`fetch_url`, and see the refiner below) are **never silenced by the ceiling**:
the approval prompt is the exfiltration defense. Only an explicit
`auto_approve_tools` name — the operator's conscious act — overrides that.

## Per-call refiners (`send_email_recipient@v1`)

Some tools carry a `risk_refiner` id on their row. A refiner may **only lower**
a single call below its band, at approval time, when it can *prove* the call
is safe; it can never raise the band, and any call it cannot prove holds the
ceiling (deny-safe).

`send_email` (band `outreach`) declares `send_email_recipient@v1` (operator
ruling dm#244): a send to **the registered operator's own address**
auto-approves; a send to any other recipient asks.

- The operator address arrives as `_runtime.operator_email` — a
  **model-unwritable** key the gateway injects from the account record at run
  start (one source of truth; payload-supplied values are dropped).
- The refiner unions **all** recipient fields (`to`/`cc`/`bcc`); **every**
  recipient must equal the operator address (normalized strip + NFC +
  lowercase — no confusable/IDN folding, so a homoglyph domain never matches).
- Deny-safe at every gap → **ask**: no operator email configured (the feature
  is simply off — the email is optional), empty/unresolved recipients, a
  wrapper-nested argument shape, a display-name/group token
  (`Operator <op@self.com>` is compared verbatim, never bracket-parsed), any
  parse failure, or a version-skewed runtime with the refiner unregistered.

Until the operator email is configured, `send_email` stays in the
require-approval set: self and others both ask.


==============================================================================
# FILE: docs/tools-comms.md
==============================================================================

# Communication tools (`comms` toolset)

AbstractRuntime’s AbstractCore integration can expose an optional `comms` toolset (email, WhatsApp, Telegram). These tools are executed as **durable tool calls** via `EffectType.TOOL_CALLS`:
- tool requests/results are recorded in the **ledger** (`src/abstractruntime/core/models.py`)
- execution is controlled by the configured `ToolExecutor` (`src/abstractruntime/integrations/abstractcore/tool_executor.py`)

This document covers what is implemented in this repo: **toolset gating + wiring**. Provider credentials/config are defined by **AbstractCore tools**.

Implementation pointers (this repo):
- toolset gating: `src/abstractruntime/integrations/abstractcore/default_tools.py`
- tool execution: `src/abstractruntime/integrations/abstractcore/tool_executor.py`

## Enable (opt-in)

The `comms` toolset is disabled by default. Enable it via env vars (checked by `default_tools.comms_tools_enabled()`):

- `ABSTRACT_ENABLE_COMMS_TOOLS=1` (enable email + WhatsApp + Telegram)
- `ABSTRACT_ENABLE_EMAIL_TOOLS=1` (email only)
- `ABSTRACT_ENABLE_WHATSAPP_TOOLS=1` (WhatsApp only)
- `ABSTRACT_ENABLE_TELEGRAM_TOOLS=1` (Telegram only)

## Discover what gets enabled

```bash
python - <<'PY'
from abstractruntime.integrations.abstractcore.default_tools import list_default_tool_specs
comms = [s for s in list_default_tool_specs() if s.get("toolset") == "comms"]
print([s.get("name") for s in comms])
PY
```

## Wire into a runtime (local tool execution)

```python
import os

from abstractruntime.integrations.abstractcore import MappingToolExecutor, create_local_runtime
from abstractruntime.integrations.abstractcore.default_tools import get_default_tools

os.environ["ABSTRACT_ENABLE_COMMS_TOOLS"] = "1"

tool_executor = MappingToolExecutor.from_tools(get_default_tools())
rt = create_local_runtime(provider="ollama", model="qwen3:4b", tool_executor=tool_executor)
```

Notes:
- Install Runtime with `pip install abstractruntime`; AbstractCore tool integration and the MCP worker entry point are part of the base remote-light install.
- In untrusted deployments, prefer passthrough tools so a host/worker boundary approves and executes tool calls (`PassthroughToolExecutor` in `src/abstractruntime/integrations/abstractcore/tool_executor.py`).
- For local bridge-owned delivery flows, `ApprovalToolExecutor` can auto-run the Telegram send tools while requiring approval for email, WhatsApp, unknown tools, and write/command-style tools by default.
- Separate from the durable `TOOL_CALLS` path, Runtime also exposes **host wrappers** for operator-owned email and Telegram surfaces:
  - email helpers on `get_abstractcore_host_facade(runtime)` and `abstractruntime.integrations.abstractcore.comms_facade`
  - Telegram lifecycle/send wrappers in `abstractruntime.integrations.abstractcore.telegram_facade`
  - read/bootstrap helpers stay host-local and do not create run history by themselves
  - if an outbound send belongs to a run, prefer the durable run facade:
    `get_abstractcore_run_facade(runtime).send_email(...)` /
    `send_telegram_message(...)`
  - if that durable child run pauses for approval or passthrough execution, resume it via
    `get_abstractcore_run_facade(runtime).resume_tool_calls(...)`

## Credentials/config (provided by AbstractCore)

The actual comms tools live in AbstractCore:
- email + WhatsApp: `abstractcore.tools.comms_tools`
- Telegram: `abstractcore.tools.telegram_tools`

AbstractRuntime does **not** store secrets in run state. Secrets should be supplied as environment variables in the **process that executes the tool calls**.

Practical starting points (provided by AbstractCore; see `pyproject.toml` for the minimum supported version):
- Email:
  - `ABSTRACT_EMAIL_ACCOUNTS_CONFIG=/path/to/emails.yaml` (YAML/JSON config), or `ABSTRACT_EMAIL_{IMAP,SMTP}_*` env vars
  - Passwords are resolved indirectly via `*_PASSWORD_ENV_VAR` (default: `EMAIL_PASSWORD`)
  - Repo template (this repo): `emails.config.example.yaml` (static examples: OVH + Gmail)
- WhatsApp (Twilio):
  - defaults use `TWILIO_ACCOUNT_SID` and `TWILIO_AUTH_TOKEN`
- Telegram:
  - transport selection via `ABSTRACT_TELEGRAM_TRANSPORT` (`tdlib` default, or `bot_api`)
  - bot token default env var: `ABSTRACT_TELEGRAM_BOT_TOKEN`

## Security and privacy notes

- Tool calls and results are durable: message bodies, recipients, and response metadata may be persisted in the ledger and/or checkpoint vars.
- Keep secrets out of tool arguments; prefer env-var resolution. Even when a tool accepts `*_env_var` parameters, those should be **names**, not secret values.
- Treat run storage and ledgers as sensitive when enabling comms tools.

## See also

- `integrations/abstractcore.md` — AbstractCore wiring (`LLM_CALL`, `TOOL_CALLS`)
- `provenance.md` — tamper-evident ledger


==============================================================================
# FILE: docs/entity-runtime.md
==============================================================================

# Per-entity runtime (homes, leases, visit durability)

How AbstractRuntime hosts a **summoned entity**: one self-contained home
directory per entity, one `Runtime` per home, one writer at a time, and
visit turns that survive process restarts. This page documents the runtime
half of the entity-topology consensus plan ("1 gateway + N runtimes");
the door half (auth, stamps, serving) is AbstractGateway's.

## The home is the unit

An entity home is one directory holding the whole life:

```
entities/castor/
  manifest.json            # entity_id (birth marker — never a lookup key)
  spark.yaml               # the attested seed
  substrate.yaml           # operator's ONE mind choice (provider+model)
  memory.sqlite3           # graph + journal (abstractmemory)
  home.sqlite3             # diary book + command inbox
  runtime_castor.sqlite3   # run store + ledger (THIS page)
  artifacts/               # verbatims and run artifacts
  workspace/               # the entity's own files
  .writer_lease            # writer lease (see below)
```

**Copying the directory moves the whole life** — including pending runs,
durable waits, and commitments. Nothing at rest references the door's
address (relocation-stable keys).

## One writer per directory: the lease (`storage/lease.py`)

The lease is a **generic one-writer-per-directory mechanism** (entity homes
are its first consumer; project workplaces are the designed second — the
2026-07-10 vocabulary sign-off re-homed it from `identity/` to `storage/`).
It arbitrates *processes*, never principals: visitors' write-rights are
denied structurally at the deposit gate regardless of the lease.

Four writer windows exist for a home: a visit host, the own-time loop's
day, the dream window, and maintenance passes. Exactly one may hold the
directory at a time:

```python
from abstractruntime.storage.lease import acquire_directory_lease, DirectoryLeaseHeld, read_directory_lease

with acquire_directory_lease(home_dir, holder="visit-host", session_id="visit-1"):
    ...  # the directory is yours for this window

read_directory_lease(home_dir)   # {"holder": ..., "pid": ..., "held": True/False}
```

- Mechanics: `flock(LOCK_EX|LOCK_NB)` on `<dir>/.writer_lease`. The flock
  is the truth; the file's JSON metadata (holder kind, pid, acquired_at,
  session/run id) is diagnostics for refusal messages and "who holds
  Castor?".
- Refusal raises `DirectoryLeaseHeld` **naming the incumbent** — loud,
  never a silent wait. The own-time loop treats a held home as a *yield*
  back to its gate; the home-direct chat CLI refuses to double-summon.
- A crashed holder releases with its process (kernel drops the flock with
  the fd). A **copied** directory carries stale lease bytes but no lock —
  `read_directory_lease` answers `held` by a non-destructive flock probe,
  never by trusting metadata.
- Release truncates the file to a released record; it never unlinks
  (unlink races a concurrent acquirer onto a dead inode).
- The one-lease relay rule: cross-home delivery acquires ONE home's lease
  at a time, never two.

## One Runtime per entity (`identity/entity_runtime.py`)

```python
from abstractruntime.identity.entity_runtime import open_entity_runtime

ert = open_entity_runtime(home_dir, extra_handlers={EffectType.LLM_CALL: my_llm_handler})
run_id = ert.runtime.start(workflow=visit_workflow, vars={}, session_id="visit-1")
ert.runtime.tick(workflow=visit_workflow, run_id=run_id)
ert.close()   # checkpoints WAL so the directory is copy-clean
```

- The run store + ledger live in `runtime_<slug>.sqlite3` **inside the
  home** (slug = the directory name, the registry key).
- Effect handlers are the home's own seam + diary handlers (strict entity
  posture): `MEMORY_*` writes land in the home's graph, `DIARY_*` in the
  book, artifacts in the home's store. Handlers are RAW here — the gateway
  wraps door-served instances with stamp verification.
- `extra_handlers` lets the host add `LLM_CALL`/`TOOL_CALLS` at
  composition. Attempting to shadow a home handler **raises**: identity
  effects route through the home, never a host override.
- Every host `LLM_CALL` handler is automatically wrapped with the
  **act-only dereference** (below) — the privacy boundary cannot be
  forgotten.

## Act-only content: references at rest, words only in flight (`identity/act_only.py`)

Diary words must never rest outside the book. When an entity reads its own
diary mid-visit, the **durable transcript carries a typed reference**, not
the words:

```json
{"$act_only": {"tool": "diary_read", "entry_id": "diary_ab12", "gist": "one bounded line"}}
```

- The ref is the tool message's entire `content` (exact JSON, one top-level
  key). Detection is parse-based, never regex — and only on
  `role == "tool"` messages, which hosts append: a visitor pasting
  ref-looking JSON into their message stays inert text.
- At send time the wrapped `LLM_CALL` handler resolves refs **through the
  run's own `DIARY_READ` handler** into a wire copy (in-place content
  substitution; message identity untouched). The original payload — what
  the ledger and run store hold — keeps the ref.
- An unresolvable ref fails the effect **loud and non-retryable**; the
  provider never receives a degraded payload.
- Consequence for audits: like media `{"$artifact": ...}` refs, the wire
  payload is not byte-reconstructible from the ledger alone —
  reconstruction re-resolves refs through doors that enforce authority.

## Durable visit waits: event + deadline

A visit parks on the visitor's next message; the park may carry an idle
deadline (no reaper daemons):

```python
Effect(type=EffectType.WAIT_EVENT, payload={
    "wait_key": "visitor_input",
    "until": "2026-07-11T09:00:00+00:00",          # optional deadline (UTC-normalized)
    "details": {"kind": "visitor_message"},         # self-describing for clients
}, result_key="_temp.resume")
```

- An event resume before the deadline wins (payload lands in `result_key`).
- Past the deadline, `tick()` (or the scheduler's due-scan —
  deadline-carrying EVENT waits join `list_due_wait_until` on every store
  backend) resolves the wait with `{"timed_out": true}` so the workflow
  routes to its close path.
- The ledger wait record carries `wait_key`, `until`, and `details`
  together, so clients can render a chat composer for visit waits and the
  deadline without new transport.

## Tests

`tests/test_directory_lease.py`, `tests/test_entity_runtime.py`,
`tests/test_act_only_dereference.py`, `tests/test_wait_event_deadline.py` —
including cross-process lease exclusion, waits-traveling-on-home-copy, the
ledger-keeps-the-ref end-to-end pin, and the three-backend due-scan.


==============================================================================
# FILE: docs/mcp-worker.md
==============================================================================

# MCP worker (`abstractruntime-mcp-worker`)

AbstractRuntime ships an MCP worker that exposes AbstractCore toolsets over MCP (JSON-RPC) via:
- stdio (default)
- HTTP (optional)

Entry point:
- CLI script: `abstractruntime-mcp-worker` (`pyproject.toml`)
- implementation: `src/abstractruntime/integrations/abstractcore/mcp_worker.py`

## Install

```bash
pip install abstractruntime
```

## Run (stdio)

Choose toolsets explicitly (comma-separated):

```bash
abstractruntime-mcp-worker --toolsets files,web,system
```

Toolsets come from `get_default_toolsets()` (`src/abstractruntime/integrations/abstractcore/default_tools.py`). If comms tools are enabled, you can also expose `comms` (`docs/tools-comms.md`).

## Run (HTTP)

```bash
abstractruntime-mcp-worker --transport http --toolsets files,system --host 127.0.0.1 --port 8765
```

For anything beyond localhost, enable auth:

```bash
export ABSTRACT_WORKER_TOKEN="..."
abstractruntime-mcp-worker --transport http --toolsets files,system --http-require-auth
```

Optional origin allowlist (when clients send an `Origin` header):

```bash
abstractruntime-mcp-worker --transport http --toolsets files --http-allow-origin http://localhost:3000
```

## Security notes

- Exposing `system` tools can execute commands; treat the worker as privileged.
- Prefer stdio transport over an authenticated channel (e.g., SSH) when possible.

## See also

- `integrations/abstractcore.md` — tool executors and default toolsets


==============================================================================
# FILE: docs/evidence.md
==============================================================================

# Evidence capture

AbstractRuntime can record **provenance-first evidence** for selected “external boundary” tools (web + process execution). Evidence is stored durably as:
- a small index entry in `RunState.vars["_runtime"]["memory_spans"]`
- an artifact-backed payload (so checkpoints stay JSON-safe and bounded)

Implementation pointers:
- recorder: `src/abstractruntime/evidence/recorder.py`
- capture hook: `Runtime._maybe_record_tool_evidence(...)` (`src/abstractruntime/core/runtime.py`)
- retrieval helpers: `Runtime.list_evidence(...)` / `Runtime.load_evidence(...)` (`src/abstractruntime/core/runtime.py`)

## When evidence is recorded

Evidence capture runs best-effort after a successful `EffectType.TOOL_CALLS` step:
- only for tool names in `DEFAULT_EVIDENCE_TOOL_NAMES` (`web_search`, `fetch_url`, `execute_command`)
- only when an `ArtifactStore` is configured on the runtime (`Runtime(..., artifact_store=...)`)

If evidence capture fails, the runtime records a warning under `vars["_runtime"]["evidence_warnings"]` and continues execution.

## How to inspect evidence

```python
evidence = rt.list_evidence(run_id)
for e in evidence:
    print(e.get("tool_name"), e.get("created_at"), e.get("evidence_id"))

payload = rt.load_evidence(evidence_id="...")  # loads from ArtifactStore
```

## Storage and privacy

- Evidence payloads can include fetched page text or command stdout/stderr; treat artifacts and ledgers as sensitive.
- Secrets should never be passed as tool arguments (arguments are ledger-recorded). Prefer env-var resolution in tool implementations.

## See also

- `provenance.md` — tamper-evident ledger chain
- `architecture.md` — where evidence fits (runtime-owned, artifact-backed)


==============================================================================
# FILE: docs/snapshots.md
==============================================================================

# Snapshots (bookmarks)

A **snapshot** is a named, searchable checkpoint of a run state.

Motivation:
- debugging (“return to a known-good state”)
- observability (“inspect state at time T”)
- manual experimentation (“branch from snapshot later”)

Implementation: `src/abstractruntime/storage/snapshots.py`

## Data model

A snapshot stores:
- `snapshot_id`, `run_id`, optional `step_id`
- `name`, `description`, `tags`
- timestamps
- `run_state` (as a JSON dict)

## Stores

Included stores:
- `InMemorySnapshotStore` (tests/dev)
- `JsonSnapshotStore` (file-per-snapshot)

Search (MVP):
- filter by `run_id`
- filter by single `tag`
- substring match in `name` / `description`

## Restore semantics

Restoring a snapshot is a **host-level** operation:
1. load a snapshot from `SnapshotStore`
2. write `snapshot.run_state` back into your configured `RunStore`

Compatibility note:
- snapshot restore cannot guarantee safety if the workflow spec/node code has changed since the snapshot was taken.

## See also

- `architecture.md` — how snapshots fit with RunStore/LedgerStore/ArtifactStore


==============================================================================
# FILE: docs/provenance.md
==============================================================================

# Provenance (tamper-evident ledger)

AbstractRuntime’s ledger is an append-only journal of `StepRecord` entries. For audit/debug workflows, you can add **tamper-evidence** via a hash chain:
- each record carries `prev_hash` + `record_hash`
- modifications/reordering become detectable when you verify the chain

Implementation pointers:
- model fields: `src/abstractruntime/core/models.py` (`StepRecord.prev_hash`, `StepRecord.record_hash`, `StepRecord.signature`)
- hash-chain decorator + verifier: `src/abstractruntime/storage/ledger_chain.py`

## What is implemented (v0.4.9)

- `HashChainedLedgerStore(inner_store)` — wraps any `LedgerStore` to compute hashes on append
- `verify_ledger_chain(records)` — validates the chain and returns a verification report

Example:

```python
from abstractruntime import Runtime, WorkflowSpec
from abstractruntime.storage import InMemoryLedgerStore, InMemoryRunStore
from abstractruntime.storage.ledger_chain import HashChainedLedgerStore, verify_ledger_chain

ledger = HashChainedLedgerStore(InMemoryLedgerStore())
rt = Runtime(run_store=InMemoryRunStore(), ledger_store=ledger)

# ... run workflows ...

records = rt.get_ledger(run_id="...")  # list[dict]
report = verify_ledger_chain(records)
print(report.get("ok"), report.get("errors"))
```

## What is intentionally not implemented (yet)

- cryptographic signatures (non-forgeability)
- key management / delegation / revocation

Those belong in an optional extra (e.g., `abstractruntime[crypto]`) once the design is finalized.

## See also

- `architecture.md` — ledger as the source of truth
- `evidence` capture: `src/abstractruntime/evidence/recorder.py` (stores external-boundary evidence as artifacts + index)


==============================================================================
# FILE: docs/workflow-bundles.md
==============================================================================

# WorkflowBundles (`.flow`)

A **WorkflowBundle** is a portable distribution unit for VisualFlow JSON workflows:
- bundle format: zip file with `manifest.json`, `flows/*.json`, optional `assets/*`
- portability comes from shipping **VisualFlow JSON**, not `WorkflowSpec` (which contains Python callables)

Implementation pointers:
- manifest model: `src/abstractruntime/workflow_bundle/models.py`
- pack/unpack helpers: `src/abstractruntime/workflow_bundle/packer.py`, `src/abstractruntime/workflow_bundle/reader.py`
- on-disk registry: `src/abstractruntime/workflow_bundle/registry.py`
- compiler: `src/abstractruntime/visualflow_compiler/*`

The compiler also handles current VisualFlow authoring conveniences such as multi-entry execution fan-in. When a node has multiple incoming `exec-in` routes and per-route input overrides, the compiler lowers them into internal `join_exec` and `path_mux` nodes so the bundle remains portable and the runtime behavior stays explicit.

## Bundle layout

Minimal bundle:

```
manifest.json
flows/<flow_id>.json
```

Optional:

```
assets/<name>
```

## Packing a bundle

```python
from abstractruntime.workflow_bundle import pack_workflow_bundle

pack_workflow_bundle(
    root_flow_json="flows/root.json",
    out_path="out/my_bundle.flow",
    bundle_id="my_bundle",
    bundle_version="0.1.0",
)
```

`pack_workflow_bundle(...)` is stdlib-only and validates that referenced subflows exist in `flows_dir` (defaults to the root file’s directory).

## Reading a bundle

```python
from abstractruntime.workflow_bundle import open_workflow_bundle

b = open_workflow_bundle("out/my_bundle.flow")
print(b.manifest.bundle_id, b.manifest.bundle_version)
print([ep.flow_id for ep in b.manifest.entrypoints])
```

## Registry (installed bundles)

`WorkflowBundleRegistry` is a host-side convenience layer for storing and resolving `.flow` bundles from a directory:
- default directory resolution: `default_workflow_bundles_dir()` (`src/abstractruntime/workflow_bundle/registry.py`)
- resolve `bundle_id[@version]` and entrypoints (`resolve_bundle`, `resolve_entrypoint`)

Default directory resolution checks `ABSTRACTFRAMEWORK_WORKFLOWS_DIR`, then AbstractFlow authoring env names, then `./flows/bundles/`, then `~/.abstractframework/workflows/`. Hosts with Gateway-specific flow settings should pass `bundles_dir` explicitly or translate them to the shared framework env name before constructing the registry.

## VisualFlow multi-entry fan-in

Visual authoring tools may connect more than one execution edge into the same target `exec-in` pin. For example, a first prompt can enter a node from `on_flow_start`, while a later loop can re-enter the same node with a different prompt produced by the previous turn.

Store two metadata fields on the target node:
- `entryRoutes`: ordered execution entries. Each route has a stable `key`, `sourceNodeId`, and `sourceHandle`.
- `inputRouteOverrides`: per-input route overrides. Shape: `pinId -> routeKey -> {sourceNodeId, sourceHandle}`.

Minimal target-node fragment:

```json
{
  "id": "ask",
  "type": "ask_user",
  "data": {
    "pinDefaults": {"prompt": "start"},
    "entryRoutes": [
      {"key": "start::exec-out", "sourceNodeId": "start", "sourceHandle": "exec-out"},
      {"key": "ask::exec-out", "sourceNodeId": "ask", "sourceHandle": "exec-out"}
    ],
    "inputRouteOverrides": {
      "prompt": {
        "ask::exec-out": {"sourceNodeId": "ask", "sourceHandle": "response"}
      }
    }
  }
}
```

Compiler behavior:
- incoming exec edges are rerouted through an internal `join_exec` node
- overridden pins are routed through internal `path_mux` nodes
- the selected route is persisted in run state, so pause/resume and file-store restarts keep the same input selection
- stale metadata is rejected when `entryRoutes` no longer matches the incoming exec edges

Authoring guidance:
- use the default route key `${sourceNodeId}::${sourceHandle}` unless your editor needs a custom stable key
- keep route keys unique per target node
- use one normal data edge or `pinDefaults` for the fallback value, then add `inputRouteOverrides` only for routes that need a different value

## See also

- `architecture.md` — VisualFlow → WorkflowSpec compilation path


==============================================================================
# FILE: docs/manual_testing.md
==============================================================================

# Manual testing

This guide is a small set of **manual smoke tests** you can run to verify the durable runtime loop (start/tick/wait/resume), scheduler resumption, and persistence.

## Prerequisites

From the repo root:

```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
python -m pip install -e .
```

## Test 1: Zero-config hello world

```python
from abstractruntime import create_scheduled_runtime, StepPlan, WorkflowSpec


def greet(run, ctx):
    name = run.vars.get("name", "World")
    return StepPlan(node_id="greet", complete_output={"message": f"Hello, {name}!"})


workflow = WorkflowSpec(
    workflow_id="hello",
    entry_node="greet",
    nodes={"greet": greet},
)

sr = create_scheduled_runtime()
run_id, state = sr.run(workflow, vars={"name": "Alice"})

print(state.status.value)
print(state.output)

sr.stop()
```

Expected:
- status is `completed`
- output contains `{"message": "Hello, Alice!"}`

---

## Test 2: Ask user (pause + resume)

```python
from abstractruntime import create_scheduled_runtime, Effect, EffectType, StepPlan, WorkflowSpec, RunStatus


def ask_name(run, ctx):
    return StepPlan(
        node_id="ask",
        effect=Effect(
            type=EffectType.ASK_USER,
            payload={"prompt": "What is your name?"},
            result_key="user_input",
        ),
        next_node="greet",
    )


def greet(run, ctx):
    name = run.vars.get("user_input", {}).get("text", "Unknown")
    return StepPlan(node_id="greet", complete_output={"greeting": f"Hello, {name}!"})


workflow = WorkflowSpec(
    workflow_id="ask_and_greet",
    entry_node="ask",
    nodes={"ask": ask_name, "greet": greet},
)

sr = create_scheduled_runtime()
run_id, state = sr.run(workflow)
assert state.status == RunStatus.WAITING
print(state.waiting.prompt)

state = sr.respond(run_id, {"text": "Bob"})
assert state.status == RunStatus.COMPLETED
print(state.output)

sr.stop()
```

Expected:
- first run blocks with `status=waiting` and a `prompt`
- after `respond`, run completes with a greeting

---

## Test 3: Wait until (scheduler auto-resume)

```python
from datetime import datetime, timedelta, timezone
import time

from abstractruntime import create_scheduled_runtime, Effect, EffectType, StepPlan, WorkflowSpec, RunStatus


def schedule_task(run, ctx):
    until = (datetime.now(timezone.utc) + timedelta(seconds=2)).isoformat()
    return StepPlan(
        node_id="schedule",
        effect=Effect(type=EffectType.WAIT_UNTIL, payload={"until": until}),
        next_node="execute",
    )


def execute_task(run, ctx):
    return StepPlan(node_id="execute", complete_output={"ok": True})


workflow = WorkflowSpec(
    workflow_id="scheduled_task",
    entry_node="schedule",
    nodes={"schedule": schedule_task, "execute": execute_task},
)

sr = create_scheduled_runtime(poll_interval_s=0.2)
run_id, state = sr.run(workflow)
assert state.status == RunStatus.WAITING
print("waiting until:", state.waiting.until)

for _ in range(20):
    time.sleep(0.2)
    state = sr.get_state(run_id)
    if state.status == RunStatus.COMPLETED:
        break

print(state.status.value, state.output)
sr.stop()
```

Expected:
- run first blocks with `wait_reason=until`
- within a few seconds, scheduler resumes and the run completes

---

## Test 4: Persistence (survive restart)

```python
import tempfile
from pathlib import Path

from abstractruntime import (
    create_scheduled_runtime,
    Effect,
    EffectType,
    StepPlan,
    WorkflowSpec,
    JsonFileRunStore,
    JsonlLedgerStore,
    RunStatus,
)


def ask(run, ctx):
    return StepPlan(
        node_id="ask",
        effect=Effect(type=EffectType.ASK_USER, payload={"prompt": "Continue?"}, result_key="answer"),
        next_node="done",
    )


def done(run, ctx):
    answer = run.vars.get("answer") or {}
    text = answer.get("text") if isinstance(answer, dict) else None
    return StepPlan(node_id="done", complete_output={"answer": text})


workflow = WorkflowSpec(workflow_id="persistent_wf", entry_node="ask", nodes={"ask": ask, "done": done})
data_dir = Path(tempfile.mkdtemp())

# Session 1: start + block
sr1 = create_scheduled_runtime(run_store=JsonFileRunStore(data_dir), ledger_store=JsonlLedgerStore(data_dir))
run_id, state = sr1.run(workflow)
assert state.status == RunStatus.WAITING
sr1.stop()

# Session 2: “restart”, reload + resume
sr2 = create_scheduled_runtime(
    run_store=JsonFileRunStore(data_dir),
    ledger_store=JsonlLedgerStore(data_dir),
    workflows=[workflow],  # re-register
)
state = sr2.get_state(run_id)
assert state.status == RunStatus.WAITING
state = sr2.respond(run_id, {"text": "yes"})
assert state.status == RunStatus.COMPLETED
print(state.output)
sr2.stop()
```

Expected:
- run id remains valid after a restart
- ledger and checkpoint files exist under `data_dir`

---

## Test 5: Find waiting runs

```python
from abstractruntime import create_scheduled_runtime, Effect, EffectType, StepPlan, WorkflowSpec, WaitReason, RunStatus


def wait_for_event(run, ctx):
    return StepPlan(node_id="wait", effect=Effect(type=EffectType.WAIT_EVENT, payload={"wait_key": f"event_{run.run_id[:8]}"}))


workflow = WorkflowSpec(workflow_id="event_wf", entry_node="wait", nodes={"wait": wait_for_event})

sr = create_scheduled_runtime()
ids = [sr.run(workflow)[0] for _ in range(3)]

waiting = sr.find_waiting_runs()
waiting_events = sr.find_waiting_runs(wait_reason=WaitReason.EVENT)

print("waiting:", len(waiting), "events:", len(waiting_events))
assert all(r.status == RunStatus.WAITING for r in waiting_events)

sr.stop()
```

Expected:
- at least 3 waiting runs are listed
- filtering by `WaitReason.EVENT` works

## Run the automated tests

```bash
python -m pytest -q
```

Expected: the test suite passes.

Some integration tests depend on optional local services or packages (for example a configured Ollama model or `lancedb`). In lean development environments, run the focused unit tests for your change first, then run the full suite in a fully provisioned integration environment before release.

## See also

- `getting-started.md` — first steps
- `../examples/README.md` — runnable examples
- `architecture.md` — where these behaviors come from


==============================================================================
# FILE: docs/adr/README.md
==============================================================================

# Architectural Decision Records (ADRs)

ADRs document significant architectural decisions made during AbstractRuntime development. They explain *why* certain approaches were chosen, not *what* was built (that's in the backlog).

## Why ADRs Matter

When you ask "why is it designed this way?", the answer is in an ADR. ADRs are:
- **Immutable**: Once accepted, they are not edited (only superseded by new ADRs)
- **Historical**: They capture the context and constraints at decision time
- **Educational**: They help new contributors understand the architecture

## Index

| ID | Title | Status | Date | Summary |
|----|-------|--------|------|---------|
| 0001 | [Layered Coupling with AbstractCore](0001_layered_coupling_with_abstractcore.md) | Accepted | 2025-12-11 | Kernel stays dependency-light; AbstractCore integration is opt-in |
| 0002 | [Execution Modes](0002_execution_modes_local_remote_hybrid.md) | Accepted | 2025-12-11 | Support local, remote, and hybrid execution topologies |
| 0003 | [Provenance Hash Chain](0003_provenance_tamper_evident_hash_chain.md) | Accepted | 2025-12-11 | Tamper-evident ledger first; cryptographic signatures deferred |
| 0004 | [Runtime Owns Run-Scoped Media Execution Truth](0004_runtime_owns_run_scoped_media_execution_truth.md) | Accepted | 2026-05-20 | Hosts must route run-scoped media execution through Runtime |
| 0005 | [Runtime Owns AbstractCore Host Discovery Queries](0005_runtime_owns_abstractcore_host_discovery_queries.md) | Accepted | 2026-05-20 | Hosts should ask Runtime for Core discovery/catalog snapshots |
| 0006 | [Runtime Owns Durable AbstractCore Bloc Prompt-Cache Control](0006_runtime_owns_durable_abstractcore_bloc_prompt_cache.md) | Accepted | 2026-05-20 | Hosts should use Runtime for durable bloc/KV controls and binding-aware execution |
| 0007 | [Runtime Relays Core-Owned Model Residency Truth](0007_runtime_relays_core_owned_model_residency_truth.md) | Accepted | 2026-05-21 | Runtime reports loaded state only from AbstractCore residency truth |

## Relationship to Backlog

ADRs explain *why*. Backlog items explain *what* and *how*.

| ADR | Related Implementation |
|-----|------------------------|
| 0001 | `backlog/completed/005_abstractcore_integration.md` |
| 0002 | `backlog/completed/005_abstractcore_integration.md` |
| 0003 | `backlog/completed/007_provenance_hash_chain.md`, `backlog/planned/008_signatures_and_keys.md` |
| 0004 | `backlog/completed/023_truthful_local_media_residency_boundaries.md`, `backlog/completed/024_runtime_owned_run_scoped_media_execution.md` |
| 0005 | `backlog/completed/026_runtime_host_discovery_facade_for_core_catalogs.md` |
| 0006 | `backlog/completed/027_runtime_durable_bloc_prompt_cache_facade.md` |
| 0007 | `backlog/completed/0035_model_residency_provider_truth_for_local_http_clients.md`, `backlog/proposed/0036_local_media_residency_bridge_to_core_residency.md` |

## Adding New ADRs

When making a significant architectural decision:
1. Create `docs/adr/NNNN_short_title.md`
2. Use the template: Status, Context, Decision, Consequences
3. Set status to "Accepted" once the decision is final
4. If superseding an old ADR, update the old one's status to "Superseded by NNNN"


==============================================================================
# FILE: examples/README.md
==============================================================================

# AbstractRuntime Examples

Runnable examples demonstrating AbstractRuntime capabilities.

## Quick Start

```bash
cd examples
python 01_hello_world.py
```

## Examples

| Example | Description | Dependencies |
|---------|-------------|--------------|
| 01_hello_world.py | Minimal workflow with zero-config | None |
| 02_ask_user.py | Pause for user input, resume with response | None |
| 03_wait_until.py | Schedule a task for later | None |
| 04_multi_step.py | Multi-node workflow with branching | None |
| 05_persistence.py | File-based storage, survive restart | None |
| 06_llm_integration.py | LLM call with AbstractCore | abstractcore, ollama |
| 07_react_agent.py | Full ReAct agent with tools | abstractcore, abstractagent, ollama |

## Requirements

Examples 1-5 only require abstractruntime:
```bash
pip install abstractruntime
```

Examples 6-7 use Runtime's base AbstractCore integration:
```bash
pip install abstractruntime
# Example 07 also requires AbstractAgent (separate package/repo).
# Also requires Ollama running locally with qwen3:4b-instruct-2507-q4_K_M
```

## Running Examples

Each example is self-contained. Run directly:

```bash
python 01_hello_world.py
```

For interactive examples (02, 05), follow the prompts.


==============================================================================
# FILE: SECURITY.md
==============================================================================

# Security Policy

## Reporting a vulnerability

Please report security issues **privately**.

Preferred channel:
- Use **GitHub Security Advisories** / the repository’s “Report a vulnerability” feature (private).

Include as much of the following as you can:
- affected versions (from `pyproject.toml` / `CHANGELOG.md`)
- impact and realistic attack scenario
- minimal reproduction steps or proof-of-concept
- environment details (OS, Python version, storage backend used)

## Coordinated disclosure

- Do not open public issues/PRs for security vulnerabilities.
- Avoid data exfiltration, service disruption, or destructive testing; keep verification to the minimum needed.

## Non-security bugs

If you are unsure whether an issue is security-related, prefer reporting it privately first.


==============================================================================
# FILE: CONTRIBUTING.md
==============================================================================

# Contributing to AbstractRuntime

Thanks for your interest in contributing!

AbstractRuntime is a **durable workflow runtime** (interrupt → checkpoint → resume) with an append-only execution ledger.

## Quick start (dev setup)

Prereqs: **Python 3.10+**.

Recommended (workspace checkout): develop inside the [AbstractFramework](https://github.com/lpalbou/AbstractFramework) workspace.  
The test bootstrap (`tests/conftest.py`) will auto-wire sibling projects on `sys.path` (e.g., `abstractcore/`, `abstractmemory/`, `abstractsemantics/`, `abstractflow/`).

```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install -U pip

# Full dev install (runtime + docs/test tooling)
python -m pip install -e ".[test,docs]"

python -m pytest -q
```

Inside the AbstractFramework workspace, prefer `python -P -m pytest tests -q`: `-P` keeps the current folder off
`sys.path`, so a workspace folder named like a package (for example `abstractflow/`) cannot shadow it.

If you cloned **only** this repo (without the AbstractFramework workspace), make sure the sibling packages above are importable (install them or clone them next to this repo) before running the full test suite.

## Repo map (source of truth)

- Public exports: `src/abstractruntime/__init__.py` (keep this consistent with `docs/api.md`)
- Core kernel (durable semantics): `src/abstractruntime/core/`
- Durability backends: `src/abstractruntime/storage/`
- Driver loop (in-process): `src/abstractruntime/scheduler/`
- Runtime integrations: `src/abstractruntime/integrations/`
- Tests: `tests/`

Docs entrypoints:
- `README.md` → `docs/getting-started.md`
- Docs index: `docs/README.md`
- Architecture: `docs/architecture.md`

## Change guidelines

### Code

- Preserve durability invariants: values stored in `RunState.vars` must stay JSON-serializable (`src/abstractruntime/core/models.py`).
- Add/adjust tests for new behavior (see `tests/`).
- If you touch effect semantics, update `docs/architecture.md` and ensure handlers and models stay aligned.

### Documentation

Docs should be **user-facing**, **actionable**, and anchored to code (prefer referencing `src/...` paths for claims).

When behavior changes, update:
- `docs/api.md` (public API surface + imports)
- `docs/getting-started.md` (onboarding examples)
- `docs/architecture.md` (semantics/invariants)
- `CHANGELOG.md` (user-visible changes)

List every `docs/*.md` page in `docs/README.md`. Keep `llms.txt` (the hand-curated index) and
`llms-full.txt` in step with the documentation in the same change. `llms-full.txt` is generated:
run `python scripts/generate_llms_full.py` after editing any page it includes, and
`python scripts/generate_llms_full.py --check` to confirm it is current (it exits 1 when the file
is stale). Add a page to the script's `DOCUMENTS` list when it joins the core set.

## Releases

- Bump `version` in `pyproject.toml`
- Add a dated section to `CHANGELOG.md` (Keep a Changelog format)


==============================================================================
# FILE: CODE_OF_CONDUCT.md
==============================================================================

# Code of Conduct

## Our Standard

This project is maintained as a professional software collaboration. Contributors, maintainers, and users are expected
to keep discussions respectful, technically focused, and welcoming to people with different backgrounds and experience
levels.

Examples of expected behavior:

- Use clear, constructive language when giving feedback.
- Assume good faith while still asking for evidence and reproducible details.
- Keep disagreements focused on the code, docs, design, or release process.
- Respect privacy and do not publish private contact details, credentials, logs, or user data.

Examples of unacceptable behavior:

- Harassment, threats, insults, or discriminatory language.
- Sustained off-topic disruption of issues, pull requests, or discussions.
- Publishing private information without explicit permission.
- Pressuring maintainers or contributors to bypass safety, security, or release checks.

## Reporting

Report conduct concerns privately to the maintainer contact listed in the package metadata or through the repository
owner's GitHub profile. Include the relevant links, screenshots, or context when possible.

Maintainers may remove comments, close threads, block accounts, or restrict repository access when needed to protect the
project and its contributors.


==============================================================================
# FILE: ACKNOWLEDGMENTS.md
==============================================================================

# Acknowledgments

AbstractRuntime is designed to pair with the wider Abstract ecosystem and the open-source Python tooling community.

This project depends on (and is shaped by) the following libraries.
The canonical dependency list lives in `pyproject.toml`.

## Runtime dependencies (core install)

- **abstractsemantics** (`>=0.0.5`) — structured schema registry support (declared in `pyproject.toml`, used in `src/abstractruntime/integrations/abstractmemory/effect_handlers.py` and VisualFlow execution wiring).
- **AbstractMemory** (`>=0.3.0`) — TripleStore models and store contract used by Runtime's `MEMORY_KG_*` effects. Durable/vector backend dependencies such as LanceDB remain selected by hosts.

## Runtime integrations
- **abstractcore** (`abstractcore[remote,tools,vision,voice,audio,music]>=2.15.1`) — LLM, tools, media, and capability integration used by the base `abstractruntime` install; the `apple` and `gpu` extras select `abstractcore[all-apple]` / `abstractcore[all-gpu]` at the same floor (declared in `pyproject.toml`, implementation under `src/abstractruntime/integrations/abstractcore/*`, docs: `docs/integrations/abstractcore.md`).
  - The AbstractCore integration uses **httpx** for remote mode (`src/abstractruntime/integrations/abstractcore/llm_client.py`) and **pydantic** for structured validation (`src/abstractruntime/integrations/abstractcore/effect_handlers.py`). These are provided by AbstractCore’s dependency set.
  - The `tools` extra of AbstractCore backs the base Runtime toolset and the `abstractruntime-mcp-worker` entry point.
- **RestrictedPython** (`>=7.0`, core install) — sandbox for VisualFlow “Code” nodes and pin expressions (`src/abstractruntime/visualflow_compiler/visual/code_executor.py`).
- **pypdf** and **reportlab** — VisualFlow `Read PDF` / `Write PDF` nodes in the base install.

## Build & test tooling

- **hatchling** — build backend (`pyproject.toml` `[build-system]`).
- **pytest** — test runner (`pytest.ini`, `tests/`).

And thanks to everyone who reports bugs, discusses design tradeoffs, and contributes improvements.

See also: `LICENSE`, `CONTRIBUTING.md`.


==============================================================================
# FILE: ROADMAP.md
==============================================================================

# AbstractRuntime Roadmap

## Current status

What changed in each release is in [CHANGELOG.md](CHANGELOG.md); planned work is tracked in the [AbstractFramework backlog](https://github.com/lpalbou/abstractframework/tree/main/docs/backlog).

AbstractRuntime provides a durable workflow kernel plus optional integrations:
- durable execution: `Runtime.start/tick/resume`, explicit `WaitState` (`src/abstractruntime/core/runtime.py`)
- append-only ledger (`StepRecord`) + persistent stores (JSON/JSONL, SQLite) (`src/abstractruntime/storage/*`)
- built-in scheduler (`Scheduler`, `ScheduledRuntime`) (`src/abstractruntime/scheduler/*`)
- snapshots/bookmarks (`src/abstractruntime/storage/snapshots.py`)
- tamper-evident hash-chained ledger (`src/abstractruntime/storage/ledger_chain.py`)
- artifacts + offloading for large payloads (`src/abstractruntime/storage/artifacts.py`, `src/abstractruntime/storage/offloading.py`)
- retries/idempotency hooks (`src/abstractruntime/core/policy.py`)
- VisualFlow compiler + WorkflowBundles (`src/abstractruntime/visualflow_compiler/*`, `src/abstractruntime/workflow_bundle/*`)
- AbstractCore integration for `LLM_CALL` / `TOOL_CALLS` (`docs/integrations/abstractcore.md`)

## Longer-term (not scheduled)

- distributed scheduling primitives (beyond in-process polling)
- workflow versioning/migration patterns for long-lived runs and snapshot restore
- stronger reproducibility contracts for replays (workflow snapshotting + run history bundles)


==============================================================================
# FILE: docs/README.md
==============================================================================

# Documentation

This folder contains **user-facing docs** (how to use AbstractRuntime) and **maintainer docs** (ADRs/backlog).

If you are new: read `getting-started.md` → `api.md` → `architecture.md`.

## Ecosystem

AbstractRuntime is part of the wider AbstractFramework ecosystem:
- AbstractFramework umbrella: [lpalbou/AbstractFramework](https://github.com/lpalbou/AbstractFramework)
- AbstractCore (LLM + tools): [lpalbou/abstractcore](https://github.com/lpalbou/abstractcore)

In this repo, the AbstractCore wiring lives under `src/abstractruntime/integrations/abstractcore/*` and is documented in `integrations/abstractcore.md`.

## Start here

- `../README.md` — install + quick start
- `getting-started.md` — first steps (recommended)
- `api.md` — public API surface (imports + pointers)
- `architecture.md` — how the runtime is structured (with diagrams)
- `troubleshooting.md` — symptom-oriented setup, runtime, and integration fixes
- `proposal.md` — design goals and scope boundaries

## Guides

- `faq.md` — common questions (recommended)
- `manual_testing.md` — manual smoke tests and how to run the test suite
- `artifacts.md` — Runtime artifact identity, descriptors, provenance, catalog search, and access stats
- `integrations/abstractcore.md` — wiring `LLM_CALL` / `TOOL_CALLS`, the `config_facade` models/engines/host-jobs passthroughs for hosts, cached sessions with per-session prompt-cache listing/clearing, host memory snapshots, model-residency listings, locks and model switching, live token streaming, workspace-scoped tools, durable bloc prompt-cache control, media inputs, generated media outputs, video progress events, and tool approval waits via AbstractCore
- `tools-comms.md` — enabling the optional comms toolset (email/WhatsApp/Telegram)
- `tool-approval.md` — tool risk tiers, the run-policy rank ceiling, and per-call refiners (the `send_email` self-recipient rule)
- `api.md#workflowbundles-flow-and-visualflow-distribution` — VisualFlow compiler APIs, media nodes, and document nodes (`read_pdf` / `write_pdf` / `write_docx`)
- `api.md#sessions-and-history` — session history replay: one window, the most recent 50,000 tokens of whole turns (`HISTORY_REPLAY_MAX_TOKENS`), returned as a `ReplayedHistory` whose `.report` records what was replayed and dropped
- `api.md#run-history-bundle-export-portable-replay-artifact` — run history bundles, including `resolved_actions` summaries for cross-client capability replay

## Features (reference)

- `automations.md` — automations: a workflow run on a schedule or on request as a durable controller run; triggers (`schedule@1`, `manual@1`, entry-point sources), commands, independent or growing context, discussions forked at any occurrence (own writable workspace, the automation's workspace mounted read-only), tool approval, typed waits, notifications and retries, and the storage guarantees (create-if-absent, per-run lock, turn roots, session kinds, one writer process per store)
- `entity-runtime.md` — per-entity runtimes for summoned entities: homes, the one-writer lease, act-only diary privacy (`$act_only` refs), durable visit waits with deadlines
- `evidence.md` — artifact-backed evidence capture for external-boundary tools
- `mcp-worker.md` — MCP worker CLI (`abstractruntime-mcp-worker`)
- `snapshots.md` — snapshot/bookmark model and stores
- `provenance.md` — tamper-evident hash-chained ledger
- `limits.md` — runtime-aware `_limits` namespace and APIs
- `workflow-bundles.md` — `.flow` bundle format, VisualFlow distribution, and multi-entry fan-in metadata

## Maintainers

- `../CHANGELOG.md` — release notes
- `../CODE_OF_CONDUCT.md` — contributor conduct expectations
- `../CONTRIBUTING.md` — how to build/test and submit changes
- `../SECURITY.md` — responsible vulnerability reporting
- `../ACKNOWLEDGMENTS.md` — credits
- `../ROADMAP.md` — current status and longer-term direction
- `adr/README.md` — architectural decisions (why)
- `backlog/README.md` — implemented and planned work items (what/how)
