Experimental · Disabled by default

Reusable Runbooks

Define a repeatable operational checklist once, run it against the terminal you can see, approve every action, and leave with a durable report of what was already compliant, what changed, and what remains unresolved.

Start Here

From Package to Checked List

Native v1 targets one active, visible POSIX terminal. That can be a local shell or an SSH/container session you established before starting the run.

  1. 01
    Enable the experiment Open Settings → Runbooks and enable Runbooks. It is off by default while the format is v1alpha1.
  2. 02
    Choose an example or import a package Start with one of the included macOS assessments, or select a local folder whose root contains exactly one runbook.vrun.yaml. VTerminal validates imported packages before registering them.
  3. 03
    Review the preflight Choose the terminal, confirm its observed target, enter only non-secret inputs, review declared capabilities, and confirm how much output this run keeps.
  4. 04
    Run beside the terminal Follow the live checklist, approve actions, resolve pauses, review the runbook definition mid-flight, add comments, then inspect or export the final report from History.
Definitions and inputs must not contain secrets.

Inputs, operator comments and evidence metadata are retained in the report. Redaction is a best-effort safety net, not permission to place passwords, tokens, private keys or credentials in a Runbook.

Model credentials configured in Settings → Models stay in macOS Keychain or Windows Credential Manager and use the same backend-only provider path as ordinary AI requests. A token-bearing self-hosted provider requires HTTPS, except for localhost or loopback HTTP.

Create Without YAML

Build a Runbook in VTerminal

Choose New in the Library to open the deterministic, offline builder. Drafts autosave in app-managed storage and remain resumable across restarts. They cannot be run until the final package passes the same strict validation as an import.

  1. 01
    Basics Set the stable ID, semantic version, title, description, tags, failure policy, declared network or root capability, and the absolute paths the runbook writes to.
  2. 02
    Inputs Add non-secret string, integer, boolean, path, or enum inputs with typed defaults.
  3. 03
    Steps Order shell or manual checks. Shell inputs use explicit VRUN_* environment mappings and compliant/non-compliant exit codes. Tick Remediate when this check fails to add an apply that fixes it and a verify that proves it worked.
  4. 04
    Review and publish Inspect the generated YAML and validation findings, then publish the app-managed package into the Library.
Apply and verify travel together.

A step that changes something must prove it worked, so the two phases are added and removed as a pair. A runbook with an apply and no verify is rejected at publish time. The wizard still generates no agent, goal, or Ansible actions: those change who decides whether a step succeeded, so they are read as YAML before they run. macOS 13+ and Linux choices add a locked stopping target guard; Any adds no guard and requires portable commands.

Published wizard projects remain editable. Changed publications require a strictly greater semantic version; unchanged publication is a no-op. Removing a Library registration keeps its draft for republishing. Discarding a published draft detaches wizard editing but leaves the published source and historical snapshots intact.

Turn What You Just Did Into a Runbook

New → Generate with AI authors a draft from a plain-language requirement and, optionally, the terminal session where you already did the work by hand. Install and configure something once, attach that tab, and the model turns the commands you ran into steps: what you ran becomes the apply, and the way you would tell it already happened becomes the check and the verify.

  1. 01
    Describe it Say what the runbook should do. Requirements alone are enough; the terminal is optional context, not a requirement.
  2. 02
    Choose what you share Pick the session, untick any command you do not want to send, and edit the assembled transcript directly. That box is the payload, verbatim.
  3. 03
    Review before saving The result opens on Review as an ordinary draft, validated, with every field editable. Nothing reaches the Library until you publish it.
Generated commands earn no extra trust.

The model writes a draft, never a running command: each line is still approved in your terminal when the runbook runs. Output is constrained to shell and manual actions by the draft format itself, and every generated document passes the same validation and secret scan as a hand-written one. A credential echoed out of a transcript is refused, not published. Attaching a terminal is unavailable when terminal context is switched off in Settings → AI, and for any tab that has run a Runbook.

Ready to Run

Included Packages

Four packages ship with VTerminal. The three macOS assessments below target macOS 13 and newer and are assessment-only: they use no model actions and declare no network, privilege or write capability. Each stops on an unsupported target and, on a supported Mac, continues through independent findings to produce a useful full report. A fourth, ansible-localhost-example, demonstrates the Ansible executor against a local inventory and does declare one write.

macOS Security Posture

Checks FileVault, SIP, Gatekeeper, the application firewall, and automatic critical and system-data updates. It is the first included Runbook selected when no valid selection exists.

View package →

macOS Developer Workstation Health

Checks Xcode Command Line Tools, the macOS SDK, Git, native versus Rosetta-translated shell execution, configurable free space, and optional Homebrew availability.

View package →

macOS Backup & Storage Readiness

Checks configurable free space, the Time Machine destination and latest-backup age, APFS, and optional local snapshots.

View package →

Included examples stay recoverable.

They carry an Included with VTerminal badge. Removing one only hides it from the Library; it does not delete the package or its run history. Choose Restore examples to bring all hidden included sources back.

Own Format

A Small, Versioned YAML Package

A package is deliberately local and constrained. The definition uses a stable Runbook ID, a semantic content version and ordered steps with durable IDs.

my-runbook/
├── runbook.vrun.yaml   required
├── README.md           optional
└── ansible/            optional; used by ansible.playbook actions
apiVersion: runbooks.veviad.com/v1alpha1
kind: Runbook

metadata:
  id: ensure-jq-installed
  version: 1.0.0
  title: Ensure jq is installed
  tags: [linux, software]

spec:
  target:
    kind: active-terminal

  declaredCapabilities:
    network: true
    privilege: root
    writes: [/usr/bin, /usr/local/bin]

  defaults:
    onFailure: pause

  steps:
    - id: jq-installed
      title: jq is installed
      required: true
      check:
        uses: shell
        with:
          command: "command -v jq >/dev/null 2>&1"
        outcomes:
          compliantExitCodes: [0]
          noncompliantExitCodes: [1]
      apply:
        uses: agent
        instructions: |
          Install jq with the target's package manager.
          Make the smallest safe change.
      verify:
        uses: shell
        with:
          command: "command -v jq >/dev/null 2>&1"
        passExitCodes: [0]

Validation rules that matter

  • Every step requires check. If apply is present, verify is mandatory.
  • Shell actions are one inline command of at most 4,096 characters, with no newlines, control characters or heredocs.
  • Typed inputs are non-secret and enter shell actions only through an explicit VRUN_* environment mapping. There are no arbitrary template expressions.
  • Unknown fields/actions, duplicate IDs, aliases, custom tags, merge keys, includes, multiple YAML documents and invalid semantic versions are rejected.
  • Symlinks, traversal and package-root escapes are rejected. Native v1 never uploads or executes package scripts in a remote terminal.
A live definition cannot drift.

Run creation stores the original YAML, canonical JSON and a SHA-256 digest of each. Refreshing an edited package affects a future run, never an active or historical one.

Portable Packages

Export, Edit, Version and Import

Export runbook in the Library creates a complete, import-ready folder named runbook-<id>-v<version>. The folder contains the exact validated YAML, optional README, and allowed ansible/ tree from the registered source.

  1. 01
    Export from the Library Choose a destination that does not already contain the generated folder. VTerminal revalidates source digests and refuses drift, symlinks, path escapes, merge or overwrite.
  2. 02
    Edit the exported files Change the YAML or README in an ordinary editor. Keep metadata.id when the operational purpose is the same, and retain step IDs for controls whose meaning is unchanged.
  3. 03
    Bump the semantic version Update metadata.version whenever requirements, behavior, inputs or step meaning changes.
  4. 04
    Import the edited folder Exported included examples import as normal user sources. The recoverable bundled original remains a separate source.

Registrations are path-based. Importing the same folder refreshes that source; importing a copy at another path creates a separate user source.

Package export and report export serve different jobs.

Export runbook creates reusable source files from the Library. Export report in History creates report.json, report.md, and eligible evidence for one completed run; that evidence bundle cannot be imported as a Runbook.

Execution Model

Check → Apply → Verify

Steps execute sequentially. A check explicitly distinguishes a compliant exit code, a non-compliant exit code and an execution error. Success is never inferred from an apply command alone.

Already compliant

A compliant check skips apply and completes the step as already_compliant.

Changed and verified

A non-compliant check may run an approved apply, but only a separate passing verify completes it as remediated_verified.

Assessment only

Omit apply for a control that should be assessed and reported, not changed automatically.

Paused, not replayed

A failure can pause for retry, skip, timestamped waiver or stop. Mutations are never retried blindly.

Only already_compliant and remediated_verified render as checked. A required skipped, waived, blocked, failed, non-compliant or unknown step produces a completed run with exceptions rather than a false success.

Adaptive Remediation

State the Goal, Not the Commands

Installing Docker is apt-get on Debian, dnf on RHEL and pacman on Arch. A fixed command list needs one runbook per distribution. Handing the whole job to a model swaps that for a worse problem: the model decides both what to do and whether it worked.

A goal splits those apart. You state what must be true and the exact conditions that prove it. The model picks the commands. The engine runs the conditions and decides.

- id: docker-running
  title: Docker Engine is installed and running
  goal:
    intent: |
      Docker Engine is installed from the distribution's own
      repository and its daemon is running and enabled at boot.
    checks:
      - command: "command -v docker"
        expect: [0]
      - command: "docker info >/dev/null 2>&1"
        expect: [0]
  constraints:
    maxCommands: 12
    privilege: root
  apply:
    uses: agent
    instructions: |
      Install Docker Engine using this distribution's package manager.

There is no check and no verify: the goal serves as both, so one condition is written once instead of twice. Goal conditions are ordinary commands in your visible terminal and carry the same shell_observed assurance as any other. Nothing claims a more trustworthy executor. What changed is who reads the result.

Discover the target first

spec.context.discover runs a few probes once, before the first step, each individually approved. Their output is shown to every model phase as data, never instructions, so one runbook can serve Debian, RHEL and Arch.

Bounds that narrow

constraints cap commands, wall clock and model rounds, and can refuse network access or privilege escalation. Every field narrows; none can widen what you already allow.

Refused before the card

A proposal that breaks a bound is refused before an approval card is drawn, so it costs the model a round rather than costing you a click. The refusal returns to the model, which can then try something allowed.

Nothing implicit

context.inputs names which input values the model may see and context.priorSteps whether it learns what earlier steps did. Both are opt-in per step.

Best-effort bounds, not a sandbox.

They read command text. They cannot see through a script the model wrote in an earlier step, an alias in your dotfiles, or python -c. They narrow what a careless model does; they do not contain a hostile one.

Operator Control

Actions, Approvals and the Trust Boundary

shell

Runs one approved inline command in the bound visible terminal and observes its exit status and output.

agent

Gives bounded Markdown instructions to the current model, plus whatever the step's goal and context allow it to know. Each command it proposes receives its own approval, and the model phase itself is approved separately; that card names the goal and the enforced bounds before the model has them.

manual

Collects an operator outcome, required comment and optional evidence note for a human-only control.

ansible.playbook

Native. It invokes a user-installed ansible-runner as the explicit local controller, binds approval to exact project and inventory digests, and retains structured per-host outcomes. Check and verify are preview-only; apply still requires subsequent verification.

What approval means in native v1

  • Every visible-terminal shell action requires explicit approval, including checks and verification. Apply always requires it.
  • Model phases are opaque and must be approved one step at a time.
  • If you prefer, use Acknowledge and approve all remaining steps in live mode to continue through later approvals. It waits for each approved command to finish before the next approval appears, so a step that takes minutes is normal. It stops at the first operator pause, manual step, finished run or invalid command, and Stop auto-approve ends it at any point.
  • Abort run stops an active run and returns you to the Library. It cannot prove the process stopped or undo a change already made; the active step is reported unknown.
  • This flow is different from the agent panel’s Full setting and can only advance runbook approvals.
  • On-device models are reported as local; cloud and configured remote providers are additionally reported as networked.
  • Each approval shows the immutable target and asks you to attest that the visible row is a POSIX prompt on the intended host/container, and that its shell, functions, aliases and PATH are trusted.
  • A run binds to one terminal and observed local/SSH/container context. Target or session drift pauses execution, and two runs cannot drive the same terminal concurrently.
The interactive shell is inside the trust boundary.

A visible PTY is not a deterministic sandbox: aliases, exported functions, startup files, PATH shims and an already privileged remote shell can change what a textual command does. Shell results are therefore reported as shell_observed. A compromised interactive shell is outside v1's trust model.

A timeout is unknown; it is not an automatic retry. Cancellation sends SIGINT to the owned foreground job but cannot prove detached work stopped or undo a mutation. After an app restart, an active run is interrupted and must be explicitly rebound before resuming.

Durable Results

Evidence and Canonical Reports

Successful, exceptional, failed and cancelled runs all finalize a canonical report.json. The human-readable report.md is generated only from that report data, so the two views cannot invent different outcomes.

none

Keep attempt and outcome metadata, but no output body.

tail · default

Keep a redacted 8 KiB output tail for each attempt.

full

Keep the tail plus a redacted artifact capped at 1 MiB per attempt, readable from the report with View recorded output.

Export report

Export JSON, Markdown and eligible evidence from History to a folder you choose. This is not a reusable package export.

Settings → Runbooks → Record terminal output sets the floor: Never, As the runbook asks (the default, honouring spec.audit.recordOutput), or Always, in full. Preflight offers only levels at or above it, so a single run can be raised but never lowered, and the clamp is applied in Rust rather than in the picker.

What a recording is and is not.

Redaction catches recognisable credential shapes, not every secret. Capture reads what the terminal still holds in scrollback, so a chatty command yields an honest tail rather than a complete transcript, and truncation is marked. An artifact's digest is re-verified on every read: one altered since the run is reported unreadable rather than shown. Recorded output is kept until you delete the run, and report evidence is never sent to a model to improve prose.

Reports include definition IDs, versions and digests; target and input data; checklist states; attempts; approvals; edited-command deviations; evidence metadata; agent summaries; comments; exceptions and unresolved risks. Removing a package registration does not erase its historical runs.

Step and executive summaries use the engine-fixed outcomes and bounded structured evidence. They are deterministic by default, so report evidence is not sent to a model merely to improve the prose. A summary returned by an explicitly approved agent phase is retained.

Credential-shaped content and private-key blocks are redacted before display, persistence and export, and truncation is explicit. Detection is necessarily best-effort: arbitrary secrets without a recognizable name or shape cannot be found reliably.

Native

Ansible Runner, Explicitly Controlled

ansible.playbook executes through a user-installed ansible-runner; there is no fallback to ansible-playbook or a generic shell action. Every approval shows and binds the exact project and inventory digests, and reports retain structured outcomes per host. Check and verify run with --check --diff, while a successful apply must still pass verification. Runner is launched as a local controller without inheriting the visible terminal shell; its inventory may target remote hosts.