Warp-style command blocks and a multi-step AI agent, without the
bulk. Local models run in-process with Metal on Mac,
or Vulkan with automatic CPU fallback on Windows. No Ollama and no
daemon to babysit. Cloud models sit behind the same interface.
Published 11 Sept 2026·114 downloads across
all releases
macOS Apple Silicon + Windows 11 x64 preview
Tauri 2 + Rust
On-device by default
GPL-3.0
Actual app
A real VTerminal session with demo-only service data. No interface mockups.
Why VTerminal
Three Positions the Other AI Terminals Don't Take.
Modern AI terminals tend to be Electron apps that phone home for
every completion. VTerminal is a native Tauri app that treats your
shell, your machine and your consent as the things that must not
be replaced.
01
Instead of a reimplemented shell
Your Shell, Not a Lookalike
A real zsh -il shell on macOS or Bash in the
default WSL2 distribution on Windows, each over its native
PTY. vim, htop, tmux and ssh run in the real shell,
not a reimplementation.
02
Instead of cloud round trips
On-Device by Default
The default model is a 9B GGUF running inside the app process
via llama.cpp, with Metal on Mac or Vulkan and automatic CPU
fallback on Windows. Nothing leaves the machine unless you
pick a cloud model. There is no daemon to maintain.
03
Instead of an agent with free rein
Nothing Runs Without You
In a session you are driving, every command a model writes
passes an approval gate. A scheduled action runs only what
you armed it for in advance, and skips and records
everything else. Natural-language suggestions are
inserted into your prompt, never executed. Restored
tabs never reconnect on their own.
The Terminal
A Fast Terminal First. The AI Is What It Grew.
Shell integration marks every command as a block, so the app
knows what you ran, where, and how it ended. That is also
exactly the context the model needs.
01You Type a Command
02The Shell Marks It via OSC 133
03Output Streams Into It
04The Exit Code Closes It
Real Shell, Real TTY
portable-pty runs zsh -il on macOS or
Bash in the default WSL2 distribution on Windows, rendered by
xterm.js 6 with a WebGL renderer. Full TUI support, and saved SSH
hosts with one-click connect.
Command Blocks
Exit-code badges, copy the command or just its output, re-run, or
attach a block to the AI as context. Positions come from live
terminal markers, so they never drift.
Flow-Controlled Output
The reader thread blocks above a watermark and the UI
acknowledges every 256 KB, so cat-ing a
gigabyte file won't balloon memory.
Tabs, Search, Palette
Tabs, in-terminal search, persistent command history, and a
command palette for actions and model switching. Use Command on
macOS or Ctrl+Shift on Windows.
SSH-Aware
Nested sessions are detected by command shape. While you're on a
remote host, local paths and git branch are withheld from the
model rather than quietly misreported.
Six Themes
Veviad Developer UI (default), Veviad UI, Midnight, Nord,
Solarized Dark and Light, each with a matched terminal ANSI
palette, not just recoloured chrome.
The AI
Every Way In. One Keystroke Each.
The same Provider trait covers in-process llama.cpp,
Anthropic, OpenAI, Mistral and any OpenAI-compatible server you
run yourself. Every feature below works the same whichever you
pick.
Suggest ⌘I / Ctrl+Shift+I
Describe the goal, or type # at an empty prompt, then
and get a command inserted into your prompt, ready to read before
you press enter.
Explain & Fix
One click on a failed block streams a diagnosis of what went
wrong, plus a corrected command you can inspect and run.
Ask ⌘J / Ctrl+Shift+J
A chat panel with your blocks, their output and your files as
context. Resizable, and it keeps its proportion when you resize
the window.
Agent Mode
Multi-step runs that propose commands, execute them in your
visible terminal, and read the real output before
deciding what to do next.
Reasoning Effort
off → low → medium → high → max, per model, showing
only the rungs a model actually accepts because a wrong value
is a 400, not a downgrade.
Images & Files
Drag, paste or pick. An optional on-device vision sidecar
transcribes screenshots, so even a non-vision chat model can work
with them.
Standing Instructions
Conventions to keep, tools to prefer, the language to answer in.
Three fields — all AI, Agent only, Chat and Ask — added to the
built-in instructions rather than replacing them. They tune how
the model works; they cannot grant a permission or approve a
command.
Scheduled Actions
A saved sequence of commands and prompts, run on a schedule
against a local shell or a saved SSH host. Experimental and off
by default. It fires only while VTerminal is open — no tray
icon, no launch agent, no background process — and every step is
recorded, including what it wanted to do and was not allowed to.
Chat Workspace
A conversation with no terminal attached, kept in its own
threads with their own history, model and MCP servers. It can
search Knowledge, but it has no run_command: the
terminal stays Agent-only.
Local and Remote Together
Pair a local tab with a live ssh tab and drive both
from one conversation. Each target keeps its own approval mode,
and the run stops if either tab drifts from the machine it was
bound to.
Model Context Protocol
Bring Your Own Tools. Keep Every Call Visible.
Connect several remote or local MCP servers, choose them per chat,
and expose only the tools that conversation needs. MCP tools work in
Ask, Agent and the Chat workspace, while terminal commands remain
Agent-only. See the
MCP setup and security guide.
Several Servers per Chat
Select multiple servers for a conversation. Defaults apply to
new chats, while every chat can keep its own server and tool
choices.
Remote Streamable HTTP
Connect HTTPS endpoints using JSON or request-scoped SSE.
VTerminal negotiates protocol versions and supports compatible
legacy initialization lifecycles.
Sandboxed Local stdio
Run npx, uvx, Docker, or an exact
command on macOS and Windows through WSL2. Local MCP stays off
unless the mandatory sandbox self-test passes.
OAuth and Header Authentication
OAuth 2.1 with PKCE, bearer tokens, and custom secret headers
are built in. Credentials remain in the operating-system vault,
and exported configuration is redacted.
Tool-Level Control
Tool discovery is paginated and cached. Each server is a compact
group in the chat picker, where individual tools can be enabled
or disabled for that conversation.
Approval by Default
Calls show the server, exact tool, and complete arguments. Allow
once, deny, or remember that exact tool until its server
configuration or schema changes.
Knowledge
Local Documents and Qdrant. One Search Surface.
Attach local SQLite buckets and permission-filtered Qdrant
collections to the same Ask or Agent request. VTerminal creates
the right query embedding for every bucket, searches each source
independently, and fuses ranked results without comparing
incompatible raw scores.
Actual app
Local embedding models are selected and managed inside VTerminal.
One-click local embeddings
Download, verify, load and test a pinned model from Settings →
Knowledge, with progress, cancellation, RAM and disk estimates,
and license acceptance where required. No Python, conversion,
compilation or arbitrary model files.
Qwen3-Embedding 0.6B Q8 (recommended)
Qwen3-Embedding 4B Q4_K_M
Qwen3-Embedding 8B Q4_K_M
EmbeddingGemma Q8 (768 dimensions)
Multilingual E5 Base Q8 (visible, release-gated)
Multilingual E5 Large Q8 (visible, release-gated)
Multilingual E5 Base and Large remain unavailable until signed,
checksum-pinned Veviad Q8 GGUFs pass multilingual parity tests
against the official Sentence Transformers models.
Qdrant collections are buckets
Add a Qdrant database REST URL (normally port 6333) and granular
database key. VTerminal uses REST only, lists managed collections
that key may see, and keeps attachment read-only unless the
credential allows document or collection management.
Mix local and remote buckets in one request
The attachment picker shows only exact runnable profiles
Upload, replace, edit or delete individual documents
Managed metadata is shared by every VTerminal client
Shared source identities and staged revisions coordinate concurrent clients
Source-qualified citations keep every result traceable
Managed buckets require Qdrant 1.16+. Their immutable
metadata.vterminal contract stores the exact profile,
vector and payload schema in Qdrant itself. Unmarked collections
stay hidden instead of asking every client to recreate a fragile
local field mapping. Local bindings from before the managed
contract remain under Advanced as read-only search sources. Qdrant stores extracted chunks, metadata and vectors, not
original binaries.
Profiles fingerprint the model, revision, dimensions and exact
query/document transforms; changing one creates a new profile
and requires re-embedding.
UI-first ingestion
Pick files or folders, preview estimated chunks and cloud
transfer, then follow persistent Extract → Chunk → Embed → Upload
progress, even while the document list is collapsed. Jobs are
queued immediately, run automatically, resume safely after
interruption, and failed jobs can be retried without selecting the
source again or duplicating documents.
Keyword-only local buckets remain usable
OpenAI and Mistral are the only guided cloud embedders
Anthropic is chat-only; it does not offer embeddings
Ollama and LM Studio embedding probes live under Advanced
Chat, vision and embedding models have separate jobs: chat writes
answers, vision transcribes attached images, and embeddings map
documents and queries into a retrieval vector space.
Automation and advanced storage
The signed standalone vterminal-docs CLI installs
from the UI into ~/.local/bin on macOS. The Windows
installer places it in VTerminal's managed user bin directory,
and Settings reports or repairs that copy. It shares profiles,
model cache, connections, chunking
and job records with the app for repeatable terminal automation.
List profiles; list and test saved connections
List, create and delete buckets
List, ingest, replace and delete documents; search buckets
JSON output, stdin and safe Ctrl-C
Qdrant 1.18+ TurboQuant bits4 as an opt-in sidecar
TurboQuant is advanced and off by default. VTerminal confirms the
saved configuration from Qdrant before updating the UI. It keeps
the original vectors, so it can be disabled; bits2, bits1.5 and bits1 are
advanced choices that trade recall for lower memory use and a
different search-speed profile. The
CLI handles text, structured text, page JSON and text-layer PDFs;
OCR-required files are directed to the UI.
Experimental · Disabled by Default
Repeat the Operation. Keep the Evidence.
Reusable Runbooks turn a versioned YAML checklist into a visible
check → apply → verify workflow. Use one to assess a
server baseline, remediate it with approval, or install software
without losing the operator decisions and output behind the
result.
Import a local package rooted at runbook.vrun.yaml.
Each run snapshots the original YAML, canonical JSON and both
digests, so editing a package cannot rewrite live or historical
execution.
The Visible Terminal Is the Target
Native steps run in the terminal you can see, including an SSH
or container session you opened yourself. Every shell action
requires a fresh approval and target attestation; agent
permission modes never carry over.
Canonical, Auditable Reports
Completed, exceptional, failed and cancelled runs produce a
canonical report.json. Human-readable
report.md is derived from it, with approvals,
deviations, comments and captured evidence. Where a step kept a
full artifact, you can open it from the report.
State the Goal, Not the Commands
A step can declare what must be true and the exact
conditions that prove it. The model picks commands for the
distribution it actually finds; the engine runs those conditions
and decides. Per-step bounds cap commands, time and rounds, and
can refuse network access or privilege escalation.
Read how goal-directed steps work.
Native v1 is intentionally bounded.
Enable it in Settings → Runbooks. Inputs are
non-secret, the interactive shell remains inside the trust boundary,
and ansible.playbook runs through a user-installed
ansible-runner with digest-bound approval and structured
per-host outcomes. Check and verify stay preview-only, and apply
still requires verification.
Safety
How Command Execution Is Gated.
This matters more in a terminal than anywhere else, so it is worth
being precise about what is enforced and what is not.
Two-Axis Classification
Per-Session Permission Mode
Every Step Recorded
Your Edits Are Yours
API Keys Never Reach the UI
Classified on Two Axes
Every proposed command is checked for whether it is read-only and
whether it reaches the network independently, since a fetch
that writes no file still pulls unreviewed content into a loop
that proposes shell commands. Both are shown when a card is required.
Arming It Is the Authorization
For a session you are driving, the permission mode is per
session, never persisted, and never inherited by a new tab.
Reads and Smart retain approval gates, Auto runs standard
commands but protects sensitive operations, and Full runs every
executable command without approval. Deny rules and disabled
capabilities still block execution.
A scheduled action is the one exception, and it is bounded.
Full is not offered to it at all. The arming is tied to
that action's target, steps and attachments, and editing any of
them resets it. Attaching a knowledge bucket or an MCP server
caps it at reads. Anything the mode does not cover is skipped
and recorded, never left waiting.
In Your Visible Terminal
For an agent run you start, approved commands run in the tab
you're looking at, not a hidden subprocess. You see exactly what
ran, and it runs wherever that tab is, including over
ssh.
A scheduled action runs in the background by default, where
there is no tab to watch — so every step is recorded instead:
the command, its exit code and a capped extract of its output.
Driving a real terminal tab is a separate opt-in.
Web Access Is Withheld, Not Asked
With web access off, the model's fetch tooling is never sent in
the first place, and network-shaped commands are refused before
an approval card is even drawn.
Edited Means Yours
A command you rewrite before approving is treated as your own
text on your own gesture, not as model output. It is deliberately not
re-classified.
Nothing Runs at Launch
Restored sessions never replay anything. An ssh tab
you had open offers Reconnect; it does not reconnect
itself. A scheduled action set to catch up on a missed
occurrence runs at most once, and only after the app has
settled — never as part of starting up.
This Is a Safety Rail, Not a Sandbox
The command classifier cannot see through a script the agent
wrote in an earlier step, an alias in your dotfiles, a
base64-decoded one-liner, or python -c. It is
documented as best-effort in the app itself and should not be
relied on as a security boundary. What actually enforces "no
internet" for a capable model is the withheld tool, not the
string match.
Models
One Interface. On-Device, Cloud, or Your Own Server.
Chat models write and reason, the optional vision sidecar transcribes
images, and embedding models power Knowledge retrieval. Each has a
curated catalog and a separate runtime role rather than one model
being treated as interchangeable with another.
RUNS IN-PROCESS
DEFAULT
On-Device
GGUF weights running inside the app via llama.cpp with Metal on
macOS or Vulkan with automatic CPU fallback on Windows. Downloads
are resumable and range-checked, with live progress, speed, ETA,
cancel and retry directly on each chat or vision model card.
Multi-Token Prediction accelerates compatible models automatically,
verifies every draft with the target, and safely falls back to
standard decoding. Gains vary by model, prompt, context and hardware.
Qwen3.5 2B, an ultrafast option for laptops with 8 GB or more RAM
Qwen3.5 4B / 9B
Qwen3.6 27B
Gemma 4 E2B / E4B / 31B
Automatic MTP acceleration for all seven local chat models
Prompts from each model's own chat template
Qwen3.5 9B is the default, sized to run on a 32 GB M1 Pro.
BRING YOUR OWN KEY
Cloud
Frontier models behind the same trait, with per-model effort
mapping onto each vendor's own parameter.
Claude Haiku 4.5, Sonnet 5, Opus 5
GPT-5.6 Luna, Terra, Sol
Mistral Small 4, Magistral Medium, Large 3
Transient 429s and 5xxs retry with backoff
Provider keys, Hugging Face tokens and remote-server tokens stay
in macOS Keychain or Windows Credential Manager (service com.veviad.terminal).
They never round-trip to the UI or settings.json.
ON YOUR OWN HARDWARE
Self-Hosted
Point VTerminal at any OpenAI-compatible server, press
Test, and pick which of the served models to
expose.
Ollama, LM Studio, llama.cpp server
vLLM, LiteLLM
Per-server tokens, optional
Raw <think> traces split out
Hosts are only contacted behind an explicit Test and are never probed
in the background. A token-bearing connection requires HTTPS,
except for localhost or loopback HTTP.
ALSO ON-DEVICE
OPTIONAL
Vision Sidecar
A second local model loaded beside the chat model, purely to
transcribe the images you attach.
PaddleOCR-VL 1.6
Qwen3-VL 4B / 8B
Decodes greedily, so a transcript is reproducible
Works with a cloud chat model
Transcribed text is fenced and labelled as data. A screenshot is
attacker-controllable by construction.
Switching models is one keystroke from the
⌘K / Ctrl+Shift+K palette, and
every reply keeps the name of the model that wrote it.
Download
Pick Your Platform. Start in Minutes.
The macOS build is signed and notarized. Windows 11 is available
as an unsigned preview for WSL2 and Bash. Both installers include
on-device inference with automatic hardware acceleration.
Apple SiliconWindows 11 x64 previewNo app telemetryGPL-3.0
Preview status: the installer is not Authenticode signed.
Windows may show an unknown publisher. Verify its checksum
before choosing More info → Run anyway.
Issues and pull requests are welcome. Please open an issue before
starting substantial work so the approach can be agreed first. If you
intend to contribute regularly, get in touch: a licensing agreement
may be needed to keep future relicensing possible.