Arrow Documentation
Arrow is a Windows desktop platform for building, running, and scheduling intelligent automation entirely on local hardware.
You record any interaction once, compose it into a visual workflow graph, and execute it deterministically — with local AI, local vision, local OCR, and local speech. No cloud, no telemetry, no per-token fees.
- Visual automation — workflows ("chains") are node graphs, not scripts.
- Deterministic execution — the graph is a symbolic program; AI is used only at the seams.
- Agent platform — Agent Mode turns chains into tools for a local model.
- Fully on-device — bundled GGUF models, ONNX vision, PaddleOCR, Piper TTS, Vosk STT.
- Composable primitives — approvals are Input nodes plus branch routing; memory is Context nodes plus SQLite.
Capability map
| Capability | What it does |
|---|---|
| Node-graph editor | Author, edit, expand, and run chains. |
| Desktop recorder | Capture clicks, typing, scrolls, drags, clipboard, screenshots, and UI elements. |
| Web recorder | DOM-level browser capture via CDP; selector-based sessions. |
| Execution engine | Dependency-gated scheduler loop over typed nodes. |
| Desktop replay | Element-first (UIA) + template-matching cascade, OCR conditions. |
| Web replay | Native-first CDP ladder with a durable shared browser profile. |
| Local AI | llama.cpp GGUF and Ollama; chat, generate, and embed endpoints. |
| Semantic retrieval | ComoRAG probe-driven retrieval, embeddings, Graph RAG. |
| Vision & OCR | Screen understanding, ONNX detection, PaddleOCR. |
| Speech | Piper TTS and Vosk STT for voice mode. |
| Agent mode | System-chain-driven agent; chat overlay, phone client. |
| Orchestrator | Goal loop over mini-brain chains; trace + synthesis. |
| Memory | In-run context, SQLite persistence, RAG, consolidation. |
| Scheduling | Once, daily, weekly, interval; in-process runs. |
| RDP sandbox | Run automation inside an isolated Windows session. |
| Agent export | Compile a system chain + assets into a standalone .exe. |
High-level architecture
Arrow is built around three pillars:
- Editor — the graph editor and all runtime surfaces: typed nodes, dialogs, chain library, scheduler UI, agent chat.
- Player — the execution engine and all input/output machinery: interpreter, desktop recorder, web recorder/replay, and grounding.
- AI — local inference and retrieval: a FastAPI server over llama.cpp and Ollama, embeddings, retrieval, and the Laya decision engine.
Runtime topology
- Launcher — bootstraps the GUI, or runs a
.pyargument directly (how the sandbox agent and exported agents are spawned). - GUI process — PyQt5 main window; hosts the editor, overlays, scheduler, recorder, and (by default) the FastAPI server.
- llama.cpp subprocess — the bundled
llama-server.exe(default port 8081) serves GGUF inference. - Vision subprocess — multimodal inference via
llama-mtmd-cliwith a GGUF + projector pair. - Laya engine — an embedded decision daemon, lazy-launched and released after each burst; every consumer falls back to LLM paths on failure.
- Exported agents — standalone executables that embed the app, the system chain, and transitive chains; they run with
ARROW_DIRECT_ENGINE=1.
Node catalog
| Node | Role |
|---|---|
| Sequence | Replay a recorded desktop sequence. |
| Web Sequence | Replay a recorded browser session. |
| Conditional | Branch on image / OCR / wait / web conditions. |
| Input | Value injection, questions, decision routing, media acceptance. |
| LLM | Local inference: reasoning, tools, vision, structured output, TTS. |
| Chain Import | Run another chain as a sub-workflow or tool. |
| Form Filler | Document-grounded multi-field form completion. |
| Code | Custom Python with JSON in/out. |
| Context | Persistent cross-chain memory. |
| Handle | Vision-based click/scroll/zoom grounding. |
| MCP | Model Context Protocol server tools. |
| Output | Expose results to the overlay / outside the chain. |
| Orchestrator | Goal loop over mini-brain chains. |
Chain metadata includes name, description, collection (e.g. System, Tools, Brains), and is_default (marks the default system chain).
Wiring patterns
Record, branch, act. A recorded sequence feeds a conditional; the chosen branch continues.
LLM tool router. An LLM node picks one of several tool chains wired to its tools port.
Orchestrator. An orchestrator routes a goal across mini-brain chains and deterministic chains.
Ports & data flow
Every node exposes named ports. Execution flows along exec edges (base ports such as output → input); values flow along data edges (data, true/false, trace, files, and so on). A data edge never triggers execution; an exec edge never carries values.
- Port values are stored per chain and exposed as
node_<id>_<port>variables. - Chain-to-chain state rides explicit variables such as
_chain_input_context; there is no hidden global state. - Branch-capable nodes own
true/falseoutput ports. - A node whose exec edges land only on a router input port (
llm.tools,orchestrator.brains) is flaggedtool_providerand never runs in the main loop — its router invokes it.
Branching & loops
- Branch sweep — when a branch-capable node resolves, the unchosen branch is marked skipped so merges never stall.
- Loop-backs — an Input node re-resolves on every pass, which is how iterative workflows are built with pure topology.
- Decision re-evaluation — decision and routed questions are re-evaluated per loop pass.
Execution engine
Semantics
- Dependency-gated scheduling — a node runs only when every incoming edge it needs is settled.
- Tool-provider gating — chain-import nodes wired to a tool consumer run only when selected.
- Stop semantics — a single stop flag is polled everywhere; ESC aborts a run immediately.
- Fail-closed — unparseable model answers route to a configured default branch; never a crash, never a silent wrong turn.
- Serialized model access — LLM calls share one managed engine and a serialization lock; local inference runs without timeouts.
Recording
Desktop recorder
Captures, alongside each action, the cropped image of the click target, the Windows UIA element it landed on (ControlType + bounding box), and precise coordinates. Rapid scrolls are grouped into one action; typed text is buffered into one action; clipboard shortcuts are recognized.
Web recorder
Records a CDP session on a real, visible browser. Recordings are stored as web sequences: element-driven steps with selector identities, never absolute coordinates. Identity rules favor stable attributes (data-*, ids, text containment) and treat repeated items as entities.
Playback & grounding
Desktop replay — two independent locators
- Element-first (UIA) — resolve the recorded element's live bounding box, surviving moved windows, theme/DPI changes, and scrolled lists.
- Template matching — multi-scale OpenCV matching within region constraints and confidence thresholds.
- OCR conditions — PaddleOCR powers text conditions in Conditional nodes.
- Human-like behavior — jittered click targets, curved mouse paths, variable timing.
Web replay
One handler class per action type, each running a deterministic ladder: native CDP input first, then in-page JS dispatch, then a locator retry ladder, then reconnect-and-retry. Replay is visible-browser by default with a single durable shared Chrome profile.
Anti-detection & human-like input
Playback reads like a person, not a bot: jittered click targets, curved mouse paths, variable timing, native CDP input (with an in-page JS fallback), and a durable shared browser profile. Web replay never captures your physical cursor.
Inference & local API
| Engine | Transport | Notes |
|---|---|---|
| llama.cpp | llama-server.exe, port 8081 | GGUF models, GPU offload (Vulkan), embedding mode. |
| Ollama | HTTP :11434 | Alternative provider, selected by name. |
| llama-mtmd-cli | subprocess per inference | Multimodal vision (GGUF + mmproj projector). |
The engine is authoritative about context size — the app caps requests to the model's training context.
Local API surface (FastAPI)
| Endpoint | Purpose |
|---|---|
/health, /models | Liveness and model inventory. |
/chat, /generate | Unified chat/generation to the active provider. |
/embeddings | Embedding vectors. |
/llamacpp/* | Direct llama.cpp lifecycle + inference. |
Retrieval & Laya
Semantic retrieval
- ComoRAG — probe-driven retrieval: intent, targeted probes, retrieve, then synthesize against verbatim chunks.
- Graph RAG — surfaces relevant past knowledge to reasoning nodes and records execution fragments.
Laya decision engine
An embedded typed-decision model (noul / choice / score questions answered in one encoder pass) used by default for the orchestrator picker, input-node decisions, retrieval fact curation, and form-filler field judging. It runs as a lazy-launched daemon and is always optional — every consumer falls back to the LLM path on failure. LOOPER_LAYA=off forces LLM-only paths.
Vision & speech
| Layer | Technology |
|---|---|
| Screen VLM | LFM2.5-VL 450M + projector (Handle node scene understanding). |
| Object detection | ONNX YOLO-based detector for UI element detection. |
| OCR | PaddleOCR (det/rec/cls) for text conditions and documents. |
| Document extraction | PyPDF2 (PDF), python-docx (Word), plain text. |
| TTS | Piper ONNX voices. |
| STT | Vosk offline speech-to-text for voice mode. |
Agent mode — system chains & tools
Agent mode has no hardcoded behavior. It requires a system chain (a chain with collection: "System") that defines the entire cognitive architecture: Input nodes receive the query, LLM nodes reason, Context nodes remember, Chain Import nodes act as tools, Output nodes surface the answer. Multiple system chains = multiple named agents.
- Dispatch — selects the default system chain and runs it through the same engine as manual playback.
- Chains-as-tools — any chain connected to an LLM's tools port is selectable at runtime; tool descriptions are derived at call time.
- Stepped protocol — the router picks a tool by name, then each agent-modifiable Input node extracts its own single value. This is why a 3B model suffices.
- Ask-user — any Input node with a prompt pauses the run and asks through the richest available channel.
Orchestrator
The default system chain ships with an Orchestrator node as its root: several mini-brain chains are wired to its brains port, and the orchestrator loops toward a goal — it scopes one step at a time and asks the picked specialist brain to run it.
- Three inputs to every decision — the goal, what was done (step result + trace), and the world now (observed state digest).
- Picker — one strictly-parsed token (worker number,
DONE,ASK), or the Laya engine with zero parsing. - Verification — a per-step verdict judges whether the step accomplished what it was asked.
- Guards — step cap, no-progress detection, and a stop flag between steps.
- Synthesis — a final short LLM pass turns the trace into a plain-language answer.
Deterministic chains
The orchestrator's second input port carries deterministic chains: orchestrator-free compositions that answer a known request end to end with no routing inference. One rule decides which kind a chain is — does it contain an orchestrator, directly or through its imports?
- A verified run can be frozen into a learned chain (composition +
routing.examples), wired to thechainsport. - Self-improving — learned chains accumulate on the
chainsport, so the agent's library grows with every verified run. - Partial fits are assembled from existing artifacts into one continuous chain and run as a single dispatch.
- The gate asks the step question per directive, judged against the current directive with a confidence floor plus object-name alignment.
Input nodes & persistent goals
Input node v2
- Decision mode — a natural-language criterion evaluated with one strictly-contracted model turn or the Laya noul evaluator.
- Question modes — free text, yes/no, or a choice list; answers are normalized against the options.
- Accepted media — text, images, and documents (PDF/Word/txt) ride the
dataport.
Persistent goals
With use_goal_ledger enabled, a blocked goal parks itself (blocked port) and resumes on the next activation — wake-driven persistence, not a busy daemon. Schedules create the wakes.
Memory & context
- In-run state — named-port values plus a turn-based history per chain.
- Persistence — Context nodes write append-only turns to SQLite (
data/context.db) with per-(chain, node) rolling caps. - Retrieval — Graph RAG surfaces relevant past knowledge; ComoRAG consolidates large sources into grounded facts.
- Inspection — the agent settings dialog renders a goal-relationship graph with reviewable, correctable, clearable memory.
Specialized nodes
Form Filler
Document-grounded multi-field completion: field discovery on live pages, knowledge-file lookups, per-field validation, repair loops, and a review gate. Choice controls are resolved by the Laya semantic picker, which refuses unclear picks so the field is asked instead of guessed.
Handle
Combines VLM scene understanding with LLM reasoning to issue grounded click/scroll/zoom/key actions on screen content that lacks selectors.
MCP
Connects Model Context Protocol servers as workflow nodes; discovered tools are exposed to the router model as individually selectable actions.
Code
Inline or file-based Python with JSON arguments and output variables. The Code Node Studio embeds an Aider-style iterative local-agent loop — an integrated coding agent that writes, edits, and debugs Python inside the graph.
Scheduling, sandbox & remote access
Scheduler
Runs any chain on a schedule — once, daily, weekly, interval — with enable/disable, next-run display, and run-now. Schedules persist to schedules.json and execute in-process.
RDP sandbox
Automation runs inside an isolated Windows session via a local HTTP agent, so the desktop you are using stays untouched. Chains can opt in per sequence/chain-import.
Remote access
The agent chat is served over LAN or Tailnet with a scannable QR code — a phone can drive runs, answer ask-user questions, and stop executions.
Export & build
Standalone agent export
Collects every chain referenced transitively (including orchestrator brains), bundles chains/sequences/models, and builds a one-folder executable. Exported agents run directly on the local engines (ARROW_DIRECT_ENGINE=1) — no FastAPI dependency, fully offline.
Application build (MSI)
The build pipeline stages dependencies (llama.cpp binaries, Chrome, model files), runs PyInstaller, generates a WiX source, and produces a per-user MSI installer.
Configuration & data
Environment variables
| Variable | Purpose |
|---|---|
API_PORT, AUTO_START_API | Local API control (default 8000 / true). |
OLLAMA_HOST, OLLAMA_PORT | Ollama endpoint. |
LOOPER_LAYA | off forces all Laya consumers onto LLM paths. |
LOOPER_ORCH_STATE | off disables the orchestrator's per-step state digest. |
LOOPER_LEARN | off disables freezing runs into learned chains. |
ARROW_DIRECT_ENGINE | 1 in exported agents: direct engines, no API calls. |
Data locations
| Data | Location |
|---|---|
| Chains | chains/*.json (learned: chains/learned/*.json) |
| Sequences | sequences/*.json, web_sequences/*.json |
| Context memory | data/context.db (SQLite) |
| Schedules | schedules.json |
| Goal ledger | data/agent_goals.json |
| Models / voices | data/piper_voices/, data/vosk_models/ |
Engineering principles
- Neuro-symbolic execution — the graph is a deterministic program; AI has a fallback or fail-closed default everywhere.
- Small-model discipline — stepped protocols and single-token contracts for 1–4B models.
- Named-port contracts — values move only along explicit edges.
- Runtime-derived descriptions — tool descriptions are generated at call time, never persisted stale.
- Fail-closed defaults — new capabilities ship default-off; unparseable answers route to defaults.
- Entity-first grounding — selectors and UIA elements over absolute coordinates.
- No timeouts on local inference — interruption is a stop-flag concern.
- Serialized, on-demand model lifecycle — no preload at startup; burst-aware cleanup.
- Approvals and memory are primitives — composed by users, not hardcoded policy.
- Offline and telemetry-free — nothing leaves the machine.
- Inspectable and reversible — logs, trace ports, blocked states; ESC always aborts.
Licensing
Arrow is open source. Use it, modify it, and share it freely — every feature and node type, no execution limits.
- All features included: visual workflows, agent mode, the orchestrator, vision, OCR, speech, and local AI.
- Runs fully offline — no activation servers, no telemetry.
- Export agents as standalone executables straight from the source.
- The packaged MSI installer (WiX) is the recommended path for end users; the source is the recommended path for contributors.