Arrow Documentation

Arrow is a Windows desktop platform for building, running, and scheduling intelligent automation entirely on local hardware.

You record any interaction once, compose it into a visual workflow graph, and execute it deterministically — with local AI, local vision, local OCR, and local speech. No cloud, no telemetry, no per-token fees.

  • Visual automation — workflows ("chains") are node graphs, not scripts.
  • Deterministic execution — the graph is a symbolic program; AI is used only at the seams.
  • Agent platform — Agent Mode turns chains into tools for a local model.
  • Fully on-device — bundled GGUF models, ONNX vision, PaddleOCR, Piper TTS, Vosk STT.
  • Composable primitives — approvals are Input nodes plus branch routing; memory is Context nodes plus SQLite.

Capability map

CapabilityWhat it does
Node-graph editorAuthor, edit, expand, and run chains.
Desktop recorderCapture clicks, typing, scrolls, drags, clipboard, screenshots, and UI elements.
Web recorderDOM-level browser capture via CDP; selector-based sessions.
Execution engineDependency-gated scheduler loop over typed nodes.
Desktop replayElement-first (UIA) + template-matching cascade, OCR conditions.
Web replayNative-first CDP ladder with a durable shared browser profile.
Local AIllama.cpp GGUF and Ollama; chat, generate, and embed endpoints.
Semantic retrievalComoRAG probe-driven retrieval, embeddings, Graph RAG.
Vision & OCRScreen understanding, ONNX detection, PaddleOCR.
SpeechPiper TTS and Vosk STT for voice mode.
Agent modeSystem-chain-driven agent; chat overlay, phone client.
OrchestratorGoal loop over mini-brain chains; trace + synthesis.
MemoryIn-run context, SQLite persistence, RAG, consolidation.
SchedulingOnce, daily, weekly, interval; in-process runs.
RDP sandboxRun automation inside an isolated Windows session.
Agent exportCompile a system chain + assets into a standalone .exe.

High-level architecture

Arrow is built around three pillars:

  • Editor — the graph editor and all runtime surfaces: typed nodes, dialogs, chain library, scheduler UI, agent chat.
  • Player — the execution engine and all input/output machinery: interpreter, desktop recorder, web recorder/replay, and grounding.
  • AI — local inference and retrieval: a FastAPI server over llama.cpp and Ollama, embeddings, retrieval, and the Laya decision engine.
flowchart TB subgraph UI["Editor"] A["Node graph editor"] B["Chain library + scheduler + agent chat"] end subgraph Engine["Player"] C["Workflow interpreter"] D["Desktop + web recorder"] E["Grounding (UIA, templates, OCR, vision)"] end subgraph AI["AI"] F["FastAPI local API"] G["llama.cpp + Ollama"] H["Embeddings + ComoRAG + Graph RAG"] I["Laya decision engine"] end subgraph Store["Persistence"] J["chains JSON"] K["sequences + web sequences"] L["context.db (SQLite)"] end A --> C B --> C C --> D C --> E C --> F C --> H C --> I F --> G C --> J C --> L

Runtime topology

  • Launcher — bootstraps the GUI, or runs a .py argument directly (how the sandbox agent and exported agents are spawned).
  • GUI process — PyQt5 main window; hosts the editor, overlays, scheduler, recorder, and (by default) the FastAPI server.
  • llama.cpp subprocess — the bundled llama-server.exe (default port 8081) serves GGUF inference.
  • Vision subprocess — multimodal inference via llama-mtmd-cli with a GGUF + projector pair.
  • Laya engine — an embedded decision daemon, lazy-launched and released after each burst; every consumer falls back to LLM paths on failure.
  • Exported agents — standalone executables that embed the app, the system chain, and transitive chains; they run with ARROW_DIRECT_ENGINE=1.
flowchart LR L["Launcher"] --> G["GUI process"] G --> A["FastAPI server :8000"] G --> M["llama-server subprocess :8081"] G --> V["llama-mtmd-cli (vision)"] G --> S["Sandbox agent (RDP)"] G -.-> Laya["Laya decision daemon"] G --> X["Exported agent exe"]

Node catalog

NodeRole
SequenceReplay a recorded desktop sequence.
Web SequenceReplay a recorded browser session.
ConditionalBranch on image / OCR / wait / web conditions.
InputValue injection, questions, decision routing, media acceptance.
LLMLocal inference: reasoning, tools, vision, structured output, TTS.
Chain ImportRun another chain as a sub-workflow or tool.
Form FillerDocument-grounded multi-field form completion.
CodeCustom Python with JSON in/out.
ContextPersistent cross-chain memory.
HandleVision-based click/scroll/zoom grounding.
MCPModel Context Protocol server tools.
OutputExpose results to the overlay / outside the chain.
OrchestratorGoal loop over mini-brain chains.

Chain metadata includes name, description, collection (e.g. System, Tools, Brains), and is_default (marks the default system chain).

Wiring patterns

Record, branch, act. A recorded sequence feeds a conditional; the chosen branch continues.

flowchart LR S["Sequence (recorded)"] --> C{"Conditional (image / OCR)"} C -- true --> L["LLM (reason)"] C -- false --> F["Sequence (fallback)"] L --> O["Output"] F --> O

LLM tool router. An LLM node picks one of several tool chains wired to its tools port.

flowchart LR I["Input (query)"] --> L["LLM (tool router)"] L -- tools port --> T1["Chain Import (web search)"] L -- tools port --> T2["Chain Import (fill form)"] L --> O["Output"]

Orchestrator. An orchestrator routes a goal across mini-brain chains and deterministic chains.

flowchart LR G["Input (goal)"] --> O["Orchestrator"] O -- brains port --> B1["Brain chain (research)"] O -- brains port --> B2["Brain chain (write)"] O -- chains port --> D["Deterministic chain"] O --> Out["Output + trace"]

Ports & data flow

Every node exposes named ports. Execution flows along exec edges (base ports such as output → input); values flow along data edges (data, true/false, trace, files, and so on). A data edge never triggers execution; an exec edge never carries values.

  • Port values are stored per chain and exposed as node_<id>_<port> variables.
  • Chain-to-chain state rides explicit variables such as _chain_input_context; there is no hidden global state.
  • Branch-capable nodes own true / false output ports.
  • A node whose exec edges land only on a router input port (llm.tools, orchestrator.brains) is flagged tool_provider and never runs in the main loop — its router invokes it.
flowchart LR I["Input node"] -- "exec edge" --> L["LLM node"] I -- "data edge (value)" --> L L -- "output (exec)" --> O["Output node"] L -. "node_X_data variable" .-> C["Code node"]

Branching & loops

  • Branch sweep — when a branch-capable node resolves, the unchosen branch is marked skipped so merges never stall.
  • Loop-backs — an Input node re-resolves on every pass, which is how iterative workflows are built with pure topology.
  • Decision re-evaluation — decision and routed questions are re-evaluated per loop pass.
flowchart TB Start["Chain start"] --> Build["Build typed graph"] Build --> Loop["Scheduler loop"] Loop --> Gate{"Any node ready?"} Gate -- yes --> Exec["Execute next ready node"] Exec --> Branch{"Branch-capable node?"} Branch -- yes --> Sweep["Route true / false, skip unchosen branch"] Branch -- no --> Pub["Publish port outputs + variables"] Sweep --> Pub Pub --> Loop Gate -- "none left / stopped" --> Done["Chain complete"]

Execution engine

Semantics

  • Dependency-gated scheduling — a node runs only when every incoming edge it needs is settled.
  • Tool-provider gating — chain-import nodes wired to a tool consumer run only when selected.
  • Stop semantics — a single stop flag is polled everywhere; ESC aborts a run immediately.
  • Fail-closed — unparseable model answers route to a configured default branch; never a crash, never a silent wrong turn.
  • Serialized model access — LLM calls share one managed engine and a serialization lock; local inference runs without timeouts.

Recording

Desktop recorder

Captures, alongside each action, the cropped image of the click target, the Windows UIA element it landed on (ControlType + bounding box), and precise coordinates. Rapid scrolls are grouped into one action; typed text is buffered into one action; clipboard shortcuts are recognized.

Web recorder

Records a CDP session on a real, visible browser. Recordings are stored as web sequences: element-driven steps with selector identities, never absolute coordinates. Identity rules favor stable attributes (data-*, ids, text containment) and treat repeated items as entities.

Playback & grounding

Desktop replay — two independent locators

  • Element-first (UIA) — resolve the recorded element's live bounding box, surviving moved windows, theme/DPI changes, and scrolled lists.
  • Template matching — multi-scale OpenCV matching within region constraints and confidence thresholds.
  • OCR conditions — PaddleOCR powers text conditions in Conditional nodes.
  • Human-like behavior — jittered click targets, curved mouse paths, variable timing.
flowchart TB Click["Click action"] --> UIA{"UIA element recorded?"} UIA -- yes --> Resolve["Resolve live element"] Resolve -- verified --> Native["Click real element"] UIA -- no --> Tmpl["Template matching"] Resolve -- miss --> Tmpl Tmpl -- found --> Human["Human-like click"] Tmpl -- miss --> Fallback["Fallback ladder"] Fallback -- found --> Human Fallback -- exhausted --> Coords["Recorded coordinates"]

Web replay

One handler class per action type, each running a deterministic ladder: native CDP input first, then in-page JS dispatch, then a locator retry ladder, then reconnect-and-retry. Replay is visible-browser by default with a single durable shared Chrome profile.

Anti-detection & human-like input

Playback reads like a person, not a bot: jittered click targets, curved mouse paths, variable timing, native CDP input (with an in-page JS fallback), and a durable shared browser profile. Web replay never captures your physical cursor.

Inference & local API

EngineTransportNotes
llama.cppllama-server.exe, port 8081GGUF models, GPU offload (Vulkan), embedding mode.
OllamaHTTP :11434Alternative provider, selected by name.
llama-mtmd-clisubprocess per inferenceMultimodal vision (GGUF + mmproj projector).

The engine is authoritative about context size — the app caps requests to the model's training context.

Local API surface (FastAPI)

EndpointPurpose
/health, /modelsLiveness and model inventory.
/chat, /generateUnified chat/generation to the active provider.
/embeddingsEmbedding vectors.
/llamacpp/*Direct llama.cpp lifecycle + inference.

Retrieval & Laya

Semantic retrieval

  • ComoRAG — probe-driven retrieval: intent, targeted probes, retrieve, then synthesize against verbatim chunks.
  • Graph RAG — surfaces relevant past knowledge to reasoning nodes and records execution fragments.

Laya decision engine

An embedded typed-decision model (noul / choice / score questions answered in one encoder pass) used by default for the orchestrator picker, input-node decisions, retrieval fact curation, and form-filler field judging. It runs as a lazy-launched daemon and is always optional — every consumer falls back to the LLM path on failure. LOOPER_LAYA=off forces LLM-only paths.

Vision & speech

LayerTechnology
Screen VLMLFM2.5-VL 450M + projector (Handle node scene understanding).
Object detectionONNX YOLO-based detector for UI element detection.
OCRPaddleOCR (det/rec/cls) for text conditions and documents.
Document extractionPyPDF2 (PDF), python-docx (Word), plain text.
TTSPiper ONNX voices.
STTVosk offline speech-to-text for voice mode.

Agent mode — system chains & tools

Agent mode has no hardcoded behavior. It requires a system chain (a chain with collection: "System") that defines the entire cognitive architecture: Input nodes receive the query, LLM nodes reason, Context nodes remember, Chain Import nodes act as tools, Output nodes surface the answer. Multiple system chains = multiple named agents.

  • Dispatch — selects the default system chain and runs it through the same engine as manual playback.
  • Chains-as-tools — any chain connected to an LLM's tools port is selectable at runtime; tool descriptions are derived at call time.
  • Stepped protocol — the router picks a tool by name, then each agent-modifiable Input node extracts its own single value. This is why a 3B model suffices.
  • Ask-user — any Input node with a prompt pauses the run and asks through the richest available channel.
sequenceDiagram participant U as User participant O as Chat overlay participant R as Chain router participant S as System chain participant L as LLM node participant T as Tool chain U->>O: query O->>R: handle request R->>S: run system chain S->>L: reasoning turn L->>T: tool picked by name T-->>L: tool result L-->>S: final answer S-->>O: output + rating O-->>U: answer

Orchestrator

The default system chain ships with an Orchestrator node as its root: several mini-brain chains are wired to its brains port, and the orchestrator loops toward a goal — it scopes one step at a time and asks the picked specialist brain to run it.

  • Three inputs to every decision — the goal, what was done (step result + trace), and the world now (observed state digest).
  • Picker — one strictly-parsed token (worker number, DONE, ASK), or the Laya engine with zero parsing.
  • Verification — a per-step verdict judges whether the step accomplished what it was asked.
  • Guards — step cap, no-progress detection, and a stop flag between steps.
  • Synthesis — a final short LLM pass turns the trace into a plain-language answer.
flowchart TB G["Goal arrives"] --> W["Observe world (page + window)"] W --> P{"Picker: next step"} P -- worker --> Sc["Scope one instruction"] Sc --> R["Run mini brain"] R --> V["Collect result + verify"] V --> Guard{"Stop or progress?"} Guard -- continuing --> W Guard -- cap --> Synth["Synthesize answer from trace"] P -- DONE --> Synth P -- ASK --> Q{"User channel?"} Q -- yes --> V Q -- no --> B["Blocked, park goal"] Synth --> Out["Output + trace"]

Deterministic chains

The orchestrator's second input port carries deterministic chains: orchestrator-free compositions that answer a known request end to end with no routing inference. One rule decides which kind a chain is — does it contain an orchestrator, directly or through its imports?

  • A verified run can be frozen into a learned chain (composition + routing.examples), wired to the chains port.
  • Self-improving — learned chains accumulate on the chains port, so the agent's library grows with every verified run.
  • Partial fits are assembled from existing artifacts into one continuous chain and run as a single dispatch.
  • The gate asks the step question per directive, judged against the current directive with a confidence floor plus object-name alignment.

Input nodes & persistent goals

Input node v2

  • Decision mode — a natural-language criterion evaluated with one strictly-contracted model turn or the Laya noul evaluator.
  • Question modes — free text, yes/no, or a choice list; answers are normalized against the options.
  • Accepted media — text, images, and documents (PDF/Word/txt) ride the data port.

Persistent goals

With use_goal_ledger enabled, a blocked goal parks itself (blocked port) and resumes on the next activation — wake-driven persistence, not a busy daemon. Schedules create the wakes.

Memory & context

  • In-run state — named-port values plus a turn-based history per chain.
  • Persistence — Context nodes write append-only turns to SQLite (data/context.db) with per-(chain, node) rolling caps.
  • Retrieval — Graph RAG surfaces relevant past knowledge; ComoRAG consolidates large sources into grounded facts.
  • Inspection — the agent settings dialog renders a goal-relationship graph with reviewable, correctable, clearable memory.
flowchart TB N["Workflow nodes"] --> S["Port store (in-run state)"] S --> D["context.db (SQLite)"] D --> R["Graph RAG"] D --> C["ComoRAG consolidation"] R --> L["LLM nodes"] C --> L

Specialized nodes

Form Filler

Document-grounded multi-field completion: field discovery on live pages, knowledge-file lookups, per-field validation, repair loops, and a review gate. Choice controls are resolved by the Laya semantic picker, which refuses unclear picks so the field is asked instead of guessed.

Handle

Combines VLM scene understanding with LLM reasoning to issue grounded click/scroll/zoom/key actions on screen content that lacks selectors.

MCP

Connects Model Context Protocol servers as workflow nodes; discovered tools are exposed to the router model as individually selectable actions.

Code

Inline or file-based Python with JSON arguments and output variables. The Code Node Studio embeds an Aider-style iterative local-agent loop — an integrated coding agent that writes, edits, and debugs Python inside the graph.

Scheduling, sandbox & remote access

Scheduler

Runs any chain on a schedule — once, daily, weekly, interval — with enable/disable, next-run display, and run-now. Schedules persist to schedules.json and execute in-process.

RDP sandbox

Automation runs inside an isolated Windows session via a local HTTP agent, so the desktop you are using stays untouched. Chains can opt in per sequence/chain-import.

Remote access

The agent chat is served over LAN or Tailnet with a scannable QR code — a phone can drive runs, answer ask-user questions, and stop executions.

Export & build

Standalone agent export

Collects every chain referenced transitively (including orchestrator brains), bundles chains/sequences/models, and builds a one-folder executable. Exported agents run directly on the local engines (ARROW_DIRECT_ENGINE=1) — no FastAPI dependency, fully offline.

Application build (MSI)

The build pipeline stages dependencies (llama.cpp binaries, Chrome, model files), runs PyInstaller, generates a WiX source, and produces a per-user MSI installer.

Configuration & data

Environment variables

VariablePurpose
API_PORT, AUTO_START_APILocal API control (default 8000 / true).
OLLAMA_HOST, OLLAMA_PORTOllama endpoint.
LOOPER_LAYAoff forces all Laya consumers onto LLM paths.
LOOPER_ORCH_STATEoff disables the orchestrator's per-step state digest.
LOOPER_LEARNoff disables freezing runs into learned chains.
ARROW_DIRECT_ENGINE1 in exported agents: direct engines, no API calls.

Data locations

DataLocation
Chainschains/*.json (learned: chains/learned/*.json)
Sequencessequences/*.json, web_sequences/*.json
Context memorydata/context.db (SQLite)
Schedulesschedules.json
Goal ledgerdata/agent_goals.json
Models / voicesdata/piper_voices/, data/vosk_models/

Engineering principles

  • Neuro-symbolic execution — the graph is a deterministic program; AI has a fallback or fail-closed default everywhere.
  • Small-model discipline — stepped protocols and single-token contracts for 1–4B models.
  • Named-port contracts — values move only along explicit edges.
  • Runtime-derived descriptions — tool descriptions are generated at call time, never persisted stale.
  • Fail-closed defaults — new capabilities ship default-off; unparseable answers route to defaults.
  • Entity-first grounding — selectors and UIA elements over absolute coordinates.
  • No timeouts on local inference — interruption is a stop-flag concern.
  • Serialized, on-demand model lifecycle — no preload at startup; burst-aware cleanup.
  • Approvals and memory are primitives — composed by users, not hardcoded policy.
  • Offline and telemetry-free — nothing leaves the machine.
  • Inspectable and reversible — logs, trace ports, blocked states; ESC always aborts.

Licensing

Arrow is open source. Use it, modify it, and share it freely — every feature and node type, no execution limits.

  • All features included: visual workflows, agent mode, the orchestrator, vision, OCR, speech, and local AI.
  • Runs fully offline — no activation servers, no telemetry.
  • Export agents as standalone executables straight from the source.
  • The packaged MSI installer (WiX) is the recommended path for end users; the source is the recommended path for contributors.