Live System

Local Vector Agent

Zero-cost, zero-internet AI tool calling. Powered by Ollama + local vector search over 20 indexed tools.

20
Tools Indexed
0
API Cost
<4s
Avg Response
100%
Local

How It Works

Every message flows through a 5-step pipeline. The local agent sits at priority 15, catching all non-test messages before they hit the OpenAI-powered production path.

Step 1
Embed
nomic-embed-text
~5ms local
→
Step 2
Vector Search
Cosine similarity
20 tool docs
→
Step 3
Route
3-tier confidence
scoring
→
Step 4
LLM Decide
qwen2.5:7b
JSON tool call
→
Step 5
Execute + Reply
Run tool, format
natural response

Three-Tier Confidence Routing

The vector similarity score determines how aggressively we use tools. This prevents the LLM from hallucinating tool calls on casual chat.

ScoreTierBehaviorExample
> 0.70 High Confidence LLM is told to call the tool with reasonable defaults. No hesitation. "Generate an auth link for Nathan" → create_auth_link
0.60 – 0.70 Medium Confidence LLM sees candidates but is told to only call if the intent is clear. "What's happening on X about AI?" → grok_x_search
< 0.60 No Match Skip tools entirely. Pure conversational reply via Ollama. "Hey Alfred, how are you?" → Natural chat response

Server Architecture

Messages flow through a hot-reloadable plugin pipeline. Each plugin declares a priority and a match function. The first match wins.

1 noise_filter Reject spam, empty messages, delivery receipts
10 fast_reply OpenAI-powered — catches "test" prefix messages only
15 local_agent ★ This one — Ollama + vector DB, handles all other messages
20 vector_tools OpenAI + vector DB (production iMessage path, never reached on webchat)
90 cursor_agent Fallback — sends coding tasks to Cursor IDE

Why This Architecture

🧠

Local LLM (qwen2.5:7b)

7B parameter model via Ollama. Fast enough for routing decisions (~1-2s), small enough to run alongside other services. Handles JSON output mode for reliable tool call generation.

📐

Vector-First Routing

Embeddings narrow 20 tools to 1-5 candidates before the LLM sees anything. This means the LLM gets rich context per tool instead of shallow descriptions of all tools.

🔌

Hot-Reload Plugins

Plugins are re-imported if their file's mtime changes. Edit local_agent.py, save, and the next message uses the new code — no server restart needed.

🌐

Proxy Architecture

The webchat page is served through the main server (port 8787, Cloudflare), which proxies /api/chat to the toolcall server (port 8788). Single domain, no CORS.

💰

Zero API Cost

Embedding (nomic-embed-text) and reasoning (qwen2.5:7b) both run on Ollama locally. The only external calls are when tools themselves need APIs (Grok, Twilio, etc.).

🔄

Reuses Existing Executor

Tool execution shares the same executor as fast_reply.py — all 20 MCP tools work identically. No duplication of tool integration code.

Live Test Results

Every test below was run against the live local agent. Tool calls execute real tools, not mocks.

✓

Auth Link Generation

"Can you generate an auth link for Nathan?"
→ Vector matched create_auth_link at 75%. Ollama called it correctly. Returned a live auth URL.
4.0s · Tool executed · Real auth token generated
✓

X/Twitter Search

"Search X for the latest news about AI agents"
→ Vector matched grok_x_search at 69%. Ollama built the query, called Grok API, summarized results about AI agent market reaching $183B.
24.1s · Tool executed · Real Grok API call
✓

User Nudge (Dry Run)

"Can you send nudges to inactive users? Just a dry run please"
→ Vector matched user_nudge at 80%. Ollama set dry_run=true. Scanned all users, generated personalized messages without sending.
14.8s · Tool executed · Ollama generated nudge messages
✓

YouTube Transcript (Error Handling)

"Get me the transcript for this YouTube video: https://youtube.com/..."
→ Vector matched youtube_transcript at 74%. Tool called, returned connection error. Ollama formatted a helpful error message.
4.6s · Graceful error handling
✓

Pure Conversation (No Tools)

"Hey Alfred, what are you up to?"
→ No tools above 0.60 threshold. Skipped tool routing entirely. Ollama replied conversationally: "Not much, just keeping an ear out for users like you!"
1.4s · No tools invoked · Pure chat
✓

Vector Search Accuracy

"send a message to Nathan" — inline unit test
→ Top match: send_imessage (0.737), followed by send_sms_group_message (0.690), send_twilio_sms (0.669). Correct ranking.
<2s · 20 tools searched · nomic-embed-text

20 Indexed Tools

All tools are embedded with nomic-embed-text and stored in database/tool-vectors.json. Each entry includes the full tool doc (summary, parameters, examples, triggers) for rich vector matching.

Test the Local Agent

This chat connects directly to the local vector agent via /api/chat. All tool calls are real — auth links, searches, nudges, etc. Try anything.

Alfred — Local Agent

Checking...
Hey! I'm Alfred's local agent — fully powered by Ollama and vector search. Ask me to do something or just chat. I have access to 20 tools.
Generate an auth link for Nathan Search X for AI agent news Show me inactive users (dry run) Hey, how are you?