Zero-cost, zero-internet AI tool calling. Powered by Ollama + local vector search over 20 indexed tools.
Every message flows through a 5-step pipeline. The local agent sits at priority 15, catching all non-test messages before they hit the OpenAI-powered production path.
The vector similarity score determines how aggressively we use tools. This prevents the LLM from hallucinating tool calls on casual chat.
| Score | Tier | Behavior | Example |
|---|---|---|---|
| > 0.70 | High Confidence | LLM is told to call the tool with reasonable defaults. No hesitation. | "Generate an auth link for Nathan" → create_auth_link |
| 0.60 – 0.70 | Medium Confidence | LLM sees candidates but is told to only call if the intent is clear. | "What's happening on X about AI?" → grok_x_search |
| < 0.60 | No Match | Skip tools entirely. Pure conversational reply via Ollama. | "Hey Alfred, how are you?" → Natural chat response |
Messages flow through a hot-reloadable plugin pipeline. Each plugin declares a priority and a match function. The first match wins.
7B parameter model via Ollama. Fast enough for routing decisions (~1-2s), small enough to run alongside other services. Handles JSON output mode for reliable tool call generation.
Embeddings narrow 20 tools to 1-5 candidates before the LLM sees anything. This means the LLM gets rich context per tool instead of shallow descriptions of all tools.
Plugins are re-imported if their file's mtime changes. Edit local_agent.py, save, and the next message uses the new code — no server restart needed.
The webchat page is served through the main server (port 8787, Cloudflare), which proxies /api/chat to the toolcall server (port 8788). Single domain, no CORS.
Embedding (nomic-embed-text) and reasoning (qwen2.5:7b) both run on Ollama locally. The only external calls are when tools themselves need APIs (Grok, Twilio, etc.).
Tool execution shares the same executor as fast_reply.py — all 20 MCP tools work identically. No duplication of tool integration code.
Every test below was run against the live local agent. Tool calls execute real tools, not mocks.
create_auth_link at 75%. Ollama called it correctly. Returned a live auth URL.grok_x_search at 69%. Ollama built the query, called Grok API, summarized results about AI agent market reaching $183B.user_nudge at 80%. Ollama set dry_run=true. Scanned all users, generated personalized messages without sending.youtube_transcript at 74%. Tool called, returned connection error. Ollama formatted a helpful error message.send_imessage (0.737), followed by send_sms_group_message (0.690), send_twilio_sms (0.669). Correct ranking.
All tools are embedded with nomic-embed-text and stored in database/tool-vectors.json.
Each entry includes the full tool doc (summary, parameters, examples, triggers) for rich vector matching.
This chat connects directly to the local vector agent via /api/chat.
All tool calls are real — auth links, searches, nudges, etc. Try anything.