A standardized system for building, indexing, and discovering AI tools across any number of nodes. Wide codebase, zero merge conflicts, vector-powered discovery.
Every tool is a standalone Python file that runs by itself. No shared imports, no base classes, no framework. An LLM on any node can build a new tool from scratch without understanding the rest of the repo.
THE PROBLEM WITH DEEP CODEBASES
══════════════════════════════════════════════════════════════
Traditional repo:
┌────────────────────────────────────────────────────────────┐
│ shared_utils.py ←──┐ │
│ base_tool.py ←──┤── Every tool imports these │
│ config.py ←──┤── Change one → break all │
│ auth_helper.py ←──┘── Merge conflicts everywhere │
│ │
│ tool_a.py ──imports──► 5 shared modules │
│ tool_b.py ──imports──► 3 shared modules │
│ tool_c.py ──imports──► 7 shared modules │
│ │
│ Result: LLMs struggle. Humans struggle. Merges break. │
└────────────────────────────────────────────────────────────┘
This repo:
┌────────────────────────────────────────────────────────────┐
│ │
│ tool_a.py ──imports──► stdlib only ✓ │
│ tool_b.py ──imports──► stdlib only ✓ │
│ tool_c.py ──imports──► stdlib only ✓ │
│ │
│ Each file is a universe unto itself. │
│ Copy it to an empty directory → it still works. │
│ 100 agents build 100 tools in parallel → zero conflicts. │
│ │
└────────────────────────────────────────────────────────────┘
The insight: LLMs are really good at building something from scratch in a single file. They're less good at performing heart surgery on a repo full of dependencies, shared components, and cross-file imports. We lean into the strength.
2. The Discovery Pipeline
Tools are discovered by meaning, not by name. A user says "I need an auth link for Nisha" and the system finds create_auth_link with zero token cost — no need to stuff 19 tool definitions into every GPT call.
💬
User Message
"I need an auth link for Nisha"
🧮
Embed
Ollama nomic-embed-text (local, free)
🔍
Vector Search
Cosine similarity vs 19 tool docs
📋
Rich Context
Top 5 tool docs loaded into GPT
⚡
Execute
GPT calls the tool, returns result
Why Vector Search Instead of Stuffing All Tools?
Approach
Tokens per call
Latency
Scales to
Stuff all 19 tools as OpenAI functions
~3,000 tokens
Included in API call
~50 tools before context window issues
Vector search → top 5 with rich docs
~800 tokens (5 relevant only)
+80ms (local embed)
1,000+ tools with no degradation
The magic: Each tool doc is written with natural language descriptions, example queries, and trigger scenarios. The embedding captures the meaning of what the tool does. So "check my DMs" matches social_inbox_check at 0.61 similarity even though those words never appear in the tool name.
3. The Standard Format
Every tool is two files: the Python tool itself and its JSON documentation for the vector database.
File 1: The Tool — tools/my-tool.py
#!/usr/bin/env python3
"""
my_tool — One-line summary of what this tool does.
WHEN TO USE: ◄── Agents read this section
- User asks for X or wants to Y to decide if this tool
- "Example user message that triggers this" is the right one.
- Another scenario where this is the right tool
DESCRIPTION: ◄── This gets embedded in
Longer explanation of what the tool does, the vector database.
how it works, what API it connects to. Write for SEMANTIC
Include synonyms and related concepts. MATCHING, not humans.
PARAMETERS: ◄── Maps to CLI args
--name (required) Person's name and JSON schema.
--format (optional) Output format, default: text
RETURNS: ◄── What callers expect.
JSON: {"result": "...", "error": null}
EXAMPLES: ◄── Copy-paste runnable.
python3 tools/my-tool.py "value" --json
DEPENDENCIES: ◄── Almost always "None."
None beyond Python 3.10+ stdlib.
"""
import argparse, json, os, sys # ◄── stdlib only
API_KEY = os.environ.get("MY_API_KEY", "") # ◄── config from env
def my_tool(name: str, format: str = "text") -> dict:
"""Pure logic. No printing. No sys.exit.""" # ◄── testable function
return {"result": f"Hello {name}"}
def main():
parser = argparse.ArgumentParser(...)
# ...
if args.json: # ◄── machine-readable output
print(json.dumps({"result": ..., "error": None}))
if __name__ == "__main__":
main()
File 2: The Doc — tools/my-tool.tool.json
This is what the vector database indexes. Write it for semantic matching — the embedding of this doc is compared against user messages via cosine similarity.
{
"tool_name": "my_tool", ◄── snake_case, unique
"summary": "One sentence, shown in search.", ◄── brief
"description": "2-4 sentences. Include synonyms, ◄── this gets EMBEDDED
related concepts, the kinds of questions that so write for meaning,
should trigger this tool.", not for humans
"when_to_use": [ ◄── natural language
"User wants to X", trigger scenarios
"User says 'do Y for Z'"
],
"example_queries": [ ◄── literal things
"I need to X for Nisha", users might say
"Can you check the Y?"
],
"parameters": { ◄── maps to OpenAI
"name": { function schema
"type": "string",
"required": true,
"description": "Person's name"
}
},
"returns": "What the caller gets back.",
"file": "tools/my-tool.py", ◄── path from root
"standalone": true, ◄── always true
"dependencies": [] ◄── pip packages, if any
}
4. The Tool Builder
build-tool.py — The Meta-Tool
Takes a natural language description and does everything automatically:
📝
Describe
"A tool that checks domain WHOIS info"
🔎
Dupe Check
Vector search for existing tools ≥0.65
🏗️
Scaffold
.py + .tool.json pre-filled by local LLM
🤖
Cursor Task
Detailed instructions for implementation
Usage
# Just check if something similar already exists
$ python3 tools/build-tool.py "Send a text message" --check-only
SIMILAR TOOLS FOUND:
send_imessage (0.741) — Send iMessage/SMS via BlueBubbles
send_twilio_sms (0.691) — Send SMS via Twilio number
create_sms_group(0.675) — Create group MMS thread
# Scaffold a genuinely new tool
$ python3 tools/build-tool.py "Look up WHOIS info for a domain" --name domain-lookup
FILES CREATED:
tools/domain-lookup.py ← Python file, template-filled
tools/domain-lookup.tool.json ← Vector DB documentation
database/tool-docs/domain-lookup.json ← Copy for indexer
CURSOR AGENT TASK:
────────────────────────────────────────────────
Implement the tool at tools/domain-lookup.py.
READ tools/TOOL-BUILDING-GUIDE.md FIRST...
(full implementation instructions for Cursor)
# Full auto: scaffold + hand to Cursor immediately
$ python3 tools/build-tool.py "Fetch podcast RSS feeds" --auto-build --json
What this solves: Every new tool starts from the same template, with the same format, in the same directory. The vector doc is pre-generated. The implementation instructions reference the guide. No agent invents its own format. No duplicates sneak in.
5. Design Rules
📦
Self-Contained
Each file works if copied to an empty directory. No internal imports. Duplicate small helpers rather than creating shared dependencies.
🔑
Env-Configured
No hardcoded URLs, API keys, ports, or paths. os.environ.get("KEY", "default") for everything. Every node has its own .env.
🔄
Idempotent
Running a tool twice with the same inputs produces the same result or safely skips. No "first-run-only" behavior.
📊
JSON Output
--json flag always outputs {"result": ..., "error": ...}. No extra logging to stdout. Use stderr for debug.
🛡️
Defensive
Timeouts on all network calls. UTF-8 safe. Truncate huge outputs. Check required env vars early with clear error messages.
🚫
Non-Interactive
No input(), no prompts, no stdin reads. Agents can't type. Everything through CLI args and env vars.
File Structure
tools/
├── TEMPLATE.py ← Copy this to start any new tool
├── TEMPLATE.tool.json ← Copy this for the vector DB doc
├── TOOL-BUILDING-GUIDE.md ← Strategy guide for agents
├── build-tool.py ← Meta-tool: dupe check + scaffold + Cursor task
├── build-tool.tool.json ← Its own vector DB doc
├── requirements.txt ← Shared pip deps (keep minimal)
│
├── my-tool-name.py ← A tool
└── my-tool-name.tool.json ← Its documentation
database/
├── tool-docs/*.json ← All tool docs (indexed by ingest-tools.py)
├── tool-vectors.json ← Pre-computed embeddings (JSON fallback)
├── ingest-tools.py ← Embeds tool docs → vector store (fast)
└── reindex-vectors.py ← Full codebase reindex (376 files, ~3 min)
Legacy tools (pre-standard):
├── mcp-server/src/*.py ← Will migrate to tools/ over time
└── scripts/*.py ← Will migrate to tools/ over time
6. Current Tool Library
create_auth_link
Generate one-click authenticated access links for any page
send_imessage
Send iMessage or SMS via BlueBubbles / Messages.app
bb_history
Pull real iMessage conversation history for any contact
send_twilio_sms
Send SMS via Twilio phone number (green bubble)
create_sms_group
Create group MMS thread between 2+ people
send_sms_group_message
Message an existing Twilio group thread
list_sms_groups
List all active Twilio group conversations
ask_ollama
Run prompts through local LLMs — zero API cost
grok_chat
Chat with Grok (xAI) for analysis and reasoning
grok_web_search
Live web search via Grok with citations
grok_x_search
Search X/Twitter posts, trends, and public sentiment
ask_delphi
Query Delphi AI clones (Pace Morby, Nick Fisher)
youtube_transcript
Fetch captions from any YouTube video
read_link
Read any URL as clean text via Jina Reader
hivemind
Agent-to-agent communication (Luna, Jarvis, Q)
fb_group_member_requests
Pull pending Facebook group members + Q&A answers
social_inbox_check
Check unread DM counts on Facebook + Instagram
run_cursor_agent
Hand off complex tasks to the Cursor IDE agent
build_tool
Scaffold new tools with the standard format
7. Vector Database
Infrastructure
Component
Tech
Details
Database
Postgres 17 + pgvector 0.8.1
Local, HNSW index for sub-1ms search. 2,842 rows (2,824 codebase + 18 tool docs).
Embeddings
Ollama nomic-embed-text
768 dimensions, runs locally on Apple Silicon. ~40-80ms per embed. Zero API cost.
Fallback
database/tool-vectors.json
Pre-computed embeddings as JSON. Works without Postgres. Same search quality.
Plugin
plugins/vector_tools.py
Priority 20 in server_toolcall.py. Embeds user message → finds tools → GPT decides.
Search Quality (real results)
Query Top Match Similarity
───────────────────────────────────── ────────────────── ──────────
"I need an auth link for Nisha" create_auth_link 0.755
"What did Pace text me?" bb_history 0.672
"Search the web for AI news" grok_web_search 0.707
"Check the Facebook group requests" fb_group_member_req 0.715
"Send a message to Nathan" send_imessage 0.712
"Build me a tool that checks stocks" build_tool 0.648
The right tool lands at #1 every time. This is why the natural language docstring matters so much — it's the primary signal for vector matching. A well-written tool doc with good when_to_use and example_queries is the difference between 0.55 and 0.75 similarity.
8. Guide for Agents
Before Building a New Tool
Read the guide:tools/TOOL-BUILDING-GUIDE.md
Check for duplicates:
python3 tools/build-tool.py "what you want to build" --check-only
If anything comes back ≥0.65, extend it instead of building new.
Scaffold:
python3 tools/build-tool.py "description of your tool" --name my-tool
This creates the .py, .tool.json, and Cursor task automatically.
Implement the TODO in the generated Python file. Keep it self-contained.
Test:
python3 tools/my-tool.py --help # should print usage
python3 tools/my-tool.py <args> --json # should print valid JSON
Index:python3 database/ingest-tools.py
Verify:python3 mcp-server/src/vector-search.py "a query it should match"
The promise: If every agent follows this process, 100 nodes can build 100 tools in parallel. Every tool is a single file. Every tool has a rich vector doc. No merge conflicts. No duplicate tools. No broken imports. The library grows cleanly forever.