This page is intentionally “living”: every new discovery + fix gets appended here so Matt always has one canonical reference.
This is a comprehensive checklist of actions that improve foundations without changing Alfred’s code. It focuses on visibility, backup/safety, deployment hygiene, workspace legibility, and reducing “unknown unknowns”.
Progress + notes are stored in your browser (localStorage). Use export/import to move progress between devices/browsers.
techteam over LAN./p/upload), but the running agent-orchestrator container appears to be on an older build that does not include the upload routes the UI expects (/api/upload, /api/uploads, /uploads).oc-bot-refactor), cloud/services/agent-orchestrator/main.py does define those upload routes — but the container’s baked /app/main.py does not contain them. That strongly suggests “source has the fix, but it was never rebuilt/deployed”.403 for gated/unknown paths before routing, so the page incorrectly concludes /api/upload “exists” and then POSTs to a guaranteed-fail path.alfredhottakes.com) + a voice server (voice.alfredhottakes.com) + an API server (api.alfredhottakes.com). That’s flexible, but it increases “page points at wrong backend” failure modes.The transcript you shared shows multiple different failure modes that look like “memory issues”, but are usually a mix of: message delivery quirks, split services, and the model confabulating when it lacks ground truth.
+1831..., 1831..., tel:+1831..., etc.| Hostname | Local target | Purpose | Notes |
|---|---|---|---|
alfredhottakes.combot.alfredhottakes.com |
localhost:8887 |
Docker container agent-orchestrator (ports 8887 → 8787) |
/p/sitemap is public (200). Many API endpoints are gated (403) unless authenticated. |
voice.alfredhottakes.com |
localhost:8788 |
Uvicorn app (malli/server.py) |
/docs is open (200) but it’s not the pages/upload host. |
api.alfredhottakes.com |
localhost:8042 |
Another Uvicorn app (main:app) |
Not serving /p/* pages; appears to be API-only. |
~/.cloudflared/config.yml - alfredhottakes.com → http://localhost:8887 - bot.alfredhottakes.com → http://localhost:8887 - voice.alfredhottakes.com → http://localhost:8788 - api.alfredhottakes.com → http://localhost:8042
compose working dir: oc-bot-refactor/cloud/services/agent-orchestrator/ compose files: docker-compose.yml docker-compose.override.yml env file: oc-bot-refactor/.env
- oc-bot-refactor/local → /workspace (so host local/pages is container /workspace/pages) - cloud/services/agent-orchestrator/data → /app/data - oc-bot-refactor/logs → /logs - oc-bot-refactor/local/context → /workspace/context - plus campaign + malli mounts
oc-bot-refactor/local/pages/upload.html
This is the “make the repo readable again” section. Alfred currently has two separate repos on disk:
oc-bot-refactor (newer orchestrator architecture) and oc-bot (older/legacy stack).
The biggest source of confusion is that production-facing services pull from both.
oc-bot in Cursor)oc-bot-refactor, Cursor will not automatically show sibling folders like oc-bot unless you add them to the workspace.voice.alfredhottakes.com is a host-run Uvicorn process whose working directory is /Users/techteam/oc-bot (so oc-bot is actively used today).alfredhottakes.com is the Docker agent-orchestrator built from /Users/techteam/oc-bot-refactor.oc-bot-refactor.AGENT_NAME=Alfred, domain variables, page titles, and UI copy.oc-bot-refactor → alfred) only after deployment is stable; optionally keep a symlink for compatibility.agent-orchestrator if it’s already wired everywhere; “pretty names” can be layered on in UI/docs./Users/techteam/oc-bot-refactor (new / canonical)
Size snapshot (approx): dashboard/ ~173MB, local/ ~129MB, logs/ ~17MB. Empty/placeholder files found: ~75 (most are virtualenv artifacts under local/services/.venv).
Common root files you’ll see in Cursor (oc-bot-refactor): - .env / .env.example env config for services (keep; make sure secrets aren’t committed) - README.md / CONTEXT.md / MALLI_ROADMAP.md documentation (keep) - package.json / package-lock.json dashboard deps (keep) - tools.json / mcp-servers.json / memory.json agent/tool registry + memory (keep; consider consolidating docs) If a root file is empty or mysterious: - treat it as suspect until you know who reads it (search references before deleting) - prefer moving runtime state into local/context or service data dirs, not root.
| Path | What it is | Actively used? | Recommendation | Why / how (later) |
|---|---|---|---|---|
cloud/services/agent-orchestrator/ |
Main Alfred portal/orchestrator FastAPI service (Docker image build context). | Yes (serves alfredhottakes.com). |
KEEP (core). | Make this the single owner of pages + uploads + auth. When we “fix Alfred”, most changes should land here. |
cloud/services/agent-orchestrator/main.py |
The orchestrator app entrypoint (huge; owns routes, auth middleware, pipeline wiring). | Yes (source-of-truth), but deployed container may be behind. | CHANGE (deploy hygiene). | Ensure the running container’s /app/main.py matches this file (rebuild/redeploy later). Add a version stamp endpoint to detect drift. |
cloud/services/agent-orchestrator/pipeline/ |
Core pipeline modules: auth, identity, history, send, tools, tasks, etc. | Yes. | KEEP. | This is the “brains” of the orchestrator. Prefer refactors here over spawning new services. |
cloud/services/agent-orchestrator/portal/ |
HTML portal UI (login/dashboard pages served by orchestrator). | Yes. | KEEP. | Central place for “operator UX”. If pages feel confusing, improve these instead of adding more ad-hoc pages. |
cloud/services/agent-orchestrator/data/ |
Stateful storage: identity sqlite/dbs, visitors, prompt audits, processed GUIDs. | Yes. | LEAVE ALONE (but document). | Do not delete casually. Later: make a clear backup/rotation policy; consider moving to a dedicated state directory outside the repo tree. |
local/ |
Host-mounted workspace data + pages + modes + local services. | Yes (mounted into Docker as /workspace). |
KEEP (but organize). | Think of local/ as “runtime workspace”, not product code. Later: tighten which subfolders are “state” vs “config” vs “UX pages”. |
local/pages/ |
Public/operator HTML pages (includes upload.html). |
Yes. | CHANGE (align contracts). | Keep pages, but remove brittle endpoint auto-detection and align to the deployed API surface. |
local/context/ |
Stateful conversations + attachments + runner progress. | Yes (core bot memory). | LEAVE ALONE (but fence it). | Do not delete casually. Later: ensure it’s gitignored + backed up; optionally relocate to a dedicated data volume outside repo. |
local/services/cursor-runner/ |
“Hardened runner” service (runs agent jobs; orchestrator prefers it). | Yes (called by orchestrator). | KEEP. | Good separation: runner can be sandboxed and scaled later. |
local/services/agent-server/ |
Decision-tree / fallback agent service. | Yes (configured fallback). | KEEP (but consider deprecating later). | If cursor-runner is stable, you may phase this out to reduce moving parts. |
local/services/rv-park-finder/ |
Standalone “RV Park Finder” FastAPI microservice (port 8042). | Yes (serves api.alfredhottakes.com). |
MOVE / ISOLATE. | It’s not part of Alfred’s basic UX; it can consume resources. Later: move to cloud or separate host; keep the API stable. |
local/services/.venv/ |
Checked-in local Python virtualenv for host-run services. | No (not “source of truth”; it’s generated). | DELETE (cleanup) once we confirm nothing is hard-pinned to it. | This is a major repo-noise source and explains many “empty files”. Later: remove the venv dir, ensure it’s ignored, and reinstall from requirements.txt. |
dashboard/ |
Next.js dashboard frontend. | Unknown (exists in compose, but not mapped in current Cloudflare ingress). | KEEP (for now). | Later: either properly publish it under a hostname or remove it from the default “prod” compose to reduce noise. |
whatsapp-bridge/ |
WhatsApp Web bridge service (Baileys). | Unknown (present in compose; not confirmed live). | KEEP (for now). | Later: only run in environments where WhatsApp is actively needed; otherwise disable to reduce surface area. |
logs/ |
Runtime logs directory mounted into containers. | Yes. | KEEP (but rotate). | Add log rotation + size caps; don’t keep unbounded logs inside the repo forever. |
.venv-*, .pytest_cache/, __pycache__/, tmp/, graphify-out/ |
Local caches/build artifacts. | No (not required for runtime correctness). | DELETE (safe cleanup). | Later: remove locally + ensure ignored. This is the fastest “workspace sanity” win with minimal risk. |
/Users/techteam/oc-bot (legacy / still powering Twilio)
Size snapshot (approx): imessage-bridge/ ~888MB, installer/ ~147MB, vercel/ ~156MB, malli/ ~35MB. Empty/placeholder files found: ~433 (overwhelmingly in imessage-bridge/.venv).
| Path | What it is | Actively used? | Recommendation | Why / how (later) |
|---|---|---|---|---|
malli/server.py |
Twilio voice/SMS gateway + Malli persona/queue endpoints. | Yes (serves voice.alfredhottakes.com). |
KEEP (temporarily). | This is “live legacy”. Later: migrate the Twilio surface into oc-bot-refactor (or at least unify identity/auth + upload conventions). |
imessage-bridge/, messages-bridge/ |
Older bridge stacks (iMessage/other messaging surfaces). | Unknown (not confirmed in current Cloudflare ingress). | ARCHIVE / DEPRECATE (if unused). | Later: confirm whether any Alfred surfaces still depend on them; if not, move to an archive/ folder or separate repo to reduce cognitive load. |
imessage-bridge/.venv/ |
Committed Python virtualenv inside the legacy repo. | No (generated). | DELETE (cleanup) once confirmed unused. | This is the single biggest “clutter amplifier” (hundreds of empty REQUESTED/py.typed marker files). Replace with requirements.txt + recreate on demand. |
.pycache/ |
Unusual cache tree that mirrors absolute paths (e.g. .pycache/Users/...). |
No (generated). | DELETE (safe cleanup). | It’s almost certainly a cache directory produced by a nonstandard pycache prefix config. Later: delete and add ignore rules; check env like PYTHONPYCACHEPREFIX. |
vercel/ |
Frontend build/deploy artifacts. | Unknown (likely generated). | DELETE (cleanup) if not actively used. | If nothing is serving from it, it’s pure noise. Later: delete and ensure build output dirs are ignored. |
node_modules/, dist/, .pnpm-store/ |
JS build artifacts/deps. | Maybe (depends on whether legacy frontends are used). | DO NOT TOUCH blindly. | Later: if legacy web frontend isn’t used, remove + regenerate on demand; otherwise keep but ensure it’s not tracked in git. |
pages/ |
Legacy HTML pages for the old stack. | Unknown. | ARCHIVE (if unused). | Later: extract any still-valuable content into oc-bot-refactor/local/pages and deprecate the rest. |
database/, data/, agent_data/, logs/ |
Legacy state directories. | Possibly. | LEAVE ALONE until proven safe. | Later: identify which state is still referenced by running processes; migrate it into the new state layout if needed. |
local/context, cloud/services/agent-orchestrator/data, legacy oc-bot state dirs. Back up before any reorg.tmp/, __pycache__/, stray venvs, build outputs — the stuff that never should have been in the “mental model”.oc-bot-refactor (or at minimum unify config/auth and keep it as a clearly named “gateway” service).oc-bot-refactor:
uploads API, where files land, how pages call APIs, how auth works, where context lives.oc-bot/malli/server.py into oc-bot-refactor as a dedicated service,
or keep it as a separate service but with a versioned interface and shared identity/auth conventions.
When you say “three servers”, you’re seeing three different FastAPI apps exposed through Cloudflare under three hostnames.
They are not redundant copies; each has a different responsibility. The confusing part is they also span two different repos
(oc-bot-refactor and oc-bot), which strongly suggests a migration was mid-flight.
| Public hostname | Local port | Running process | What it does (evidence) | Redundancy |
|---|---|---|---|---|
alfredhottakes.combot.alfredhottakes.com |
8887 |
Docker agent-orchestrator (container port 8787) |
Pages + auth + orchestration layer. Docker Compose config pins:
PAGES_ROOT=/workspace/pages, CONTEXT_ROOT=/workspace/context.
This is the “main” Alfred web surface people hit.
|
Not redundant; it’s the portal/orchestrator. |
voice.alfredhottakes.com |
8788 (bound to 127.0.0.1) |
python -m uvicorn malli.server:app |
OpenAPI title: “Alfred Twilio Pipeline”. Exposes Twilio SMS/voice endpoints and Malli persona endpoints
(e.g. /twilio/sms, /twilio/voice/*, /malli/persona, queue/followup runners).
The codebase for this process lives under /Users/techteam/oc-bot (older repo) and loads a hot-reloaded pipeline.json from that directory.
|
Not redundant; it’s the Twilio/voice/SMS gateway. It is architecturally “separate” and currently split from orchestrator. |
api.alfredhottakes.com |
8042 (bound to 0.0.0.0) |
python -m uvicorn main:app |
OpenAPI title: “RV Park Finder”. Endpoints are deal-sourcing specific:
/rv-parks/search, /brokers/search, /ingest/run.
Process cwd is oc-bot-refactor/local/services/rv-park-finder.
|
Not redundant; it’s a separate product microservice. It is optional to Alfred’s “chat/upload” UX but competes for machine resources. |
Alfred’s orchestrator is configured to call two more local services (not currently exposed via Cloudflare in the ingress shown above):
cursor-runner (port 8791): OpenAPI title “Cursor Runner” with /run, /cancel, /health, and a pages route. This is the “hardened runner” the orchestrator prefers (PREFER_CURSOR_RUNNER=1).agent-server (port 8792): OpenAPI title “Agent Server” with conversation + mode routing endpoints (/modes, /pipeline/reload, /run, etc.). This appears to be a fallback/decision-tree server.oc-bot while orchestrator is running from oc-bot-refactor is a classic “mid-migration split-brain” smell. It’s likely Nathan was partway through moving capabilities into the orchestrator pattern.oc-bot vs oc-bot-refactor): fixes/feature additions can land in the “wrong” codebase and never reach the public surface you’re using.403 in ways that make naive “is this route available?” checks unreliable (this directly breaks the current upload page’s auto-detection logic).alfredhottakes.com/p/sitemap returns 200 (public)./api/pages and /api/page-search exist but return 403 without auth (expected; these back the sitemap’s search).403 (gated) or 404 (not implemented) depending on URL + auth state.oc-bot-refactor/cloud/services/agent-orchestrator/main.py, but the running container’s baked /app/main.py does not contain those routes. So from the public site’s perspective, the upload endpoints are “missing”.oc-bot-refactor/local/pages/upload.html) posts to one of /api/upload, /api/uploads, or /uploads and expects list/download/delete helpers. Until the orchestrator container is rebuilt/redeployed (and ACLs aligned), uploads will keep failing.
If you were given a page that posts to /uploads (legacy) or /api/upload (Jarvis-style),
it will fail unless the alfredhottakes.com backend implements that exact endpoint in the deployed container.
Today, it appears the deployed container is behind the source repo, so “not found / error” is expected even if the repo contains the code.
This is the “why Alfred trips over basics” section. The short version: Jarvis is closer to a single, coherent monolith,
while Alfred is an orchestrator + multiple microservices with a legacy sidecar (the Twilio pipeline in oc-bot)
and at least one unrelated product service (RV Park Finder). That’s not inherently bad, but it requires stricter contracts and better observability.
agent-orchestrator (Docker) which then calls helper services (cursor-runner, agent-server), while Twilio voice/SMS is a separate app and RV Park Finder is another separate app.oc-bot and oc-bot-refactor looks accidental/migration-in-progress./Users/techteam/oc-bot (legacy). Orchestrator + runner services live under /Users/techteam/oc-bot-refactor (new).alfredhottakes.com, voice.alfredhottakes.com, api.alfredhottakes.com route to different apps.403 early for non-whitelisted paths, which can happen before the router decides if a route exists.local → /workspace (pages, modes, context) can change immediately with host edits./app inside the container image and only changes after a rebuild/redeploy.Organized by subsystem. Each item is fixable later; we are not executing changes right now.
alfredhottakes.com (Docker orchestrator), voice.alfredhottakes.com (voice Uvicorn), api.alfredhottakes.com (API Uvicorn).cloudflared tunnel ... run processes for his primary tunnel config, plus multiple BlueBubbles-managed cloudflared tunnel --url localhost:1234 processes.403 without an authenticated identity.403 and assume “broken”./p/upload public with strong limits, or generate one-tap auth links that redirect into the upload page./api/pages (currently gated 403 without auth), but /p/sitemap itself is public./api/pages public-read (safe fields only) or embed the page list at build time.POST /api/upload to store arbitrary files and returns saved_to.agent-orchestrator container’s baked /app/main.py does not contain /uploads / /api/upload(s) routes (even though the UI expects them).oc-bot-refactor/cloud/services/agent-orchestrator/main.py does define /uploads / /api/upload(s) routes, which implies the container image is behind source (needs rebuild/redeploy later).403 for unauthenticated/gated paths even when the route doesn’t exist, so the UI “locks onto” a non-existent endpoint.POST /api/upload + GET /api/uploads + download/delete helpers) to the agent-orchestrator app./app/main.py matches the repo version.oc-bot-refactor/cloud/services/agent-orchestrator/docker-compose.yml (+ override).oc-bot-refactor/local → /workspace (so /workspace/pages is host local/pages).cloud/services/agent-orchestrator/data → /app/data (durable app data).oc-bot-refactor/logs → /logsoc-bot-refactor/local/context → /workspace/contextpython-multipart. So the likely root cause is deployment drift (container build behind source) plus route/ACL alignment, not “multipart isn’t installed”.403 before routing for paths that aren’t whitelisted/public, even if they don’t exist./api/capabilities endpoint (public-read) or make the upload UI use a single known endpoint only..env..env, the harder it is to rotate/limit blast radius; it also increases “oops, logged env” risk./p/upload page claims “Up to 2 GB per file, 7-day retention”..docx (extract text, store, index, attach to a job/conversation)./p/sitemap public, but /p/upload (and upload APIs) are not reliably usable without auth + correct whitelisting.upload explicitly public with rate limits + per-link tokens, or provide a dedicated “upload link generator” that creates one-tap authenticated sessions.oc-bot-refactor/local tree into /workspace, so runtime behavior can change with host file edits.POST /api/upload on Alfred’s orchestrator + matching /p/upload UI.upload to PUBLIC_PAGE_SLUGS (rate-limited) or provide one-tap auth redirect to upload./api/capabilities contract).cloudflared ... run processes for the same config).POST /api/upload and the UI posts to that exact endpoint.Next I will: keep expanding this list (still read-only), and attach exact “fix diffs” (what I would change, where) without deploying them until you say go.