Alfred — Audit Notes (Living)

This page is intentionally “living”: every new discovery + fix gets appended here so Matt always has one canonical reference.

last updated: 2026-04-17 (read-only audit; no Alfred changes)

Comprehensive no-code stabilization checklist (interactive)

This is a comprehensive checklist of actions that improve foundations without changing Alfred’s code. It focuses on visibility, backup/safety, deployment hygiene, workspace legibility, and reducing “unknown unknowns”.

loading…
Export / Import checklist progress
localStorage

Progress + notes are stored in your browser (localStorage). Use export/import to move progress between devices/browsers.

Export
Import

Executive summary

Why Alfred sometimes says things that aren’t true (and “echo back my texts” fails)

The transcript you shared shows multiple different failure modes that look like “memory issues”, but are usually a mix of: message delivery quirks, split services, and the model confabulating when it lacks ground truth.

1) “I replied earlier but you didn’t receive it”

2) “Not yet — you’re Unknown” even for the same number

3) The big one: Alfred “claims it patched / restarted / wired things”

4) “Echo back my last texts” failures

What to do about it (no-code mitigations)

What’s running on Alfred (high signal)

HostnameLocal targetPurposeNotes
alfredhottakes.com
bot.alfredhottakes.com
localhost:8887 Docker container agent-orchestrator (ports 8887 → 8787) /p/sitemap is public (200). Many API endpoints are gated (403) unless authenticated.
voice.alfredhottakes.com localhost:8788 Uvicorn app (malli/server.py) /docs is open (200) but it’s not the pages/upload host.
api.alfredhottakes.com localhost:8042 Another Uvicorn app (main:app) Not serving /p/* pages; appears to be API-only.

Ground-truth snapshot (collected via SSH; no changes made)

Cloudflare tunnel ingress (Alfred)

~/.cloudflared/config.yml
- alfredhottakes.com  → http://localhost:8887
- bot.alfredhottakes.com → http://localhost:8887
- voice.alfredhottakes.com → http://localhost:8788
- api.alfredhottakes.com → http://localhost:8042

Docker Compose source of truth (agent-orchestrator)

compose working dir:
  oc-bot-refactor/cloud/services/agent-orchestrator/
compose files:
  docker-compose.yml
  docker-compose.override.yml
env file:
  oc-bot-refactor/.env

Container mounts (agent-orchestrator)

- oc-bot-refactor/local → /workspace
  (so host local/pages is container /workspace/pages)
- cloud/services/agent-orchestrator/data → /app/data
- oc-bot-refactor/logs → /logs
- oc-bot-refactor/local/context → /workspace/context
- plus campaign + malli mounts

Upload UI file location (host)

oc-bot-refactor/local/pages/upload.html

Workspace architecture map (Alfred) — what everything is + what to do with it

This is the “make the repo readable again” section. Alfred currently has two separate repos on disk: oc-bot-refactor (newer orchestrator architecture) and oc-bot (older/legacy stack). The biggest source of confusion is that production-facing services pull from both.

0A) Cursor vs Finder vs “what’s running” (why you don’t see oc-bot in Cursor)

0) First principle (recommended cleanup goal)

0B) Naming goals (make it “Alfred” without breaking macOS)

1) Repo A: /Users/techteam/oc-bot-refactor (new / canonical)

Size snapshot (approx): dashboard/ ~173MB, local/ ~129MB, logs/ ~17MB. Empty/placeholder files found: ~75 (most are virtualenv artifacts under local/services/.venv).

Why there are “loose files” at the repo root (and what to do)
Common root files you’ll see in Cursor (oc-bot-refactor):
- .env / .env.example      env config for services (keep; make sure secrets aren’t committed)
- README.md / CONTEXT.md / MALLI_ROADMAP.md  documentation (keep)
- package.json / package-lock.json           dashboard deps (keep)
- tools.json / mcp-servers.json / memory.json  agent/tool registry + memory (keep; consider consolidating docs)

If a root file is empty or mysterious:
- treat it as suspect until you know who reads it (search references before deleting)
- prefer moving runtime state into local/context or service data dirs, not root.
PathWhat it isActively used?RecommendationWhy / how (later)
cloud/services/agent-orchestrator/ Main Alfred portal/orchestrator FastAPI service (Docker image build context). Yes (serves alfredhottakes.com). KEEP (core). Make this the single owner of pages + uploads + auth. When we “fix Alfred”, most changes should land here.
cloud/services/agent-orchestrator/main.py The orchestrator app entrypoint (huge; owns routes, auth middleware, pipeline wiring). Yes (source-of-truth), but deployed container may be behind. CHANGE (deploy hygiene). Ensure the running container’s /app/main.py matches this file (rebuild/redeploy later). Add a version stamp endpoint to detect drift.
cloud/services/agent-orchestrator/pipeline/ Core pipeline modules: auth, identity, history, send, tools, tasks, etc. Yes. KEEP. This is the “brains” of the orchestrator. Prefer refactors here over spawning new services.
cloud/services/agent-orchestrator/portal/ HTML portal UI (login/dashboard pages served by orchestrator). Yes. KEEP. Central place for “operator UX”. If pages feel confusing, improve these instead of adding more ad-hoc pages.
cloud/services/agent-orchestrator/data/ Stateful storage: identity sqlite/dbs, visitors, prompt audits, processed GUIDs. Yes. LEAVE ALONE (but document). Do not delete casually. Later: make a clear backup/rotation policy; consider moving to a dedicated state directory outside the repo tree.
local/ Host-mounted workspace data + pages + modes + local services. Yes (mounted into Docker as /workspace). KEEP (but organize). Think of local/ as “runtime workspace”, not product code. Later: tighten which subfolders are “state” vs “config” vs “UX pages”.
local/pages/ Public/operator HTML pages (includes upload.html). Yes. CHANGE (align contracts). Keep pages, but remove brittle endpoint auto-detection and align to the deployed API surface.
local/context/ Stateful conversations + attachments + runner progress. Yes (core bot memory). LEAVE ALONE (but fence it). Do not delete casually. Later: ensure it’s gitignored + backed up; optionally relocate to a dedicated data volume outside repo.
local/services/cursor-runner/ “Hardened runner” service (runs agent jobs; orchestrator prefers it). Yes (called by orchestrator). KEEP. Good separation: runner can be sandboxed and scaled later.
local/services/agent-server/ Decision-tree / fallback agent service. Yes (configured fallback). KEEP (but consider deprecating later). If cursor-runner is stable, you may phase this out to reduce moving parts.
local/services/rv-park-finder/ Standalone “RV Park Finder” FastAPI microservice (port 8042). Yes (serves api.alfredhottakes.com). MOVE / ISOLATE. It’s not part of Alfred’s basic UX; it can consume resources. Later: move to cloud or separate host; keep the API stable.
local/services/.venv/ Checked-in local Python virtualenv for host-run services. No (not “source of truth”; it’s generated). DELETE (cleanup) once we confirm nothing is hard-pinned to it. This is a major repo-noise source and explains many “empty files”. Later: remove the venv dir, ensure it’s ignored, and reinstall from requirements.txt.
dashboard/ Next.js dashboard frontend. Unknown (exists in compose, but not mapped in current Cloudflare ingress). KEEP (for now). Later: either properly publish it under a hostname or remove it from the default “prod” compose to reduce noise.
whatsapp-bridge/ WhatsApp Web bridge service (Baileys). Unknown (present in compose; not confirmed live). KEEP (for now). Later: only run in environments where WhatsApp is actively needed; otherwise disable to reduce surface area.
logs/ Runtime logs directory mounted into containers. Yes. KEEP (but rotate). Add log rotation + size caps; don’t keep unbounded logs inside the repo forever.
.venv-*, .pytest_cache/, __pycache__/, tmp/, graphify-out/ Local caches/build artifacts. No (not required for runtime correctness). DELETE (safe cleanup). Later: remove locally + ensure ignored. This is the fastest “workspace sanity” win with minimal risk.

2) Repo B: /Users/techteam/oc-bot (legacy / still powering Twilio)

Size snapshot (approx): imessage-bridge/ ~888MB, installer/ ~147MB, vercel/ ~156MB, malli/ ~35MB. Empty/placeholder files found: ~433 (overwhelmingly in imessage-bridge/.venv).

PathWhat it isActively used?RecommendationWhy / how (later)
malli/server.py Twilio voice/SMS gateway + Malli persona/queue endpoints. Yes (serves voice.alfredhottakes.com). KEEP (temporarily). This is “live legacy”. Later: migrate the Twilio surface into oc-bot-refactor (or at least unify identity/auth + upload conventions).
imessage-bridge/, messages-bridge/ Older bridge stacks (iMessage/other messaging surfaces). Unknown (not confirmed in current Cloudflare ingress). ARCHIVE / DEPRECATE (if unused). Later: confirm whether any Alfred surfaces still depend on them; if not, move to an archive/ folder or separate repo to reduce cognitive load.
imessage-bridge/.venv/ Committed Python virtualenv inside the legacy repo. No (generated). DELETE (cleanup) once confirmed unused. This is the single biggest “clutter amplifier” (hundreds of empty REQUESTED/py.typed marker files). Replace with requirements.txt + recreate on demand.
.pycache/ Unusual cache tree that mirrors absolute paths (e.g. .pycache/Users/...). No (generated). DELETE (safe cleanup). It’s almost certainly a cache directory produced by a nonstandard pycache prefix config. Later: delete and add ignore rules; check env like PYTHONPYCACHEPREFIX.
vercel/ Frontend build/deploy artifacts. Unknown (likely generated). DELETE (cleanup) if not actively used. If nothing is serving from it, it’s pure noise. Later: delete and ensure build output dirs are ignored.
node_modules/, dist/, .pnpm-store/ JS build artifacts/deps. Maybe (depends on whether legacy frontends are used). DO NOT TOUCH blindly. Later: if legacy web frontend isn’t used, remove + regenerate on demand; otherwise keep but ensure it’s not tracked in git.
pages/ Legacy HTML pages for the old stack. Unknown. ARCHIVE (if unused). Later: extract any still-valuable content into oc-bot-refactor/local/pages and deprecate the rest.
database/, data/, agent_data/, logs/ Legacy state directories. Possibly. LEAVE ALONE until proven safe. Later: identify which state is still referenced by running processes; migrate it into the new state layout if needed.

3) “How do we clean this up without breaking progress?” (recommended approach)

  1. Inventory what is actually live: for each public hostname, record the PID/container + working directory + git commit hash (where possible).
  2. Mark state directories as sacred: local/context, cloud/services/agent-orchestrator/data, legacy oc-bot state dirs. Back up before any reorg.
  3. Delete only safe artifacts first: caches, tmp/, __pycache__/, stray venvs, build outputs — the stuff that never should have been in the “mental model”.
  4. Then consolidate repo split-brain: migrate Twilio gateway into oc-bot-refactor (or at minimum unify config/auth and keep it as a clearly named “gateway” service).

4) How to merge toward “one optimized Alfred” (without intense architectural changes yet)

  1. Define the canonical “contracts” in oc-bot-refactor: uploads API, where files land, how pages call APIs, how auth works, where context lives.
  2. Bring Twilio/voice into the canonical repo in the least disruptive way: either move oc-bot/malli/server.py into oc-bot-refactor as a dedicated service, or keep it as a separate service but with a versioned interface and shared identity/auth conventions.
  3. Isolate unrelated product microservices (e.g. RV Park Finder) so Alfred’s core UX doesn’t degrade when product jobs spike.
  4. Only then consider deeper refactors (single-hostname “front door”, collapsing services, cloud migration, etc.).

What are the “three servers” on Alfred? (purpose + redundancy assessment)

When you say “three servers”, you’re seeing three different FastAPI apps exposed through Cloudflare under three hostnames. They are not redundant copies; each has a different responsibility. The confusing part is they also span two different repos (oc-bot-refactor and oc-bot), which strongly suggests a migration was mid-flight.

A) Public surfaces (Cloudflare hostnames → local ports)

Public hostnameLocal portRunning processWhat it does (evidence)Redundancy
alfredhottakes.com
bot.alfredhottakes.com
8887 Docker agent-orchestrator (container port 8787) Pages + auth + orchestration layer. Docker Compose config pins: PAGES_ROOT=/workspace/pages, CONTEXT_ROOT=/workspace/context. This is the “main” Alfred web surface people hit. Not redundant; it’s the portal/orchestrator.
voice.alfredhottakes.com 8788 (bound to 127.0.0.1) python -m uvicorn malli.server:app OpenAPI title: “Alfred Twilio Pipeline”. Exposes Twilio SMS/voice endpoints and Malli persona endpoints (e.g. /twilio/sms, /twilio/voice/*, /malli/persona, queue/followup runners). The codebase for this process lives under /Users/techteam/oc-bot (older repo) and loads a hot-reloaded pipeline.json from that directory. Not redundant; it’s the Twilio/voice/SMS gateway. It is architecturally “separate” and currently split from orchestrator.
api.alfredhottakes.com 8042 (bound to 0.0.0.0) python -m uvicorn main:app OpenAPI title: “RV Park Finder”. Endpoints are deal-sourcing specific: /rv-parks/search, /brokers/search, /ingest/run. Process cwd is oc-bot-refactor/local/services/rv-park-finder. Not redundant; it’s a separate product microservice. It is optional to Alfred’s “chat/upload” UX but competes for machine resources.

B) Internal helper services (called by orchestrator)

Alfred’s orchestrator is configured to call two more local services (not currently exposed via Cloudflare in the ingress shown above):

C) So… is having 3 public apps “necessary”?

D) Why this topology can make Alfred feel “lobotomized” or flaky

Why your upload UI is failing

Observed behavior

Practical implication

If you were given a page that posts to /uploads (legacy) or /api/upload (Jarvis-style), it will fail unless the alfredhottakes.com backend implements that exact endpoint in the deployed container. Today, it appears the deployed container is behind the source repo, so “not found / error” is expected even if the repo contains the code.

Core architectural differences: Jarvis vs Alfred (what’s intentional vs accidental)

This is the “why Alfred trips over basics” section. The short version: Jarvis is closer to a single, coherent monolith, while Alfred is an orchestrator + multiple microservices with a legacy sidecar (the Twilio pipeline in oc-bot) and at least one unrelated product service (RV Park Finder). That’s not inherently bad, but it requires stricter contracts and better observability.

1) Monolith vs orchestrator graph

2) Repo split-brain (highest-signal “testing grounds” smell)

3) Hostname/port split (and cookie/auth implications)

4) Security posture differences (403-before-routing)

5) Mutable runtime vs release-driven runtime

6) Resource contention / unrelated workloads

Comprehensive architecture differences + redundancies + fixable issues

Organized by subsystem. Each item is fixable later; we are not executing changes right now.

1) Service topology (single-server vs multi-surface)

2) Cloudflare Tunnel routing complexity

3) Auth model and operational friction

4) Pages system and page index

5) Upload pipeline gap (root cause of DOCX pain)

6) Container/volume layout (where state actually lives)

7) Dependency/runtime drift

8) Route existence vs auth ambiguity

9) Redundancy and failure modes

10) Secrets + configuration sprawl

11) Upload UX claims vs infra realities

12) Storage + ingestion are not defined for DOCX

13) Public pages allowlist mismatch

14) Observability gaps

15) Coupling between projects (oc-bot-refactor vs agent-orchestrator)

16) Security hardening opportunities

Fix queue (doable later; not executed now)

  1. Standardize uploads: implement POST /api/upload on Alfred’s orchestrator + matching /p/upload UI.
  2. Make it usable: either add upload to PUBLIC_PAGE_SLUGS (rate-limited) or provide one-tap auth redirect to upload.
  3. Remove upload endpoint auto-detect: make the UI call only the canonical endpoint (or consult a /api/capabilities contract).
  4. Unify “front door”: reduce split surfaces by proxying pages+uploads behind one hostname.
  5. Observability: add a status page that prints hostname→port mapping + last errors + upload destinations.
  6. Runtime hardening: prevent restarts from switching Python/env unexpectedly; verify multipart/upload dependencies at startup.
  7. Tunnel process hygiene: converge to one managed tunnel runner (avoid multiple simultaneous cloudflared ... run processes for the same config).
  8. Large-file readiness: implement size limits + chunking (or direct-to-object-storage uploads) before promising multi-GB uploads publicly.

Notes on “why Jarvis works better” (early signals)

Next update slot

Next I will: keep expanding this list (still read-only), and attach exact “fix diffs” (what I would change, where) without deploying them until you say go.