Why Ollama matters for cost: At high message volumes, every GPT-4o-mini call costs $0.01–0.03. Ollama on Apple Metal is free. A bot that handles 1,000 messages/day saves ~$10–30/day in API costs vs cloud inference. At scale this is the dominant cost driver.
Apple ID management: Apple occasionally requires re-authentication (2FA prompts, iCloud sign-in renewal). In a headless rack setup, this is handled via Screen Sharing into the specific Mac user account. Plan for ~1 hour/month of operational overhead per 10 bots.
EC2 and iMessage: AWS does offer Mac instances (mac1.metal / mac2.metal) but they require a 24-hour minimum dedicated host allocation at $1.30–1.80/hour = $936–1,296/month minimum. This is never cost-competitive with a physical Mac mini. EC2 Linux is the correct cloud path; Mac hardware belongs in a rack.
Cursor credit burn: Each agent invocation consumes Cursor credits. At high user volumes (many simultaneous conversations), this becomes the dominant cost. The long-term solution is replacing Cursor with a direct Claude/Anthropic API agent — same capability, usage-based pricing instead of seat pricing.