How to build it
Build one loop first.
The safest way to copy the system is not to automate everything. Pick one painful workflow, make it visible, verify it, then widen the pattern.
Phase plan
Build sequence.
- Phase 1: Inventory. List inputs, systems of record, recurring reports, approval points, and painful handoffs.
- Phase 2: Choose one workflow. Good first choices are leasing follow-up, daily maintenance triage, owner brief prep, or delinquency review.
- Phase 3: Run dry. Let the system classify and draft, but do not let it send, dispatch, pay, delete, or change official records.
- Phase 4: Add receipts. Every run should leave proof: source data, draft, decision, approval, final action, and log.
- Phase 5: Add the scheduler. Put recurring judgment work on an agent schedule. Put boring syncs and health checks in system services.
- Phase 6: Add memory. Write lessons into the playbook so the next run uses the current truth instead of old assumptions.
- Phase 7: Expand lanes. Add the next workflow only after the first one is boring, visible, and easy to turn off.
Stack build order
Put the pieces in this order.
- Models
- Pick model roles first: OpenAI/Codex-style coding support, Claude-style reasoning, fast classification, and retrieval/search.
- Control
- Add a Hermes-style control plane: agents, tools, gateway, dashboard, approvals, logs, and memory startup.
- Records
- Connect systems of record: AppFolio or your property platform, Google Sheets/Docs, accounting exports, and a task manager if you use one.
- Comms
- Add communication APIs: Telegram for approvals, Gmail for inbox signals, ElevenLabs/Twilio for voice/SMS intake, and business-text reminders if needed.
- Schedules
- Use cron for recurring agent judgment and LaunchAgents/system services for background polling, syncs, watchdogs, backups, and drains.
- Proof
- Use browser automation, logs, receipts, screenshots, returned records, and dashboards to prove each run worked.
Templates
Copy these worksheets.
Workflow card
Input: Where does this start?
Owner: Which lane owns the next step?
Context: What data is needed?
Draft: What should the system prepare?
Verify: What proof says it worked?
Approve: Who must say yes?
Log: Where does the result live?
Risk card
External impact: Could this message or action affect someone else?
Money impact: Could this move cash, invoices, charges, or owner reporting?
Legal impact: Could this affect notices, lease status, collections, or compliance?
Data impact: Could this expose private records?
Rollback: How do we turn it off?
Minimum viable version
What to build first.
Start with one dashboard, one lane, one recurring dry-run job, one approval rule, and one memory note. That is enough to prove the pattern without creating a fragile automation maze.
Model layer
AI models are split by job.
| Model/API category | What it does here | How another team should copy it |
|---|---|---|
| OpenAI / Codex | Coding, implementation, site edits, local verification, browser checks, and structured tool work. | Use a coding-oriented model for repository edits, scripts, tests, deployments, and repeatable tooling. |
| Claude / Claude CLI | Long-form planning, architecture review, audits, policy reasoning, and operator-facing analysis. | Use a strong reasoning model for ambiguous business logic, audits, and complex operating decisions. |
| Fast classification models | Message triage, extraction, routing, summarization, and repeated low-risk ops checks. | Use cheaper or faster models for high-frequency classification and extraction after the rules are clear. |
| Retrieval/search layer | Finds prior decisions, runbooks, tenant-safe context, and historical lessons before a workflow acts. | Build search over docs and logs before adding automation. The system should remember what it already learned. |
Control layer
Hermes is the operator hub.
Gateway
One always-on local gateway that connects all profiles, tools, Telegram lanes, and scheduled jobs.
Agents
Nine lanes: the conductor (main/control), work, Lyra (leasing and resident intake), finance, collections, ops, memory, maintenance, and a Grok-powered tweeter lane.
Tools
Narrow access to shell, files, browser automation, Google tools, Vercel deploys, GitHub, local scripts, and app-specific commands.
Approvals
Human gates sit in Telegram and the operator loop. Work orders, sensitive messages, legal/financial actions, and destructive edits stay approval-owned.
APIs and software
What sits around Hermes.
| Software/API | Role in the system | Copy pattern |
|---|---|---|
| Telegram Bot API | Primary notification, approval, alert, and handoff channel across specialist lanes. | Use one visible human channel first. Every risky workflow should have a clear approval surface. |
| Google Workspace APIs | Gmail signals, Sheets queues, Docs/Drive references, and calendar-style context. | Use Workspace as the first queue and source-of-truth surface if your team already lives there. |
| AppFolio | Property system of record for residents, work orders, vendors, occupancy, delinquency, and reports. | Connect to the portfolio system of record. Read first, draft second, approve before writes. |
| ElevenLabs | Lyra voice AI: inbound call handling, intent collection, knowledge-base answers, and call analysis. | Start with inbound-only voice intake. Let it collect context and hand off before it promises outcomes. |
| Twilio | Phone number plus live two-way texting around customer communication. | Keep compliance and opt-in requirements outside the model. Let the phone/SMS layer enforce them. |
| iMessage | Business-text digests and urgent-message scans. | Keep specialized message scanners separate from the main agent loop so they can fail independently. |
| Kanban handoffs | Hermes's built-in board for passing work between lanes. | Give every handoff an owner, so work never falls between lanes. |
| Cloudflare Workers | Thin webhook/edge API pattern used for worker-style glue and public-facing handoffs when useful. | Put small public API surfaces at the edge; keep private decision logic and secrets out of the public docs. |
| Vercel | Hosts this public explanation site. | Use a public site to explain the pattern, not to expose private machinery. |
| Browser automation | Used where a system lacks a clean API or visual verification is necessary. | Prefer APIs first. Use browser automation for legacy portals and proof-gathering. |
| Node.js + Python | Local scripts, sync jobs, parsers, browser tooling, reports, and small service glue. | Keep scripts small, named, logged, and easy to run by hand before scheduling them. |
Schedules
Cron and LaunchAgents do different jobs.
Hermes scheduled jobs
Recurring agent-owned jobs: ops daily digest, occupancy, tenant directory, lease renewals, daily rent roll, weekly delinquency and collections snapshot, work-order history, unit-turn digest, business-text digests and urgent scans, Lyra QA and improvement drafts, wiki synthesis/curation/verification, daily logs, weekly reviews, backups, and a morning newsletter.
macOS LaunchAgents
Only a handful: the gateway, the 5-minute email pipeline, the daily AppFolio pull, the weekly Lyra KB sync, a Telegram delivery re-sender, and the wiki file watcher. Everything else moved into Hermes scheduled jobs.
The rule is simple: if the job needs judgment, put it in the agent scheduler. If it mostly checks, syncs, drains, watches, or keeps something alive, make it a boring service.
Current job families
What the recurring work covers.
The real setup has 32 scheduled jobs. Publicly, they are best understood as job families instead of exact private schedules.
| Job family | What it does | Why it exists |
|---|---|---|
| Wiki synthesis + curation | Updates the shared planning wiki and keeps recent decisions searchable. | Stops the agents from working from stale notes. |
| Lyra QA (twice daily) | Reviews customer-service calls and texts and flags quality or routing issues. | Keeps the voice agent useful and accountable. |
| Daily log writer | Writes the day into a durable record. | Makes tomorrow's work smarter than today's memory. |
| Weekly failure review | Looks for repeated failures and turns them into rules or fixes. | Prevents the same mistake from repeating quietly. |
| Property sync | Pulls property-system data like work orders, occupancy, tenants, and updates. | Keeps dashboards and follow-ups based on current records. |
| Work-order sheets | The daily AppFolio pull turns work-order changes into per-vendor sheets. | Turns maintenance status into visible follow-up. |
| Ops daily brief | Summarizes property operations and open work. | Gives the operator a morning control panel. |
| Occupancy report | Checks which units or assets are occupied, vacant, or changing. | Supports leasing and owner visibility. |
| Weekly collections snapshot | Summarizes late-rent totals and how fresh the data is. | Turns accounting data into reviewable insight. |
| Weekly delinquency + daily rent roll | Tracks past-due balances and follow-up. | Keeps money-risk work from slipping. |
| Business-text digests and urgent scans | Digests business texts three times a day and scans for urgent ones every 15 minutes in the daytime. | Keeps communication follow-up moving. |
| Lease renewal pipeline | Tracks renewals and next-step follow-up. | Turns lease timing into tasks before it becomes urgent. |
| Tenant directory refresh | Refreshes tenant/contact context. | Helps agents find the right person and property context. |
| Work-order history sync | Maintains longer work-order history. | Supports reporting, vendor review, and pattern detection. |
Always-on families
What LaunchAgents cover.
Six Mac background services keep the always-on pieces running.
Gateway
Keep the single Hermes gateway alive so every profile, tool, and Telegram lane can talk to each other.
Inbox
Poll Gmail every 5 minutes and run the leasing drafter so communication inputs keep moving into the operating loop.
Property and Lyra sync
Run the daily AppFolio pull and the weekly Lyra knowledge-base sync.
Reliability
Re-send failed Telegram deliveries and watch the wiki for changes. The other watchdogs (credentials, browser profile, backups, missed jobs) now run as Hermes scheduled jobs.
Lyra
Customer service is a workflow, not just a voice bot.
- Call
- A prospect, resident, vendor, or other caller reaches the intake line.
- Voice AI
- ElevenLabs gives Lyra a natural conversation layer. Lyra asks questions and identifies intent.
- Tools
- Scoped backend tools check availability, property info, tenant lookup, tenant verification, emergency alerts, caller history, tour requests, and warm transfer; the call log is written automatically after the call.
- Knowledge
- Maintained property notes and customer-service rules tell Lyra what she may say and what she must hand off.
- Queue
- The call becomes a summary, queue item, alert, or follow-up request.
- Human
- Maintenance dispatch, legal/payment/account issues, emergencies, uncertain answers, and sensitive follow-up stay human-owned.
Parts list
The pieces, in plain English.
| Piece | Simple meaning | Portfolio example | How to build it |
|---|---|---|---|
| Input layer | All the places work arrives. | Tenant call, leasing email, owner question, vendor invoice, delinquency report. | Make one list of sources. For each source, name the owner and what proof should be captured. |
| Router | The traffic controller. | A message gets classified as leasing, maintenance, finance, collections, or review. | Start with simple rules. Later add AI classification once the rules are clear. |
| Operating lane | A named owner for a kind of work. | Leasing lane, work-order lane, finance lane, collections lane, executive-review lane. | Give each lane allowed tools, blocked actions, escalation rules, and a dashboard view. |
| Agent job | A recurring task that needs judgment. | Summarize delinquency risk, prepare owner brief, review leasing follow-ups. | Run it on a schedule, but require a receipt and human review before action. |
| System service | A boring background job. | Sync tasks, check inbox, rotate logs, verify backup, detect config drift. | Move repetitive checks outside the reasoning layer so the AI is used for judgment, not plumbing. |
| Knowledge base | The written source of truth. | Property notes, leasing policies, emergency rules, vendor instructions, team runbooks. | Keep it in simple files or docs. Sync it into the tools that need it. |
| Memory | The system's durable lessons. | “This vendor needs photos before dispatch” or “this report must reconcile to owner statement.” | After every important run, write the lesson in a place future runs can search. |
| Tool connection | A narrow doorway into another system. | Property software, CRM, accounting, calendar, inbox, messaging, task manager. | Give each connection the least power it needs. Avoid broad admin access when read-only is enough. |
| Approval gate | The stop sign before impact. | Work order dispatch, tenant message, owner update, payment action, legal escalation. | Define the exact approval phrase or button. Log who approved and what was approved. |
| Verification receipt | Proof that a run did what it claimed. | Returned record count, screenshot, sync receipt, sent-message ID, dashboard link. | Do not call a workflow done until it leaves proof a person can inspect. |
| Dashboard | A simple view of current state. | Open approvals, failed jobs, leasing follow-ups, unresolved work orders, cash-risk items. | Start with one page or table. Show what changed, what is blocked, and what needs a decision. |
| Rollback path | The way back if something goes wrong. | Disable a new automation and return to manual review. | Before turning a workflow on, write down how to turn it off. |
Rule of thumb
Use AI for judgment, not blind action.
Good AI work
Classifying messages, summarizing context, drafting replies, comparing reports, finding exceptions, preparing decisions, and explaining what changed.
Keep human-owned
Sending sensitive messages, dispatching costly work, approving payments, changing legal status, deleting data, publishing externally, or overriding policy.
Battle-tested
The prompts I actually use.
These are the ones that ship — pulled directly from my Claude behavioral file, the wiki schema, and the scheduled jobs that run every day. Copy what fits, adjust for your work.
Operator behavior
Make Claude (or Codex) actually useful.
This is the pattern I put at the top of any session that involves building something. It changes the AI from “eager assistant” to “careful collaborator.”
Before you write code or take action:
1. State your assumptions explicitly. If uncertain, ask.
2. If multiple interpretations exist, present them — don't pick silently.
3. For non-trivial work (3+ steps or architectural decisions),
propose an approach and wait for confirmation before editing.
4. If a fix feels hacky, pause and ask "is there a more elegant way?"
5. If something goes sideways, STOP and re-plan — don't keep pushing.
Style:
- Terse. State results directly. No trailing summaries of what
you just read or did.
- Define success criteria. Verify before claiming done.
- After edits to runnable code: run tests, type-check, or lint
as appropriate. Don't claim done based on "it looks right."
- Reconcile against LIVE state (files, processes, deployed code),
not narrative reports.
- On correction passes: retract plainly, don't rationalize.
Don't add features, refactor, or introduce abstractions beyond
what the task requires. Three similar lines is better than a
premature abstraction.Two-AI review
When you want a real second opinion.
This is the prompt I use when Claude has just produced something I'm about to ship — before sending it through Codex for review (or vice versa).
Read the diff/output above. Don't validate it — review it.
Specifically:
- Are there assumptions baked in that could be wrong?
- Is there an edge case the code doesn't handle?
- Is there a simpler way that I'm missing?
- Is anything overcomplicated for the actual task?
- If you were reviewing this for a senior engineer to ship,
what would you push back on?
Be specific. Quote line numbers. If you'd ship as-is, say
so plainly. Don't soften your critique to be polite.Weekly review
The Friday afternoon pattern.
My memory agent runs this every Friday at 5 PM. The result lands in my Telegram thread and gets appended to the planning wiki.
Pull the last 7 days of:
- decisions made (from the decisions log)
- things that broke (from the ops Telegram channel)
- cron failures (from launchd logs)
- any new project pages or playbooks
For each one, answer:
1. What was the actual outcome? (vs. the intended outcome)
2. Is there a pattern with previous weeks?
3. What rule or playbook should be updated as a result?
4. What's worth telling future-me even if it seems obvious now?
Be honest about what didn't work. Don't summarize for the
sake of summarizing — if a week was uneventful, say so in
one line.
Append the output to wiki/log.md with this week's date.Wiki ingest
Turn raw notes into durable knowledge.
When I dump a chunk of session transcripts or rough notes into my wiki's raw/ folder, this prompt processes them into structured project pages.
Read all new files in raw/ that I haven't ingested yet.
For each one, extract:
- Decisions made (date them, link to source line)
- Risks identified (separate from decisions)
- Action items with explicit owners
- Facts worth remembering (gotchas, working recipes, gotchas about gotchas)
Then for each affected project page in wiki/projects/:
- Update Current State if it changed
- Append to Decisions section (don't overwrite history)
- Add new risks to Risks and Blockers
- Update Next Checkpoint
Rules:
- Never delete historical decisions unless I tell you to.
- If new content contradicts old content, KEEP BOTH and
mark the contradiction. Add "superseded by" if applicable.
- Mark uncertain claims as "Status: tentative".
- Cite source files for every non-trivial claim.
When done, append a timestamped entry to wiki/log.md
summarizing what changed.Pre-commit safety
Keep the AI from doing something destructive.
This is the explicit guardrail block I include in any session that has shell or git access. Pulled directly from my CLAUDE.md.
Destructive Action Guardrails:
- Never `rm -rf`, `git reset --hard`, force-push, drop tables,
or kill shared processes without explicit confirmation.
- Never skip hooks (--no-verify), bypass signing, or disable
safety checks as a shortcut.
- Never commit .env, credentials, or token-bearing files.
- If state looks unfamiliar (stray files, unknown branches),
investigate — don't delete.
- Before running destructive operations, consider whether there
is a safer alternative. Only use destructive operations when
they are truly the best approach.
- Approval in one context doesn't extend to the next. Confirm
before each new destructive action.Project planning
Get to a real plan before you write a line of code.
For anything bigger than a 30-minute task, I use this to force Claude into planning mode first.
Don't write code yet. Help me plan.
The task: [DESCRIBE IT]
Walk me through:
1. What this actually needs to do (re-state in your own words)
2. The 3-5 main steps, in order
3. Each step's:
- Files that need to change
- What can go wrong
- How to verify it worked
4. Risks I'm not thinking about
5. Questions you have for me before starting
If the task feels under-specified, ask me clarifying
questions instead of guessing.
After I confirm the plan, you can start.Lyra-style intake
The system prompt for a voice agent.
A simplified version of what runs Lyra. Adapt the property-specific bits for your domain.
You are [NAME], an inbound assistant for [BUSINESS].
Your job:
- Answer the call warmly. Identify yourself once.
- Understand what the caller actually needs (not what they
first say — listen for intent).
- Use the tools available to look up real information.
- For anything sensitive (account data, money, legal),
verify the caller's identity before sharing details.
- For anything urgent (emergency, safety), use the
emergency_alert tool immediately, then keep talking
to the caller.
- For anything you can't resolve, get the details and
promise a callback. Don't make commitments on behalf
of the business.
What you DON'T do:
- Make pricing exceptions.
- Promise outcomes you can't deliver.
- Send messages or take actions the user can't see.
- Continue trying to "solve" when the caller wants a human.
If the caller is frustrated, acknowledge it directly. Don't
keep offering options when they want escalation.
Always log the call summary at the end with: who, what,
intent, urgency, what I told them, what's outstanding.Original prompts (still useful)
Generic build prompts.
The original generic prompts from earlier versions of this playbook — useful for setting up a new system from scratch.
You are helping me build a local-first workflow operating system for a portfolio management team.
Explain everything in plain language.
The system should have:
- AI model roles for coding, reasoning, classification, retrieval, and review.
- A local control plane like Hermes with agents, tools, approvals, logs, and memory.
- Operating lanes for customer intake, leasing, maintenance, finance, collections, reliability, and memory.
- Software/API connections for our systems of record, inboxes, task tools, voice/SMS tools, dashboards, and documentation.
- Cron-style recurring jobs for agent-owned judgment work.
- LaunchAgent/system-service-style jobs for background polling, syncs, backups, watchdogs, and cleanup.
- Human approval gates before any send, dispatch, payment, legal step, official-record change, deletion, or public action.
- Receipts and logs for every run.
Do not design silent automation.
Start with one workflow in dry-run mode.
Ask me questions until you can produce a build plan, a risk list, and the first workflow card.Discovery
Map the current business.
Input inventory
Prompt: Interview me about every place work enters our portfolio management team. Build a table with: source, example message, owner, current system of record, urgency, privacy risk, and what proof would show the item was handled.
Workflow pain map
Prompt: Help me find the 10 recurring workflows that waste the most time or create the most missed follow-up risk. Rank them by business value, automation risk, data availability, and ease of starting in dry-run mode.
Lane design
Prompt: Turn our workflows into operating lanes. For each lane, define purpose, owner, allowed tools, blocked actions, approval rules, dashboard needs, and first three recurring checks.
Software map
Prompt: Create a public-safe software map for our workflow OS. Separate systems of record, communication tools, task tools, AI/model tools, schedulers, dashboards, and documentation/memory.
Stack prompts
Ask for the real pieces.
Model plan
Prompt: Design our model layer using separate roles for coding, long-form reasoning, fast classification, extraction, retrieval/search, and final review. Explain where OpenAI/Codex-style tools and Claude-style tools fit.
Hermes-style control plane
Prompt: Design a Hermes-style control plane for our team. Include agents, gateway/router, tools, approvals, Telegram or Slack handoff, dashboard, logs, memory, scheduled jobs, and background services.
API/software matrix
Prompt: Build a matrix for Google Workspace, AppFolio or our property system, ElevenLabs or voice AI, Twilio or SMS, a task manager, a wiki, Telegram or approval channel, Vercel or public docs, and Cloudflare Workers or edge API glue.
Schedule split
Prompt: Separate this workflow into scheduled jobs, LaunchAgents/system services, manual approvals, and dashboards. Explain why each piece belongs in that layer.
Build prompts
Design the first loop.
- Dry run
Prompt:Design a dry-run version of this workflow. It may read data, classify, summarize, and draft next steps, but it may not send messages, dispatch work, change records, move money, or delete anything.- Receipts
Prompt:Define the receipt this workflow must produce. Include source records checked, decision made, draft output, approval status, final action, and where the log should live.- Approval
Prompt:Write a human approval policy for this workflow. Specify what needs approval, who can approve, what phrase or button counts as approval, and what should happen if approval is missing.- Schedule
Prompt:Decide whether this workflow belongs in an agent schedule or a background service. If it needs judgment, keep it agent-owned. If it mostly checks or syncs, make it a boring service.- Memory
Prompt:After this run, write a short memory note: what changed, what was proven, what failed, what should be checked next time, and what rule should be updated.
API prompts
Connect software without making a mess.
API boundary
Prompt: For this software connection, define the minimum API access needed. Separate read-only lookup, draft creation, queue writing, official record changes, sends, and deletes. Recommend the safest starting permission.
System of record
Prompt: Identify the system of record for this workflow. Tell me which fields should be read from it, which should never be overwritten automatically, and what human approval is required before any write.
Service choice
Prompt: Decide whether this integration should use a direct API, browser automation, spreadsheet queue, webhook, or manual upload. Compare reliability, privacy risk, setup cost, and auditability.
Failure handling
Prompt: Design failure handling for this API. Include timeout behavior, retry limits, stale-data warnings, fallback mode, human alert, and the receipt that proves no external action was taken.
Lyra-style prompt set
Build customer-service intake.
Voice intake scope
Prompt: Design an inbound customer-service agent for our portfolio team. It can answer common questions, collect context, classify intent, and create a handoff summary. It cannot make final promises, send sensitive account details, approve work, or make legal/financial decisions.
Handoff summary
Prompt: Create the perfect human handoff summary for a customer call. Include caller, property/account, intent, urgency, facts collected, missing information, recommended next step, risk level, and whether approval is required.
Knowledge base
Prompt: Build the first knowledge base outline for customer-service intake. Include property facts, policies, emergency rules, leasing answers, maintenance categories, escalation rules, and what the agent must never answer alone.
Safety review
Prompt: Review this customer-service workflow for risks: privacy, wrong promises, legal exposure, payment/account sensitivity, emergency handling, unclear ownership, and missing logs. Return fixes before launch.