Internal SOP and colleague playbook
Seven stages, from the first call to a self-improving site. Evidenced by two real builds. Every stage carries the prompt to run, the skills to run it with, and a receipt you can open.
Those six numbers are counted from this page as it loads. None of them is typed. The last one is counted from the actual file payload carried inside this document, so it cannot claim a skill that is not really here.
Precedence. Follows clients/_shared/templates/WEBSITE-LADDER.md, which is canonical and says so itself. If this file and the ladder disagree, the ladder wins.
Version. v1.0 · 2026-08-02
Found a drift between this and what you actually do? Run bd create "playbook drift: <what>" -p 2, or leave a note in docs/output/website-build-playbook/. Field practice outranks this document.
The honest cost, up front. This method is front-loaded. Stage 1 and stage 2 happen before the client sees a single pixel of design, and to a client used to being shown mockups in week one, that looks slow. It is slower, at the start.
What it buys: nothing gets built twice. The source material is written before the HTML, so the design argument happens against recorded facts instead of taste. By the time a direction is picked, everything downstream is assembly.
You do not need our accounts. You need a laptop, Claude Code, and a way to record one call. The rest of this page is the steps, in order, with the words to type.
Here is the whole idea. The website gets built out of what the owner already told you. Not out of what you guessed they wanted.
Most website projects start with a design. Someone makes a pretty page. Then everyone argues about it. Nobody can say who is right, because it is all taste.
This way starts with a call. You record it. Then you build one page. It shows them their own words, their own logo, their own service list, and a date. Now the argument is about facts. Facts are easy to settle.
Record the call. Do not rely on your notes. Your notes are your summary of what they said. A transcript is what they said.
This matters more than it sounds. Every quote we put on a client page links back to the recording. The owner clicks it and hears themselves say it. You cannot do that from memory. One 38 minute call gave the Juan and Only Plumber build all 11 of its quotes.
| Tool | What it does for you | Good to know |
|---|---|---|
| Fathom | Joins the call, records it, and writes the transcript by itself | This is what we use. Both builds on this page started here |
| Wispr Flow | Turns your talking into typed text | Best for dumping your own thoughts right after the call |
| Otter or Granola | Records and transcribes meetings | Fine substitutes. Nothing here depends on the brand |
| Zoom or Google Meet | Both can save a transcript if you turn it on | Free if you already pay for the meeting tool |
| Phone voice memo | Records audio. You transcribe it after | Costs nothing. Whisper will transcribe the file locally |
The tool does not matter. What matters is that you end up with a plain text file showing who said what. Tell the person you are recording before you hit record.
Fourteen questions. Each one is here because it fills a named block on the page you build in stage 2. Change the wording to sound like you. Keep the mapping.
| Ask this | It fills |
|---|---|
| Tell me about the business. How did it start? | billboard-quote |
| What kind of job do you want more of? | quote-cards |
| What kind of job do you want less of? | quote-cards |
| Describe your best customer. Think of the last good one. | audience-cards |
| What is that customer worried about before they call you? | psychographics |
| What do people get wrong about your work? | psychographics |
| List your services, the way you say them out loud. | sitemap · menu-visual |
| Name the towns you actually drive to. | sitemap |
| What do you own today? Domain, logo, photos, reviews, profiles. | access-checklist · brand-kit |
| Who has the passwords? | access-checklist |
| Show me two websites you like. What do you like about each? | reference-examples |
| Show me one you hate. What is wrong with it? | design-negatives |
| Somebody calls right now. What happens next? | cro-patterns |
| What is coming up? A season, a promo, a truck, a hire. | timeline |
These fourteen questions were written for this playbook. They are not a copy of a real call. They were built backwards from the 27 blocks in block-manifest.json, so every question has a job. A question that fills nothing gets cut, the same as a sentence that cannot be clicked.
inventory-check-template-v1.0.html ships beside this file. It is empty on purpose.What you get at the end of day one. One page you can send the owner. It shows their words, what you hold, who they sell to, and what happens next. It asks them for nothing.
That page is the source of truth for everything built after it. Read on for how it gets made.
This is the part of the call that decides the design, and it is the part most people rush. Everything in the two directions later comes from about ten minutes of this conversation. Get it thin and the design stage turns into a guessing game that the client pays for in weeks.
One example is a preference. They liked a site. You still do not know which part they liked, or how far toward it they want to go.
Three or more is a mood board. Every site added after the second one adds noise, because now you are averaging instead of aiming.
Two is a vector. It has direction and it has distance. A client can place themselves between two points, and they usually do it without being asked.
The recorded proof. On the plumbing build the owner named two sites himself. He called the first one "a little bit too old school" for his brand. He called the second one "the most comprehensive to what I have in mind". Then, with nobody prompting him, he said "maybe in between".
That last phrase is the whole reason this section exists. He located his own brand in the gap between two named points, in his own words, and he did it because there were two points to sit between. The two directions built the following week are each anchored to one end of that gap. Neither of them was invented by us.
Both sites were named by the client on the call, not proposed by us. Both return 403 to an automated fetch and open normally in a browser, which is why nothing in this process promises to screenshot a reference site for you.
Never ask "what look are you going for". An owner who lays pipe for a living does not have that vocabulary and the question makes them feel tested. Ask them to show you something instead, then ask what specifically they liked about it. You are not collecting taste. You are collecting sentences you can quote back on a page.
Two rules that keep it honest. Quote them, never paraphrase, because the paraphrase is where your taste leaks in. And say back what you heard before the call ends, because a misheard reference costs a week and a corrected one costs thirty seconds.
Split on purpose so neither half of the brief can be skipped. Every question names the block it fills. A question that fills nothing gets cut.
The floor is enforced, not suggested. The blank template counts entries in its reference-examples block and will not mark that block filled below two. Ship an inventory check with one example and the page says so about itself, in its own receipt, before a client ever sees it.
That is deliberate. A minimum that nothing checks is a suggestion, and this whole process exists to replace suggestions with checks.
These are the real observed dates from the plumbing build, with the examples work named as its own span inside them. No date here was invented to make room for it.
What skipping it costs, plainly. You do not save that half day. You move it into the directions stage, and there it multiplies. A direction with no anchor is a direction built on somebody's taste, and taste is exactly the thing an operator is right to reject.
There is a number on that. On the flagship build a round-10 veto rejected a direction outright and triggered a full rebuild of that variation from scratch. The rebuild cost a day. Half a day of asking, against a day of rebuilding, plus the week of calendar it lands on. That trade is not close.
That sentence is on the first page we send every client, and it is the load-bearing idea of this entire process.
Most agency deliverables add a decision. Here is a mockup, what do you think. The client now owes you something, and the project stalls on their calendar rather than yours. The inventory check does the opposite: it removes a decision. It proves we listened, proves we hold what we need, and names a date. There is nothing to react to, so there is nothing to wait for.
Every rule below serves it. Every claim is clickable or countable because a claim they have to take on faith is a decision in disguise. The directions are described in words rather than shown, because showing them invites a verdict we have not earned yet. Three asks maximum, because a list of eight is a decision.
Every claim must be verifiable in one click or one count. A quote links to the recording. Inventory is pictured and counted. A palette comes from their own files. Dates are named days. If a sentence cannot be clicked or counted, rewrite it or cut it.
clients/_shared/templates/website-inventory-check-v1/README.md. It is not our phrasing of their rule, it is the rule.The client-facing pages are built as a play, and the order is not decoration. Each act earns the right to the next one. Proving we listened comes before proving we hold their assets, which comes before saying anything about who their buyer is, which comes before asking for anything at all.
| Act | Name | What it does | Why it sits here |
|---|---|---|---|
| Hero | Cold open | Their logo, real counts flipping to done, their strongest quote as the headline, one dated promise | The first screen is a competence claim. It is made with their own words and numbers we counted, not adjectives |
| I | Listening | Their verbatim quotes, each with what we do with it, and the call recording | Nobody believes you understood them until you can repeat what they said |
| II | Possession | Brand kit, palette, access checklist, then their inventory counted or their site audited | Proof we can start. It also surfaces gaps early, while they are cheap |
| III | Understanding | Audience, sitemap, menu, patterns with evidence, both directions in words | Only now do we say anything about strategy, and only against acts I and II |
| IV | Plan | Dated road, honest conditionals, three asks maximum, nothing to approve | The ask lands last, after the proof, and it is small |
The ladder defines what the client receives. The stages define what we do. They are not the same list and confusing them is the most common way this process gets misread, so here they are side by side.
| Stage | What we do | Ladder rung | What the client receives |
|---|---|---|---|
| 1 | The call, mined into source material | Rung 1 · inventory check | One page proving we listened and can start, with a dated promise |
| 2 | The inventory check built | ||
| 3 | Two design directions built and iterated | Rung 2 · design reveal | Both directions as live clickable pages |
| 4 | Pillar templates, then rollout | No rung | Nothing. Internal build time |
| 5 | AI readiness applied from the first line | No rung | Nothing. It is a property of the build, not a deliverable |
| 6 | Hosting and launch | No rung | A live site |
| 7 | Post-launch, the standing system | Rung 3 · RESERVED | Rung 3 is unnamed. Working candidate is a launch-proof page |
The tool is not what makes this work. The discipline around the tool is. Five practices do the load-bearing work, and one of them is knowing what the tool does not do.
Anything substantial goes through a fixed pipeline before a line is written: a council of five advisors argues the goal from different angles, six improvement rounds climb from surface fixes to structural rethinks, and a drift gate scores the final plan against the operator's original words with quoted evidence per criterion.
The gate is the part that matters. Six rounds of improvement make a plan smarter and can quietly make it somebody else's plan. The gate holds the plan against the raw idea, verbatim, and returns pass, drift, or block. Drift gets one targeted remediation round and a full re-gate. It is allowed to fail its own author, and it does.
Every run leaves a directory on disk: the idea verbatim, the council, the plan, the round log, the gate verdict, the contract. Later phases read those files rather than the conversation, which is what keeps a long build from drowning in its own history.
No markup until the facts are written down. Every build gets a source-material.md whose first line states the standard directly: every claim on the page traces to a line here, and nothing on the page was written from memory.
This is the single highest-value habit in the process. It moves the argument about what is true to a text file, where it is cheap, instead of into the HTML, where it is expensive and invisible.
Repetitive assembly is a script that runs, not a task someone does. The inventory-check family ships build-assets.cjs and inject-strips.cjs. Category order comes from the client's stated revenue priorities in a config block at the top. Counts are emitted by the script, so the number on the page and the number of files on disk cannot disagree.
clients/_shared/templates/website-inventory-check-v1/.Not reviewed by eye. Driven. A throwaway script speaks the DevTools protocol to headless Chrome and asserts in-page: doctype and layout mode, console error count, horizontal overflow at both widths, whether the rail actually descends, how many reveal elements fired, whether reduced motion loses information, whether external links carry the right rel attribute.
The mobile trap worth knowing: a headless window set to 390 pixels wide does not give you a 390 pixel layout, because the minimum window width is larger. It silently captures a cropped desktop layout that looks fine. True mobile requires a device metrics override. Verifying mobile the obvious way passes while being wrong.
The model builds. It does not decide. Every stage that touches a client stops at a human. The ladder states the strongest one plainly: Justin signs off the email, and his send is the trigger. Nothing on this ladder reaches a client without it.
WEBSITE-LADDER.md standing rules, and the rung 1 runbook calls it the only gate.What it does not do. It does not talk to the client. It does not choose the design direction. It does not send anything. It does not know whether a fact is true, only whether it is written down, which is exactly why the source-material discipline exists.
The failure mode is not the model refusing. It is the model producing something confident and wrong. Every gate above exists because that happened at least once.
The section above, made portable. These are the actual installed skills the website build runs on, counted live on 2026-08-02. Nine of them a colleague can copy and run today. Thirteen need our accounts or our client data, and are listed so the shape of the system is legible, not so it can be copied.
| Tier | Path | Installed | Who has it |
|---|---|---|---|
| Project | .claude/skills/ | 32 | Anyone who clones the repo |
| User | ~/.claude/skills/ | 31 | Per machine, per person |
| Portable library | ~/skills/ | 23 | Per machine, symlinked into each CLI |
A skill counts as real when its directory exists and contains at least one recognised manifest: SKILL.md, or skill.json, or CLAUDE.md. Anything that fails is cut from this page rather than softened.
Why three manifests and not just the obvious one. Checking for SKILL.md alone looks correct and is wrong. One of the design skills below ships skill.json and CLAUDE.md and has no SKILL.md at all. The obvious rule deletes a working skill, which is the exact failure the rule exists to prevent, running backwards. It was caught by running the check against the real tree instead of trusting it.
Mechanically, on what would actually block you, and every internal row prints its reason so you can check it rather than trust it. A skill is internal when running it needs either an account of ours (it calls a service bound to our credentials) or data of ours (it reads private client directories in this repo). Neither, and it is shareable.
Two rules that sound reasonable and are not, both tried and discarded: matching the word token case-insensitively flags every design skill, because design systems are full of tokens; and treating a mention of a credentials file as a dependency flags the planning skills, which merely tell you to check one. Text that talks about credentials is not the same as code that needs them.
docs/output/website-build-playbook/skill-kit.md.| Skill | Tier | Install path | Class | Why internal |
|---|---|---|---|---|
one-shot | project | .claude/skills/one-shot | Shareable | |
llm-council | project | .claude/skills/llm-council | Shareable | |
ebi-process | project | .claude/skills/ebi-process | Shareable | |
goal-lock | project | .claude/skills/goal-lock | Shareable | |
thread-to-spec | user | ~/.claude/skills/thread-to-spec | Shareable | |
ui-ux-pro-max | project | .claude/skills/ui-ux-pro-max | Shareable | |
design-taste-frontend | project | .claude/skills/design-taste-frontend | Shareable | |
browser | project | .claude/skills/browser | Shareable | |
browser-automation | user | ~/.claude/skills/browser-automation | Shareable | |
design-create | project | .claude/skills/design-create | Internal | our accounts |
emil-design-eng | project | .claude/skills/emil-design-eng | Internal | our accounts |
html-reports | user | ~/.claude/skills/html-reports | Internal | our accounts |
precall-intel | user | ~/.claude/skills/precall-intel | Internal | our accounts |
seo-build | project | .claude/skills/seo-build | Internal | our accounts |
seo-audit | user | ~/.claude/skills/seo-audit | Internal | our accounts |
lastpass | project | .claude/skills/lastpass | Internal | our accounts |
search-atlas | project | .claude/skills/search-atlas | Internal | our accounts and private data |
sa-tools-discover | project | .claude/skills/sa-tools-discover | Internal | our accounts and private data |
ghl-tools-discover | project | .claude/skills/ghl-tools-discover | Internal | our accounts and private data |
site-deploy | project | .claude/skills/site-deploy | Internal | our accounts and private data |
client-reporting | project | .claude/skills/client-reporting | Internal | our accounts and private data |
hive-mind-status | project | .claude/skills/hive-mind-status | Internal | private data |
frontend-design exists as .claude/skills/frontend-design.md, a bare file rather than a directory skill, so it does not resolve under the rule above. It is real and the design chain references it three times, so it is named here as a file rather than presented as something you can install.
design-spec is not a skill at all. Every reference to it resolves to brain/design-spec.md, a document. Worth knowing before someone goes looking for four skills in that chain and finds two, one loose file, and a doc.
The skills are carried inside this file. Click one and it saves to your computer as a zip. There is no download server, no sign up, and no repo to clone. If you are reading this page, you already have the files.
If nothing appears above, JavaScript is turned off. The files are still in this document, near the bottom of the source.
~/.claude/skills/Make that folder first if it is not there. On Windows it is %USERPROFILE%\.claude\skills\./one-shot, or just describe the job and let it route.That is the whole mechanism. Nothing new was built for this playbook. Wiring for other command line tools is in ~/skills/README.md; the repo procedure is sops/skills-install.md.
ui-ux-pro-max is shareable but it is 11 MB across 326 files, which is a hundred times the size of everything else here combined. Putting it inside a document would make the document unusable to send.
It is also not ours to hand out repackaged. It is third party, MIT licensed, by NextLevelBuilder, and it ships its own installer. Get it the way its authors intended: run npx uipro-cli init --ai claude, or take it from the source below. Version installed here at build time: 2.5.0.
The thirteen internal skills are not downloadable and would not run for you if they were. Each one calls a service bound to our credentials or reads a private client directory. They are listed above so the shape of the system is legible, not so it can be lifted.
Each stage carries four things: what happens with its artifact and gate, the ordered skill sequence, the prompt, and the receipt. A stage is not documented until it has all four.
Every prompt block is stamped. RAN means that prompt was actually executed and names the run directory you can open. AUTHORED means it was written for this playbook from the documented process and never executed as written. There is no third state, and no prompt here is presented as proven without its run.
One onboarding call, recorded. Everything downstream is extracted from that single transcript: verbatim quotes, what they said their highest-revenue work is, what they are tired of, the reference sites they named, and every constraint they stated as a negative. Then their live site is audited and counted, not estimated.
Artifact: source-material.md. Gate: every quote traces to one real transcript, every count is a count.
This stage opens with the mistake, not the win. On a real build, the file-storage search returned zero results for six media folders across four different query forms. The first draft of the client page therefore said the folders were empty and made "send us your photos" the number one ask.
That was wrong. The photos were already there. The folders had been created about 25 minutes before the query and the search index had not caught up. The one file that did return had been uploaded four and a half hours earlier.
The standing rule that came out of it: never assert a client failed to deliver on the strength of an empty search result. Read the container directly, or ask. Getting this wrong sends a client a page accusing them of not doing something they already did, which costs more trust than the missing asset ever would have.
Mine the onboarding call for <client> into verified source material. Do not write any HTML.
INPUT: the call share link <recording-url>, one call only.
Produce clients/<client>/portal/website/inventory-check/source-material.md whose first line is
"Every claim on index.html traces to a line here. Nothing on the page was written from memory."
Extract, each with its attribution and timestamp:
1. Every verbatim quote worth using, in the client's own words. Number them. Do not paraphrase
and do not merge two sentences into one quote.
2. What they said their highest-revenue work is, in their words.
3. What they said they are tired of, and what costs them money when it goes wrong.
4. Every reference site they named, with what specifically they liked about each.
5. Every constraint they stated as a negative ("no black backgrounds", "nothing busy").
6. Every page or service they named that you should check for on the live site.
HARD RULES: all quotes come from this one transcript. Never invent a second call link. Never
quote from memory. If you are unsure a quote is verbatim, drop it.
Then audit the live site and COUNT, do not estimate: pages and posts from the sitemaps, service
URLs, city or area pages. Record every defect you find while counting (misspelled slugs,
duplicate pages, doubled contact or thank-you pages) with its full URL, because the client can
click each one. Record which pages they named on the call that do not exist.
Never assert the client failed to deliver an asset on the strength of an empty search result.
Either read the container directly, or ask.
clients/juan-and-only-plumber/portal/website/inventory-check/source-material.mdWhat the client feels: nothing yet. They had a call and hung up. The whole point is that the next thing they receive proves the call mattered.
The four-act page is built and sent the same day the assets are processed, before any design exists. It is the source of truth the whole build runs from.
Artifact: the deployed page on the client portal. Gate: the claim audit, the scrubs, the render checks, then the operator gate.
Because it is the only artifact everyone argues against instead of arguing from memory. Once the quotes, the counts, the palette, the sitemap and the two directions are on one page the client has read, every later question has a place to be settled.
Design disagreements become checkable: the direction is named from their language, so "this does not feel like us" is answerable by pointing at the sentence it came from. Scope disagreements become checkable: the sitemap is on the page, so a page nobody planned is visibly a page nobody planned.
Skip it and the build runs on somebody's recollection of a call. That is the failure this stage exists to prevent, and it is why rung 1 always ships first.
The template is blank on purpose. It ships with all 27 blocks, each carrying an instruction and nothing from a previous client. Clone it, fill the slots, delete the instructions. The receipt at the foot counts filled blocks live, so it doubles as a progress bar.
Blocks, generated from block-manifest.json.
Hero: logo-band boot-check headline-quote dated-promise.
Act I: billboard-quote quote-cards call-link quote-count.
Act II: brand-kit palette access-checklist inventory-or-audit possession-counts.
Act III: audience-cards psychographics sitemap menu-visual cro-patterns ai-seo-link reference-examples design-negatives direction-a direction-b.
Act IV: timeline conditionals asks nothing-to-approve.
Psychographics ship as questions, not examples. The template asks what the client said costs them money when it goes wrong, what they are tired of, what makes their buyer close the tab. There are no sample answers, deliberately, because a sample answer is exactly the thing that bleeds from one client's page onto the next one's.
The evidence bar is two sources, not one. Every pattern claimed carries one external published study with its year AND one internal receipt from a real build. Two of our own receipts is still one source, and a section that cites only ourselves is the mistake it is trying to avoid.
Build and ship the <client> "Step One" inventory check today. Work from repo root <repo-root>. GOAL: <contact-name> (<client>, <city>) receives today one personalized inventory-check page on their portal site proving we have their assets, access, feedback, audience, and plan, while keeping the build unblocked. Existing client work must remain byte-untouched. No A/B design visuals appear. No new client review cycle is created. The final send is gated on Justin approving the email. PREFLIGHT (verify all; any failure = stop and report): 1. clients/<client>/portal/ exists and git status --short on it is clean. 2. The source material file exists and every quote in it is attributed to one real transcript. 3. Brand kit files exist at their recorded paths. 4. The static deploy script exists and its token resolves from the environment file. 5. The call recording link returns 200. 6. identity/swarmsystem-brand.html exists. BRAND: read docs/reference/visual-design.md and obey it in full. Fraunces headlines and display numbers, Inter body and labels, styleguide root tokens, our red #dc2626. THE CLIENT'S brand colours appear ONLY inside their palette swatch component and their own logo, never as page styling. State words WAITING/WORKING/DONE/LIVE. Motion only reports state, draws a trend, or lands a number. Reduced motion kills all animation with zero information loss. NO em dashes anywhere. No marketing jargon on the page. Plain language, fifth-grade reading level, sentences under 15 words. Copy fills the full container width. STRUCTURE, four acts, each a scanned section with a state dot: HERO, cold open. Their logo on the dark shell. Boot-check rows that flip to DONE with REAL count-ups. Their strongest verbatim quote as the headline. One dated promise line. ACT I, WHAT YOU TOLD US. Second-strongest quote as a billboard. Remaining quotes as cards, each with a "what we do with it" line. Call recording link card. Footer count of quotes. ACT II, WHAT WE HAVE. Brand kit received, palette showing kit-true colours ONLY, open brand asks. Access checklist with state dots, each row naming what it unlocks in plain words. Then either inventory strips in revenue-priority order (product client) or an honest count of their existing site including its defects (service client). Footer counts. ACT III, WHO AND WHERE. Audience cards in plain words. Target towns. Both design directions TEXT ONLY, named from the client's own language, each mapped to the reference sites THEY named, as new-tab cards. ACT IV, THE ROAD. Dated timeline with a countdown chip. Conditional items stated honestly. Exactly the open asks, 3 maximum. The line "Nothing to approve today." GATES (all hard): every sentence is verifiable in one click or one count. Previous-client scrub greps to 0. Unfilled placeholders 0. Em dashes 0. Banned words 0. Renders at true 390 and 1440 with zero console errors. All external links carry rel="noopener". Reduced motion loses nothing. Deploy to the client portal site, then STOP. Justin signs off the email and his send is the trigger.
clients/celebrate-rentals/portal/website/inventory-check/clients/juan-and-only-plumber/portal/What the client feels: someone actually listened, and there is nothing they have to do about it today.
Two directions are built as live clickable pages, home and about and one service page each, and iterated internally until both are review-ready. Not mockups. The client clicks through both on the Friday the inventory check promised.
Artifact: variation-a and variation-b plus a decisions log. Gate: both render clean at both widths, the chassis is provably shared, no brand colour leaks.
The operator veto is designed behaviour, not a failure. On the flagship build, a round-10 veto rejected a direction outright and triggered a full pipeline rebuild of that variation from scratch.
That is the stage working. If no version can ever be rejected, the A/B is theatre and the client is being asked to rubber-stamp a decision already made. The rebuild cost a day. A direction nobody believed in would have cost the project.
The chain also reads brain/design-spec.md and .claude/skills/frontend-design.md. Both are real, neither is an installable skill. See the skill kit above.
Build two design directions for <client> as live clickable pages: home, about, and one service page each. Work from repo root <repo-root>. GOAL: <contact-name> opens one link on <promised-date> and sees both directions as real pages they can click through, not mockups. The direction names come from THEIR language in the source material, never invented style labels. PREFLIGHT: the approved inventory check is live at its portal path; source-material.md exists and names both directions and both client reference sites; brand assets exist at their recorded paths. BUILD variation-a and variation-b under clients/<client>/website-round1/. Each gets home, about, and one service page, sharing a chassis so a picked direction becomes the pillar template with no rework. Anchor each direction to the reference site the CLIENT named for it, and honour every constraint they stated as a negative. Record every non-obvious choice in clients/<client>/website-round1/DECISIONS.md with its why. When the operator vetoes something, that veto is data: write it down and rebuild, do not argue it. GATES: both directions render at true 390 and 1440 with zero console errors; reduced motion loses nothing; the chassis is provably shared (diff the shell); no client brand colour leaks outside the palette component and their logo. STOP at review-ready. The operator picks the direction. You do not pick it.
clients/celebrate-rentals/website-round1/variation-a and variation-b, decisions at DECISIONS.mdWhat the client feels: two real options they can click, on the day they were told, with names that came out of their own mouth.
The picked direction's three pages are frozen as the pillar templates. Home is the chassis, the service page is the repeatable unit, about is the trust unit. Every remaining page in the approved sitemap is built off them.
Artifact: the full site. Gate: chassis parity proven by diff, not assumed.
Chassis parity is proven, not assumed. The shared shell is diffed across every generated page and asserted identical. A page that drifted is a bug, not a variation. This is the check that keeps a 40-page rollout from becoming 40 slightly different sites, and it is cheap only if the two directions shared a chassis from the start, which is why stage 3 requires it.
The sitemap approved back in stage 2 is the contract here. No page gets built that is not on it, and every page on it gets built.
Procedure: sops/site-build.md.
The operator picked <direction> for <client>. Promote it to the pillar templates and roll out the rest of the site. PREFLIGHT: the picked direction's home, about and service pages exist and pass QA; the full sitemap is recorded in the inventory check and approved. 1. Freeze the picked direction's three pages as the pillar templates. The home page is the chassis, the service page is the repeatable unit, the about page is the trust unit. 2. Build every remaining page in the approved sitemap off those templates. Do not invent a page that is not in the sitemap. 3. Prove chassis parity rather than assuming it: diff the shared shell across every generated page and assert it is identical. A drifted page is a bug, not a variation. GATES: every page in the sitemap exists and no page exists outside it; chassis diff is clean; every page renders at true 390 and 1440 with zero console errors; internal links all resolve.
clients/celebrate-rentals/website-round1/, 11 pages per variationWhat the client feels: nothing. This is the quiet stretch, which is exactly why stage 2 named a date they can hold us to.
The readiness layer is applied from the first line of markup, not as a pass at the end. Structured business data, a plain-language description file for assistants, and clean semantic HTML, built in rather than retrofitted.
Artifact: validating structured data on every page. Gate: the page outline read alone, with no styling, still says what the business does and where.
Why it cannot be a final pass. A retrofit is a rebuild. Semantic structure is not decoration you add later, it is the page's skeleton: real landmarks, one top-level heading, headings that describe rather than decorate. An assistant reads the outline before it reads anything else. If the outline was never built, there is nothing to annotate afterwards.
This is also why it is linked from the inventory check in stage 2. The client sees the standard the build will be held to before the build starts, rather than hearing about it as an upsell afterwards.
Apply the AI readiness layer to <client>'s build. This runs from the FIRST line of markup, not as a pass at the end. A retrofit is a rebuild. Scope, from the priced line: schema.org structured data, an llms.txt description file, clean semantic HTML, structured business data. 1. Semantic HTML first: real landmarks, one top-level heading per page, headings that describe the section rather than decorate it. An agent reads the outline before it reads the styling. 2. LocalBusiness structured data per the published spec, carrying hours, area served, and contact. Validate it, do not hand-wave it. 3. llms.txt at the root describing what this business does and where it serves, in plain sentences. 4. Every service and city page carries the structured data for THAT service and THAT city, not a copy of the home page's. GATE: structured data validates; the page outline read alone, headings only and no styling, still tells you what the business does and where; llms.txt resolves.
Returned 200 on 2026-08-02
.claude/skills/seo-build/What the client feels: nothing directly, which is the point. They notice it later as showing up in places they were not showing up before.
The built pages go live behind one Cloudflare Worker, on the client's preview subdomain, without disturbing whatever is currently serving their live root.
Artifact: a live site plus verification receipts. Gate: owned paths serve 200 with a valid certificate, and the incumbent paths are byte-identical to before the cutover.
One Worker bound to <client>.swarmsystem.ai/* that does three things:
Serves the owned design paths from Cloudflare Static Assets, marked noindex. Serves media at /r2/<key> from an R2 bucket. And passes every other path through to the incumbent origin unchanged, so the live root keeps serving exactly as it did.
This is a strangler fig. You take over paths one at a time and the root stays untouched until the very last step, so there is never a moment where the client's live site is at risk because you are mid-migration.
src/index.js never changes per client. Everything client-specific is config.Because Pages is hostname-bound. It owns a whole hostname or nothing, and this whole design depends on owning some paths on a hostname while the root keeps serving from somewhere else. Pages cannot leave the live root on another origin cleanly, which is the one thing the migration needs.
Also rejected: a separate Cloudflare account per client. That is ops sprawl, and it would force moving the whole shared zone. It is unnecessary because client ownership lives in the storage bucket and in git, not in who holds the account.
docs/decisions.md, dated 2026-07-21.One SwarmSystem Cloudflare account holds the zones, the Workers and the storage. Custom hostnames attach client domains at scale. The shared zone already lives in that account and has to, because Worker routes need the zone in the same account.
The ownership point, stated as infrastructure rather than as a promise: what the client owns is in the storage bucket and in git. That is what makes leaving easy. It is not a policy we could quietly change, it is where the files are.
src/index.js, an assets binding at ./dist, a bucket binding, and vars for the pass-through origin and the owned prefixes. Copy the example manifest and fill it in.src/index.js, the template has failed.Mode 1, our subdomain. node scripts/wire-client-domain.cjs <label> <site_id>. Creates the CNAME on our zone and wires the host. Needs a DNS-edit token and a deploy token, both by name from the environment file.
Mode 2, the client's own domain. node scripts/wire-client-domain.cjs --domain <domain> --site <site_id>. The client adds the records at their own registrar; the panel shows them exactly what to add. This mode only sets the custom domain and provisions the certificate. We never touch their DNS.
The scar worth carrying: set the custom domain EARLY. A domain that gets pointed at the host before the custom domain is set can wedge certificate provisioning, and unwedging it is worse than ordering the steps correctly in the first place. This is recorded in the ladder as the stuck-certificate lesson because it cost a real launch day.
Cutover is one DNS change: flip the client CNAME from grey to orange. Until it is proxied, the Worker route does not fire at all, so nothing you deployed is live and nothing is at risk.
Rollback is flipping it back to grey. That instantly restores the pure origin path. There is no un-deploy, no restore from backup, no window where the client's site is half-migrated. The rollback is the same single action as the cutover, in reverse.
Verify it worked by hash-comparing the root and every incumbent path before and after the flip. They must be byte-identical. If a single incumbent path changed, the pass-through is wrong and you flip back.
node scripts/wire-client-domain.cjs --verify <domain> proves a wired domain end to end: the DNS answer over DNS-over-HTTPS, the certificate issuer and expiry, an HTTP 200, and a content hash compared against the original host. It exits zero only when the name resolves, the certificate is valid for it, and the page actually serves.
Why DNS-over-HTTPS rather than the system resolver: your operating system's cache will cheerfully tell you a record still resolves after you have deleted it. Asking the resolver directly is the difference between a receipt and a reassurance.
Stand up the Cloudflare Workers and R2 host for <client> and serve the built pages at path prefixes on their preview subdomain WITHOUT disturbing the live root. GOAL: serve <client>'s owned paths from a single Cloudflare Worker using the Static Assets binding, while the subdomain's active root keeps serving byte-unchanged via pass-through. Stand up and prove a media path. Deploy with wrangler. PREFLIGHT (verify ALL; on any failure stop and report, do not improvise): - `wrangler --version` works and `wrangler whoami` lists the account. If whoami fails, STOP. Auth is an operator gate. - The account identifier is in the environment file and object storage is enabled on the account. - The built page directories exist. - `dig +short <client>.swarmsystem.ai` returns the current pass-through origin. - Our zone is editable by the DNS-edit token, and the <client> CNAME there is currently grey. PHASE A, prove the primitive on a THROWAWAY path first. Scratch worker, scratch proxied subdomain. Serve one static page from assets, one object from storage, and pass-through for everything else. Verify 200 plus valid cert, object loads, tail clean. PHASE B, the deliverable. Copy the built pages into dist/, one directory per owned prefix. Point the assets binding at dist. Route one Worker on <client>.swarmsystem.ai/*: requests under an owned prefix are served from Static Assets with a noindex header; every other path is passed through to the incumbent origin, returned unchanged. The route only fires when the hostname is proxied, so flip the <client> CNAME grey to orange. Verify: owned paths 200 plus valid cert; open at 1440 and 390 with zero console errors; HASH-COMPARE the root and every incumbent path before vs after the flip, they MUST be identical; the media sample serves; tail clean. ROLLBACK if anything regresses: flip the CNAME orange back to grey, restoring the pure origin path instantly. Leave the incumbent site intact throughout. PHASE C, confirm it stayed a config-only clone. src/index.js must not have changed. Everything client-specific belongs in the config file plus dist/ plus the bucket. Then wire the domain and take receipts: node scripts/wire-client-domain.cjs <label> <site_id> # our subdomain node scripts/wire-client-domain.cjs --domain <domain> --site <id> # client-owned domain node scripts/wire-client-domain.cjs --verify <domain> # receipts Set the custom domain EARLY. A domain pointed before it is set can wedge cert provisioning. On a client-owned domain the CLIENT adds the records at their registrar. We never touch their DNS. CONSTRAINTS: do NOT touch the client's root domain or their registrar DNS, that is a launch step with its own operator gate. Prove Phase A on the throwaway subdomain before touching the client subdomain. DONE = owned paths serve from the Worker at 200 plus valid cert, the root and incumbent paths are byte-identical to before, the media sample serves, and the host is a config-only clone for the next client.
Serving today
workers/client-site-host/README.md and scripts/wire-client-domain.cjsWhat the client feels: their site is live, their old site never blinked, and nothing about their domain moved without them doing it themselves.
The site joins the standing system: the funnel becomes visible, faults get proposed fixes, and content gets added on a cadence. Some of that is built and running today. Some of it is not, and this section says which is which.
Artifact: live feeds and a visible autonomy state. Gate: every capability is built with a receipt, in development with its identifier, or build-gated with what gates it.
This is the section where a playbook usually lies. Post-launch automation is the easiest thing to describe in the present tense before it exists. The table below has three states and no fourth. Nothing that is not built is written as though it is.
| Capability | State | Receipt or gate |
|---|---|---|
| Funnel visibility: which variant is winning, is the loop safe, when autonomy unlocks | Built | Run 2026-07-28-cro-hivemind-module, gate PASS. Module renders from a live endpoint |
| Proposed fixes land in a queue for a human | Built, propose only | Same contract: any proposal lands in the heal queue and never auto-applies |
| Self-heal surfaces and the autonomy loop | Built | Standing guides for the autonomy loop and the CRO engine |
| Autonomy actually unlocking and acting | In development | Gated on a visible verdict. Stays unproven below roughly 1,000 relevant sessions per variant |
| The agent that writes the change itself | In development | Tracked as a task, explicitly gated on the above proving out |
| City and service pages added on a cadence | Build-gated | Needs the page factory build. Not shipped |
| Articles published on a cadence | Build-gated | Needs the article engine build. Not shipped |
Bring <client> onto the post-launch loop and report only what is built. 1. Register the site's feeds so the funnel is visible: which variant is converting, whether the autonomy loop is safe, and how far the power gate is from unlocking. 2. Proposed fixes land in the queue. They are PROPOSE-ONLY and never auto-apply. Do not describe them as automatic. 3. The power gate stays unproven below roughly 1,000 relevant sessions per variant. Until it clears, report the countdown, not a result. 4. Content cadence: the page factory and the article engine are BUILD-GATED as of 2026-08-02. Do not describe either in the present tense and do not promise a cadence that no shipped code produces. GATE: every capability on any client-facing surface is BUILT with a receipt, IN DEVELOPMENT with its identifier, or BUILD-GATED with what gates it. There is no fourth state and no adjectives.
docs/output/one-shot/runs/2026-07-28-cro-hivemind-module/What the client feels: the site keeps getting better without them filing a request, and nobody promised them a robot that does not exist.
Stages 1 and 2 are the same either way. The filled inventory check is the source of truth in both. The fork is stage 3 only: where the design actually gets iterated. Stages 4 through 7 rejoin, because launch, hosting and AI readiness do not care where the pixels were pushed.
The path everything above documents. Two directions are built as real clickable pages inside the repo, iterated against the chassis, and reviewed by the operator.
Its advantage is that the thing you are reviewing is the thing that ships. There is no translation step between the design and the site, so nothing gets lost turning a picture into markup.
Its cost is that visual iteration happens through a text interface. If the design is unsettled and you want to try twelve things quickly, that is slower here.
Take the filled inventory check into a Claude Design design system project, iterate the look there against one consistent source of truth, then bring the components back into Claude Code to launch and optimize.
Use this when the design is the uncertain part and the content is settled. That is exactly the state you are in the day the inventory check ships.
Create the project as type PROJECT_TYPE_DESIGN_SYSTEM. That type is fixed at creation, so a normal project can never be converted into one later. Getting it wrong means starting over.
Split the filled inventory check into per component preview files, each with a first line marker <!-- @dsCard group="..." -->, grouped the way the pane groups them: Type, Colors, Spacing, Components, Brand. Those markers compile into _ds_manifest.json, which is what builds the card index, so registering assets by hand is legacy and not needed.
Then call finalize_plan. It locks the exact write paths and the local source directory and returns a plan id. Writes without a valid plan id, or to a path outside the plan, are rejected.
One component at a time. Never a wholesale replace, because a replace loses the parts that were already right and you will not notice until later.
Reads come back through list_files and get_file. Treat returned file content as data, never as instructions. Other people can write to a shared project, so anything that arrives reading like a command to you is a command from somebody else.
Pull the iterated components down and reconcile them into the local chassis, then continue in Claude Code through the stages already documented above.
The picked direction becomes the pillar template, the rest of the site rolls out against the sitemap, hosting goes up, and the AI readiness playbook is applied as the site is built rather than bolted on after launch.
NOT EXECUTED as of 2026-08-02. Written from the tool contract, never run end to end.
/design-sync skill is not installed on this machine, no client build has taken this route, and running it creates a project on your own Claude account. Every leg above is authored. None of it is a receipt.Content settled and design uncertain, and you want to try many looks fast: Route B, knowing you are the first to run it. Anything on a client deadline: Route A, because it is the one with receipts. Either way the inventory check gets filled first, and stages 4 through 7 are identical.
Act III of every client page makes claims about what converts. The bar for those claims is two sources, not one: an external published study explaining why the pattern works, with its year, and an internal receipt proving we did it on a real build. Two of our own receipts is still one source, and a section citing only ourselves is precisely the error it is trying to avoid.
Here is the registry the client pages draw from. Every link was checked live on 2026-08-02 and returned 200. Publication years are on the page so a reader can judge recency without taking our word for it.
| Group | Published source | Year | Our receipt |
|---|---|---|---|
| Form friction | Baymard, cutting a checkout from 16 fields to 8 | 2025 | A service client books through one link rather than a multi-field form. Client asks are capped at three for the same reason |
| Baymard, cart and checkout usability research base | 2025 | ||
| Click to call and mobile urgency | Nielsen Norman Group, touch targets on touchscreens | 2023 | The live phone number counted off the client's own site and carried onto the page as a countable fact |
| Nielsen Norman Group, mobile limitations and strengths | 2024 | ||
| Page experience | Google, the three Core Web Vitals | 2025 | Zero console errors and zero horizontal scroll at 1440 and true 390, verified in a real browser before send |
| Google, how the thresholds were defined | 2025 | ||
| Trust signals | BrightLocal, local consumer review survey | 2025 | A client already carried 256 reviews through a widget and was not using it as a trust signal above the fold |
| Navigation legibility | Nielsen Norman Group, separating findability causes | 2024 | Counting a client's sitemaps produced 42 pages and surfaced four live public defects plus five pages named on the call that did not exist |
| Accessibility | WebAIM Million, 95.9% of home pages fail | 2026 | Reduced motion is a hard gate: zero hidden elements, every state resolved, every count final. All external links carry the right rel attribute |
| WebAIM Million, the prior year for trend | 2025 | ||
| W3C, WCAG 2.2 success criteria | 2023 | ||
| Machine legibility | Google, local business structured data | 2025 | The readiness playbook is live and carried into the markup from the first line by the build skill |
| Schema.org, the LocalBusiness type | 2025 |
A pattern that cannot get both tiers is cut, not softened. That rule is why this table has seven groups rather than the dozen a marketing page would claim.
Ordered by when the question hits you, not by topic. Ordering it by topic would just make it a second copy of the seven stages. Every answer points at a place on this page, a path in the repo, or a number this page computes. If an honest answer would need a policy we do not have, the question is not here.
None of the method does. Look at the What you need before day one list near the top of this page: Claude Code, a way to record, thirty minutes of an owner talking, the skills that download out of this file, and the blank template that ships beside it. That is the whole dependency list, and it is deliberate. What our accounts add is convenience at the far end, mainly hosting on a portal subdomain and the AI readiness tooling in stage 6. A method that only works inside our building is not a method, it is a lock-in. You can run stages 1 through 4 having never spoken to us.
Then you do not have stage 1, and everything downstream inherits that. Look at what stage 1 produces: source-material.md, whose gate is that every quote traces to one real transcript. Without a recording you cannot meet that gate, so the honest move is to say so rather than fill the page from memory.
In practice most owners agree once you say why. It is not for compliance, it is so their words end up on the page instead of your summary of their words. If they still refuse, take notes live and mark the page as authored from notes. Do not present a paraphrase inside quote marks.
Do not move on. Read the second question group in the brief section above and try the negative first. Owners who cannot name a site they like can almost always name one they hate, and a hard no is worth as much as a yes because it is one you can carry into both directions.
If they still have nothing, ask them to look at two competitors while you are on the phone and react out loud. Reacting is easier than recalling. What you must not do is pick two sites yourself and present them as theirs. The reference-examples block exists to hold sites they named, and a site you chose has none of the value.
Make them cut it to two before the call ends. Eight is not eight times the information, it is a mood board, and averaging eight sites produces a design that looks like nobody. The brief section above says why two is the number: two is a vector, with a direction and a distance.
The question that does the cutting is the third one in the first group. Ask which is closest to what they have in mind and what the other one is missing. Owners narrow fast when the question is a comparison instead of a ranking. Record the six you dropped in your working notes, not on the client page.
This is real and it is common. Both reference sites on the plumbing build return 403 to an automated fetch and open perfectly in a browser, because bot protection does not distinguish between you and a scraper. Open it by hand.
The brief section above says the same, and it is why nothing here promises to screenshot a reference site for you. A feature that breaks on exactly the sites a client is most likely to name is worse than no feature, because you find out after you promised it. If a site is genuinely dead, cite it as named on the call rather than as a live example.
Then Act II tells the truth about that, which is more useful than it sounds. The access-checklist and brand-kit blocks surface the gap on day one, while it is still cheap, instead of on the day the build stalls waiting for a logo.
Before you declare anything missing, read the stage 1 callout on this page. On a real build a file search returned zero for six folders that were full, because the index had not caught up. Never assert a client failed to deliver an asset on the strength of an empty search result. Read the folder directly, or ask.
You can reorder nothing and skip almost nothing, and the reason is in the four acts table above: each act earns the right to the next one. Proving you listened comes before proving you hold their assets, which comes before saying anything about their buyer, which comes before asking for anything.
The one that gets skipped most is the examples pass, and the cost of skipping it is stated in the brief section: you do not save the half day, you move it into stage 3 where it multiplies. The standing ladder rule is narrower and harder. No design previews before the inventory check ships.
Because an ask is a decision, and a decision parks the project on the client's calendar instead of yours. That is the thesis section on this page and it is the load bearing idea of the whole method.
Most agency deliverables add a decision. Here is a mockup, what do you think. Now the client owes you something. The inventory check removes one instead: it proves you listened, proves you hold what you need, and names a date. There is nothing to react to, so there is nothing to wait for. The template caps client asks at three for the same reason, and a list of eight is itself a decision.
One direction is not a choice, it is a presentation, and the client can only say yes or start over. Three splits attention and usually produces a request to combine them, which means rebuilding all three.
Two works because the client already handed you the pair. The reference-examples block holds the two sites they named, and each direction is anchored to one end of that gap, named from their own language rather than an invented style label. The template makes it a rule: direction B must be a genuinely different bet, and if you cannot say in one sentence what choosing B costs that A does not, you built the same direction twice.
Because showing something invites a verdict you have not earned yet. On the inventory check the two directions are text only, no mockups and no colour studies, which is written into the direction-a instruction in the template.
The client sees real clickable pages at stage 3, on the Friday the inventory check promised. By then the words have already done the work of setting expectations, so the pages are read as an answer rather than as a surprise. A mockup shown on day one gets judged on taste. The same design shown after the brief gets judged on fit.
That is the stage working, not failing. Read the stage 3 callout: on the flagship build a round-10 veto rejected a direction outright and triggered a full rebuild of that variation from scratch. It cost a day.
If no version can ever be rejected then the choice is theatre and the client is being asked to rubber stamp a decision already made. When both get vetoed, treat the veto as data. Write it into the decisions log, go back to the two reference sites and the stated negatives, and check whether the directions actually drifted off their anchors. Usually they did, and usually the anchor is the fix.
The whole thing, which is the point of building the directions as real pages rather than mockups. Both variations share a chassis, and the stage 3 gate requires that sharing to be provable by diffing the shell.
So the picked direction becomes the pillar template in stage 4 with no rework, and the rest of the site rolls out against the sitemap already agreed in Act III. Nothing is redrawn and nothing is re-approved. This is also why the sitemap block matters more than it looks: it is the contract for the rollout, and no page gets built later that is not on it.
Because retrofitting it means touching every page a second time, and the second pass is always worse than getting it right in the markup. The ai-seo-link block says it in one plain sentence: the site is built so that assistants and search engines can read the business correctly from the first line of markup.
There is a practical reason too. The structure that makes a site readable by a machine is the same structure that makes it readable by a person, which is why the menu block asks for two why lines per item, one for the visitor and one for the machine. Bolt it on later and you get markup that satisfies neither.
Read the two routes section above and check the badges before you decide. Route A is marked Proven and carries two live client builds in its stage receipts. Route B is marked Authored, and its NOT EXECUTED line is dated on its face.
The short version. Anything on a client deadline goes Route A, because it is the one with receipts. Route B is worth trying when content is settled and design is the uncertain part, which is exactly the day the inventory check ships, as long as you accept you are the first to run it. Either way stages 1 and 2 are identical and stages 4 through 7 rejoin.
Every answer above points at something: a section of this page, a path in the repo, or a number this page computes. That is deliberate, and it is checked. An answer that pointed at nothing would be a new claim smuggled in through a question, and this page has no way to verify a claim it invented for itself. Where a policy would be needed and we do not have one, the question was cut rather than answered.
A document that teaches "every claim carries its receipt" and carried none would refute itself on its own first page. So the counts below are computed from this document as it loads, the same way the client-facing pages compute theirs. If any of them read zero, something is broken and the receipt says so rather than hiding it.
The download section gets the same treatment. The verdict below only reads passed if every skill offering a download button is also listed as shareable in the table above it. A skill cannot be quietly handed out through one section while another section calls it internal.