Pre-launch — full task kanban: every tracked task and subtask
updated just now · generated 2026-09-12T03:47:06Z by muse-code
24 components · 4 worlds · 17 recipes · 146 CI steps green
The site can take test payments but not real ones, and runs on a temporary database. Five handoffs need you: the real payment keys, photo storage scope, the production database, rotated secrets, and a second AI provider so the system never depends on one. Everything else keeps moving meanwhile.
Real money and real customer data need your accounts, not mine.
0 of 6 subtasks · 0%
Owner-gated list per ARCHITECTURE.md honest gaps + STATUS.md since-6-Sept + CLOUDFLARE.md owner section (endpoint + invoice.paid later + STRIPE_SECRET_KEY rotation before first live key).
ARCHITECTURE.md — honest gaps with the owner-gated listCLOUDFLARE.md — owner section: secret, endpoint, rotationRisk: Test-mode billing must never be mistaken for live billing again — that exact failure already happened once.
Four decisions are yours, not the system's: what the three prices are and whether there is a trial; approving the photo-cleanup test on twenty real photos; hiring an accountant review and a solicitor review of the shop legal pages; and confirming the draft prices on the marketing page are right.
Prices, approvals and professional reviews cannot be invented by the builder.
0 of 5 subtasks · 0%
P1-03 remainder (tiers are two not three, no trial — inventing them would be fabrication); T0-09 (human approval non-optional); P7-03 (nothing starts before sign-off); LIBRARY-v3 Part 18 order step 6; CLOUDFLARE.md (landing pricing owner-unconfirmed draft).
docs/archive/PLATFORM-BUILD-TASKS.md — P1-03, T0-09, P7-03 remaindersWaiting on something external
Two standard website scans — speed-over-time tracking and a security baseline scan — cannot run on this machine because it has no Docker. They are documented and waiting for a machine that has it. Nothing else is held up by this.
Documented-not-proven beats quietly skipped.
0 of 2 subtasks · 0%
T0-10: documented-not-proven, no Docker daemon here. Re-evaluate where Docker exists. Self-hosted LanguageTool graduated to T0-02 (proven separately).
docs/archive/PLATFORM-BUILD-TASKS.md — T0-10 recordTwo industry-standard scans — one that replays the site on a slow connection, one that probes for common attacks — both documented but never run, because they need a tool this machine does not have. Re-evaluate on any machine that does.
Claiming performance and security without running the scanners is theatre.
0 of 2 subtasks · 0%
T0-10 documented-not-proven: no Docker daemon here. Same constraint as BL-T010.
docs/archive/PLATFORM-BUILD-TASKS.md — T0-10 definitionThe system can already fetch real, credited stock photos that pass every check — proven live. But nothing calls that from the photo upload page yet, so every tenant photo is still uploaded by hand. The next step wires the fetch into the page with the photographer's credit shown.
A shop with no photos gets good ones in one click instead of an empty frame.
0 of 5 subtasks · 0%
render/stock.py retrieve() proven live 2026-09-11 (3/3 slots, 0 errors, pinned IDs, imgix server-side crops). Zero callers in platform/server.py. Last two subtasks are the P2-01 remainder (fact→query richness, aesthetic filter).
render/stock.py — retriever, live-proven, zero callersplatform/server.py — asset page where it plugs inThe design advice the system gives comes from books, guides and example collections. Each one needs its licence checked and its lessons written down in one place, so nobody has to wonder where an idea came from.
Unattributed sources become untraceable decisions.
0 of 5 subtasks · 0%
Commerce-agents row claims skills/design extraction but the skill text has no trace of it; Vercel row is UNVERIFIED with extraction planned into a skills/qa that does not exist; refero has no registry row; motion has no consulted-list; Lucide picked per DECISIONS + installed, zero catalogue call sites.
registry/registry.json — 142 source rows with licence stateskills/design/references/ — where extraction records liveOnce ten sites build from ten descriptions, they go side by side. If a designer would call them one system's output, the variety machinery is not working and gets fixed before anything else. This is the test the whole catalogue exists to pass.
Ten same-looking sites are one template, not a generator.
0 of 3 subtasks · 0%
SPEC step 8, unblocked by SP-7's ten (contact sheet exists in evidence). M-1..M-5 are the variety mechanisms; the universal-blocker rule already fired once here and was fixed (adjacent-gaps pairs pass, Q-1). Note for the run: 7 of the 10 shipped bloom — the pressure is real.
render/variety.py — K-008 grammar rulesThe question flow exists and a seven-question creation flow with a downloadable brief runs on the site. What is missing is proving a stranger with no context can finish it, and the adaptive questioning the spec asks for. Until then the front door is impressive but unproven.
A builder nobody can start is not a product.
2 of 4 subtasks · 50%
SPEC step 10 per Part 5. render/interview.py derives the next question from Required data + ladder (controls exist). /build skeleton landed 2026-09-11 (seven questions, photography choice, taste sliders, world choice, brief download; build-check 15/15, site-check 12/12). S-05 interview rebuild (anonymous start, six adaptive questions) still parked-phase.
render/interview.py — ladder-driven next questionmarketing/build-check.py — /build flow checks 15/15The spec lists ten sentences that together mean the system is finished — a stranger publishing in six minutes, ten distinct sites, three photoless but presentable, readable by everyone, fast, unbreakable, private, and honest. One holds today. Nine do not.
This is the definition of finished; everything else is commentary.
1 of 10 subtasks · 10%
Only the honesty check holds (sim billing stripped, ceca793). WCAG: console surfaces + 10 events pages pass axe/a11y gates; generated-site keyboard nav unproven. Six-minute and 90-second figures unmeasured. Isolation: tenant-scoping controls exist per surface, no whole-system proof.
KOZTO-SPEC.md — Part 17, the ten sentencesScreenshots that vote: the system photographs its own pages and compares them against approved originals, so a visual break fails the build instead of reaching customers. It works for the console today. Generated sites, the thing customers see, are not covered yet.
Eyes in CI catch what assertions cannot describe.
1 of 4 subtasks · 25%
T0-01: vrt/console.spec.mjs shipped 2026-09-07 (4 baselines, 4/4 in 5.5s). No BackstopJS/Percy/Argos at this scale, by decision.
vrt/console.spec.mjs — console visual baselinesvrt/budgets.json — project perf/a11y budgetsA checker reads every customer-facing sentence for banned tones and broken formatting, and it already keeps the site clean. What remains is running its bigger self-hosted engine as a service, so the check stays sharp as the words grow.
Threat-toned copy ships when nobody is watching the words.
2 of 3 subtasks · 67%
T0-02: platform/copycheck.py + controls, CI steps; ladder ** balance mirrored; implementation-noun denylist; content/worlds/meta swept clean; proven twice (public API + self-hosted 6.6).
platform/copycheck.py — tone + formatting lintThe JavaScript side is clean — the one bad package was removed outright and the audit is empty. The Python side still needs its own audit, plus a scanner that proves no secret ever lands in the code.
A leaked key in history is a breach, not a bug.
1 of 3 subtasks · 33%
T0-04 done 2026-09-07 for npm. Roadmap keeps the node-vibrant pin for the harvest pipeline.
package.json — dependency pinsEvery page has speed and quality budgets, and both the console and a generated page currently hold them — two real failures were found and fixed proving the wiring bites. The box stays unchecked until the budgets are a standing gate rather than a good run.
Performance rots silently without a number that fails the build.
3 of 5 subtasks · 60%
T0-05: vrt/budgets.json + vrt/lh-ci.sh + vrt/lh-check.py + 3/3 controls; Lighthouse 13.4.1; measured baseline 93/100/100/100. Box unchecked in source — standing gate, not a good run, closes it.
vrt/lh-check.py — budget checkervrt/budgets.json — the budgetsThe checker proved itself by finding 58 errors in old output and zero in new output, and it caught a real swallowed-tag regression the same day. It still needs wiring into CI, which needs a machine allowed to fetch its runner.
Invalid HTML breaks assistive technology first.
2 of 3 subtasks · 67%
T0-06: .htmlvalidate.json + single-quote normalization; CSP meta stays double-quoted as the one documented exception.
.htmlvalidate.json — project-tuned configSmall vision models were proven to invent things that are not there, so the system only trusts them for coarse signals from a fixed allowlist, enforced by a test. The upgrade path — a flagship vision model under provenance rules — needs a provider key.
A model that invents a business name cannot QA anything.
2 of 3 subtasks · 67%
T0-08: moondream invented 'Karoos Web Hosting'; eyevotes-negatives #6 asserts analyze() emits exactly the allowlist. Blocked subtask waits on OW-1's provider key.
vrt/eyevotes.py — deterministic signals + allowlistThe console is watched by deterministic image analysis with per-surface budgets, all passing — and the checker once caught a real math bug before it shipped. Generated pages, the surfaces customers see, still need their budgets, plus type census and fill-rate tracking.
The motivating void numbers came from generated pages, not the console.
2 of 6 subtasks · 33%
T0-11: vrt/eyevotes.py byte-identical re-runs, budgets json, ci.sh wiring, 4/4 echoing controls. T0-12's void-injected control is #5.
vrt/eyevotes.py — void/palette/edge analysisThe front door works and is checked on every build, but it is still hand-written code rather than output of the site generator. A generator that cannot build its own front door has not proven itself. The footer and navigation also need finishing.
Eat your own cooking, or admit the recipe is incomplete.
1 of 3 subtasks · 33%
P1-01 partial 2026-09-07: live site count from DB, pricing from billing, categories from spines; tests 147, negatives 47/47. Done-when (root 200, not 302) holds.
platform/server.py — hand-authored homepage (the problem)marketing/site-check.py — marketing checks 12/12Prices show live from the billing system on the homepage and their own page. But the business only has two tiers, and the page needs three plus a trial marker — those do not exist yet, and inventing them would be fabrication. This waits on your pricing decision.
A price the business never set is a lie with a checkout button.
1 of 3 subtasks · 33%
P1-03 partial 2026-09-08: per-build price from the ledger debit, plain-terms translation per request; tests 218, negatives 60/60. Blocked subtasks are OW-2's pricing decision.
billing/ledger.py — plan catalogue + per-build debitFour real example sites with live links and an honesty header saying they are demonstrations, not customer stories. Still missing: a home-services example whose deploy 404s, an automatic check that links stay alive, and customer results — which do not exist, so none are claimed.
One invented testimonial ends trust permanently.
2 of 5 subtasks · 40%
P1-05 partial 2026-09-07: curated SHOWCASE list; tests 160, negatives 49/49. Home-services needs a real deploy (Phase C).
marketing/site-check.py — marketing checks incl. /workUploaded photos already obey the slot contract — ratios, subject points, labelled holes. What remains is the same enforcement when the generator itself picks crops and focal points, rather than a human uploading. The tenant path is proven; the generation path is not.
Two paths, one contract, or the contract is decoration.
2 of 3 subtasks · 67%
P2-02: component-declared ratios and focal roles. Step 6 (IM-6) closed the tenant side; generation-side selection is NX-2/T1-adjacent work.
render/imagery.py — crop derivation from recordsrender/assets.py — slot contract auditUploads already warn when the shape does not fit the slot, and that warning once caught real confusion. Still missing: warnings for dark, blurry or tiny photos — which need pixel reading the offline build does not have.
A blurry photo that ships is a support ticket with your name on it.
2 of 3 subtasks · 67%
P2-03 partial 2026-09-08: save_asset(..., want=) surfaces build refusal at upload; tests 220, negatives 61/61. Pillow is not in offline CI.
platform/generate.py — save_asset shape warningUploading several photos at once already files them into the right empty slots by shape. Still missing: spotting logos, cutting out backgrounds, and pulling a business's existing Google photos — the first two need pixel tools, the last needs a key.
Tenants arrive with messy photo folders, not studio shoots.
1 of 4 subtasks · 25%
P2-04 partial 2026-09-08: match_photos pure, one photo per slot, unmatched refused with shapes named; tests 223, negatives 63/63. GBP import is OW-1-adjacent (key).
platform/generate.py — match_photos batch filingEvery site already shows its health — missing answers, missing photos, build state — with links to the right screen. What remains is fixing holes in one click from the health screen itself, which needs the photo search and generation seams first.
A hole you can see but not fix is a chore list, not a product.
2 of 3 subtasks · 67%
P2-07 partial 2026-09-08: GEN.site_health(), per-card lines, workspace metrics strip; tests 219, negatives 60/60. One-click waits on NX-2.
platform/generate.py — site_health() + workspace rollupQuestions already come one screen at a time with progress, hints and saving halfway through. Still missing: the animated transitions between them, and questions that narrow based on earlier answers — the ladder has no conditional questions yet, so there is nothing true to show.
Onboarding is seduction; a form is paperwork.
1 of 3 subtasks · 33%
P3-01 partial 2026-09-08: tests 184, strings asserted in the real page (no clean mid-interview demo site for a screenshot).
platform/server.py — interview flow screensThe lint already bans threatening tones and caught a real bug where the ban itself unfataled every recipe. What remains is the actual rewrite — every question, hint and error in a human voice — which needs a writer, not a linter.
A lint keeps words legal; only a writer makes them kind.
2 of 3 subtasks · 67%
P3-02 partial 2026-09-08: copycheck check_tone 8/8 controls; ladder fatality now explicit; mutants 85/85. Remainder needs a writer (owner-adjacent).
platform/copycheck.py — tone lint that guards the rewritePhoto folders and website links already import through hardened, tested paths. Still missing: Google profiles, social presence, PDF menus and product spreadsheets — each needs an external service or parser the offline build does not have.
Retyping forty products is where onboarding dies.
2 of 6 subtasks · 33%
P3-03 partial 2026-09-08: dashboard URL import shares the hardened fetcher parse layer (facts_from_html); both sides screenshot-verified.
platform/generate.py — public_fetch + facts_from_htmlEach design already shows its real section lineup with every slot marked filled or missing, plus the exact photo shopping list to finish it. Still missing: full-bleed renders instead of lineups, and a single switchable view — both need structured data most tenants have not supplied yet.
The reveal moment sells the product; a lineup is a packing list.
2 of 4 subtasks · 50%
P3-05 partial 2026-09-08: _direction_preview shares the demo_preview layer so tenant and stranger can never disagree; proven by the P1-02 0/15 sweep that full renders need a real interview. Tests 217.
platform/generate.py — direction previews + shopping listsThe dashboard already shows real numbers, a launch checklist, next actions and site health worst-first. Still missing: headlines written for the specific trade — which needs per-category event data that does not exist yet.
A generic dashboard is a mirror; a trade-aware one is an advisor.
2 of 3 subtasks · 67%
P4-02 shipped 2026-09-07 + health 2026-09-08; tests 232. Publish + rollback audited (site.publish/site.rollback rows).
platform/server.py — dashboard + health panelSite cards already show status, category and actions on desktop and mobile. Still missing: photographic thumbnails of each site, which need a thumbnail service, and real business names for signups instead of demo-seed naming.
Cards without pictures are a list wearing a costume.
1 of 3 subtasks · 33%
P4-03 shipped 2026-09-07, screenshot-verified desktop + mobile, test-covered. Thumbnails wait on Phase C.
platform/server.py — site card gridThe SEO machinery is done — script-free business markup in every page head, wired into real builds. Everything else in settings is still greenfield: domains, search preview, languages, integrations, team, sessions and the danger zone.
Settings is where trust is configured or lost.
1 of 4 subtasks · 25%
P4-05 partial 2026-09-08: JSON-LD rejected (script-src 'none' forbids scripts); seo controls 10/10; tests 188. Canonical from url_live, skip-not-guess without one.
render/seo.py — SEO layer in every built page headFour worlds exist and each new one must prove it is genuinely different, not a recolour. Twelve are needed before any launch claim, each passing the validators and each authored against a real business with real photography.
Three worlds cannot produce visual variety; the interview ends up doing design's job.
4 of 12 distinct worlds in catalogue
not yet confirmed by the ownerP6-01: distinct in axes not recolours; seeded structural variation designed into the catalogue FIRST (spec Part 10). Worlds today: obsidian, ledger, bloom, luminous. Queue earns recipes; worlds are authored, not queued.
worlds/validate.py — contrast/gamut/coherence validatorsThe tax position is written down with the questions for an accountant, and the order is clear: country first, then tax lines, then currencies. Nothing starts before sign-off, because invented tax rates would be fabrication.
Wrong tax is a liability the business carries, not us — unless we invented it.
1 of 4 subtasks · 25%
P7-03 partial 2026-09-08: facts today are sim processor, no tax, no country field. Multi-currency needs P2-04 locale infra. Note: LIBRARY-v3 Part 8 now governs the tenant-facing VAT truth — reconcile with this position at build time.
JURISDICTION.md — tax position + accountant question listTwo automated checkers already scan every console surface and every generated test page with zero violations — and they have caught real bugs. What remains is the human checklist no script can check: real keyboard journeys, real screen readers, real judgment.
Automated checks catch markup; humans catch meaning.
3 of 4 subtasks · 75%
P8-01 partial 2026-09-08: a11y.py + axe-ci.mjs + controls proving the scanner bites; first run caught an unlabeled field. Generated pages enumerated from site-audit output.
platform/a11y.py — token-contrast gatevrt/axe-ci.mjs — axe-core scanThe cap machinery is armed — one function for levels, warnings before limits, daily alerts to owners. But no model call has ever passed through it, so the gate guards nothing yet. It gets wired the moment the first live call flows.
An unwired cap is a hope, not a limit.
1 of 3 subtasks · 33%
P8-03 partial 2026-09-08: tests 146 billing, negatives 72/72, mutants 87/87. Mail/uplink via outbox; ownerless scopes audit and say so.
billing/ledger.py — spend caps + alertingThe read path was measured at three sizes and the first break — unindexed scans — was fixed with fourteen indexes, proven flat with plan assertions guarding it. Documented next: one shared database connection blocking writers, no background job queue, and operator-wide scans needing their own pass.
Scale breaks where measurement stops.
2 of 5 subtasks · 40%
P8-04 partial 2026-09-08: inbox unread 0.03→0.81ms linear, now 0.02ms flat (40x at 10k); tests 233. Margin-per-tier and support load need real traffic.
platform/migrations/0012_indexes.sql — the fourteen FK indexesRouting already picks providers from a config table of measured results, and the first live scoreboard sends two tasks to the fast model and the rest to the safe stub. What remains is measuring a second provider — one source is an anecdote, two is routing.
Single-provider routing is a preference with a table.
2 of 3 subtasks · 67%
P10-02: render/routing.py + routing.json; route() picks best observed pass/total at current PROMPT_VERSION; empty board routes stub out loud. Recording half: PROMPT_VERSION=1, site.json provenance, job manifests stamped. Honest gap: worker stamp lines have no mutation owner.
render/routing.py — observed-quality routerAfter the shared product contract and base shop come the categories that break weak abstractions — food, clothing and digital first — then the services, tax truth, legal pages, fulfilment, failure states, feeds and the long tail. Each with its own tests that prove refusal refuses.
The recipe book is a specification; this is its construction schedule.
0 of 9 subtasks · 0%
LIBRARY-v3 Part 18.1 order steps 2–10 (step 11 compliance gates fold into each). Gated by RC-1's contract. Interview consequence: top-six refusal-first questions per detected recipe.
recipes/LIBRARY-v3.md — the order, Part 18.1Tenant uploads with messy backgrounds get clean cutouts — proven working on a normal computer in 26 seconds a photo, no paid services. Before it touches real shops it must prove itself on 20 typical photos, and a person signs off, not a script.
Product photos on grey bedroom backgrounds do not sell.
0 of 3 subtasks · 0%
T0-09 open: proven operational (900x700 to RGBA in 26s CPU-only, zero keys), never productised.
docs/archive/PLATFORM-BUILD-TASKS.md — T0-09 definitionShops without their own photos get good stock matched to their category, colours and section — with the photographer's credit rendered wherever the licence demands it. Needs stock-service keys first.
A builder that cannot fill a photo slot cannot finish a site.
0 of 3 subtasks · 0%
P2-01 open. Step-6 T1-1 built the live-stock lookup path; this is the full integration. Stock API keys via OW-1.
docs/archive/PLATFORM-BUILD-TASKS.md — P2-01 definitionWhere no photo exists at all, the builder can generate one from the business facts and the design world's palette — the tenant picks one of three, every one labelled as machine-made, every one costed against the budget.
A missing hero image blocks a launch; a labelled generated one unblocks it.
0 of 3 subtasks · 0%
P2-05 open. Generation provider path is CF-1 C7 (Workers AI); cost hooks are S-09.
docs/archive/PLATFORM-BUILD-TASKS.md — P2-05 definitionEvery photo a shop owns lives in one searchable library — editable descriptions, warnings before deleting a photo a page uses, licence records, and fast modern formats served from the CDN.
Uploads without a library become a pile nobody dares delete.
0 of 5 subtasks · 0%
P2-06 open. Upload + batch-match paths exist (P2-04); this is the library around them.
platform/server.py — current upload/asset endpointsdocs/archive/PLATFORM-BUILD-TASKS.md — P2-06 definitionEvery sentence on a built site carries its origin — tenant-written, imported, sample, machine-drafted or template — sample content renders visibly watermarked, and publishing stays blocked until provenance is clean. Praise without a name never ships.
A site that cannot say where its words came from cannot be trusted.
0 of 4 subtasks · 0%
P3-04 open. Model stamping exists (P10-02 provenance); this extends it to every origin.
render/prompt.py — PROMPT_VERSION + stamp contractdocs/archive/PLATFORM-BUILD-TASKS.md — P3-04 definitionThe dashboard, inbox, billing and settings all speak the same visual language as the sites they manage — today's beige browser-default chrome deleted, and every gap found along the way (tables, charts, notifications) becomes a proper authored component.
A console that looks nothing like its product teaches distrust.
0 of 3 subtasks · 0%
P4-01 open. P6-03 (thirty console components) is the component half of the same work.
platform/server.py — console surfaces todaydocs/archive/PLATFORM-BUILD-TASKS.md — P4-01 definitionShop owners rearrange their site in a real editor — the left panel shows the true page structure, the middle previews instantly, the right shows only the selected element's genuine fields with character counts before limits hit.
Without an editor every change is a support ticket.
0 of 4 subtasks · 0%
P5-01 open. IR work starts early per the phase plan; device-toggle precedent exists on the reveal page (P3-06).
docs/archive/PLATFORM-BUILD-TASKS.md — P5-01 definitionOwners edit a draft while visitors see the published site — a persistent indicator says which is which, one click publishes, and a complete rollback is instant when a change misfires.
Editing the live site is how shops break at lunchtime.
0 of 3 subtasks · 0%
P5-02 open. Publish/rollback audit rows exist (P4-02 site.publish/site.rollback); this is the draft layer above them.
docs/archive/PLATFORM-BUILD-TASKS.md — P5-02 definitionOwners can ask the machine to redo a section, a page or the whole site — and see exactly what would change before accepting. Hand edits are never silently destroyed by a regeneration.
Regeneration without a diff is a slot machine that eats work.
0 of 3 subtasks · 0%
P5-03 open. Revise-path provider contract (call()) exists via T0-07; this is the product surface.
docs/archive/PLATFORM-BUILD-TASKS.md — P5-03 definitionOwners type 'make the hero bigger' and get a previewed change to accept or reject — and the validation layer proves the machine cannot produce a structurally broken page no matter what it is asked.
Natural-language editing without validation is a footgun.
0 of 2 subtasks · 0%
P5-04 open. Builds on P5-03 diff flow + the bind gate (render/gate.py).
docs/archive/PLATFORM-BUILD-TASKS.md — P5-04 definitionOwners see their actual words and photos live in every design world and switch between them in seconds — the moment the product's generative bet becomes tangible.
Switching designs is the demo that sells the product.
0 of 2 subtasks · 0%
P5-05 open. Direction previews (P3-05) are the static precedent; this is the live switch.
docs/archive/PLATFORM-BUILD-TASKS.md — P5-05 definitionMenus, galleries, validated forms, maps, booking, restaurant menus, schedules, testimonials, pricing toggles and announcement bars — each working without JavaScript first and enhanced with a little, all clean under the strict content policy.
Static sections cannot take bookings or enquiries.
0 of 2 subtasks · 0%
P6-02 open. Console is zero-JS by CSP (P4-08); the island runtime is the tenant-side answer.
docs/archive/PLATFORM-BUILD-TASKS.md — P6-02 definitionTables, charts, drawers, pickers, palettes and the rest — each studied from outside references but rebuilt against Kozto tokens. Outside libraries feed authoring only; they never enter a generation prompt.
P4-01's console rebuild has nothing to build with until these exist.
0 of 2 subtasks · 0%
P6-03 open. Golden queue assigns component work (golden/queue.py); build only what it asks for.
golden/queue.py — component ask, in rank orderdocs/archive/PLATFORM-BUILD-TASKS.md — P6-03 definitionEach design world declares how it moves — easing, durations, stagger — and the island runtime reads it, so switching worlds changes how the site moves, not just how it looks.
Motion is half of a design language; static tokens are the other half.
0 of 3 subtasks · 0%
P6-04 open. Blocked in practice on P6-02 (no runtime yet).
docs/archive/PLATFORM-BUILD-TASKS.md — P6-04 definitionNo component, world or check lands without the metadata that describes it and a control proving the check can fail — echoing what it caught. This is the standing rule that keeps the whole catalogue honest.
Rule 1 of the project: a validator that has never failed is not known to work.
0 of 1 subtasks · 0%
P6-05 open as a checklist item; practised throughout (mutate.py 154/154, golden queue discipline).
docs/archive/PLATFORM-BUILD-TASKS.md — P6-05 definitionNew owners start from real generated sites — pinned world-by-category-by-seed outputs, one click to launch, every one switchable and regenerable with its seed labelled. Frozen hand-made templates are banned: they rot. Pasting a liked URL extracts style signals, never content.
Starting from blank loses strangers; starting from frozen lies to them.
0 of 4 subtasks · 0%
P9-01 open, redefined by the gap essay section 1 (frozen templates banned). Needs P5-05 switching + P3-04 quarantine.
docs/archive/PLATFORM-GAP-ESSAY.md — gallery redefinition, section 1docs/archive/PLATFORM-BUILD-TASKS.md — P9-01 definitionCategory, location and competitor-alternative pages generated by the pipeline itself — real pages for real searches, not a sitemap of thin content.
Discovery compounds; paid acquisition does not.
0 of 3 subtasks · 0%
P9-02 open. From week 12 per the phase plan; needs generation at scale first.
docs/archive/PLATFORM-BUILD-TASKS.md — P9-02 definitionBusiness-profile sync, domain registration and email on the tenant's own domain — the surfaces that stop the product donating its value to other companies.
Every review or email on someone else's domain is churn fuel.
0 of 3 subtasks · 0%
P9-03 open. Mail/domain halves live in CF-1 (C5/C6).
docs/archive/PLATFORM-BUILD-TASKS.md — P9-03 definitionDocumented keys, scopes and rate limits plus webhook docs — and marketplace foundations with attribution and revenue share on catalogue metadata.
A platform without an API is a cul-de-sac.
0 of 2 subtasks · 0%
P9-04 open. Internal JSON API surface exists (OPENAPI.md); this is the public v1.
OPENAPI.md — internal API surface todaydocs/archive/PLATFORM-BUILD-TASKS.md — P9-04 definitionEvery release regenerates Kozto's own marketing site through the pipeline — and a failure there blocks the launch. The product eats what it serves.
A generator the team will not run on itself is not ready.
0 of 2 subtasks · 0%
P10-03 open. P1-01 (hand-authored homepage) is what the loop replaces.
docs/archive/PLATFORM-BUILD-TASKS.md — P10-03 definitionEvery build starts from one document with a fixed structure, produced by a single top-tier model call with a cached prefix — validated, retried, then template-fallback, and stored with its inputs, seed, model and cost.
Reproducibility starts with knowing what was asked.
0 of 4 subtasks · 0%
S-01 open. K-009 zones as implementation.
docs/archive/PLATFORM-BUILD-TASKS.md — S-01 definitionEvery configuration pins its module at a version; published versions can never change or disappear — deprecation warns but never removes. Template-conformance joins the existing gate while the readable typechecked source stays readable.
A site that changes under its owner was never theirs.
0 of 3 subtasks · 0%
S-02 open.
docs/archive/PLATFORM-BUILD-TASKS.md — S-02 definitionVersions and briefs live in the database, every build stores its seed and inputs, and publishing is a pointer swap — so rollback is instant.
Rollback that takes an hour is not rollback.
0 of 3 subtasks · 0%
S-03 open. Generalises the P5-02 draft story to the data layer.
docs/archive/PLATFORM-BUILD-TASKS.md — S-03 definitionOwners start without an account, get a magic link later, answer six adaptive questions with thin honest progress, every field saved as typed — including palette and reference questions, with machine guesses flagged as guesses.
Signup walls lose the strangers the demo just won.
0 of 3 subtasks · 0%
S-05 open. Replaces the current ladder flow (P3-01) once designed.
docs/archive/PLATFORM-BUILD-TASKS.md — S-05 definitionWhile the site builds, owners see a stage checklist and a live iframe — no spinner, no fake percentage, honest lines about slowness, one button after a beat. Closing the tab never kills the build; at 180 seconds it ships degraded, honestly.
Fake progress destroys more trust than slowness.
0 of 4 subtasks · 0%
S-06 open. Extends the P3-06 progress/reveal screens.
marketing/build-check.py — current build-flow checksdocs/archive/PLATFORM-BUILD-TASKS.md — S-06 definitionA prompt bar with four routed intents, double-click inline text editing, an 8-second undo chip and safe concurrent edits — plus a restorable version list with generated summaries, and a pre-flight that warns without ever blocking.
Power without undo is a trap; gates without override are a wall.
0 of 4 subtasks · 0%
S-07 open. The editor half needs P5; the pre-flight half extends render/qa.py.
docs/archive/PLATFORM-BUILD-TASKS.md — S-07 definitionBooking, menus, enquiries, portfolios, loyalty and newsletters ship as packs under their own directory — each declaring its modules, data, routes, integrations and plan tier — while the knowledge layer stays untouched.
Real shops need bookings and menus, not just pages.
0 of 3 subtasks · 0%
S-08 open. Runs on the P6-02 island runtime once it exists.
docs/archive/PLATFORM-BUILD-TASKS.md — S-08 definitionEvery build records its tokens by tier, cache-hit rate, images, wall clock and degradations — and a cache-hit rate under 80% is filed as a bug. A $0.40/90-second budget is enforced, with top-tier calls never silently downgraded.
Unmetered generation is how margins die quietly.
0 of 3 subtasks · 0%
S-09 open. Tenant/global caps precedent exists (P8-03); this is per-build metering.
docs/archive/PLATFORM-BUILD-TASKS.md — S-09 definitionTwo providers per tier with a circuit breaker and a recorded serving provider — and customer input delimiter-wrapped at all three model contact points, uploads sanitised, tenant-scoped agent calls, rate limits honoured.
One provider is a single point of failure; raw input is an open door.
0 of 3 subtasks · 0%
S-10 open. Seam exists (T0-07); routing exists (P10-02); this is failover + hardening.
render/routing.py — routing table to extend with failoverdocs/archive/PLATFORM-BUILD-TASKS.md — S-10 definitionThe console design system — palette, type scale, grid, motion rules and the forbidden-tells list — with the two commercial typefaces licensed first and fallbacks specified.
The console cannot outgrow its typefaces.
0 of 2 subtasks · 0%
S-11 open. Precedes P4-01 in practice (the system comes before the rebuild).
docs/archive/PLATFORM-BUILD-TASKS.md — S-11 definitionLanding to publish in six minutes; a weekly ten-brief designer test from seed-system day; accessibility conformance; and three resilience drills — kill any agent, prove isolation, prove no lost work.
A release process that cannot prove itself cannot be trusted.
0 of 3 subtasks · 0%
S-12 open. Dogfood loop (P10-03) is its canary.
docs/archive/PLATFORM-BUILD-TASKS.md — S-12 definitionSessions in the edge store with existing expiry behaviour, rate limits on login, signup and enquiry, bot verification on signup and public forms, and managed firewall rules plus API shielding.
The stranger-facing front door gets attacked first.
0 of 4 subtasks · 0%
C2 open. Turnstile skill available; needs the C0 account first.
docs/archive/PLATFORM-BUILD-TASKS.md — C2 definitionThe Generate button enqueues instead of blocking; a durable workflow runs the stages with retries; a per-site object streams progress over sockets — today's progress screen as architecture. Heavy Python stays in containers, off the request path, and every consumer is idempotent like the ledger.
Inline generation falls over exactly when it matters.
0 of 5 subtasks · 0%
C4 open: wiring the producer without the generation port would be decoration, so it stays open until the port.
wrangler.jsonc — queue producer + Progress DO declareddocs/archive/PLATFORM-BUILD-TASKS.md — C4 definitionProgrammatic custom hostnames with automatic certificates — collapsing the whole domain epic to an API integration plus UI. The permanent domain comes first because it is a performance task: it kills the tunnel tax.
Shops whose domains live elsewhere churn elsewhere.
0 of 3 subtasks · 0%
C5 open. Needs the C0 account first.
docs/archive/PLATFORM-BUILD-TASKS.md — C5 definitionTransactional mail on the email service, tenant-domain mail on routing, feature flags, scheduled health/hole/reconcile/eval jobs, learning-loop events flowing from day one, and an error pipeline feeding the runbooks.
A platform that cannot email, schedule or learn is a demo.
0 of 5 subtasks · 0%
C6 open. Outbox + flags UI exist locally (P3-06, P4-06); this is the edge wiring.
docs/archive/PLATFORM-BUILD-TASKS.md — C6 definitionText, image, embedding and vision inference with no external key for v1 — gateway caching, fallbacks and per-tenant metering feeding the cost remodel and model routing. No model name or fixed price assumed anywhere.
The second provider the routing layer is waiting for.
0 of 3 subtasks · 0%
C7 open. Routing (P10-02) needs observed quality from two sources; this is the second.
render/routing.py — routing table awaiting a second boarddocs/archive/PLATFORM-BUILD-TASKS.md — C7 definitionUnit economics re-derived on requests, CPU time, database, storage and inference — and today's advisory caps become enforced limits with pre-cap alerts. Blocks paid traffic until it lands.
Paid traffic on unmetered costs is a blank cheque.
0 of 2 subtasks · 0%
C8 open and blocking paid traffic. Pre-cap alerting exists locally (P8-03); enforcement is edge work.
docs/archive/PLATFORM-BUILD-TASKS.md — C8 definitionData residency decided per store before the first EU tenant — control plane, objects, files, analytics — and the tunnel and the box retired only after the parity board is fully green.
Residency decided after the first EU tenant is a breach with a date.
0 of 2 subtasks · 0%
C9 open. Last step of the strangler by design.
docs/archive/PLATFORM-BUILD-TASKS.md — C9 definitionYour recipe book is filed and the pause is lifted. The first build from it is the shared product contract plus the base shop recipe, proven by one product that passes and one that is correctly refused. The 17 existing shop folders use the older taxonomy and get reconciled, not rebuilt.
Shops are the revenue surface; the recipe book is the specification they build from.
1 of 5 subtasks · 20%
LIBRARY-v3 Part 18.1 order step 1. Required fields carry _never_generated and provenance; fulfilmentMethods non-empty; sustainabilityClaims need a basis. 41 categories = 32 physical + 8 service + B2B overlay.
recipes/LIBRARY-v3.md — the v3 authoring source, 19 partsrecipes/validate.py — existing recipe checksRisk: The 17 existing recipe folders use the older taxonomy; building v3 alongside without reconciling forks the layer.
Publishing, test payments with disputes and refunds, dunning emails and invoices all work in test mode with honest records. What remains is the only part that matters: one real money round-trip on live keys, which needs your Stripe setup.
Nobody can trade until real money moves once.
3 of 5 subtasks · 60%
P7-01 (tests 116, negatives 67/67, mutants 81/81), P7-02 (tests 137, negatives 71/71, mutants 84/84), P7-04 voice lint. No auto-recharge by design; operator-run sweep documented. Per-jurisdiction tax stays P7-03.
billing/ledger.py — idempotent ledger + capsplatform/dunning.py — dunning sweep + grace + downgradeThe admin surface already finds tenants, shows their money and members, carries an undismissable impersonation banner, inspects failed builds without leaking paths, and toggles feature flags from the browser. Still missing: catalogue coverage, work queue, eval trends, costs, moderation and health scoring.
Operating blind means every incident becomes an excavation.
4 of 8 subtasks · 50%
P4-06 partial 2026-09-08: tests 249/251, inbox tests 46. Progress page renders its own banner (bypasses _page). Form posts accepted on the flags endpoint (own CSRF-channel gate). Step-7 finding 2026-09-12: tenant-seeded collections missing — sitemap recipes (events-promoter) cannot build from tenant input because nothing writes content._structured; the empty-collection refusal is correct and stays until the seeding path exists (SPEC 12.5: site management belongs to the site).
platform/server.py — admin screens + flags + inspectorplatform/inbox.py — inbox + moderation-adjacent surfacesThe strangler plan is underway: Workers in front, tunnelled origin behind, route-by-route parity-proven cutover. The ten tasks C0–C9 below hold the checklists and states; this card holds the live narrative. Watch the 493ms tunnel tax vanish route by route in Lighthouse numbers.
The 493ms tunnel tax dies route by route, with proof at each step.
not yet confirmed by the ownerStates live on C0–C9; narrative here. Live facts 2026-09-11: remote control DB created + migrated + verified, job queue + R2 photo bucket created, edge worker deployed with all four bindings, read ports landed (upload, words, site actions, build screens, generate producer). Known gaps: per-statement D1 atomicity, Progress DO broadcast unverified, Containerfile unbuilt, tunnel serving until cutover. First live hits failed TLS handshake (alert 40) on the brand-new hostname — propagation, retry not reroute.
edge/worker.js — the strangler workerplatform/d1.py — env-selected D1 transportCLOUDFLARE.md — runbook + cutover orderRisk: Cutover without per-route parity proof repeats the tunnel era with new branding.
The operator console finds any tenant, shows their money, members, sites and audit trail, and acts as them only under a banner that cannot be dismissed without stopping. Feature flags toggle live; failed builds explain themselves in an inspector.
Operating blind — or impersonating invisibly — is how platforms abuse.
4 of 7 subtasks · 57%
P4-06 partial 2026-09-08: search/detail/impersonation/inspector/flags shipped (tests 251). Remainder is operator analytics + moderation.
platform/server.py — admin routesdocs/archive/PLATFORM-BUILD-TASKS.md — P4-06 definitionBreadcrumbs on every site screen, a real dark mode, reduced-motion respect, and secrets shown once as selectable text. The remaining wishlist — optimistic UI, undo, shortcuts, autosave — needs console JavaScript, and the console is deliberately script-free; that is an architecture decision to revisit, not a gap to patch.
Polish is the part strangers feel first.
2 of 3 subtasks · 67%
P4-08 partial 2026-09-08: server-rendered slice shipped (tests 238). JS-dependent remainder waits on the zero-JS architecture decision.
platform/server.py — breadcrumbs + dark tokensdocs/archive/PLATFORM-BUILD-TASKS.md — P4-08 definitionTen sites from ten briefs must look undesigned-by-committee. The first mechanism is live — built pages are scanned against their own world's banned list (pure black, square corners, monospace body). Still to come: seeded selection, anti-brief constraints, palettes from uploads, and a similarity check against 500 sites.
A generator whose output is attributable to one system has no moat.
1 of 5 subtasks · 20%
S-04 partial 2026-09-08: qa.py check 6b live, controls 36/36, mutants 95/95. Open ruling: obsidian nav backdrop-filter vs its glassmorphism ban (1 WARN).
render/qa.py — check 6b anti-pattern enforcementdocs/archive/PLATFORM-BUILD-TASKS.md — S-04 definitionThe Cloudflare account project with preview and production environments, current compatibility, logging from the first deploy, secrets in Workers secrets — never the repo — and a preview URL per change. The local config slice is done; the account-side remainder waits on you.
Deployments without a foundation are snowflakes.
1 of 5 subtasks · 20%
C0 partial 2026-09-08 (Option A slice vs installed wrangler 4.130.0 schema). Remainder needs owner account access (OW-1).
wrangler.jsonc — bindings + compat + observabilitydocs/archive/PLATFORM-BUILD-TASKS.md — C0 definitionThe console frontend served from the edge with API routes proxied to the origin at first. Cutover reads first — dashboard, inbox, settings — then writes, generation last, each step proven byte-identical against the Python oracle before it counts.
Cutover without per-route proof is a migration-shaped outage.
3 of 5 subtasks · 60%
C1 partial 2026-09-08 (Option A slice: first two reads). Parity harness is owner-run (needs npx).
edge/worker.js — ported reads + passthroughdocs/archive/PLATFORM-BUILD-TASKS.md — C1 definitionThe control plane on the edge database with the tenant stamped on every row and enforced by the query builder, per-tenant coordination in memory-objects rather than business tables, and content-addressed files in object storage. All thirteen migrations ran unmodified and behaved identically.
A tenant who can read another tenant's rows is a breach, not a bug.
4 of 7 subtasks · 57%
C3 partial 2026-09-08 (Option A slice +2 mutants). Remote DB live 2026-09-11.
platform/d1.py — env-selected D1 transportdocs/archive/PLATFORM-BUILD-TASKS.md — C3 definitionThis page. It now lists every tracked task and subtask with derived progress, instead of nine summary cards. Open any card for the plain-English version, expand it for subtasks, evidence and technical detail. Tell me if anything on it is wrong — correcting the board is the fastest way to keep it honest.
You should never have to read the code to know the state of the project.
4 of 5 subtasks · 80%
status/board.json -> status/build.py (stdlib, <1s) -> status/dist/index.html. build.py refuses a done card without evidence and verifies every files[].path exists. Facts measured at build time. Kanban expansion at owner request 2026-09-12: subtask schema + derived x-of-y percentages (computed, never estimated).
status/board.json — the data, maintained per sessionstatus/build.py — renderer + evidence gatedocs/status-board-brief.md — the owner brief, filed verbatimstatus/build.py358 lines · renderer with the evidence gate


The system kept its own shopping list: every business tested against every world, failures ranked. It asked for one reassurance component — faq-accordion was built, process-steps joined the luminous grammar, the adjacent-gaps rule was fixed — and the list emptied: 176 of 176 pairs bound with zero gaps and zero failures.
The queue, not a guess, decides what gets built next.
4 of 4 subtasks · 100%
Closed 2026-09-12: 176/176 bound (100%), 0 catalogue gaps, 0 failed. faq-accordion authored (serves all four worlds); process-steps added to luminous grammar; adjacent-gaps rule fixed to runs of 3+; 56 ladder eval rows rebased to reassure-served reality, each triple-asserted. 17 distinct spines, 68 effective problems.
golden/queue.py — the ranked ask, measured livecomponents/faq-accordion/meta.json — the built component declarationworlds/luminous.json — grammar gains process-stepspython3 golden/queue.py176 pairs, 176 bound, 0 gaps, 0 failed2026-09-12
The product milestone, shipped: ten real websites from ten plain descriptions across ten recipes and three worlds — all ten with no customer photos, no stock and no generated images, every failure recovered honestly instead of invented around, all ten screenshotted at both widths with the report in the log.
Nothing else matters until this works. It is the product.
6 of 6 subtasks · 100%
SPEC step 7 closed 2026-09-12. Tenant inputs verified verbatim (mechanical assertion, 0 nonverbatim); CTA labels, form-field trios and href '#' disclosed as tenant-chosen chrome. Refusals observed: gesture-carrier decline with reroute-or-name, events-promoter empty-collection refusal (kept as evidence, collections gap filed on SP-14). Adjacent-gaps rule fixed to runs of 3+ along the way (17/17 controls, mutant-proven).
render/page.py — component binding and slot fillplatform/generate.py — the product path the ten ran throughrender/variety.py — adjacent-gaps fix the run forced










Risk: The photoless path is the hardest case and the most common one.
The specification itself is filed in the repo and every step since has been built against it. The STOP reports and the rules that govern the work all cite it by section.
A build order without a saved spec is folklore.
1 of 1 subtasks · 100%
SPEC step 0 + STOP confirmed. AGENTS.md, LOG.md Part 18 reports and this board all cite SPEC sections.
KOZTO-SPEC.md — the build order, acceptance and reporting specKOZTO-SPEC.md1013 lines · Parts 0–19 including build order and acceptance
The database the site runs on today and the cloud database it will move to were proven identical — every migration applied unchanged with the same behavior. The move itself is still to come, but the equivalence is measured, not assumed.
Migrating to a database that behaves differently is a data-loss plan.
2 of 2 subtasks · 100%
SPEC step 1. Option A parity run: schema/FK/UNIQUE/index behavior identical; platform/d1.py transport 23/23 tests (+2 mutants — the stub used to encode array rows and hid live misreads); per-statement atomicity documented as the known semantic gap.
platform/d1.py — env-selected D1 transportplatform/d1.py177 lines · transport with half-config refusal
All three model jobs were run live and scored: writing, understanding and directing pass; the revision job has one stable miss that is understood, not flaky. Routing already sends each job to whichever provider earned it.
Unmeasured model quality is a demo, not a capability.
3 of 3 subtasks · 100%
SPEC step 3 per 11.2. Comprehend's miss is genuine fees fabrication caught twice. Revise miss is verbatim-dictation decline, twice running (1.1 answered it). Single-provider — the second source is OW-1/P10-02.
evals/run.py — eval runner over the four task setsevals/run.py236 lines · runner for copy/comprehend/director/revise sets
The decision about how components become web pages is settled and recorded: static component files lowered with patterns, no clever runtime. The report went out with the recommendation and the reasons.
An undecided renderer forks every component built after it.
1 of 1 subtasks · 100%
SPEC step 4 per 7.1. Retroactive Part 18 report 5f246b0 (LOG +28). Island:none lowering stands.
render/page.py — the decided renderergit show --stat 5f246b0LOG.md +28, retroactive renderer-decision report2026-09-11
Five components were taken through the whole intake pipeline — cards and tables first — passing the sandbox gate and rendering in all three design worlds. The pipeline, not just the components, is what was proven.
Components nobody can ingest are drawings, not inventory.
2 of 2 subtasks · 100%
SPEC step 5 per 7.8. Commit be2926d: 32 files, +1679/−2. Proofs held in intake/proof with _intake ids, no catalogue merge until the queue asks.
components/ingest.py — the proven intake pipelinegit show --stat be2926d32 files, 1679 insertions, 2 deletions2026-09-11
When someone uploads a photo for the main subject, they now mark where the subject is so crops never cut heads off. The system can fetch real credited stock photos that pass every check. And where no photo exists at all, every world shows a labelled empty frame — never a fake-looking gradient, never a broken image. Proven with screenshots for all three worlds.
Photography is 60–70% of perceived quality; every tier of the supply now proves itself or stays silent.
6 of 6 subtasks · 100%
SPEC step 6 + Tier-3 wiring (commit 99da0c6). save_asset(..., focal) refuses subject slots without a point; nine-point picker; batch excludes subject slots; stock.retrieve proven live; cta-monument k-mon__figure -> k-mon__fig (old class never matched _REGION).
platform/generate.py — save_asset focal validation, assets_of threadingplatform/server.py — focal picker, edge enforcement, batch exclusionrender/stock.py — Tier-2 record with want-keyed ratiosrender/gate.py — 3f missing-asset treatment requirementcomponents/cta-monument — treatment + region renamepython3 platform/negatives.py91/91 controls fired (7 new focal controls)2026-09-11
python3 render/stock-negatives.py11/11 stock controls hold (record passes the crop gate)2026-09-11
python3 render/gate-negatives.py16/16 correct (3f catches an untreated component)2026-09-11
live Unsplash retrieve + asset audit, bloom/restaurant3/3 slots filled, 0 errors, 0 fatal; hero photo renders with credit2026-09-11





Your recipe book for shops — 41 kinds of business, what each shop must ask, and what must refuse — arrived and is saved word for word as the source the shop-building work follows. Nothing in it has been built yet; filing it was the step that unblocked the recipe pause.
The shop work was paused waiting for exactly this document.
3 of 3 subtasks · 100%
Persisted at recipes/LIBRARY-v3.md: 19 parts (0–18), taxonomy 32+8+overlay, Part 18 implementation order. Reconciliation with the 17 older dirs is RC-1's work.
recipes/LIBRARY-v3.md — the v3 authoring sourcerecipes/LIBRARY-v3.md2561 lines · all 19 parts, verified by header count
An earlier version pretended test billing was the real thing. It was deleted — thousands of lines removed — and only honest test-mode billing with a read-only live-key check remains. Live billing stays waiting on you (see the top card).
Work described as finished that was not is the failure this board exists to prevent.
2 of 2 subtasks · 100%
ceca793 Phase 1 remediation §3.1–§3.8: strip sim billing, injection boundary, rate limits, fonts, palette, docs archive.
billing/ledger.py — honest test-mode ledger that survivedgit show --stat ceca79386 files changed, 2757 insertions, 6206 deletions2026-09-10
The system talks to AI providers through one guarded door: requests go out with exactly the right shape, answers come back parsed and checked, and the tests prove the door slams shut on bad keys, timeouts and over-limit responses. Adding a second provider later means implementing the same small contract.
A provider change must not rewrite the product.
3 of 3 subtasks · 100%
T0-07 done 2026-09-07: exact request shape, Authorization override refusal, 429 retry-after passthrough, JSON.parse fenced-code recovery; +3 targeted mutants. Anthropic extended with live-scan provocation proving no real key leaks.
render/providers/openai.py — OpenAI provider behind the seamrender/providers/anthropic.py — Anthropic provider behind the seamrender/providers/openai.py159 lines · seam contract + controls
Every image analysis has a matching sabotage test — inject a blank hole, overflow the math, delete a file — and each one fires with the exact message expected. This is the layer that keeps every other checker honest.
A validator that has never failed is not known to work.
1 of 1 subtasks · 100%
T0-12 done 2026-09-07: void-injected hole fires budget control #5; int16 overflow fires #3 before the fix was found; 6/6 with echoes coverage.py can see.
worlds/fontmetrics.py — home of the sabotage-control patternworlds/fontmetrics.py114 lines · controls beside the validator
The design skill's typography claims are cross-checked by measuring real font files — x-heights, stem widths, contrast ratios — so a claim about readability is a measurement, not an opinion. Thirty-three checks plus five controls, with mutation testing proving they bite.
Two independent measures beat one confident opinion.
2 of 2 subtasks · 100%
T0-03 done 2026-09-07: 33/33 + 5/5 controls, 99/99 mutants; DELIBERATE BUG injected and caught before proclaiming safety.
worlds/fontmetrics.py — font measurement controlsworlds/fontmetrics.py114 lines · 33 checks + 5 controls
Strangers see a live demonstration in one screen — no scrolling, real content, and every number on it traceable to its source. The remaining gap is a demo driven by a genuinely empty site, which needs a real interview first.
The first screen decides whether anyone reads the second.
2 of 3 subtasks · 67%
Hmm — one todo subtask on a done card: the card's DONE claim is the shipped demo (0/15 sweep verdict: real-interview demo, not fixture-polish). The empty-site render is tracked as remainder, visibly unfinished. P1-02 shipped 2026-09-07.
platform/generate.py — demo_preview + assemble_contentplatform/generate.py1168 lines · demo layer over real content assembly
Every billing and settings sentence was checked for threatening tones and confusing jargon, with user-facing errors that explain what to do. The lint runs on every build so new threats cannot slip in.
Threat-toned billing copy reads as extortion.
3 of 3 subtasks · 100%
P1-04 shipped 2026-09-07: voice now includes explicit failure-mode entries; delinquency_banner states grace + downgrade with dates. Tests 148, negatives 48/48.
platform/copycheck.py — the lint that keeps it cleanplatform/copycheck.py258 lines · tone lint + fail-closed surfaces
A site with no photos, no direction and nothing filled in renders as declared empty inventories and named holes — never a crash, never a silent gap. Seventeen checks cover it including hostile inputs.
The empty state is the first state every tenant sees.
1 of 1 subtasks · 100%
P1-06 shipped 2026-09-07. Empty states guide to the asset page with slot/alt text; preview 7/7 against real assembled content.
platform/generate.py — zero-asset honesty pathsplatform/generate.py1168 lines · 17 zero-asset checks green
Every hole in a built page declares its own treatment — a photograph slot asks for a photograph, a map slot asks for a map — so nothing renders silent or pretends to be designed. The step-6 work later gave each treatment its own styling.
A hole with a name gets filled; a silent one ships.
2 of 2 subtasks · 100%
P2-08 shipped 2026-09-08: tests 222, negatives 62/62; later extended by IM-6 (focal, gate 3f, monument).
render/assets.py — needs/questions/audit contractplatform/generate.py1168 lines · save_asset + match_photos + questions
Tenants approve an AI-written photo description before it becomes the accessibility text — shown beside the image, one click to accept, one to rewrite by hand. Nothing machine-written reaches a visitor unreviewed.
Alt text nobody approved is a lawsuit with a timestamp.
1 of 1 subtasks · 100%
P3-02b shipped 2026-09-08: tests 210, negatives 58/58, live-DOM verified.
platform/server.py — approval UI beside uploadsplatform/server.py4731 lines · alt-quote approval step
While a site builds, visitors see honest progress with a working cancel; when it finishes, the reveal names what is missing and links straight to fixing it. No dead ends, no fake completion.
A build that lies about finishing loses the customer it just won.
3 of 3 subtasks · 100%
P3-06 shipped 2026-09-08: build-check 13/13. Coordinator runs inline in SQLite (the queue is CF-1's work).
marketing/build-check.py — build flow checks 13/13marketing/build-check.py192 lines · 13/13 progress + reveal checks
The inbox — where review flags land — is empty-state honest, plain-English, and shows only real review work. No notification theatre, no demo rows.
An inbox of fake work teaches owners to ignore the real thing.
1 of 1 subtasks · 100%
P4-04 shipped 2026-09-07: no notification center, no badge counts, no seeds; tests 198, negatives 55/55.
platform/inbox.py — review-only inboxplatform/inbox.py164 lines · review-only inbox, no theatre
Analytics, backups, dunning and the rest each have their own small test suite, and all 399 pass. Each was specified from real code, not invented — the process itself was audited twice.
Instruments nobody tests lie the loudest.
1 of 1 subtasks · 100%
P4-07 shipped 2026-09-08: analytics/backup/dunning/restore/integrity/scheduler/notifications/lifecycle/progress/flags/results, each with its own suite.
platform/analytics.py — one of the eleven slicesplatform/backup.py — timestamped backup + restoreplatform/dunning.py — dunning sweepplatform/analytics.py82 lines · one passing instrument slice
Buying, upgrading, failing, disputing and refunding all work against the test processor with idempotent records — replay the same webhook twice and the balance moves once. The sweep that charges renewals is operator-run and documented as such.
Money code without idempotency charges twice.
2 of 2 subtasks · 100%
P7-01 shipped 2026-09-08: tests 116, negatives 67/67, mutants 81/81. Sim billing stripped later (DL-1); this card is the honest test-mode that survived.
billing/ledger.py — idempotent test-mode ledgerbilling/ledger.py98 lines · ledger + reconcile + replay guards
Failed payments trigger a real sequence — emails, a visible banner, a grace period — ending in downgrade, never silent cutoff. Every run is audited with before-and-after states.
Silent cutoff is how you lose a customer and keep their money.
2 of 2 subtasks · 100%
P7-02 shipped 2026-09-08: tests 137, negatives 71/71, mutants 84/84.
platform/dunning.py — dunning sweep + grace + downgradeplatform/dunning.py168 lines · audited dunning runs
Every billing and settings surface now carries lint metadata with explicit failure modes, so new screens cannot ship without declaring how they fail. The registry has a control proving it bites.
Voice without a registry rots one screen at a time.
2 of 2 subtasks · 100%
P7-04 shipped 2026-09-08: tests 129, negatives 68/68, mutants 82/82.
platform/copycheck.py — voice metadata + registry gateplatform/copycheck.py258 lines · voice registry gate
The render quality gate was proven on a real business with seven injected faults — it caught five outright and named all seven. The remaining two taught the gate new checks the same day.
A gate tested on clean input guards nothing.
3 of 3 subtasks · 100%
P8-02 shipped 2026-09-08: render/qa.py + render/qa-negatives.py 22/22 controls incl. CI visibility/CI failure/workspace hygiene.
render/qa.py — render quality gaterender/qa.py170 lines · gate + live-fire proof
44 real businesses across every category build successfully against four design worlds — all 176 pairs bound with zero gaps and zero failures. The queue that used to name 17 gaps and 3 failures now reports an empty ask. Seventeen distinct spines carry 68 effective problems.
A corpus that hides its failures is marketing, not measurement.
4 of 4 subtasks · 100%
P10-01 + queue as of 2026-09-12: 176 pairs, 176 bound (100%), 0 catalogue gaps, 0 failed. Closed via Q-1 (faq-accordion, process-steps in luminous grammar, adjacent-gaps rule fix).
golden/queue.py — binding + ranked askpython3 golden/queue.py176 pairs, 176 bound, 0 gaps, 0 failed2026-09-12
Not starting before a paying customer needs it