Websites build with or without photographs. The binder's current ask: one reassure component and 17 catalogue gaps. Next: the first real shop recipes — waiting on credentials only you can provide.

Pre-launch — full task kanban: every tracked task and subtask

updated just now · generated 2026-09-12T00:51:33Z by muse-code

23 components · 4 worlds · 17 recipes · 146 CI steps FAILING

Waiting for you · 2

Live keys and secrets only you can provideM

The site can take test payments but not real ones, and runs on a temporary database. Five handoffs need you: the real payment keys, photo storage scope, the production database, rotated secrets, and a second AI provider so the system never depends on one. Everything else keeps moving meanwhile.

Real money and real customer data need your accounts, not mine.

0 of 6 subtasks · 0%

  • Stripe webhook secret + endpoint registered todo
  • R2 scope for tenant assets todo
  • Postgres for the control plane (or confirm D1 instead) todo
  • Rotate the leaked secrets (B-1), today regardless todo
  • Second model provider key for K-010 routing todo
  • SMTP credentials so queued mail actually sends todo
not yet confirmed by the owner

Owner-gated list per ARCHITECTURE.md honest gaps + STATUS.md since-6-Sept + CLOUDFLARE.md owner section (endpoint + invoice.paid later + STRIPE_SECRET_KEY rotation before first live key).

  • ARCHITECTURE.md — honest gaps with the owner-gated list
  • CLOUDFLARE.md — owner section: secret, endpoint, rotation

Spec: KOZTO-SPEC Part 19 (operations)

    Waiting on:
  • Stripe webhook secret + endpoint (owner)
  • R2 scope for tenant assets (owner)
  • Postgres for the control plane (owner)
  • Secret rotation B-1 (owner)
  • Second model provider key (owner)
  • SMTP credentials (owner)

Risk: Test-mode billing must never be mistaken for live billing again — that exact failure already happened once.

Started 2026-09-06

Business decisions only you can makeS

Four decisions are yours, not the system's: what the three prices are and whether there is a trial; approving the photo-cleanup test on twenty real photos; hiring an accountant review and a solicitor review of the shop legal pages; and confirming the draft prices on the marketing page are right.

Prices, approvals and professional reviews cannot be invented by the builder.

0 of 5 subtasks · 0%

  • Third pricing tier + trial marker (tiers are two, not three) todo
  • Approve the 20-photo background-removal evaluation todo
  • Commission accountant review of the VAT position (P7-03) todo
  • Commission solicitor review of GB legal templates (recipe order step 6) todo
  • Confirm draft pricing tiers on the landing page todo
not yet confirmed by the owner

P1-03 remainder (tiers are two not three, no trial — inventing them would be fabrication); T0-09 (human approval non-optional); P7-03 (nothing starts before sign-off); LIBRARY-v3 Part 18 order step 6; CLOUDFLARE.md (landing pricing owner-unconfirmed draft).

  • docs/archive/PLATFORM-BUILD-TASKS.md — P1-03, T0-09, P7-03 remainders

Spec: Build-tasks P1-03, T0-09, P7-03; LIBRARY-v3 Part 18

    Waiting on:
  • Pricing decision: third tier + trial (owner)
  • T0-09 twenty-photo approval (owner)
  • Accountant + solicitor reviews (owner)

Blocked · 2

Waiting on something external

Scans that need Docker (not available here)S

Two standard website scans — speed-over-time tracking and a security baseline scan — cannot run on this machine because it has no Docker. They are documented and waiting for a machine that has it. Nothing else is held up by this.

Documented-not-proven beats quietly skipped.

0 of 2 subtasks · 0%

  • sitespeed.io run where Docker exists blocked
  • OWASP ZAP baseline scan where Docker exists blocked
not yet confirmed by the owner

T0-10: documented-not-proven, no Docker daemon here. Re-evaluate where Docker exists. Self-hosted LanguageTool graduated to T0-02 (proven separately).

  • docs/archive/PLATFORM-BUILD-TASKS.md — T0-10 record

Spec: Build-tasks T0-10

    Waiting on:
  • Docker daemon on the build machine (environment)
Speed and security scans where Docker existsS

Two industry-standard scans — one that replays the site on a slow connection, one that probes for common attacks — both documented but never run, because they need a tool this machine does not have. Re-evaluate on any machine that does.

Claiming performance and security without running the scanners is theatre.

0 of 2 subtasks · 0%

  • sitespeed.io run where a Docker daemon exists todo
  • OWASP ZAP baseline scan where a Docker daemon exists todo
not yet confirmed by the owner

T0-10 documented-not-proven: no Docker daemon here. Same constraint as BL-T010.

  • docs/archive/PLATFORM-BUILD-TASKS.md — T0-10 definition

Spec: Build-tasks T0-10

    Waiting on:
  • Docker daemon (see BL-T010)

Next up · 72

What the binder asks for: reassure component + gapsM

The system keeps its own shopping list: every business tested against every world, and the failures ranked. Right now it asks for one new reassurance component, has 17 catalogue gaps and 3 failed pairs, and flags one rule — adjacent gaps — as wrong rather than the catalogue. This is the next build when autonomous work resumes.

The queue, not a guess, decides what gets built next.

0 of 4 subtasks · 0%

  • Build the reassure component the queue asks for (run queue, do not guess) todo
  • Fix the adjacent-gaps universal blocker (rule problem, not catalogue) todo
  • Close the 17 catalogue gaps todo
  • Fix the 3 failed pairs todo
not yet confirmed by the owner

Measured 2026-09-12: 44 businesses x 4 worlds = 176 pairs, bound 156/176 (88%), catalogue gaps 17, failed 3. Reassure ask: pricing-cards | process-steps | policy-strip | faq-accordion. 17 distinct spines, 68 effective problems.

  • golden/queue.py — the ranked ask, measured live

Spec: ARCHITECTURE.md section 7

Plug live stock photos into the upload flowM

The system can already fetch real, credited stock photos that pass every check — proven live. But nothing calls that from the photo upload page yet, so every tenant photo is still uploaded by hand. The next step wires the fetch into the page with the photographer's credit shown.

A shop with no photos gets good ones in one click instead of an empty frame.

0 of 5 subtasks · 0%

  • Call retrieve() from the asset page for empty slots todo
  • Save fetched photos with credit through save_asset todo
  • Keep the existing audit as the gate (no bypass) todo
  • Fact-to-query builder: category, world adjectives, intent, palette todo
  • Aesthetic and palette filter on results todo
not yet confirmed by the owner

render/stock.py retrieve() proven live 2026-09-11 (3/3 slots, 0 errors, pinned IDs, imgix server-side crops). Zero callers in platform/server.py. Last two subtasks are the P2-01 remainder (fact→query richness, aesthetic filter).

  • render/stock.py — retriever, live-proven, zero callers
  • platform/server.py — asset page where it plugs in

Spec: KOZTO-SPEC Part 16 step 6 (Tier-2 remainder); build-tasks P2-01

Name every design source the system learns fromS

The design advice the system gives comes from books, guides and example collections. Each one needs its licence checked and its lessons written down in one place, so nobody has to wonder where an idea came from.

Unattributed sources become untraceable decisions.

0 of 5 subtasks · 0%

  • Refero/design.md extracts + design-property index todo
  • Vercel patterns: verify licence, write the node todo
  • Commerce-agents extraction record todo
  • Motion consulted-libraries list (stays consult-only) todo
  • Lucide rule: picked, ISC, installed, unused by catalogue todo
not yet confirmed by the owner

Commerce-agents row claims skills/design extraction but the skill text has no trace of it; Vercel row is UNVERIFIED with extraction planned into a skills/qa that does not exist; refero has no registry row; motion has no consulted-list; Lucide picked per DECISIONS + installed, zero catalogue call sites.

  • registry/registry.json — 142 source rows with licence state
  • skills/design/references/ — where extraction records live

Spec: KOZTO-SPEC Part 16 step 7

Prove variety: ten sites side by sideL

Once ten sites build from ten descriptions, they go side by side. If a designer would call them one system's output, the variety machinery is not working and gets fixed before anything else. This is the test the whole catalogue exists to pass.

Ten same-looking sites are one template, not a generator.

0 of 3 subtasks · 0%

  • Lay the ten SP-7 screenshots side by side todo
  • Designer test: one system or ten makers todo
  • Fix M-1 through M-5 until it passes todo
not yet confirmed by the owner

SPEC step 8, gated on SP-7's ten. M-1..M-5 are the variety mechanisms; the universal-blocker rule (100% blocked = the rule is wrong) applies.

  • render/variety.py — K-008 grammar rules

Spec: KOZTO-SPEC Part 16 step 8

    Waiting on:
  • SP-7
Public interview that strangers can finishL

The question flow exists and a seven-question creation flow with a downloadable brief runs on the site. What is missing is proving a stranger with no context can finish it, and the adaptive questioning the spec asks for. Until then the front door is impressive but unproven.

A builder nobody can start is not a product.

2 of 4 subtasks · 50%

  • Stranger test: no-context visitor completes /build todo
  • Adaptive questions per Part 5 (ladder-driven, fatal first) todo
  • Deterministic interview slice kept (render/interview.py) done
  • Seven-question flow + brief download live (/build skeleton) done
not yet confirmed by the owner

SPEC step 10 per Part 5. render/interview.py derives the next question from Required data + ladder (controls exist). /build skeleton landed 2026-09-11 (seven questions, photography choice, taste sliders, world choice, brief download; build-check 15/15, site-check 12/12). S-05 interview rebuild (anonymous start, six adaptive questions) still parked-phase.

  • render/interview.py — ladder-driven next question
  • marketing/build-check.py — /build flow checks 15/15

Spec: KOZTO-SPEC Part 16 step 10, Part 5

Part 17 acceptance: the ten checks that mean doneL

The spec lists ten sentences that together mean the system is finished — a stranger publishing in six minutes, ten distinct sites, three photoless but presentable, readable by everyone, fast, unbreakable, private, and honest. One holds today. Nine do not.

This is the definition of finished; everything else is commentary.

1 of 10 subtasks · 10%

  • Landing to published site in under six minutes todo
  • Ten sites a designer would not call one system's output todo
  • Three of the ten have no photographs, still presentable todo
  • Every generated site passes WCAG AA + keyboard at 390px todo
  • Full build under 90 seconds within the cost ceiling todo
  • Killing any single agent mid-build still ships todo
  • Publish and rollback instant, neither loses data todo
  • No tenant reads another tenant's data on any path todo
  • The editor never loses work under any failure todo
  • No simulated, mocked or stubbed feature presented as working done
not yet confirmed by the owner

Only the honesty check holds (sim billing stripped, ceca793). WCAG: console surfaces + 10 events pages pass axe/a11y gates; generated-site keyboard nav unproven. Six-minute and 90-second figures unmeasured. Isolation: tenant-scoping controls exist per surface, no whole-system proof.

  • KOZTO-SPEC.md — Part 17, the ten sentences

Spec: KOZTO-SPEC Part 17

Visual regression for generated sitesM

Screenshots that vote: the system photographs its own pages and compares them against approved originals, so a visual break fails the build instead of reaching customers. It works for the console today. Generated sites, the thing customers see, are not covered yet.

Eyes in CI catch what assertions cannot describe.

1 of 4 subtasks · 25%

  • Console baselines cut, human-approved, comparing (4/4) done
  • Golden-site surfaces covered todo
  • CI wiring where a browser and seeded server exist todo
  • Baselines re-approved by a human, never auto-updated todo
not yet confirmed by the owner

T0-01: vrt/console.spec.mjs shipped 2026-09-07 (4 baselines, 4/4 in 5.5s). No BackstopJS/Percy/Argos at this scale, by decision.

  • vrt/console.spec.mjs — console visual baselines
  • vrt/budgets.json — project perf/a11y budgets

Spec: Build-tasks T0-01

Copy lint standing guard over every wordS

A checker reads every customer-facing sentence for banned tones and broken formatting, and it already keeps the site clean. What remains is running its bigger self-hosted engine as a service, so the check stays sharp as the words grow.

Threat-toned copy ships when nobody is watching the words.

2 of 3 subtasks · 67%

  • copycheck live in CI with echoing controls done
  • Banned-tone rules landed done
  • Self-hosted LanguageTool server deployed todo
not yet confirmed by the owner

T0-02: platform/copycheck.py + controls, CI steps; ladder ** balance mirrored; implementation-noun denylist; content/worlds/meta swept clean; proven twice (public API + self-hosted 6.6).

  • platform/copycheck.py — tone + formatting lint

Spec: Build-tasks T0-02

Dependency and secret scanning in CIS

The JavaScript side is clean — the one bad package was removed outright and the audit is empty. The Python side still needs its own audit, plus a scanner that proves no secret ever lands in the code.

A leaked key in history is a breach, not a bug.

1 of 3 subtasks · 33%

  • node-vibrant chain removed; npm audit 0 vulnerabilities done
  • pip-audit or osv-scanner pass in CI todo
  • gitleaks pass in CI todo
not yet confirmed by the owner

T0-04 done 2026-09-07 for npm. Roadmap keeps the node-vibrant pin for the harvest pipeline.

  • package.json — dependency pins

Spec: Build-tasks T0-04

Lighthouse budgets held on every pageM

Every page has speed and quality budgets, and both the console and a generated page currently hold them — two real failures were found and fixed proving the wiring bites. The box stays unchecked until the budgets are a standing gate rather than a good run.

Performance rots silently without a number that fails the build.

3 of 5 subtasks · 60%

  • Budgets committed (perf 90, a11y/SEO/BP 100, LCP/TBT/CLS) done
  • lh-ci wiring + controls in CI done
  • Two caught failures fixed (CSP frame-ancestors, LCP fonts) done
  • Console holds all budgets on every run todo
  • Generated page holds all 7 budgets on every run todo
not yet confirmed by the owner

T0-05: vrt/budgets.json + vrt/lh-ci.sh + vrt/lh-check.py + 3/3 controls; Lighthouse 13.4.1; measured baseline 93/100/100/100. Box unchecked in source — standing gate, not a good run, closes it.

  • vrt/lh-check.py — budget checker
  • vrt/budgets.json — the budgets

Spec: Build-tasks T0-05

HTML validity proven both directionsS

The checker proved itself by finding 58 errors in old output and zero in new output, and it caught a real swallowed-tag regression the same day. It still needs wiring into CI, which needs a machine allowed to fetch its runner.

Invalid HTML breaks assistive technology first.

2 of 3 subtasks · 67%

  • Project-tuned html-validate config shipped done
  • Both-directions proof (58 errors old, 0 new) + regression caught done
  • CI wiring (needs a networked runner for npx fetch) todo
not yet confirmed by the owner

T0-06: .htmlvalidate.json + single-quote normalization; CSP meta stays double-quoted as the one documented exception.

  • .htmlvalidate.json — project-tuned config

Spec: Build-tasks T0-06

Vision-model triage boundariesS

Small vision models were proven to invent things that are not there, so the system only trusts them for coarse signals from a fixed allowlist, enforced by a test. The upgrade path — a flagship vision model under provenance rules — needs a provider key.

A model that invents a business name cannot QA anything.

2 of 3 subtasks · 67%

  • TRIAGE_SIGNALS allowlist enforced in CI done
  • Local-model option closed (out, needs provider key not code) done
  • Hosted-vision upgrade path (needs provider key) blocked
not yet confirmed by the owner

T0-08: moondream invented 'Karoos Web Hosting'; eyevotes-negatives #6 asserts analyze() emits exactly the allowlist. Blocked subtask waits on OW-1's provider key.

  • vrt/eyevotes.py — deterministic signals + allowlist

Spec: Build-tasks T0-08

    Waiting on:
  • Second model provider key (owner)
Deterministic eye-votes for generated pagesM

The console is watched by deterministic image analysis with per-surface budgets, all passing — and the checker once caught a real math bug before it shipped. Generated pages, the surfaces customers see, still need their budgets, plus type census and fill-rate tracking.

The motivating void numbers came from generated pages, not the console.

2 of 6 subtasks · 33%

  • Console surfaces: budgets pass (login, dashboard) done
  • int16-overflow bug caught before shipping done
  • Generated-site surfaces covered (events 85.5%, restaurant 72.4%) todo
  • Per-world budgets todo
  • Type-size census (needs fonttools) todo
  • Image-slot fill rate todo
not yet confirmed by the owner

T0-11: vrt/eyevotes.py byte-identical re-runs, budgets json, ci.sh wiring, 4/4 echoing controls. T0-12's void-injected control is #5.

  • vrt/eyevotes.py — void/palette/edge analysis

Spec: Build-tasks T0-11

Marketing homepage generated by the pipelineM

The front door works and is checked on every build, but it is still hand-written code rather than output of the site generator. A generator that cannot build its own front door has not proven itself. The footer and navigation also need finishing.

Eat your own cooking, or admit the recipe is incomplete.

1 of 3 subtasks · 33%

  • Logged-out root returns the marketing page, checked in CI done
  • Page generated by the render pipeline (still hand-authored) todo
  • Standalone footer and nav beyond the top bar todo
not yet confirmed by the owner

P1-01 partial 2026-09-07: live site count from DB, pricing from billing, categories from spines; tests 147, negatives 47/47. Done-when (root 200, not 302) holds.

  • platform/server.py — hand-authored homepage (the problem)
  • marketing/site-check.py — marketing checks 12/12

Spec: Build-tasks P1-01

Pricing page with the settled tiersS

Prices show live from the billing system on the homepage and their own page. But the business only has two tiers, and the page needs three plus a trial marker — those do not exist yet, and inventing them would be fabrication. This waits on your pricing decision.

A price the business never set is a lie with a checkout button.

1 of 3 subtasks · 33%

  • Live plan catalogue on home + /pricing from billing sources done
  • Third tier exists in the billing layer (owner decision) blocked
  • Trial marker exists in the billing layer (owner decision) blocked
not yet confirmed by the owner

P1-03 partial 2026-09-08: per-build price from the ledger debit, plain-terms translation per request; tests 218, negatives 60/60. Blocked subtasks are OW-2's pricing decision.

  • billing/ledger.py — plan catalogue + per-build debit

Spec: Build-tasks P1-03

    Waiting on:
  • Pricing decision: third tier + trial (owner)
Case-study proofs, all realS

Four real example sites with live links and an honesty header saying they are demonstrations, not customer stories. Still missing: a home-services example whose deploy 404s, an automatic check that links stay alive, and customer results — which do not exist, so none are claimed.

One invented testimonial ends trust permanently.

2 of 5 subtasks · 40%

  • Four verified live deploys with honesty header done
  • Unpublished entries degrade to labelled holes done
  • Home-services deploy actually live (404s today) todo
  • Automated link-health check todo
  • Customer results cited (none exist — nothing invented) todo
not yet confirmed by the owner

P1-05 partial 2026-09-07: curated SHOWCASE list; tests 160, negatives 49/49. Home-services needs a real deploy (Phase C).

  • marketing/site-check.py — marketing checks incl. /work

Spec: Build-tasks P1-05

Image slots enforced end to end at generationM

Uploaded photos already obey the slot contract — ratios, subject points, labelled holes. What remains is the same enforcement when the generator itself picks crops and focal points, rather than a human uploading. The tenant path is proven; the generation path is not.

Two paths, one contract, or the contract is decoration.

2 of 3 subtasks · 67%

  • Tenant uploads obey ratios + focal roles (step 6) done
  • Holes render as labelled fixable gaps, never silent done
  • Generation-time crop and focal selection follows the contract todo
not yet confirmed by the owner

P2-02: component-declared ratios and focal roles. Step 6 (IM-6) closed the tenant side; generation-side selection is NX-2/T1-adjacent work.

  • render/imagery.py — crop derivation from records
  • render/assets.py — slot contract audit

Spec: Build-tasks P2-02

Uploads warn about bad photos, not just wrong shapesS

Uploads already warn when the shape does not fit the slot, and that warning once caught real confusion. Still missing: warnings for dark, blurry or tiny photos — which need pixel reading the offline build does not have.

A blurry photo that ships is a support ticket with your name on it.

2 of 3 subtasks · 67%

  • Shape mismatch warns at upload time next to the photo done
  • Small-size warning done
  • Darkness and blur warnings (need pixel decoders) todo
not yet confirmed by the owner

P2-03 partial 2026-09-08: save_asset(..., want=) surfaces build refusal at upload; tests 220, negatives 61/61. Pillow is not in offline CI.

  • platform/generate.py — save_asset shape warning

Spec: Build-tasks P2-03

Logo, background removal, business-profile importM

Uploading several photos at once already files them into the right empty slots by shape. Still missing: spotting logos, cutting out backgrounds, and pulling a business's existing Google photos — the first two need pixel tools, the last needs a key.

Tenants arrive with messy photo folders, not studio shoots.

1 of 4 subtasks · 25%

  • Folder-of-photos batch upload files by shape done
  • Logo detection at upload todo
  • Background removal at upload todo
  • Google Business Profile import (needs API key) todo
not yet confirmed by the owner

P2-04 partial 2026-09-08: match_photos pure, one photo per slot, unmatched refused with shapes named; tests 223, negatives 63/63. GBP import is OW-1-adjacent (key).

  • platform/generate.py — match_photos batch filing

Spec: Build-tasks P2-04

One-click fixes for every holeM

Every site already shows its health — missing answers, missing photos, build state — with links to the right screen. What remains is fixing holes in one click from the health screen itself, which needs the photo search and generation seams first.

A hole you can see but not fix is a chore list, not a product.

2 of 3 subtasks · 67%

  • Site-health core with deep links (cards + workspace rollup) done
  • Honest directions page (three-ways claim fixed) done
  • One-click fixes beyond upload links (needs stock/generate) todo
not yet confirmed by the owner

P2-07 partial 2026-09-08: GEN.site_health(), per-card lines, workspace metrics strip; tests 219, negatives 60/60. One-click waits on NX-2.

  • platform/generate.py — site_health() + workspace rollup

Spec: Build-tasks P2-07

    Waiting on:
  • NX-2
Question flow with motion and smartsM

Questions already come one screen at a time with progress, hints and saving halfway through. Still missing: the animated transitions between them, and questions that narrow based on earlier answers — the ladder has no conditional questions yet, so there is nothing true to show.

Onboarding is seduction; a form is paperwork.

1 of 3 subtasks · 33%

  • One-per-screen flow with save-and-resume + encouragement done
  • Animated transitions (needs a console JS story) todo
  • Conditional narrowing from earlier answers (needs ladder support) todo
not yet confirmed by the owner

P3-01 partial 2026-09-08: tests 184, strings asserted in the real page (no clean mid-interview demo site for a screenshot).

  • platform/server.py — interview flow screens

Spec: Build-tasks P3-01

Rewrite every word the customer readsM

The lint already bans threatening tones and caught a real bug where the ban itself unfataled every recipe. What remains is the actual rewrite — every question, hint and error in a human voice — which needs a writer, not a linter.

A lint keeps words legal; only a writer makes them kind.

2 of 3 subtasks · 67%

  • Banned-tone lint + explicit fatal cells (K-034) done
  • Six threat-toned hints rewritten done
  • Full writer rewrite of every question, hint and error todo
not yet confirmed by the owner

P3-02 partial 2026-09-08: copycheck check_tone 8/8 controls; ladder fatality now explicit; mutants 85/85. Remainder needs a writer (owner-adjacent).

  • platform/copycheck.py — tone lint that guards the rewrite

Spec: Build-tasks P3-02

Import from everywhere the business already isL

Photo folders and website links already import through hardened, tested paths. Still missing: Google profiles, social presence, PDF menus and product spreadsheets — each needs an external service or parser the offline build does not have.

Retyping forty products is where onboarding dies.

2 of 6 subtasks · 33%

  • Photo-folder batch import (tenant side) done
  • URL title/description/image import, SSRF-hardened (stranger side) done
  • Google Business Profile import todo
  • Social presence import todo
  • PDF menu and brochure import todo
  • Product spreadsheet import todo
not yet confirmed by the owner

P3-03 partial 2026-09-08: dashboard URL import shares the hardened fetcher parse layer (facts_from_html); both sides screenshot-verified.

  • platform/generate.py — public_fetch + facts_from_html

Spec: Build-tasks P3-03

Direction picker as the reveal momentM

Each design already shows its real section lineup with every slot marked filled or missing, plus the exact photo shopping list to finish it. Still missing: full-bleed renders instead of lineups, and a single switchable view — both need structured data most tenants have not supplied yet.

The reveal moment sells the product; a lineup is a packing list.

2 of 4 subtasks · 50%

  • Bound lineups with filled/missing slots per world done
  • Photo shopping list per direction done
  • Full-bleed renders (needs structured tenant data) todo
  • Switchable single view todo
not yet confirmed by the owner

P3-05 partial 2026-09-08: _direction_preview shares the demo_preview layer so tenant and stranger can never disagree; proven by the P1-02 0/15 sweep that full renders need a real interview. Tests 217.

  • platform/generate.py — direction previews + shopping lists

Spec: Build-tasks P3-05

Dashboard that knows the tradeM

The dashboard already shows real numbers, a launch checklist, next actions and site health worst-first. Still missing: headlines written for the specific trade — which needs per-category event data that does not exist yet.

A generic dashboard is a mirror; a trade-aware one is an advisor.

2 of 3 subtasks · 67%

  • Metrics strip, checklist, next-action callout, activity feed done
  • Site-health panel worst-first with deep links done
  • Recipe-driven per-category headline (needs event data) todo
not yet confirmed by the owner

P4-02 shipped 2026-09-07 + health 2026-09-08; tests 232. Publish + rollback audited (site.publish/site.rollback rows).

  • platform/server.py — dashboard + health panel

Spec: Build-tasks P4-02

Site cards with real photographic thumbnailsS

Site cards already show status, category and actions on desktop and mobile. Still missing: photographic thumbnails of each site, which need a thumbnail service, and real business names for signups instead of demo-seed naming.

Cards without pictures are a list wearing a costume.

1 of 3 subtasks · 33%

  • Card grid with chips, counts, inline publish done
  • Photographic thumbnails (needs thumbnail service / Browser Rendering) todo
  • Real signup naming (demo-seed naming only in seeds) todo
not yet confirmed by the owner

P4-03 shipped 2026-09-07, screenshot-verified desktop + mobile, test-covered. Thumbnails wait on Phase C.

  • platform/server.py — site card grid

Spec: Build-tasks P4-03

Settings beyond SEOL

The SEO machinery is done — script-free business markup in every page head, wired into real builds. Everything else in settings is still greenfield: domains, search preview, languages, integrations, team, sessions and the danger zone.

Settings is where trust is configured or lost.

1 of 4 subtasks · 25%

  • Microdata SEO layer in every built page + maybe_seo wiring done
  • Domain, DNS and TLS todo
  • Search-preview UI, localisation, integrations, forms routing todo
  • Team and invites, sessions list, danger zone todo
not yet confirmed by the owner

P4-05 partial 2026-09-08: JSON-LD rejected (script-src 'none' forbids scripts); seo controls 10/10; tests 188. Canonical from url_live, skip-not-guess without one.

  • render/seo.py — SEO layer in every built page head

Spec: Build-tasks P4-05

Twelve genuinely distinct worldsL

Four worlds exist and each new one must prove it is genuinely different, not a recolour. Twelve are needed before any launch claim, each passing the validators and each authored against a real business with real photography.

Three worlds cannot produce visual variety; the interview ends up doing design's job.

4 of 12 distinct worlds in catalogue

not yet confirmed by the owner

P6-01: distinct in axes not recolours; seeded structural variation designed into the catalogue FIRST (spec Part 10). Worlds today: obsidian, ledger, bloom, luminous. Queue earns recipes; worlds are authored, not queued.

  • worlds/validate.py — contrast/gamut/coherence validators

Spec: Build-tasks P6-01; KOZTO-SPEC Part 19

VAT and international tax, signed offM

The tax position is written down with the questions for an accountant, and the order is clear: country first, then tax lines, then currencies. Nothing starts before sign-off, because invented tax rates would be fabrication.

Wrong tax is a liability the business carries, not us — unless we invented it.

1 of 4 subtasks · 25%

  • Position + accountant question list in JURISDICTION.md done
  • Accountant sign-off (owner commission) blocked
  • Country field, tax lines, multi-currency in code order todo
  • Referral and affiliate mechanism todo
not yet confirmed by the owner

P7-03 partial 2026-09-08: facts today are sim processor, no tax, no country field. Multi-currency needs P2-04 locale infra. Note: LIBRARY-v3 Part 8 now governs the tenant-facing VAT truth — reconcile with this position at build time.

  • JURISDICTION.md — tax position + accountant question list

Spec: Build-tasks P7-03; LIBRARY-v3 Part 8

    Waiting on:
  • Accountant review (owner)
Human accessibility auditM

Two automated checkers already scan every console surface and every generated test page with zero violations — and they have caught real bugs. What remains is the human checklist no script can check: real keyboard journeys, real screen readers, real judgment.

Automated checks catch markup; humans catch meaning.

3 of 4 subtasks · 75%

  • Deterministic token-contrast gate in CI (7 pairs) done
  • axe-core scan over 8 surfaces, 0 violations done
  • Generated test pages pass both gates done
  • Human-checklist audit commissioned todo
not yet confirmed by the owner

P8-01 partial 2026-09-08: a11y.py + axe-ci.mjs + controls proving the scanner bites; first run caught an unlabeled field. Generated pages enumerated from site-audit output.

  • platform/a11y.py — token-contrast gate
  • vrt/axe-ci.mjs — axe-core scan

Spec: Build-tasks P8-01

Model-spend caps wired to a live pathM

The cap machinery is armed — one function for levels, warnings before limits, daily alerts to owners. But no model call has ever passed through it, so the gate guards nothing yet. It gets wired the moment the first live call flows.

An unwired cap is a hope, not a limit.

1 of 3 subtasks · 33%

  • caps.status() single source + pre-cap alerts + daily cadence done
  • Wiring into a live model path (no model has ever run through it) todo
  • Cost-model output metering from the first live call todo
not yet confirmed by the owner

P8-03 partial 2026-09-08: tests 146 billing, negatives 72/72, mutants 87/87. Mail/uplink via outbox; ownerless scopes audit and say so.

  • billing/ledger.py — spend caps + alerting

Spec: Build-tasks P8-03

Scale truths at 100, 1,000 and 10,000 tenantsM

The read path was measured at three sizes and the first break — unindexed scans — was fixed with fourteen indexes, proven flat with plan assertions guarding it. Documented next: one shared database connection blocking writers, no background job queue, and operator-wide scans needing their own pass.

Scale breaks where measurement stops.

2 of 5 subtasks · 40%

  • Read path measured 100/1k/10k; FK indexes shipped (migration 0012) done
  • EXPLAIN-plan regression guards on hot reads done
  • Connection-per-request + WAL before cutover (durability review) todo
  • Background job queue (builds run inline today) todo
  • Operator-wide scan pass (admin search) todo
not yet confirmed by the owner

P8-04 partial 2026-09-08: inbox unread 0.03→0.81ms linear, now 0.02ms flat (40x at 10k); tests 233. Margin-per-tier and support load need real traffic.

  • platform/migrations/0012_indexes.sql — the fourteen FK indexes

Spec: Build-tasks P8-04

Model routing on observed quality, second sourceM

Routing already picks providers from a config table of measured results, and the first live scoreboard sends two tasks to the fast model and the rest to the safe stub. What remains is measuring a second provider — one source is an anecdote, two is routing.

Single-provider routing is a preference with a table.

2 of 3 subtasks · 67%

  • Routing table as config (K-010) + prompt versioning + stamps done
  • First live board recorded (spark: copy 5/5, comprehend 3/4, director 4/4, revise 2/3) done
  • Second provider measured via routing.py record (needs key) blocked
not yet confirmed by the owner

P10-02: render/routing.py + routing.json; route() picks best observed pass/total at current PROMPT_VERSION; empty board routes stub out loud. Recording half: PROMPT_VERSION=1, site.json provenance, job manifests stamped. Honest gap: worker stamp lines have no mutation owner.

  • render/routing.py — observed-quality router

Spec: Build-tasks P10-02; K-010

    Waiting on:
  • Second model provider key (owner)
Recipe order steps 2–11: categories to legal layerL

After the shared product contract and base shop come the categories that break weak abstractions — food, clothing and digital first — then the services, tax truth, legal pages, fulfilment, failure states, feeds and the long tail. Each with its own tests that prove refusal refuses.

The recipe book is a specification; this is its construction schedule.

0 of 9 subtasks · 0%

  • Three categories end to end: food-pantry, apparel, digital-goods todo
  • service-base + classes-workshops (calendar, capacity, cancellation) todo
  • Ladder runner: product-level and booking-level refusals todo
  • Tax profile + VAT status (unregistered shown 'inc VAT' is a false statement) todo
  • Legal layer: GB templates, slot-refusal, one solicitor review todo
  • Order lifecycle + fulfilment incl. collection and local delivery todo
  • Checkout failure states (eleven defined behaviours) todo
  • Feeds + commerce manifest before the category long tail todo
  • Remaining categories in demand order + per-category negatives todo
not yet confirmed by the owner

LIBRARY-v3 Part 18.1 order steps 2–10 (step 11 compliance gates fold into each). Gated by RC-1's contract. Interview consequence: top-six refusal-first questions per detected recipe.

  • recipes/LIBRARY-v3.md — the order, Part 18.1

Spec: LIBRARY-v3 Part 18

    Waiting on:
  • RC-1
Background removal that runs on CPU with no keysM

Tenant uploads with messy backgrounds get clean cutouts — proven working on a normal computer in 26 seconds a photo, no paid services. Before it touches real shops it must prove itself on 20 typical photos, and a person signs off, not a script.

Product photos on grey bedroom backgrounds do not sell.

0 of 3 subtasks · 0%

  • Pin u2net explicitly (2.x defaults to a 1GB surprise download) todo
  • Evaluate on 20 tenant-style photos: logos, products, faces todo
  • Human approval before production, non-optional todo
not yet confirmed by the owner

T0-09 open: proven operational (900x700 to RGBA in 26s CPU-only, zero keys), never productised.

  • docs/archive/PLATFORM-BUILD-TASKS.md — T0-09 definition

Spec: Build-tasks T0-09

Stock photos that match the business, licensed properlyM

Shops without their own photos get good stock matched to their category, colours and section — with the photographer's credit rendered wherever the licence demands it. Needs stock-service keys first.

A builder that cannot fill a photo slot cannot finish a site.

0 of 3 subtasks · 0%

  • Fact-to-query builder: category, world adjectives, intent, palette todo
  • Automatic aesthetic/palette filter on results todo
  • Attribution rendered where licences require it todo
not yet confirmed by the owner

P2-01 open. Step-6 T1-1 built the live-stock lookup path; this is the full integration. Stock API keys via OW-1.

  • docs/archive/PLATFORM-BUILD-TASKS.md — P2-01 definition

Spec: Build-tasks P2-01

    Waiting on:
  • Stock API keys (see OW-1)
AI-generated images with provenance labelsM

Where no photo exists at all, the builder can generate one from the business facts and the design world's palette — the tenant picks one of three, every one labelled as machine-made, every one costed against the budget.

A missing hero image blocks a launch; a labelled generated one unblocks it.

0 of 3 subtasks · 0%

  • Prompt from facts + world axes + palette todo
  • Generate-three-pick-one flow todo
  • Costed against the cost model; provenance-labelled AI todo
not yet confirmed by the owner

P2-05 open. Generation provider path is CF-1 C7 (Workers AI); cost hooks are S-09.

  • docs/archive/PLATFORM-BUILD-TASKS.md — P2-05 definition

Spec: Build-tasks P2-05

Asset library as a real productL

Every photo a shop owns lives in one searchable library — editable descriptions, warnings before deleting a photo a page uses, licence records, and fast modern formats served from the CDN.

Uploads without a library become a pile nobody dares delete.

0 of 5 subtasks · 0%

  • Search over tenant assets todo
  • Auto alt text, editable and validated todo
  • Usage tracking with delete warnings todo
  • Licence records per asset todo
  • CDN delivery, responsive sizes, modern formats todo
not yet confirmed by the owner

P2-06 open. Upload + batch-match paths exist (P2-04); this is the library around them.

  • platform/server.py — current upload/asset endpoints
  • docs/archive/PLATFORM-BUILD-TASKS.md — P2-06 definition

Spec: Build-tasks P2-06

Provenance on every word, quarantine for samplesM

Every sentence on a built site carries its origin — tenant-written, imported, sample, machine-drafted or template — sample content renders visibly watermarked, and publishing stays blocked until provenance is clean. Praise without a name never ships.

A site that cannot say where its words came from cannot be trusted.

0 of 4 subtasks · 0%

  • Origin on every string: user/import/sample/model/template todo
  • Samples render watermarked todo
  • Publish blocked until provenance is clean todo
  • Nameless praise gets a real name or does not ship todo
not yet confirmed by the owner

P3-04 open. Model stamping exists (P10-02 provenance); this extends it to every origin.

  • render/prompt.py — PROMPT_VERSION + stamp contract
  • docs/archive/PLATFORM-BUILD-TASKS.md — P3-04 definition

Spec: Build-tasks P3-04

Console rebuilt on the shared design tokensL

The dashboard, inbox, billing and settings all speak the same visual language as the sites they manage — today's beige browser-default chrome deleted, and every gap found along the way (tables, charts, notifications) becomes a proper authored component.

A console that looks nothing like its product teaches distrust.

0 of 3 subtasks · 0%

  • Beige system-ui chrome deleted; tokens shared with worlds todo
  • Dashboard, inbox, billing, settings on the system todo
  • Gaps (tables, charts, toasts, palettes) become authored console components todo
not yet confirmed by the owner

P4-01 open. P6-03 (thirty console components) is the component half of the same work.

  • platform/server.py — console surfaces today
  • docs/archive/PLATFORM-BUILD-TASKS.md — P4-01 definition

Spec: Build-tasks P4-01

Editor with section tree, live preview, device toggleL

Shop owners rearrange their site in a real editor — the left panel shows the true page structure, the middle previews instantly, the right shows only the selected element's genuine fields with character counts before limits hit.

Without an editor every change is a support ticket.

0 of 4 subtasks · 0%

  • Left panel reflects the real IR todo
  • Centre previews instantly todo
  • Right panel: selected element's genuine fields only, pre-limit counters todo
  • Device toggle todo
not yet confirmed by the owner

P5-01 open. IR work starts early per the phase plan; device-toggle precedent exists on the reveal page (P3-06).

  • docs/archive/PLATFORM-BUILD-TASKS.md — P5-01 definition

Spec: Build-tasks P5-01

Drafts, never live, with instant rollbackM

Owners edit a draft while visitors see the published site — a persistent indicator says which is which, one click publishes, and a complete rollback is instant when a change misfires.

Editing the live site is how shops break at lunchtime.

0 of 3 subtasks · 0%

  • Persistent draft-vs-published indicator todo
  • One-click publish todo
  • Instant complete rollback todo
not yet confirmed by the owner

P5-02 open. Publish/rollback audit rows exist (P4-02 site.publish/site.rollback); this is the draft layer above them.

  • docs/archive/PLATFORM-BUILD-TASKS.md — P5-02 definition

Spec: Build-tasks P5-02

Regenerate with diff-before-applyM

Owners can ask the machine to redo a section, a page or the whole site — and see exactly what would change before accepting. Hand edits are never silently destroyed by a regeneration.

Regeneration without a diff is a slot machine that eats work.

0 of 3 subtasks · 0%

  • Regenerate at section, page and site level todo
  • Diff shown before apply todo
  • Manual edits never silently destroyed todo
not yet confirmed by the owner

P5-03 open. Revise-path provider contract (call()) exists via T0-07; this is the product surface.

  • docs/archive/PLATFORM-BUILD-TASKS.md — P5-03 definition

Spec: Build-tasks P5-03

Conversational editing that cannot break structureL

Owners type 'make the hero bigger' and get a previewed change to accept or reject — and the validation layer proves the machine cannot produce a structurally broken page no matter what it is asked.

Natural-language editing without validation is a footgun.

0 of 2 subtasks · 0%

  • Plain-language ask to validated patch to preview to accept/reject todo
  • Validation layer proves structurally-broken output impossible todo
not yet confirmed by the owner

P5-04 open. Builds on P5-03 diff flow + the bind gate (render/gate.py).

  • docs/archive/PLATFORM-BUILD-TASKS.md — P5-04 definition

Spec: Build-tasks P5-04

World switcher: your content in every design, in secondsM

Owners see their actual words and photos live in every design world and switch between them in seconds — the moment the product's generative bet becomes tangible.

Switching designs is the demo that sells the product.

0 of 2 subtasks · 0%

  • Tenant's actual content rendered live in every world todo
  • Switchable in seconds todo
not yet confirmed by the owner

P5-05 open. Direction previews (P3-05) are the static precedent; this is the live switch.

  • docs/archive/PLATFORM-BUILD-TASKS.md — P5-05 definition

Spec: Build-tasks P5-05

Ten interactive blocks on a tiny framework-free runtimeL

Menus, galleries, validated forms, maps, booking, restaurant menus, schedules, testimonials, pricing toggles and announcement bars — each working without JavaScript first and enhanced with a little, all clean under the strict content policy.

Static sections cannot take bookings or enquiries.

0 of 2 subtasks · 0%

  • Island runtime: tiny, framework-free, progressively enhanced, CSP-clean todo
  • Ten blocks: menus, galleries, forms, maps, booking, menus, schedules, testimonials, pricing, bars todo
not yet confirmed by the owner

P6-02 open. Console is zero-JS by CSP (P4-08); the island runtime is the tenant-side answer.

  • docs/archive/PLATFORM-BUILD-TASKS.md — P6-02 definition

Spec: Build-tasks P6-02

Thirty console components re-authored against our tokensL

Tables, charts, drawers, pickers, palettes and the rest — each studied from outside references but rebuilt against Kozto tokens. Outside libraries feed authoring only; they never enter a generation prompt.

P4-01's console rebuild has nothing to build with until these exist.

0 of 2 subtasks · 0%

  • Thirty components re-authored from references against our tokens todo
  • Resource library feeds authoring, never generation prompts todo
not yet confirmed by the owner

P6-03 open. Golden queue assigns component work (golden/queue.py); build only what it asks for.

  • golden/queue.py — component ask, in rank order
  • docs/archive/PLATFORM-BUILD-TASKS.md — P6-03 definition

Spec: Build-tasks P6-03

Worlds' motion grammar made executableM

Each design world declares how it moves — easing, durations, stagger — and the island runtime reads it, so switching worlds changes how the site moves, not just how it looks.

Motion is half of a design language; static tokens are the other half.

0 of 3 subtasks · 0%

  • Easing/durations/stagger declared per world todo
  • Island runtime reads the active world's grammar todo
  • Switching worlds changes motion todo
not yet confirmed by the owner

P6-04 open. Blocked in practice on P6-02 (no runtime yet).

  • docs/archive/PLATFORM-BUILD-TASKS.md — P6-04 definition

Spec: Build-tasks P6-04

    Waiting on:
  • P6-02 island runtime
Every addition ships meta, evals and negativesS

No component, world or check lands without the metadata that describes it and a control proving the check can fail — echoing what it caught. This is the standing rule that keeps the whole catalogue honest.

Rule 1 of the project: a validator that has never failed is not known to work.

0 of 1 subtasks · 0%

  • Meta + evals + negatives with every addition todo
not yet confirmed by the owner

P6-05 open as a checklist item; practised throughout (mutate.py 154/154, golden queue discipline).

  • docs/archive/PLATFORM-BUILD-TASKS.md — P6-05 definition

Spec: Build-tasks P6-05

Starting-points gallery of living instances, not templatesL

New owners start from real generated sites — pinned world-by-category-by-seed outputs, one click to launch, every one switchable and regenerable with its seed labelled. Frozen hand-made templates are banned: they rot. Pasting a liked URL extracts style signals, never content.

Starting from blank loses strangers; starting from frozen lies to them.

0 of 4 subtasks · 0%

  • Pinned generated instances, one-click launch, seed labelled todo
  • Every item world-switchable and regenerable todo
  • Reference-URL extraction: signals, never content todo
  • Sample businesses as quarantined, watermarked instances todo
not yet confirmed by the owner

P9-01 open, redefined by the gap essay section 1 (frozen templates banned). Needs P5-05 switching + P3-04 quarantine.

  • docs/archive/PLATFORM-GAP-ESSAY.md — gallery redefinition, section 1
  • docs/archive/PLATFORM-BUILD-TASKS.md — P9-01 definition

Spec: Build-tasks P9-01

Programmatic SEO generated by the engineM

Category, location and competitor-alternative pages generated by the pipeline itself — real pages for real searches, not a sitemap of thin content.

Discovery compounds; paid acquisition does not.

0 of 3 subtasks · 0%

  • Category pages generated by the engine todo
  • Location pages generated by the engine todo
  • Competitor-alternative pages todo
not yet confirmed by the owner

P9-02 open. From week 12 per the phase plan; needs generation at scale first.

  • docs/archive/PLATFORM-BUILD-TASKS.md — P9-02 definition

Spec: Build-tasks P9-02

Retention surfaces: reviews sync, domains, tenant emailL

Business-profile sync, domain registration and email on the tenant's own domain — the surfaces that stop the product donating its value to other companies.

Every review or email on someone else's domain is churn fuel.

0 of 3 subtasks · 0%

  • GBP sync todo
  • Domain registration todo
  • Email on tenant domain todo
not yet confirmed by the owner

P9-03 open. Mail/domain halves live in CF-1 (C5/C6).

  • docs/archive/PLATFORM-BUILD-TASKS.md — P9-03 definition

Spec: Build-tasks P9-03

Public API v1 with webhooks docsM

Documented keys, scopes and rate limits plus webhook docs — and marketplace foundations with attribution and revenue share on catalogue metadata.

A platform without an API is a cul-de-sac.

0 of 2 subtasks · 0%

  • Keys, scopes, rate limits + webhooks docs todo
  • Marketplace foundations: attribution + revenue share todo
not yet confirmed by the owner

P9-04 open. Internal JSON API surface exists (OPENAPI.md); this is the public v1.

  • OPENAPI.md — internal API surface today
  • docs/archive/PLATFORM-BUILD-TASKS.md — P9-04 definition

Spec: Build-tasks P9-04

Dogfood loop: marketing site regenerates every releaseS

Every release regenerates Kozto's own marketing site through the pipeline — and a failure there blocks the launch. The product eats what it serves.

A generator the team will not run on itself is not ready.

0 of 2 subtasks · 0%

  • Marketing site regenerates through the pipeline per release todo
  • Failures there are launch-blockers todo
not yet confirmed by the owner

P10-03 open. P1-01 (hand-authored homepage) is what the loop replaces.

  • docs/archive/PLATFORM-BUILD-TASKS.md — P10-03 definition

Spec: Build-tasks P10-03

    Waiting on:
  • P1-01 pipeline-generated homepage
Build Brief artifact per buildM

Every build starts from one document with a fixed structure, produced by a single top-tier model call with a cached prefix — validated, retried, then template-fallback, and stored with its inputs, seed, model and cost.

Reproducibility starts with knowing what was asked.

0 of 4 subtasks · 0%

  • One Markdown doc per build, exact section-2.2 structure todo
  • Single tier-A call, cached prefix byte-identical downstream todo
  • Brief validation with retry-then-template-fallback todo
  • Stored with inputs, seed, model, cost todo
not yet confirmed by the owner

S-01 open. K-009 zones as implementation.

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-01 definition

Spec: Build-tasks S-01

Module versioning: immutable published versionsM

Every configuration pins its module at a version; published versions can never change or disappear — deprecation warns but never removes. Template-conformance joins the existing gate while the readable typechecked source stays readable.

A site that changes under its owner was never theirs.

0 of 3 subtasks · 0%

  • Configs pin id@version; published versions immutable todo
  • Deprecation never removes todo
  • Template-conformance joins the existing gate todo
not yet confirmed by the owner

S-02 open.

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-02 definition

Spec: Build-tasks S-02

Reproducibility schema: seed, inputs, instant rollbackM

Versions and briefs live in the database, every build stores its seed and inputs, and publishing is a pointer swap — so rollback is instant.

Rollback that takes an hour is not rollback.

0 of 3 subtasks · 0%

  • Versions + briefs tables todo
  • Seed + inputs stored per build todo
  • Publish as pointer swap, rollback instant (K-006 schema) todo
not yet confirmed by the owner

S-03 open. Generalises the P5-02 draft story to the data layer.

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-03 definition

Spec: Build-tasks S-03

Interview rebuild: anonymous start, six adaptive questionsL

Owners start without an account, get a magic link later, answer six adaptive questions with thin honest progress, every field saved as typed — including palette and reference questions, with machine guesses flagged as guesses.

Signup walls lose the strangers the demo just won.

0 of 3 subtasks · 0%

  • Anonymous start, deferred magic-link account todo
  • Six adaptive questions, thin-rule progress, per-field persistence todo
  • Palette + reference question types, inferred:true flags todo
not yet confirmed by the owner

S-05 open. Replaces the current ladder flow (P3-01) once designed.

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-05 definition

Spec: Build-tasks S-05

Build screen: stage checklist, live iframe, honest slownessM

While the site builds, owners see a stage checklist and a live iframe — no spinner, no fake percentage, honest lines about slowness, one button after a beat. Closing the tab never kills the build; at 180 seconds it ships degraded, honestly.

Fake progress destroys more trust than slowness.

0 of 4 subtasks · 0%

  • Stage checklist + live iframe, no spinner/percentage todo
  • Honest slowness lines, sit-a-beat then one button todo
  • Build continues server-side on tab close todo
  • 180s hard timeout ships degraded todo
not yet confirmed by the owner

S-06 open. Extends the P3-06 progress/reveal screens.

  • marketing/build-check.py — current build-flow checks
  • docs/archive/PLATFORM-BUILD-TASKS.md — S-06 definition

Spec: Build-tasks S-06

Editor intents, version history, publish pre-flightL

A prompt bar with four routed intents, double-click inline text editing, an 8-second undo chip and safe concurrent edits — plus a restorable version list with generated summaries, and a pre-flight that warns without ever blocking.

Power without undo is a trap; gates without override are a wall.

0 of 4 subtasks · 0%

  • Prompt bar, four routed intents; double-click inline editing todo
  • 8-second undo chip; DO-serialised concurrent edits todo
  • Restorable version list with generated summaries todo
  • Pre-flight runs real checks as warnings, never blockers todo
not yet confirmed by the owner

S-07 open. The editor half needs P5; the pre-flight half extends render/qa.py.

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-07 definition

Spec: Build-tasks S-07

Runtime packs: booking, menus, enquiries and moreL

Booking, menus, enquiries, portfolios, loyalty and newsletters ship as packs under their own directory — each declaring its modules, data, routes, integrations and plan tier — while the knowledge layer stays untouched.

Real shops need bookings and menus, not just pages.

0 of 3 subtasks · 0%

  • Packs for booking, menu, enquiries, portfolio, loyalty, newsletter todo
  • Each declares modules, collections, routes, integrations, plan tier todo
  • Knowledge skills/ untouched todo
not yet confirmed by the owner

S-08 open. Runs on the P6-02 island runtime once it exists.

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-08 definition

Spec: Build-tasks S-08

    Waiting on:
  • P6-02 island runtime
Cost architecture: per-build ledger, enforced budgetsM

Every build records its tokens by tier, cache-hit rate, images, wall clock and degradations — and a cache-hit rate under 80% is filed as a bug. A $0.40/90-second budget is enforced, with top-tier calls never silently downgraded.

Unmetered generation is how margins die quietly.

0 of 3 subtasks · 0%

  • Tier mapping in config (A never downgraded) todo
  • Per-build ledger: tokens, cache-hit, images, clock, degradations todo
  • Cache-hit below 80% is a bug; $0.40/90s budget enforced todo
not yet confirmed by the owner

S-09 open. Tenant/global caps precedent exists (P8-03); this is per-build metering.

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-09 definition

Spec: Build-tasks S-09

Provider failover and prompt-injection hardeningL

Two providers per tier with a circuit breaker and a recorded serving provider — and customer input delimiter-wrapped at all three model contact points, uploads sanitised, tenant-scoped agent calls, rate limits honoured.

One provider is a single point of failure; raw input is an open door.

0 of 3 subtasks · 0%

  • Two providers per tier, circuit breaker, recorded serving provider todo
  • Delimiter-wrapped customer input at all three model contact points todo
  • SVG sanitisation on upload; tenant-scoped agent calls; 429 with Retry-After todo
not yet confirmed by the owner

S-10 open. Seam exists (T0-07); routing exists (P10-02); this is failover + hardening.

  • render/routing.py — routing table to extend with failover
  • docs/archive/PLATFORM-BUILD-TASKS.md — S-10 definition

Spec: Build-tasks S-10

Platform UI to the Studio specM

The console design system — palette, type scale, grid, motion rules and the forbidden-tells list — with the two commercial typefaces licensed first and fallbacks specified.

The console cannot outgrow its typefaces.

0 of 2 subtasks · 0%

  • Palette, type scale, grid, motion rules, forbidden-tells as console system todo
  • Commercial face licensing resolved first, fallbacks specified todo
not yet confirmed by the owner

S-11 open. Precedes P4-01 in practice (the system comes before the rebuild).

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-11 definition

Spec: Build-tasks S-11

Release gate: six minutes, ten briefs, four drillsM

Landing to publish in six minutes; a weekly ten-brief designer test from seed-system day; accessibility conformance; and three resilience drills — kill any agent, prove isolation, prove no lost work.

A release process that cannot prove itself cannot be trusted.

0 of 3 subtasks · 0%

  • Six minutes landing-to-publish todo
  • Ten-brief designer test weekly from seed-system day todo
  • WCAG AA; kill-any-agent drill; isolation proof; no-lost-work proof todo
not yet confirmed by the owner

S-12 open. Dogfood loop (P10-03) is its canary.

  • docs/archive/PLATFORM-BUILD-TASKS.md — S-12 definition

Spec: Build-tasks S-12

Auth, sessions and abuse protection on the edgeM

Sessions in the edge store with existing expiry behaviour, rate limits on login, signup and enquiry, bot verification on signup and public forms, and managed firewall rules plus API shielding.

The stranger-facing front door gets attacked first.

0 of 4 subtasks · 0%

  • Sessions in KV/D1 with existing TTLs todo
  • Rate limits on login/signup/enquiry todo
  • Turnstile on signup + public forms todo
  • WAF managed rules + API Shield todo
not yet confirmed by the owner

C2 open. Turnstile skill available; needs the C0 account first.

  • docs/archive/PLATFORM-BUILD-TASKS.md — C2 definition

Spec: Build-tasks C2

    Waiting on:
  • C0 account project
Generation on queues: durable stages, live progressL

The Generate button enqueues instead of blocking; a durable workflow runs the stages with retries; a per-site object streams progress over sockets — today's progress screen as architecture. Heavy Python stays in containers, off the request path, and every consumer is idempotent like the ledger.

Inline generation falls over exactly when it matters.

0 of 5 subtasks · 0%

  • Queue producer + Progress DO declared (broadcast contract unverified) doing
  • Generate button enqueues; Workflow runs stages durably with retries todo
  • Per-site DO fans progress over WebSockets todo
  • Heavy Python in Containers/Builds, never the request path todo
  • Idempotency on every consumer (ledger pattern generalised) todo
not yet confirmed by the owner

C4 open: wiring the producer without the generation port would be decoration, so it stays open until the port.

  • wrangler.jsonc — queue producer + Progress DO declared
  • docs/archive/PLATFORM-BUILD-TASKS.md — C4 definition

Spec: Build-tasks C4

Custom domains via Cloudflare for SaaSM

Programmatic custom hostnames with automatic certificates — collapsing the whole domain epic to an API integration plus UI. The permanent domain comes first because it is a performance task: it kills the tunnel tax.

Shops whose domains live elsewhere churn elsewhere.

0 of 3 subtasks · 0%

  • Programmatic custom hostnames + automatic TLS todo
  • Permanent domain first (kills the 493ms tax) todo
  • Domain epic collapses to API integration + UI todo
not yet confirmed by the owner

C5 open. Needs the C0 account first.

  • docs/archive/PLATFORM-BUILD-TASKS.md — C5 definition

Spec: Build-tasks C5

    Waiting on:
  • C0 account project
Mail, flags, schedules and insight pipelinesL

Transactional mail on the email service, tenant-domain mail on routing, feature flags, scheduled health/hole/reconcile/eval jobs, learning-loop events flowing from day one, and an error pipeline feeding the runbooks.

A platform that cannot email, schedule or learn is a demo.

0 of 5 subtasks · 0%

  • Transactional mail via Email Service; tenant-domain via Routing todo
  • Flags via Flagship todo
  • Cron: health checks, hole-herding, reconcile, eval trends todo
  • Learning-loop events in Analytics Engine from day one todo
  • Tail Worker error pipeline feeding the Phase 8 runbooks todo
not yet confirmed by the owner

C6 open. Outbox + flags UI exist locally (P3-06, P4-06); this is the edge wiring.

  • docs/archive/PLATFORM-BUILD-TASKS.md — C6 definition

Spec: Build-tasks C6

    Waiting on:
  • C0 account project
Inference on Workers AI behind a gatewayM

Text, image, embedding and vision inference with no external key for v1 — gateway caching, fallbacks and per-tenant metering feeding the cost remodel and model routing. No model name or fixed price assumed anywhere.

The second provider the routing layer is waiting for.

0 of 3 subtasks · 0%

  • Text, image, embedding, vision with no external key for v1 todo
  • Gateway caching, fallbacks, per-tenant metering todo
  • Metering feeds cost remodel + model routing; no fixed price assumed todo
not yet confirmed by the owner

C7 open. Routing (P10-02) needs observed quality from two sources; this is the second.

  • render/routing.py — routing table awaiting a second board
  • docs/archive/PLATFORM-BUILD-TASKS.md — C7 definition

Spec: Build-tasks C7

    Waiting on:
  • C0 account project
Cost remodel with enforced caps as configurationM

Unit economics re-derived on requests, CPU time, database, storage and inference — and today's advisory caps become enforced limits with pre-cap alerts. Blocks paid traffic until it lands.

Paid traffic on unmetered costs is a blank cheque.

0 of 2 subtasks · 0%

  • Re-derive unit economics on edge metering todo
  • Caps become enforced limits with pre-cap alerts todo
not yet confirmed by the owner

C8 open and blocking paid traffic. Pre-cap alerting exists locally (P8-03); enforcement is edge work.

  • docs/archive/PLATFORM-BUILD-TASKS.md — C8 definition

Spec: Build-tasks C8

Residency decided, tunnel decommissionedM

Data residency decided per store before the first EU tenant — control plane, objects, files, analytics — and the tunnel and the box retired only after the parity board is fully green.

Residency decided after the first EU tenant is a breach with a date.

0 of 2 subtasks · 0%

  • Residency decided per store before the first EU tenant todo
  • Cut the tunnel and decommission the box after a fully green parity board todo
not yet confirmed by the owner

C9 open. Last step of the strangler by design.

  • docs/archive/PLATFORM-BUILD-TASKS.md — C9 definition

Spec: Build-tasks C9

    Waiting on:
  • C0–C8 parity board green

In progress · 11

Shop recipes from your recipe book: contract + baseL

Your recipe book is filed and the pause is lifted. The first build from it is the shared product contract plus the base shop recipe, proven by one product that passes and one that is correctly refused. The 17 existing shop folders use the older taxonomy and get reconciled, not rebuilt.

Shops are the revenue surface; the recipe book is the specification they build from.

1 of 5 subtasks · 20%

  • product.schema.json: base contract (ids, price, tax, stock, images, variants) todo
  • commerce-base recipe folder (spine, ladder, flows, emails, guidance) todo
  • One product validates, one refuses (proven, not asserted) todo
  • Reconcile the 17 existing recipe dirs with the v3 taxonomy todo
  • Owner recipe book filed verbatim (unblocks all of the above) done
not yet confirmed by the owner

LIBRARY-v3 Part 18.1 order step 1. Required fields carry _never_generated and provenance; fulfilmentMethods non-empty; sustainabilityClaims need a basis. 41 categories = 32 physical + 8 service + B2B overlay.

  • recipes/LIBRARY-v3.md — the v3 authoring source, 19 parts
  • recipes/validate.py — existing recipe checks

Spec: LIBRARY-v3 Part 18

Risk: The 17 existing recipe folders use the older taxonomy; building v3 alongside without reconciling forks the layer.

Started 2026-09-11

Make generation work: ten sites, three photolessL

The product milestone: ten real websites from ten plain descriptions, at least three with no customer photos and no generated images, every failure recovered honestly instead of invented around, all ten screenshotted and shown to you. The photoless machinery it needs landed; the ten-site run has not happened.

Nothing else matters until this works. It is the product.

1 of 6 subtasks · 17%

  • Refusal recovery path per 6.5 (decline, gap, recover — never invent) doing
  • Photoless machinery proven (IM-6: contract, stock, honest gaps) done
  • Ten sites from ten descriptions built todo
  • At least three with no photos and no generated images todo
  • All ten screenshotted todo
  • STOP: shown to the owner with the report todo
not yet confirmed by the owner

SPEC step 7. Ten sites from ten descriptions (not businesses — descriptions). Historic Haircut failures traced to establish.subject 1:1-with-no-fallback and stage-7 copy decline gapping the gesture slot; both since addressed (IM-6, assemble_content seeding all 7 intents).

  • render/page.py — component binding and slot fill
  • render/prompt.py — the model prompt, three cache zones
  • render/direct.py — director prose for ladder questions

Spec: KOZTO-SPEC Part 16 step 7, section 6.5

Risk: The photoless path is the hardest case and the most common one.

Started 2026-09-08

Publish, payment and account — live money outstandingL

Publishing, test payments with disputes and refunds, dunning emails and invoices all work in test mode with honest records. What remains is the only part that matters: one real money round-trip on live keys, which needs your Stripe setup.

Nobody can trade until real money moves once.

3 of 5 subtasks · 60%

  • Checkout + upgrade in test mode (success, failure, dispute, refund) done
  • Dunning sequence with emails + delinquency banner done
  • Billing/settings voice lint done
  • One live Stripe round-trip observed (needs owner keys) blocked
  • Per-site admin surface (see SP-14) todo
not yet confirmed by the owner

P7-01 (tests 116, negatives 67/67, mutants 81/81), P7-02 (tests 137, negatives 71/71, mutants 84/84), P7-04 voice lint. No auto-recharge by design; operator-run sweep documented. Per-jurisdiction tax stays P7-03.

  • billing/ledger.py — idempotent ledger + caps
  • platform/dunning.py — dunning sweep + grace + downgrade

Spec: KOZTO-SPEC Part 16 step 13

    Waiting on:
  • Stripe webhook secret + endpoint (owner)

Started 2026-09-08

Per-site admin: search, inspect, flagsL

The admin surface already finds tenants, shows their money and members, carries an undismissable impersonation banner, inspects failed builds without leaking paths, and toggles feature flags from the browser. Still missing: catalogue coverage, work queue, eval trends, costs, moderation and health scoring.

Operating blind means every incident becomes an excavation.

4 of 7 subtasks · 57%

  • Tenant search + detail with ledger, members, audit done
  • Impersonation banner that cannot be dismissed without stopping done
  • Generation inspector on progress pages (no path leaks) done
  • Feature flags table + tenant overrides + form toggles done
  • Catalogue coverage + work queue views todo
  • Eval trends, cost and margin views todo
  • Moderation queue + DMCA path, health scoring, command palette todo
not yet confirmed by the owner

P4-06 partial 2026-09-08: tests 249/251, inbox tests 46. Progress page renders its own banner (bypasses _page). Form posts accepted on the flags endpoint (own CSRF-channel gate).

  • platform/server.py — admin screens + flags + inspector
  • platform/inbox.py — inbox + moderation-adjacent surfaces

Spec: KOZTO-SPEC Part 16 step 14; build-tasks P4-06

Started 2026-09-08

Phase C epic: whole product on Cloudflare (tasks C0–C9)L

The strangler plan is underway: Workers in front, tunnelled origin behind, route-by-route parity-proven cutover. The ten tasks C0–C9 below hold the checklists and states; this card holds the live narrative. Watch the 493ms tunnel tax vanish route by route in Lighthouse numbers.

The 493ms tunnel tax dies route by route, with proof at each step.

not yet confirmed by the owner

States live on C0–C9; narrative here. Live facts 2026-09-11: remote control DB created + migrated + verified, job queue + R2 photo bucket created, edge worker deployed with all four bindings, read ports landed (upload, words, site actions, build screens, generate producer). Known gaps: per-statement D1 atomicity, Progress DO broadcast unverified, Containerfile unbuilt, tunnel serving until cutover. First live hits failed TLS handshake (alert 40) on the brand-new hostname — propagation, retry not reroute.

  • edge/worker.js — the strangler worker
  • platform/d1.py — env-selected D1 transport
  • CLOUDFLARE.md — runbook + cutover order

Spec: Build-tasks Phase C; CLOUDFLARE.md

Risk: Cutover without per-route parity proof repeats the tunnel era with new branding.

Started 2026-09-08

Admin panel: tenants, impersonation, flags, inspectorL

The operator console finds any tenant, shows their money, members, sites and audit trail, and acts as them only under a banner that cannot be dismissed without stopping. Feature flags toggle live; failed builds explain themselves in an inspector.

Operating blind — or impersonating invisibly — is how platforms abuse.

4 of 7 subtasks · 57%

  • Tenant search + detail (ledger, members, sites, audit) done
  • Impersonation behind a permanent banner, stop revokes session done
  • Generation inspector on progress pages (done + failed) done
  • Feature flags with tenant overrides done
  • Catalogue coverage + work queue todo
  • Eval trends, cost/margin views todo
  • Moderation queue + DMCA path, health scoring, command palette todo
not yet confirmed by the owner

P4-06 partial 2026-09-08: search/detail/impersonation/inspector/flags shipped (tests 251). Remainder is operator analytics + moderation.

  • platform/server.py — admin routes
  • docs/archive/PLATFORM-BUILD-TASKS.md — P4-06 definition

Spec: Build-tasks P4-06

Console micro-detail: breadcrumbs, dark mode, honesty about JSM

Breadcrumbs on every site screen, a real dark mode, reduced-motion respect, and secrets shown once as selectable text. The remaining wishlist — optimistic UI, undo, shortcuts, autosave — needs console JavaScript, and the console is deliberately script-free; that is an architecture decision to revisit, not a gap to patch.

Polish is the part strangers feel first.

2 of 3 subtasks · 67%

  • Breadcrumbs on all seven site screens + dark mode + reduced-motion done
  • Shown-once secrets as selectable labelled fields done
  • Optimistic UI, undo, shortcuts, autosave, clipboard (needs console JS) todo
not yet confirmed by the owner

P4-08 partial 2026-09-08: server-rendered slice shipped (tests 238). JS-dependent remainder waits on the zero-JS architecture decision.

  • platform/server.py — breadcrumbs + dark tokens
  • docs/archive/PLATFORM-BUILD-TASKS.md — P4-08 definition

Spec: Build-tasks P4-08

Anti-sameness: five mechanisms, one liveL

Ten sites from ten briefs must look undesigned-by-committee. The first mechanism is live — built pages are scanned against their own world's banned list (pure black, square corners, monospace body). Still to come: seeded selection, anti-brief constraints, palettes from uploads, and a similarity check against 500 sites.

A generator whose output is attributable to one system has no moat.

1 of 5 subtasks · 20%

  • Build-time anti-pattern enforcement vs the world's own bans done
  • Seeded selection todo
  • Anti-brief constraint todo
  • Palette-from-uploads todo
  • Similarity-vs-500 check todo
not yet confirmed by the owner

S-04 partial 2026-09-08: qa.py check 6b live, controls 36/36, mutants 95/95. Open ruling: obsidian nav backdrop-filter vs its glassmorphism ban (1 WARN).

  • render/qa.py — check 6b anti-pattern enforcement
  • docs/archive/PLATFORM-BUILD-TASKS.md — S-04 definition

Spec: Build-tasks S-04

Wrangler foundation: project, envs, secrets, CI deployM

The Cloudflare account project with preview and production environments, current compatibility, logging from the first deploy, secrets in Workers secrets — never the repo — and a preview URL per change. The local config slice is done; the account-side remainder waits on you.

Deployments without a foundation are snowflakes.

1 of 5 subtasks · 20%

  • wrangler.jsonc: compat date current, observability on, bindings declared + schema-verified done
  • Containerfile stdlib-only (build unverified, no Docker here) doing
  • Account project + real database ids todo
  • Secrets in Workers secrets, never the repo todo
  • Preview + production envs; CI deploy + preview URL per PR todo
not yet confirmed by the owner

C0 partial 2026-09-08 (Option A slice vs installed wrangler 4.130.0 schema). Remainder needs owner account access (OW-1).

  • wrangler.jsonc — bindings + compat + observability
  • docs/archive/PLATFORM-BUILD-TASKS.md — C0 definition

Spec: Build-tasks C0

    Waiting on:
  • Owner account access (see OW-1)
Console shell on Workers, API proxied, reads firstL

The console frontend served from the edge with API routes proxied to the origin at first. Cutover reads first — dashboard, inbox, settings — then writes, generation last, each step proven byte-identical against the Python oracle before it counts.

Cutover without per-route proof is a migration-shaped outage.

3 of 5 subtasks · 60%

  • GET /v1/health + GET /v1/pixel byte-identical on 11/11 differential checks done
  • Pure core 10/10 in CI done
  • Read ports landed: upload, words, site actions, build screens, generate producer done
  • Every other route, each with its parity harness todo
  • Writes second, generation last todo
not yet confirmed by the owner

C1 partial 2026-09-08 (Option A slice: first two reads). Parity harness is owner-run (needs npx).

  • edge/worker.js — ported reads + passthrough
  • docs/archive/PLATFORM-BUILD-TASKS.md — C1 definition

Spec: Build-tasks C1

Data on D1: tenant-scoped rows, portable migrationsL

The control plane on the edge database with the tenant stamped on every row and enforced by the query builder, per-tenant coordination in memory-objects rather than business tables, and content-addressed files in object storage. All thirteen migrations ran unmodified and behaved identically.

A tenant who can read another tenant's rows is a breach, not a bug.

4 of 7 subtasks · 57%

  • Parity run: 13/13 migrations unmodified, identical behavior done
  • Env-selected transport with half-config refusal (23/23 tests) done
  • Per-statement atomicity gap documented as known done
  • Remote control DB created, migrated, schema verified done
  • tenant_id on every row + query-builder enforcement todo
  • Per-tenant coordination in DOs; IR/assets content-addressed in R2 todo
  • Remote cutover, R2 asset reads, DO coordination with generation port todo
not yet confirmed by the owner

C3 partial 2026-09-08 (Option A slice +2 mutants). Remote DB live 2026-09-11.

  • platform/d1.py — env-selected D1 transport
  • docs/archive/PLATFORM-BUILD-TASKS.md — C3 definition

Spec: Build-tasks C3

Ready for you to look at · 1

This status board, now with every taskM

This page. It now lists every tracked task and subtask with derived progress, instead of nine summary cards. Open any card for the plain-English version, expand it for subtasks, evidence and technical detail. Tell me if anything on it is wrong — correcting the board is the fastest way to keep it honest.

You should never have to read the code to know the state of the project.

4 of 5 subtasks · 80%

  • board.json + stdlib build.py with the evidence gate done
  • Seeded from real state, deployed to Pages, access stated done
  • Screenshots at 390px and 1440px attached done
  • Every tracked task + subtask listed with derived progress done
  • Owner reads the top and can say what is happening + waiting todo
not yet confirmed by the owner

status/board.json -> status/build.py (stdlib, <1s) -> status/dist/index.html. build.py refuses a done card without evidence and verifies every files[].path exists. Facts measured at build time. Kanban expansion at owner request 2026-09-12: subtask schema + derived x-of-y percentages (computed, never estimated).

  • status/board.json — the data, maintained per session
  • status/build.py — renderer + evidence gate
  • docs/status-board-brief.md — the owner brief, filed verbatim
status/build.py

358 lines · renderer with the evidence gate

Board at 1440px: headline, waiting-for-you, columns
Board at 1440px: headline, waiting-for-you, columns · 320KB
Board at 390px: stacked, waiting-for-you first
Board at 390px: stacked, waiting-for-you first · 231KB

Spec: Status board brief §1–§13

Started 2026-09-11

Done · 24

Spec saved per Part 0.1S

The specification itself is filed in the repo and every step since has been built against it. The STOP reports and the rules that govern the work all cite it by section.

A build order without a saved spec is folklore.

1 of 1 subtasks · 100%

  • KOZTO-SPEC.md saved and referenced by section done
not yet confirmed by the owner

SPEC step 0 + STOP confirmed. AGENTS.md, LOG.md Part 18 reports and this board all cite SPEC sections.

  • KOZTO-SPEC.md — the build order, acceptance and reporting spec
KOZTO-SPEC.md

1013 lines · Parts 0–19 including build order and acceptance

Spec: KOZTO-SPEC Part 16 step 0

SQLite to D1 parity proven with a real migrationM

The database the site runs on today and the cloud database it will move to were proven identical — every migration applied unchanged with the same behavior. The move itself is still to come, but the equivalence is measured, not assumed.

Migrating to a database that behaves differently is a data-loss plan.

2 of 2 subtasks · 100%

  • 13/13 migrations apply unmodified with identical behavior done
  • Env-selected transport with half-config refusal done
not yet confirmed by the owner

SPEC step 1. Option A parity run: schema/FK/UNIQUE/index behavior identical; platform/d1.py transport 23/23 tests (+2 mutants — the stub used to encode array rows and hid live misreads); per-statement atomicity documented as the known semantic gap.

  • platform/d1.py — env-selected D1 transport
platform/d1.py

177 lines · transport with half-config refusal

Spec: KOZTO-SPEC Part 16 step 1

Model capabilities verified with measured numbersM

All three model jobs were run live and scored: writing, understanding and directing pass; the revision job has one stable miss that is understood, not flaky. Routing already sends each job to whichever provider earned it.

Unmeasured model quality is a demo, not a capability.

3 of 3 subtasks · 100%

  • Copy 5/5, comprehend 3/4, director 4/4 live (spark 1.3) done
  • Revise 2/3 with the stable dictation-decline diagnosed done
  • All three reported with numbers at the STOP done
not yet confirmed by the owner

SPEC step 3 per 11.2. Comprehend's miss is genuine fees fabrication caught twice. Revise miss is verbatim-dictation decline, twice running (1.1 answered it). Single-provider — the second source is OW-1/P10-02.

  • evals/run.py — eval runner over the four task sets
evals/run.py

236 lines · runner for copy/comprehend/director/revise sets

Spec: KOZTO-SPEC Part 16 step 3, section 11.2

Started 2026-09-08

Renderer decision: static TSX plus regex loweringS

The decision about how components become web pages is settled and recorded: static component files lowered with patterns, no clever runtime. The report went out with the recommendation and the reasons.

An undecided renderer forks every component built after it.

1 of 1 subtasks · 100%

  • Recommendation reported with reasons (retroactive report) done
not yet confirmed by the owner

SPEC step 4 per 7.1. Retroactive Part 18 report 5f246b0 (LOG +28). Island:none lowering stands.

  • render/page.py — the decided renderer
git show --stat 5f246b0

LOG.md +28, retroactive renderer-decision report2026-09-11

Spec: KOZTO-SPEC Part 16 step 4, section 7.1

Ingestion pipeline proven on five componentsM

Five components were taken through the whole intake pipeline — cards and tables first — passing the sandbox gate and rendering in all three design worlds. The pipeline, not just the components, is what was proven.

Components nobody can ingest are drawings, not inventory.

2 of 2 subtasks · 100%

  • Card + table proofs through gate-sandbox done
  • Rendering proven under obsidian, bloom and luminous done
not yet confirmed by the owner

SPEC step 5 per 7.8. Commit be2926d: 32 files, +1679/−2. Proofs held in intake/proof with _intake ids, no catalogue merge until the queue asks.

  • components/ingest.py — the proven intake pipeline
git show --stat be2926d

32 files, 1679 insertions, 2 deletions2026-09-11

Spec: KOZTO-SPEC Part 16 step 5, section 7.8

Photo supply chain: subject points, live stock, honest gapsL

When someone uploads a photo for the main subject, they now mark where the subject is so crops never cut heads off. The system can fetch real credited stock photos that pass every check. And where no photo exists at all, every world shows a labelled empty frame — never a fake-looking gradient, never a broken image. Proven with screenshots for all three worlds.

Photography is 60–70% of perceived quality; every tier of the supply now proves itself or stays silent.

6 of 6 subtasks · 100%

  • Subject points: migration, validation, picker, edge + batch rules done
  • Stock records key the slot want; audit crop gate passes done
  • Gate 3f: every slot component carries its own hole treatment done
  • Monument region renamed so stamping reaches it done
  • Live stock proof: 3/3 slots, hero renders with credit done
  • Screenshots: three worlds + monument crop + live hero done
not yet confirmed by the owner

SPEC step 6 + Tier-3 wiring (commit 99da0c6). save_asset(..., focal) refuses subject slots without a point; nine-point picker; batch excludes subject slots; stock.retrieve proven live; cta-monument k-mon__figure -> k-mon__fig (old class never matched _REGION).

  • platform/generate.py — save_asset focal validation, assets_of threading
  • platform/server.py — focal picker, edge enforcement, batch exclusion
  • render/stock.py — Tier-2 record with want-keyed ratios
  • render/gate.py — 3f missing-asset treatment requirement
  • components/cta-monument — treatment + region rename
python3 platform/negatives.py

91/91 controls fired (7 new focal controls)2026-09-11

python3 render/stock-negatives.py

11/11 stock controls hold (record passes the crop gate)2026-09-11

python3 render/gate-negatives.py

16/16 correct (3f catches an untreated component)2026-09-11

live Unsplash retrieve + asset audit, bloom/restaurant

3/3 slots filled, 0 errors, 0 fatal; hero photo renders with credit2026-09-11

Obsidian, zero assets: labelled holes, no gradients
Obsidian, zero assets: labelled holes, no gradients · 34KB
Bloom, zero assets: labelled holes, legible on light
Bloom, zero assets: labelled holes, legible on light · 31KB
Luminous, zero assets: labelled holes, world-distinct
Luminous, zero assets: labelled holes, world-distinct · 30KB
Monument figure hole: dashed frame + label, full-res crop
Monument figure hole: dashed frame + label, full-res crop · 2KB
Live credited stock photo rendering in the bloom hero
Live credited stock photo rendering in the bloom hero · 41KB

Spec: KOZTO-SPEC Part 16 step 6

Started 2026-09-11

Commerce recipe book v3 received and filedS

Your recipe book for shops — 41 kinds of business, what each shop must ask, and what must refuse — arrived and is saved word for word as the source the shop-building work follows. Nothing in it has been built yet; filing it was the step that unblocked the recipe pause.

The shop work was paused waiting for exactly this document.

3 of 3 subtasks · 100%

  • Document received from the owner done
  • Persisted verbatim, all 19 parts verified by header count done
  • Recipes pause lifted on arrival done
not yet confirmed by the owner

Persisted at recipes/LIBRARY-v3.md: 19 parts (0–18), taxonomy 32+8+overlay, Part 18 implementation order. Reconciliation with the 17 older dirs is RC-1's work.

  • recipes/LIBRARY-v3.md — the v3 authoring source
recipes/LIBRARY-v3.md

2561 lines · all 19 parts, verified by header count

Spec: LIBRARY-v3 (owner-supplied)

Started 2026-09-11

Simulated billing stripped outM

An earlier version pretended test billing was the real thing. It was deleted — thousands of lines removed — and only honest test-mode billing with a read-only live-key check remains. Live billing stays waiting on you (see the top card).

Work described as finished that was not is the failure this board exists to prevent.

2 of 2 subtasks · 100%

  • Simulated processor deleted (86 files, net −3449 lines) done
  • Test-mode ledger kept honest with read-only live-key check done
not yet confirmed by the owner

ceca793 Phase 1 remediation §3.1–§3.8: strip sim billing, injection boundary, rate limits, fonts, palette, docs archive.

  • billing/ledger.py — honest test-mode ledger that survived
git show --stat ceca793

86 files changed, 2757 insertions, 6206 deletions2026-09-10

Spec: Remediation §3

Model provider seam with failing controlsM

The system talks to AI providers through one guarded door: requests go out with exactly the right shape, answers come back parsed and checked, and the tests prove the door slams shut on bad keys, timeouts and over-limit responses. Adding a second provider later means implementing the same small contract.

A provider change must not rewrite the product.

3 of 3 subtasks · 100%

  • OpenAI + Anthropic providers behind one seam done
  • Auth/timeout/rate-limit/parse controls fail correctly (13/13) done
  • Worker/upstream errors passed through honestly, never masked done
not yet confirmed by the owner

T0-07 done 2026-09-07: exact request shape, Authorization override refusal, 429 retry-after passthrough, JSON.parse fenced-code recovery; +3 targeted mutants. Anthropic extended with live-scan provocation proving no real key leaks.

  • render/providers/openai.py — OpenAI provider behind the seam
  • render/providers/anthropic.py — Anthropic provider behind the seam
render/providers/openai.py

159 lines · seam contract + controls

Spec: Build-tasks T0-07

Controls that prove the controls workS

Every image analysis has a matching sabotage test — inject a blank hole, overflow the math, delete a file — and each one fires with the exact message expected. This is the layer that keeps every other checker honest.

A validator that has never failed is not known to work.

1 of 1 subtasks · 100%

  • Six sabotage controls, all firing with exact messages done
not yet confirmed by the owner

T0-12 done 2026-09-07: void-injected hole fires budget control #5; int16 overflow fires #3 before the fix was found; 6/6 with echoes coverage.py can see.

  • worlds/fontmetrics.py — home of the sabotage-control pattern
worlds/fontmetrics.py

114 lines · controls beside the validator

Spec: Build-tasks T0-12

Font metrics as a second opinion on typeM

The design skill's typography claims are cross-checked by measuring real font files — x-heights, stem widths, contrast ratios — so a claim about readability is a measurement, not an opinion. Thirty-three checks plus five controls, with mutation testing proving they bite.

Two independent measures beat one confident opinion.

2 of 2 subtasks · 100%

  • Alphabet + x-height + contrast measurement controls done
  • Proven twice: fresh checkout and forced corruption done
not yet confirmed by the owner

T0-03 done 2026-09-07: 33/33 + 5/5 controls, 99/99 mutants; DELIBERATE BUG injected and caught before proclaiming safety.

  • worlds/fontmetrics.py — font measurement controls
worlds/fontmetrics.py

114 lines · 33 checks + 5 controls

Spec: Build-tasks T0-03

Zero-scroll demo that tells the truthM

Strangers see a live demonstration in one screen — no scrolling, real content, and every number on it traceable to its source. The remaining gap is a demo driven by a genuinely empty site, which needs a real interview first.

The first screen decides whether anyone reads the second.

2 of 3 subtasks · 67%

  • Single-viewport demo, 8 controls against the live DOM done
  • Every figure traces to ledger, query, recipes or fixtures done
  • Empty-new-site demo render (needs a real interview) todo
not yet confirmed by the owner

Hmm — one todo subtask on a done card: the card's DONE claim is the shipped demo (0/15 sweep verdict: real-interview demo, not fixture-polish). The empty-site render is tracked as remainder, visibly unfinished. P1-02 shipped 2026-09-07.

  • platform/generate.py — demo_preview + assemble_content
platform/generate.py

1168 lines · demo layer over real content assembly

Spec: Build-tasks P1-02

Billing and settings copy without threatsS

Every billing and settings sentence was checked for threatening tones and confusing jargon, with user-facing errors that explain what to do. The lint runs on every build so new threats cannot slip in.

Threat-toned billing copy reads as extortion.

3 of 3 subtasks · 100%

  • 58 billing strings swept, 2 threat-toned fixed done
  • 40 settings strings swept clean done
  • Fail-closed error surfaces (403, 404, 500, 413) done
not yet confirmed by the owner

P1-04 shipped 2026-09-07: voice now includes explicit failure-mode entries; delinquency_banner states grace + downgrade with dates. Tests 148, negatives 48/48.

  • platform/copycheck.py — the lint that keeps it clean
platform/copycheck.py

258 lines · tone lint + fail-closed surfaces

Spec: Build-tasks P1-04

Zero-asset pages render honestlyS

A site with no photos, no direction and nothing filled in renders as declared empty inventories and named holes — never a crash, never a silent gap. Seventeen checks cover it including hostile inputs.

The empty state is the first state every tenant sees.

1 of 1 subtasks · 100%

  • 17 zero-asset checks incl. oversize alt, no-null crash, hostile CSS done
not yet confirmed by the owner

P1-06 shipped 2026-09-07. Empty states guide to the asset page with slot/alt text; preview 7/7 against real assembled content.

  • platform/generate.py — zero-asset honesty paths
platform/generate.py

1168 lines · 17 zero-asset checks green

Spec: Build-tasks P1-06

Hole treatment declared per roleS

Every hole in a built page declares its own treatment — a photograph slot asks for a photograph, a map slot asks for a map — so nothing renders silent or pretends to be designed. The step-6 work later gave each treatment its own styling.

A hole with a name gets filled; a silent one ships.

2 of 2 subtasks · 100%

  • needs() with (key, want, role, required) + questions() done
  • mark_assets stamps empty regions with honest labels done
not yet confirmed by the owner

P2-08 shipped 2026-09-08: tests 222, negatives 62/62; later extended by IM-6 (focal, gate 3f, monument).

  • render/assets.py — needs/questions/audit contract
platform/generate.py

1168 lines · save_asset + match_photos + questions

Spec: Build-tasks P2-08

Photo-quote approval stepS

Tenants approve an AI-written photo description before it becomes the accessibility text — shown beside the image, one click to accept, one to rewrite by hand. Nothing machine-written reaches a visitor unreviewed.

Alt text nobody approved is a lawsuit with a timestamp.

1 of 1 subtasks · 100%

  • Alt-quote accept/rewrite beside the upload done
not yet confirmed by the owner

P3-02b shipped 2026-09-08: tests 210, negatives 58/58, live-DOM verified.

  • platform/server.py — approval UI beside uploads
platform/server.py

4731 lines · alt-quote approval step

Spec: Build-tasks P3-02b

Build progress and reveal screensM

While a site builds, visitors see honest progress with a working cancel; when it finishes, the reveal names what is missing and links straight to fixing it. No dead ends, no fake completion.

A build that lies about finishing loses the customer it just won.

3 of 3 subtasks · 100%

  • Progress page polls with cancel that actually cancels done
  • Reveal names missing ladders with fix links + brief download done
  • Gen_jobs ledger with cancelled terminal state done
not yet confirmed by the owner

P3-06 shipped 2026-09-08: build-check 13/13. Coordinator runs inline in SQLite (the queue is CF-1's work).

  • marketing/build-check.py — build flow checks 13/13
marketing/build-check.py

192 lines · 13/13 progress + reveal checks

Spec: Build-tasks P3-06

Inbox that explains itselfS

The inbox — where review flags land — is empty-state honest, plain-English, and shows only real review work. No notification theatre, no demo rows.

An inbox of fake work teaches owners to ignore the real thing.

1 of 1 subtasks · 100%

  • Review-only inbox with plain-English flag explanations done
not yet confirmed by the owner

P4-04 shipped 2026-09-07: no notification center, no badge counts, no seeds; tests 198, negatives 55/55.

  • platform/inbox.py — review-only inbox
platform/inbox.py

164 lines · review-only inbox, no theatre

Spec: Build-tasks P4-04

Console instruments that all passM

Analytics, backups, dunning and the rest each have their own small test suite, and all 399 pass. Each was specified from real code, not invented — the process itself was audited twice.

Instruments nobody tests lie the loudest.

1 of 1 subtasks · 100%

  • 11 slices, 399/399 checks green (incl. 10 each for P4-07) done
not yet confirmed by the owner

P4-07 shipped 2026-09-08: analytics/backup/dunning/restore/integrity/scheduler/notifications/lifecycle/progress/flags/results, each with its own suite.

  • platform/analytics.py — one of the eleven slices
  • platform/backup.py — timestamped backup + restore
  • platform/dunning.py — dunning sweep
platform/analytics.py

82 lines · one passing instrument slice

Spec: Build-tasks P4-07

Checkout and upgrade in test modeM

Buying, upgrading, failing, disputing and refunding all work against the test processor with idempotent records — replay the same webhook twice and the balance moves once. The sweep that charges renewals is operator-run and documented as such.

Money code without idempotency charges twice.

2 of 2 subtasks · 100%

  • Checkout + upgrade incl. failure, dispute, refund paths done
  • Dual-write + reconcile + replay idempotency done
not yet confirmed by the owner

P7-01 shipped 2026-09-08: tests 116, negatives 67/67, mutants 81/81. Sim billing stripped later (DL-1); this card is the honest test-mode that survived.

  • billing/ledger.py — idempotent test-mode ledger
billing/ledger.py

98 lines · ledger + reconcile + replay guards

Spec: Build-tasks P7-01

Dunning that downgrades honestlyM

Failed payments trigger a real sequence — emails, a visible banner, a grace period — ending in downgrade, never silent cutoff. Every run is audited with before-and-after states.

Silent cutoff is how you lose a customer and keep their money.

2 of 2 subtasks · 100%

  • Dunning sweep with emails + delinquency banner done
  • GraceDowngrade run audited with before/after done
not yet confirmed by the owner

P7-02 shipped 2026-09-08: tests 137, negatives 71/71, mutants 84/84.

  • platform/dunning.py — dunning sweep + grace + downgrade
platform/dunning.py

168 lines · audited dunning runs

Spec: Build-tasks P7-02

Billing and settings voice lintS

Every billing and settings surface now carries lint metadata with explicit failure modes, so new screens cannot ship without declaring how they fail. The registry has a control proving it bites.

Voice without a registry rots one screen at a time.

2 of 2 subtasks · 100%

  • Voice metadata on all billing/settings surfaces done
  • Registry control proves the gate bites done
not yet confirmed by the owner

P7-04 shipped 2026-09-08: tests 129, negatives 68/68, mutants 82/82.

  • platform/copycheck.py — voice metadata + registry gate
platform/copycheck.py

258 lines · voice registry gate

Spec: Build-tasks P7-04

Quality gates with live ammunitionM

The render quality gate was proven on a real business with seven injected faults — it caught five outright and named all seven. The remaining two taught the gate new checks the same day.

A gate tested on clean input guards nothing.

3 of 3 subtasks · 100%

  • 7-fault live-fire run on a real business (5 caught, 7 named) done
  • Intent-label honesty check added from the miss done
  • Prove-section ladder gating added from the miss done
not yet confirmed by the owner

P8-02 shipped 2026-09-08: render/qa.py + render/qa-negatives.py 22/22 controls incl. CI visibility/CI failure/workspace hygiene.

  • render/qa.py — render quality gate
render/qa.py

170 lines · gate + live-fire proof

Spec: Build-tasks P8-02

Golden corpus: 44 businesses boundL

44 real businesses across every category build successfully against four design worlds — 156 of 176 pairs bound, with the 17 gaps and 3 failures named and ranked by the queue rather than hidden. Seventeen distinct spines carry 68 effective problems.

A corpus that hides its failures is marketing, not measurement.

3 of 4 subtasks · 75%

  • 44 businesses x 4 worlds bound 156/176 (measured 2026-09-12) done
  • 17 gaps + 3 failures named and ranked done
  • Race-harness lanes share one gateway + policy, proven done
  • Close the gaps (see Q-1) todo
not yet confirmed by the owner

P10-01 + queue as of 2026-09-12: 176 pairs, 156 bound (88%), 17 catalogue gaps, 3 failed, universal adjacent-gaps rule blocker. One todo subtask points at Q-1 — the corpus is done, its ask is next.

  • golden/queue.py — binding + ranked ask
python3 golden/queue.py

176 pairs, 156 bound, 17 gaps, 3 failed, 1 rule blocker2026-09-12

Spec: Build-tasks P10-01; ARCHITECTURE.md section 7

Parked · 0

Not starting before a paying customer needs it

Recently changed