Actionist design grammar and preference protocol
Date: 2026-08-27 Status: current first-principles synthesis; research-only; not an implemented picker Owns: the join between P05 components, P06 preference learning, P07 token harmonization and P08 rendered shells
Why this document exists
The P06 research was useful, but it answered a narrower question than the product needs. It investigated how to learn a person's visual preference efficiently: comparison format, adaptive query selection, stopping rules, uncertainty and drift. It did not finish the design grammar that determines which palettes, font systems, spacing systems, radii, shadows and layouts are valid candidates in the first place.
That distinction matters. A sophisticated preference model over a badly defined design space learns nonsense efficiently. Conversely, a scientifically valid token catalogue without an elicitation protocol forces a non-designer to manipulate abstract values they do not understand.
The complete system is therefore:
hard constraints
-> valid design grammar
-> controlled rendered stimuli
-> adaptive preference learning
-> versioned DesignDNA
-> nearest valid token pack
-> components, shells and imported blocks re-rendered from that pack
-> later approval/rejection updates the profile
No further broad experience sprint should be promoted until this grammar and its experiments are agreed.
What P06 genuinely established
Observed or strongly supported
- Comparisons are more appropriate than abstract configuration for non-designers. Figma's radius and spacing sliders are useful escape hatches for experts, but a client often cannot state a preferred radius in pixels. Midjourney Style Creator is the closest shipped analogue: users repeatedly select visual outcomes and receive a reusable style parameter.
- Use complete realistic screens while controlling the variable under test. Content and structure remain constant within a comparison; only the targeted design dimensions vary. This preserves ecological validity without making credit assignment impossible.
- Four or five candidates per batch is a defensible starting format. Sawtooth recommends four or five items per MaxDiff set; the P06 information analysis supports four-up plus an outside option. This is a starting experimental arm, not a universal law about interface choice.
- An outside option is necessary. “None of these” prevents forced bad choices from contaminating the model. It is distinct from a tie. Three consecutive outside selections should leave the current region and re-seed rather than producing more near-duplicates.
- The query policy should be adaptive. Later comparisons should target the least-resolved dimension or the candidate set with highest expected information gain. Random comparisons waste most of the available questions.
- The stopping rule is not a fixed questionnaire length. The current supported envelope is a five-round floor, a 15-round ceiling, an always-available “this is right” exit and a confidence/information-gain stop. Sequential Gallery reported 5.36 ± 2.69 rounds to satisfaction in a six-person study; Midjourney documents stabilization around five to ten rounds and small changes past 15. These are useful anchors, not a sufficiency theorem for Actionist.
- Confidence belongs to each design dimension. “No preference about shadows” is valid information. A single percentage hides which dimensions are unresolved.
- Learn a continuous preference representation, but ship a closed valid pack. Continuous preference coordinates let a growing catalogue be re-ranked without repeating onboarding. The selected boundary artifact must still be a pre-authored or deterministically generated pack that passed all gates. Arbitrary interpolation is not allowed because it has not passed accessibility, mode and completeness checks.
- Initial stated preference is perishable. Pinterest replaced static onboarding interests with engagement-derived clusters; Netflix documents later behavior superseding initial selections. DesignDNA should be re-entrant and updated by later client approvals and rejections.
- The learner must be tested against a static default. An adaptive system is a product feature, not proof of its own effectiveness. The evaluation must compare it with a strong industry-default pack.
What P06 did not establish
- The correct list of visual dimensions.
- How many perceptually distinct levels each dimension contains.
- Which dimensions are independent or coupled.
- A scientifically privileged colour-harmony family.
- A definitive font-pairing catalogue.
- A perceptually optimal spacing base unit.
- The best radius or shadow scale.
- Whether four-up beats a scrollable gallery for Actionist clients.
- Whether 8–12 rounds outperform a simpler static pack picker in the real client workflow.
Those gaps should remain explicit. They are the work of the design grammar and its validation experiments, not reasons to discard the P06 preference-learning results.
The three-layer design
Layer 1 — validity: what the client cannot choose to break
Candidates are rejected before display when they violate the Design Validity Contract. Non-negotiable gates include:
- complete primitive -> semantic -> component token resolution;
- complete light/dark role pairing rather than RGB inversion;
- WCAG 2.2 contrast across text and applicable non-text states;
- a pinned APCA report as an additional quality signal, without claiming WCAG 3 conformance;
- monotonic, gamut-safe colour ramps;
- semantic status, focus, selection and inverse roles;
- colour-vision simulation and redundant non-colour status cues;
- available font files/weights, metric-compatible fallbacks and readable type styles;
- spacing and geometry values drawn from the declared scale;
- complete radius, border, elevation, opacity and layer resolution;
- complete interaction states and reduced-motion alternatives;
- no undeclared raw style literals in generated code.
The client chooses among valid expressions. Accessibility, internal consistency and completeness are not preference axes.
Layer 2 — grammar: the finite language of valid variation
The grammar should be expressive enough to produce materially different products but finite enough to measure, validate, explain and reproduce. Each dimension needs an identifier, candidate levels, dependencies, semantic-token mapping, preview recipe and validation rules.
G1. Canvas and mode
- mode posture: light-first, dark-first or adaptive;
- neutral temperature: cool, balanced or warm;
- canvas/surface separation: flat, subtle or pronounced;
- high-contrast variant where required.
This dimension must be resolved before fine colour evaluation because the same accent behaves differently against different canvases.
G2. Colour system
Colour is a role system, not a set of favourite hex values.
- seed or constrained brand hue;
- neutral temperature and neutral chroma;
- primary chroma/intensity;
- tonal contrast posture;
- secondary relationship family: monochromatic, analogous, complementary, split-complementary, triadic or neutral-plus-accent;
- status and focus roles;
- chart palette;
- explicit light/dark outputs.
The relationship families are a candidate taxonomy, not a claim that complementary colour is universally superior. Hue geometry alone does not establish usability: available gamut, tone, chroma, foreground pairing, surface context, cultural meaning and accessibility all intervene. Actionist should generate ramps in a pinned perceptual space such as OKLCH or HCT, map them into semantic roles, then run the validity gates. A raw colour picker can constrain the seed, but it must never directly define the shipped palette.
The first targeted research gap is to compare established seed-to-role generators and authored systems, then determine which relationship families remain perceptually distinct after semantic-role and accessibility constraints.
G3. Typography system
Typography is a coordinated role set:
- display/heading family;
- body/UI family;
- optional mono/data family;
- pairing posture: same-family, harmonious, contrasting, editorial, geometric/humanist or technical;
- type scale and heading contrast;
- available weights;
- body and heading line-height;
- letter spacing;
- reading measure and responsive policy.
The local registry's font pairings are seeds, not proof of good pairings. Every pairing must be rendered in headings, body text, labels, buttons, forms, tables and numbers. A pairing that looks good in a hero can fail in a dense operational table.
G4. Spacing and density
Spacing is a rhythm system, not a collection of independent pixel decisions:
- base unit and finite scale;
- compact, balanced or spacious density posture;
- internal group gap;
- between-group gap;
- section rhythm;
- control heights and hit targets;
- card padding, page gutters and table/list density;
- responsive scaling rules.
Material's 8dp/4dp practice and Polaris's 4px tokens establish strong conventions and practical divisibility, not controlled evidence that one base unit is perceptually optimal. Actionist should test complete coherent scales and enforce the selected scale consistently. The important visual law is relational: spacing inside a group is tighter than spacing between groups, and hierarchy should not use equal gaps everywhere.
G5. Shape and depth
- radius family: square, subtle, rounded, soft or pill-accented;
- nested-radius rule;
- border posture and weight;
- flat, bordered or elevated surface model;
- elevation count and shadow intensity;
- light-source direction and mode-specific shadow colours;
- overlays, scrims and focus rings.
Radius and shadow should initially be treated as a coupled family because both communicate object material and layer depth. The system can split them into separate dimensions only if discriminability and correlation experiments justify it.
G6. Hierarchy and layout posture
- information density and number of hierarchy levels;
- content width and reading measure;
- navigation frame;
- card/list/table composition;
- sidebar and detail-panel behavior;
- responsive collapse behavior;
- progressive disclosure posture.
This is shared with P08. P06 learns visual preference; it must not silently select an unusable shell for the workflow.
G7. Motion, icons and imagery
- motion intensity, duration and easing family;
- icon family, fill/stroke posture and stroke weight;
- photography, illustration or graphic treatment;
- image crop, radius, overlay and placeholder behavior;
- reduced-motion and asset fallbacks.
These may be skipped for operational products where they do not materially affect client value. Bespoke illustration and custom motion must not be implied when the builder cannot reproduce them.
Layer 3 — elicitation: how the user searches the grammar
The interface can support effectively unlimited exploration without generating arbitrary designs.
Each visible batch should:
- come from the valid candidate space;
- contain four or five meaningfully separated options on the targeted dimensions;
- keep content and non-target dimensions fixed or near the current posterior;
- never repeat an identical stimulus;
- expose “none,” “these are equivalent,” “show more” and “this is right” where appropriate;
- record every impression, selection, skip, response time and candidate distance;
- avoid near-duplicate batches through a diversity floor;
- update per-dimension confidence rather than one global score.
“Show more” is therefore allowed indefinitely, but it means “sample another diverse, valid batch from this region,” not “invent random colours and radii forever.” Unlimited random generation destroys experimental control, makes choices irreproducible and may show combinations that cannot ship.
Proposed round protocol
This is a staged calibration followed by an adaptive loop, not a rigid colour-to-font wizard.
R0. Constraint intake
Capture existing brand colours/fonts, logo, mandatory mode, accessibility posture, product archetype, primary device and explicit must-have/unacceptable directions. These prune the space before preference learning.
R1. Broad visual direction
Show realistic complete screens spanning high-diversity families such as restrained product, premium editorial, friendly rounded, technical dense and expressive. The result is only a prior. It does not permanently settle every visual dimension.
R2. Canvas and colour calibration
Hold structure and typography fixed. Compare neutral temperature, mode posture, chroma, contrast and relationship families using complete semantic palettes. Allow a brand seed or colour picker as an input constraint, followed by generated and validated role systems.
R3. Typography calibration
Hold palette and structure fixed. Compare complete display/body/UI systems on the same realistic screen. Include dense data, form labels and long copy rather than hero headings alone.
R4. Spacing and density calibration
Hold colour and type fixed. Compare coherent scales and density systems under realistic content stress, especially tables, lists and forms.
R5. Shape and depth calibration
Hold prior systems fixed. Compare coupled radius, border and elevation families rendered across cards, controls, menus, dialogs and overlays.
R6. Layout/context calibration
Compare the current DesignDNA across the client's actual archetype contexts: for example dense case console versus sparse client portal. Record one base profile plus context offsets rather than immediately creating unrelated profiles.
R7+. Adaptive disambiguation
Target the least-resolved dimension or highest expected-information comparison. Revisit earlier dimensions when interactions create uncertainty. Use fragments only to disambiguate a specific remaining conflict.
Final confirmation
Render the top three complete packs against representative client content and at least two important surfaces. Let the client select, hybridize only through valid grammar choices, return to a dimension or accept an industry-safe default.
What the mined repositories contribute
The P06 repository census did not produce a ready-made Actionist taste picker, but it found useful mechanisms:
| Mechanism | Strongest precedent | Actionist use |
|---|---|---|
| Preferential Bayesian model and acquisition | meta-pytorch/botorch PairwiseGP and qEUBO | Candidate production model after simpler baselines pass |
| Ask/tell experiment state | facebook/Ax | Durable adaptive trial loop |
| Simple paired-choice model | lucasmaystre/choix | Bradley-Terry baseline |
| Robust simplicity control | Copeland counting / active-ranking literature | Test whether GP complexity earns its cost |
| Visual search interaction | yuki-koyama/sequential-gallery | Adaptive gallery and satisfaction exit |
| Projected high-dimensional search | yuki-koyama/sequential-line-search, Aalto PPBO | Optional expert/fine-tuning mode |
| Experimental-design selection | idefix, BALD literature | Choose informative, non-confounded batches |
| Simulation-based sizing | skpr method | Determine rounds empirically before making claims |
| Population prior | improved aesthetic predictor | Warm-start only; learn client deviation |
| Confidence/uncertainty reference | openskill.py, Arena bootstrap methods | Per-dimension uncertainty and reporting |
| Controlled visual supply | local P05 21st corpus | Re-theme identical structures using candidate packs |
The immediate implementation should not begin with the most complicated stack. A deterministic candidate generator plus simple Bradley-Terry/Copeland baselines should be tested first. PairwiseGP/qEUBO is justified only if it improves held-out prediction, stability or round count.
Local design assets and their correct role
The laptop already contains the beginnings of the grammar:
siso-ui-baseholds the large UI corpus, palette lookup tooling, generated design-skill seeds and authored principle sections.SISO_Knowledge/design-systemdefines a component-bank structure with raw, primitive, composite, system and adapter layers.- the legacy
21st-devstore is a second complementary source-bearing component store; - the Lumelle token file demonstrates primitive and semantic colour roles plus a small radius set;
- token-pack science defines a proposed 75-requirement closed pack and machine gates A–J.
These assets are not yet one canonical design system. Palette and font seeds need deduplication, provenance, perceptual feature extraction, complete token compilation and rendered validation. Component tags need canonical entities and aliases. The local corpus is stimulus supply; it does not itself define scientific preference axes.
DesignDNA output contract
The learner should produce a reproducible record, not just “option 3 won”:
DesignDNA
identity: client, schema version, catalogue version, model version
constraints: brand seeds, must-haves, unacceptables, accessibility policy
posterior by dimension: mean/region, interval, observations, indifference state
context offsets: marketing, dashboard, portal or other proven contexts
transcript: candidate IDs, impressions, choices, skips, timing
selected pack: exact immutable pack ID/version
alternatives: second/third nearest valid packs
validation: held-out prediction, neighbour discrimination, perturbation result
freshness: created, last confirmed, decay/re-elicitation state
The selected token pack is independently versioned and contains primitive, semantic and component layers. Components and blocks consume semantic roles; they do not consume DesignDNA directly.
Experiments required before this is “nailed down”
- Local grammar inventory. Normalize the local palette, typography, spacing, radius, shadow and component seeds into one machine-readable candidate register. Report duplicates, missing roles and unsupported combinations.
- Discriminability test. Determine which levels clients can reliably distinguish. Run the ±1 perturbation test first; if adjacent levels are reliably rejected, the assumed low-resolution grammar is false.
- Correlation test. Measure whether palette/chroma, typography, density, radius and shadow can be learned independently. This decides whether rounds vary one dimension or coupled families.
- Batch-format experiment. Compare four-up plus outside, scrollable show-more, mixed four-up/pairwise and a direct-control fallback. Measure time, abandonment, held-out prediction and subjective confidence.
- Simulate-to-size. Use synthetic clients across measured noise/correlation regimes to choose the floor, ceiling and acquisition policy. Do not quote 8–12 as settled until this runs.
- Test-retest. Repeat the protocol after two weeks. Stability is more important than fitting the original clicks.
- Context transfer. Test whether one DesignDNA plus small context offsets predicts choices across marketing pages, dashboards and portals.
- Static-default A/B. Compare the elicited result with a strong industry default on actual client approval, revision count and later rejection.
- Closed-pack proof. Re-render representative P05 components and imported donor surfaces under each candidate pack and prove no raw literals or missing roles escape.
Decision and research boundary
The present evidence is sufficient to define the architecture and the next experiments. It is not sufficient to freeze the exact colour relationship catalogue, font-pairing catalogue, number of levels, round count or estimator.
The next research should be a focused design-grammar closure, not another general P06 sweep. It should investigate only:
- colour-system generation and harmony after semantic/accessibility constraints;
- typography pairing and role-system quality;
- spacing/density, shape/elevation and cross-dimension coupling;
- normalization of the actual SISO design-hub assets;
- the experiments above.
That packet should return a machine-readable grammar, direct local-asset joins, rendered stimuli and falsifiable tests. Another list of generic preference companies or repositories would add little unless it closes one of those exact gaps.
Direct evidence
- P06 research report
- P06 first-principles derivation
- P06 innovation register
- UI pick-to-spec research
- Token-pack science
- 21st corpus audit
- Local corpus join
- SISO UI Base README
- SISO Knowledge design-system README
- Lumelle tokens