P06 · Experience · Rendered from source

innovation register

Design taste and preference learner

111 lines52,779 bytessha256 2abcc64a0265
Whole-screen stimuli with only targeted knobs variedrecord 1
{
  "id": "I01",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Whole-screen stimuli with only targeted knobs varied",
  "evidence_class": "inferred",
  "source": "D-optimal / Bayesian adaptive DCE design (idefix); live demo already does this",
  "observed": "2026-08-27",
  "claim": "Render complete realistic screens but vary only the knobs the round targets, holding others at the posterior mean.",
  "limitations": "Assumes posterior mean is a sensible hold value early on.",
  "disposition": "ADOPT"
}
Content held constant across a comparisonrecord 2
{
  "id": "I02",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Content held constant across a comparison",
  "evidence_class": "inferred",
  "source": "basic experimental control",
  "observed": "2026-08-27",
  "claim": "Use the same headline/copy/imagery in all cards of a round so content preference cannot confound the token-pack signal.",
  "limitations": "Some knobs (density) interact with content length.",
  "disposition": "ADOPT"
}
P05 corpus components as re-themable controlled stimulirecord 3
{
  "id": "I03",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "P05 corpus components as re-themable controlled stimuli",
  "evidence_class": "observed",
  "source": "research/21st-corpus-audit-2026-08-27.md — 86.7% of colour-bearing CSS rules already resolve through var(--token); 87.3% carry a shadcn oklch :root block",
  "observed": "2026-08-27",
  "claim": "Use the 8,515-identity corpus as the stimulus pool: identical structure, varying token pack, re-theming is a find-and-replace on ~30 custom properties.",
  "limitations": "271 components lack previews; component SOURCE is absent from the corpus.",
  "disposition": "ADOPT"
}
Stimulus realism gradient: abstract swatches -> component -> full screenrecord 4
{
  "id": "I04",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Stimulus realism gradient: abstract swatches -> component -> full screen",
  "evidence_class": "hypothesis",
  "source": "none",
  "observed": "2026-08-27",
  "claim": "Early rounds use cheap abstract stimuli, later rounds full screens.",
  "limitations": "Risks measuring preference for the abstraction, not the design.",
  "disposition": "test"
}
Never re-show an identical stimulusrecord 5
{
  "id": "I05",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Never re-show an identical stimulus",
  "evidence_class": "inferred",
  "source": "Zajonc 1968 mere exposure",
  "observed": "2026-08-27",
  "claim": "Prevents familiarity inflating preference for repeated designs.",
  "limitations": "Constrains the generator's search near convergence.",
  "disposition": "ADOPT"
}
Log stimulus-repetition rate as a measured confoundrecord 6
{
  "id": "I06",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Log stimulus-repetition rate as a measured confound",
  "evidence_class": "inferred",
  "source": "Zajonc 1968",
  "observed": "2026-08-27",
  "claim": "Even if repeats occur, make the exposure effect measurable rather than invisible.",
  "limitations": "—",
  "disposition": "ADOPT"
}
Industry-conditioned stimulus priorsrecord 7
{
  "id": "I07",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Industry-conditioned stimulus priors",
  "evidence_class": "inferred",
  "source": "b2b-template-shelf-report.md 17-industry variant deltas",
  "observed": "2026-08-27",
  "claim": "Seed the stimulus space from an industry prior (a law firm starts in a more conservative region than a course creator).",
  "limitations": "Risks stereotyping; must remain a prior, not a constraint.",
  "disposition": "ADOPT"
}
Brand-constraint hard cutoffs before elicitationrecord 8
{
  "id": "I08",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Brand-constraint hard cutoffs before elicitation",
  "evidence_class": "observed",
  "source": "Sawtooth ACBC must-have/unacceptable mechanism (vendor manual)",
  "observed": "2026-08-27",
  "claim": "Existing brand colours/fonts become hard prunes of the stimulus space, not soft preferences; all later stimuli satisfy them.",
  "limitations": "Over-pruning can empty the space; needs a fallback.",
  "disposition": "ADOPT"
}
Dark-mode variant shown as part of the stimulusrecord 9
{
  "id": "I09",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Dark-mode variant shown as part of the stimulus",
  "evidence_class": "inferred",
  "source": "token-pack-science §3.9 — dark is not inversion",
  "observed": "2026-08-27",
  "claim": "Show each candidate in both modes since packs must ship both.",
  "limitations": "Doubles visual load per card.",
  "disposition": "test"
}
Density stress-test stimuli using realistic long copyrecord 10
{
  "id": "I10",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Density stress-test stimuli using realistic long copy",
  "evidence_class": "inferred",
  "source": "ui-pick-to-spec §6 failure mode 6",
  "observed": "2026-08-27",
  "claim": "Include a card rendered with realistically long client copy so density preference is measured under real conditions.",
  "limitations": "—",
  "disposition": "test"
}
Fragment rounds reserved for fine disambiguation onlyrecord 11
{
  "id": "I11",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Fragment rounds reserved for fine disambiguation only",
  "evidence_class": "hypothesis",
  "source": "none",
  "observed": "2026-08-27",
  "claim": "Use component fragments only when two knobs remain confounded at whole-screen scale.",
  "limitations": "Unproven that fragments resolve better.",
  "disposition": "test — this is falsifier F4"
}
Adaptive grid size by viewportrecord 12
{
  "id": "I12",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Adaptive grid size by viewport",
  "evidence_class": "observed",
  "source": "Sequential Gallery reduced 5x5 to 3x3 on a 13-inch display",
  "observed": "2026-08-27",
  "claim": "Gallery size should adapt to display, not be a fixed product constant.",
  "limitations": "—",
  "disposition": "ADOPT"
}
Stimuli drawn from the closed pack catalogue onlyrecord 13
{
  "id": "I13",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Stimuli drawn from the closed pack catalogue only",
  "evidence_class": "inferred",
  "source": "token-pack-science gates A-J",
  "observed": "2026-08-27",
  "claim": "Never show a card that is not a real, gate-passing pack — so the preview IS the deliverable.",
  "limitations": "Restricts early exploration to catalogue coverage.",
  "disposition": "ADOPT"
}
Preview must be re-rendered from spec, never the original thumbnailrecord 14
{
  "id": "I14",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Preview must be re-rendered from spec, never the original thumbnail",
  "evidence_class": "observed",
  "source": "ui-pick-to-spec §6 operational consequence",
  "observed": "2026-08-27",
  "claim": "Closes the expectation gap by construction and makes previews a free end-to-end pipeline test.",
  "limitations": "Requires the render pipeline before elicitation ships.",
  "disposition": "ADOPT"
}
Anti-prototypicality proberecord 15
{
  "id": "I15",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Anti-prototypicality probe",
  "evidence_class": "inferred",
  "source": "Tuch et al. 2012; Reber et al. 2004",
  "observed": "2026-08-27",
  "claim": "Deliberately include one low-prototypicality card per round to detect clients who want distinctive over safe.",
  "limitations": "May be systematically rejected, wasting a slot.",
  "disposition": "test"
}
Seed from client's existing website via extractionrecord 16
{
  "id": "I16",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Seed from client's existing website via extraction",
  "evidence_class": "observed",
  "source": "Canva AI extracts brand context from a public URL",
  "observed": "2026-08-27",
  "claim": "Warm-start the prior from the client's current site rather than a uniform prior.",
  "limitations": "Their current site may be exactly what they want to escape.",
  "disposition": "test"
}
Competitor-anchored stimulirecord 17
{
  "id": "I17",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Competitor-anchored stimuli",
  "evidence_class": "hypothesis",
  "source": "none",
  "observed": "2026-08-27",
  "claim": "Show packs near and far from named competitors to elicit differentiation preference.",
  "limitations": "Conflates taste with positioning.",
  "disposition": "test"
}
Stimulus set diversity floorrecord 18
{
  "id": "I18",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Stimulus set diversity floor",
  "evidence_class": "inferred",
  "source": "Negahban et al. spectral gap dependence",
  "observed": "2026-08-27",
  "claim": "Enforce a minimum pairwise distance among the 4 cards so the comparison graph stays well-connected.",
  "limitations": "Competes with least-resolved-knob targeting, which wants NEAR pairs.",
  "disposition": "ADOPT with tension noted"
}
4-up gallery with explicit outside optionrecord 19
{
  "id": "I19",
  "kind": "innovation",
  "group": "question-format",
  "name": "4-up gallery with explicit outside option",
  "evidence_class": "inferred",
  "source": "derived §3 + Sawtooth 4-5 per screen + live demo",
  "observed": "2026-08-27",
  "claim": "2.32 noiseless bits/round; ~0.91 effective at p=0.15; balances information against deliberation cost and avoids the hostility of 'worst'.",
  "limitations": "Best-worst is strictly more informative per question.",
  "disposition": "ADOPT — primary recommendation"
}
Best-worst of 4 (MaxDiff Case 1)record 20
{
  "id": "I20",
  "kind": "innovation",
  "group": "question-format",
  "name": "Best-worst of 4 (MaxDiff Case 1)",
  "evidence_class": "observed",
  "source": "Marley & Louviere 2005; Sawtooth manual",
  "observed": "2026-08-27",
  "claim": "3.58 noiseless bits/round, ~2x the 4-up rate; would cut rounds to ~3.",
  "limitations": "Roughly doubles per-screen deliberation; 'pick the worst' is hostile in a client sales context; no verified numeric information-gain multiplier exists.",
  "disposition": "test as an A/B arm — falsifier F6"
}
Attribute-level best-worst (MaxDiff Case 2)record 21
{
  "id": "I21",
  "kind": "innovation",
  "group": "question-format",
  "name": "Attribute-level best-worst (MaxDiff Case 2)",
  "evidence_class": "inferred",
  "source": "Marley, Flynn & Louviere 2008; support.BWS2",
  "observed": "2026-08-27",
  "claim": "'Which knob is most and least right on THIS design' — maps almost exactly onto per-knob credit assignment, solving the attribution problem a whole-card pick has.",
  "limitations": "Requires the client to reason about knobs explicitly, which is exactly what we said they cannot do.",
  "disposition": "test — high upside, high risk"
}
Binary pairwiserecord 22
{
  "id": "I22",
  "kind": "innovation",
  "group": "question-format",
  "name": "Binary pairwise",
  "evidence_class": "inferred",
  "source": "derived §3",
  "observed": "2026-08-27",
  "claim": "Simplest, but ~11 rounds at p=0.15 and most sensitive to near-ties.",
  "limitations": "Nearly double the rounds of 4-up.",
  "disposition": "reject as primary"
}
Three-way pairwise with a neutral middlerecord 23
{
  "id": "I23",
  "kind": "innovation",
  "group": "question-format",
  "name": "Three-way pairwise with a neutral middle",
  "evidence_class": "observed",
  "source": "Midjourney legacy Style Tuner: 'leave the middle box selected to skip the pair'",
  "observed": "2026-08-27",
  "claim": "A shipped outside-option affordance inside a pairwise format.",
  "limitations": "Fewer bits than 4-up.",
  "disposition": "reference — the affordance, not the format"
}
Full ranking of 4record 24
{
  "id": "I24",
  "kind": "innovation",
  "group": "question-format",
  "name": "Full ranking of 4",
  "evidence_class": "inferred",
  "source": "Plackett 1975; Hajek, Oh & Xu 2014",
  "observed": "2026-08-27",
  "claim": "4.58 noiseless bits, the densest format tested.",
  "limitations": "Ranking four whole screens is a heavy cognitive task; inherits IIA.",
  "disposition": "reject"
}
9-up galleryrecord 25
{
  "id": "I25",
  "kind": "innovation",
  "group": "question-format",
  "name": "9-up gallery",
  "evidence_class": "inferred",
  "source": "derived §3; Sequential Gallery used 5x5 and 3x3",
  "observed": "2026-08-27",
  "claim": "3.17 noiseless bits.",
  "limitations": "Nine whole-screen renders exceed comfortable simultaneous comparison; selection noise rises in a way the bit count does not capture.",
  "disposition": "reject for whole screens"
}
Mixed format by phase: 4-up early, pairwise laterecord 26
{
  "id": "I26",
  "kind": "innovation",
  "group": "question-format",
  "name": "Mixed format by phase: 4-up early, pairwise late",
  "evidence_class": "hypothesis",
  "source": "none",
  "observed": "2026-08-27",
  "claim": "Broad exploration wants breadth; fine disambiguation is naturally pairwise.",
  "limitations": "Format switching may confuse users and complicates the likelihood.",
  "disposition": "test"
}
Slider direct manipulation as a fallbackrecord 27
{
  "id": "I27",
  "kind": "innovation",
  "group": "question-format",
  "name": "Slider direct manipulation as a fallback",
  "evidence_class": "observed",
  "source": "Figma First Draft radius/spacing sliders",
  "observed": "2026-08-27",
  "claim": "Offer sliders to clients who DO know what they want, skipping elicitation entirely.",
  "limitations": "Most clients cannot name a radius; the slider is the competitor's approach.",
  "disposition": "ADOPT as escape hatch"
}
Sequential line search (1D slider through knob space)record 28
{
  "id": "I28",
  "kind": "innovation",
  "group": "question-format",
  "name": "Sequential line search (1D slider through knob space)",
  "evidence_class": "observed",
  "source": "Koyama et al. SIGGRAPH 2017 (MIT-licensed implementation exists)",
  "observed": "2026-08-27",
  "claim": "User picks a point on a 1D slice; reported 15-iteration budget, good by iteration 4-5 at 6D/7D.",
  "limitations": "A slider over a design continuum is harder to render than 4 discrete packs, and our output space is discrete anyway.",
  "disposition": "study"
}
Outside option as weak negative evidence on all shown cardsrecord 29
{
  "id": "I29",
  "kind": "innovation",
  "group": "question-format",
  "name": "Outside option as weak negative evidence on all shown cards",
  "evidence_class": "hypothesis",
  "source": "derived; diverges from Midjourney which makes skipping non-informative",
  "observed": "2026-08-27",
  "claim": "Preserves information from declines while keeping the re-roll honest.",
  "limitations": "Weight is a free parameter needing calibration.",
  "disposition": "ADOPT — calibrate the weight"
}
Cap consecutive outside-option selections at 3record 30
{
  "id": "I30",
  "kind": "innovation",
  "group": "question-format",
  "name": "Cap consecutive outside-option selections at 3",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Three declines in a row means the generator is in the wrong region; trigger a re-seed rather than another round.",
  "limitations": "Threshold is arbitrary pending data.",
  "disposition": "ADOPT"
}
'This is right' always-available terminal affordancerecord 31
{
  "id": "I31",
  "kind": "innovation",
  "group": "question-format",
  "name": "'This is right' always-available terminal affordance",
  "evidence_class": "observed",
  "source": "Sequential Gallery satisfaction button; mean 5.36 rounds to satisfaction",
  "observed": "2026-08-27",
  "claim": "User's own judgement is the criterion the whole exercise proxies for.",
  "limitations": "Users may satisfice early.",
  "disposition": "ADOPT"
}
Preview-before-commit on the selected packrecord 32
{
  "id": "I32",
  "kind": "innovation",
  "group": "question-format",
  "name": "Preview-before-commit on the selected pack",
  "evidence_class": "observed",
  "source": "Webflow 'Preview in Designer'",
  "observed": "2026-08-27",
  "claim": "Let the client see the pack applied to their real content before committing.",
  "limitations": "Requires the render pipeline.",
  "disposition": "ADOPT"
}
Paired-comparison tie option distinct from 'none'record 33
{
  "id": "I33",
  "kind": "innovation",
  "group": "question-format",
  "name": "Paired-comparison tie option distinct from 'none'",
  "evidence_class": "inferred",
  "source": "prefmod handles undecided/ties",
  "observed": "2026-08-27",
  "claim": "'These two are equally good' is different information from 'both are wrong'.",
  "limitations": "Adds a third response category to model.",
  "disposition": "test"
}
Confidence-weighted responses (client marks how sure they are)record 34
{
  "id": "I34",
  "kind": "innovation",
  "group": "question-format",
  "name": "Confidence-weighted responses (client marks how sure they are)",
  "evidence_class": "hypothesis",
  "source": "none",
  "observed": "2026-08-27",
  "claim": "Let the client flag a low-confidence pick so it is down-weighted.",
  "limitations": "Self-reported confidence is poorly calibrated; adds friction.",
  "disposition": "reject"
}
BALD acquisition on a GP preference modelrecord 35
{
  "id": "I35",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "BALD acquisition on a GP preference model",
  "evidence_class": "observed",
  "source": "Houlsby et al. 2011 (arXiv 1112.5745)",
  "observed": "2026-08-27",
  "claim": "Information gain expressed via predictive entropies, explicitly extended to GP preference learning; tractable.",
  "limitations": "Myopic one-step criterion.",
  "disposition": "ADOPT"
}
qEUBO acquisitionrecord 36
{
  "id": "I36",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "qEUBO acquisition",
  "evidence_class": "observed",
  "source": "Astudillo et al. AISTATS 2023; shipped in BoTorch",
  "observed": "2026-08-27",
  "claim": "One-step Bayes optimal under noise-free responses; simple regret converges at o(1/n); qEI can FAIL to converge for PBO.",
  "limitations": "Asymptotic guarantee gives no finite-sample count.",
  "disposition": "ADOPT"
}
Target the least-resolved knob each roundrecord 37
{
  "id": "I37",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Target the least-resolved knob each round",
  "evidence_class": "observed",
  "source": "live demo already does this",
  "observed": "2026-08-27",
  "claim": "Constructs cards differing on the widest-posterior axis while holding others steady.",
  "limitations": "Greedy per-knob targeting can miss interaction effects.",
  "disposition": "ADOPT"
}
Explicit exploration/exploitation schedule ('candy vs medicine')record 38
{
  "id": "I38",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Explicit exploration/exploitation schedule ('candy vs medicine')",
  "evidence_class": "observed",
  "source": "Stitch Fix Style Shuffle practice (press-reported)",
  "observed": "2026-08-27",
  "claim": "Deliberately mix near-certain-hit cards (engagement) with high-information cards (learning).",
  "limitations": "Ratio unpublished; costs information per round.",
  "disposition": "ADOPT — calibrate ratio"
}
Hierarchical Bayes shrinkage toward a population priorrecord 39
{
  "id": "I39",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Hierarchical Bayes shrinkage toward a population prior",
  "evidence_class": "observed",
  "source": "ChoiceModelR; standard conjoint practice",
  "observed": "2026-08-27",
  "claim": "Designed exactly for the few-observations-per-individual regime; the primary anti-overfitting mechanism alongside closed packs.",
  "limitations": "Needs a population of prior clients to shrink toward.",
  "disposition": "ADOPT — but cold-start it from I40"
}
Warm-start prior from a population aesthetic modelrecord 40
{
  "id": "I40",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Warm-start prior from a population aesthetic model",
  "evidence_class": "observed",
  "source": "LAION aesthetic predictor; Brochu et al. 2010 cut iterations 11.25 -> 6.5 with a learned prior",
  "observed": "2026-08-27",
  "claim": "Learn the client's DEVIATION from a population baseline rather than their taste from scratch. The largest single efficiency lever reported anywhere in the applied literature.",
  "limitations": "Population aesthetic is not per-client taste and may encode a generic bias.",
  "disposition": "ADOPT"
}
D-optimal / Bayesian-efficient stimulus set selectionrecord 41
{
  "id": "I41",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "D-optimal / Bayesian-efficient stimulus set selection",
  "evidence_class": "observed",
  "source": "idefix (R, GPL-3)",
  "observed": "2026-08-27",
  "claim": "Choose the 4 cards to maximise design efficiency under the current posterior.",
  "limitations": "GPL-3 licence; requires priors.",
  "disposition": "ADOPT-METHOD, reimplement"
}
Simulate-to-size: Monte Carlo power analysis for round countrecord 42
{
  "id": "I42",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Simulate-to-size: Monte Carlo power analysis for round count",
  "evidence_class": "observed",
  "source": "skpr (R)",
  "observed": "2026-08-27",
  "claim": "Answer 'how many rounds' EMPIRICALLY by simulating synthetic clients under a BT likelihood rather than arguing from bounds.",
  "limitations": "Power analysis is built for GLM responses; needs adapting.",
  "disposition": "ADOPT — this is how the N question should actually be settled"
}
Orthogonal-array fixed design as the passive baselinerecord 43
{
  "id": "I43",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Orthogonal-array fixed design as the passive baseline",
  "evidence_class": "observed",
  "source": "DoE.base",
  "observed": "2026-08-27",
  "claim": "The non-adaptive control our adaptive policy must beat.",
  "limitations": "—",
  "disposition": "ADOPT as control arm"
}
Projective preferential BO for high-dim knob spacesrecord 44
{
  "id": "I44",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Projective preferential BO for high-dim knob spaces",
  "evidence_class": "observed",
  "source": "AaltoPML/PPBO (MIT)",
  "observed": "2026-08-27",
  "claim": "Queries along projections, designed for human-in-the-loop high-dimensional elicitation.",
  "limitations": "Small research repo (20 stars).",
  "disposition": "study"
}
Transitivity-based pair elimination (PAPRIKA-style)record 45
{
  "id": "I45",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Transitivity-based pair elimination (PAPRIKA-style)",
  "evidence_class": "hypothesis",
  "source": "1000minds method; COULD NOT VERIFY — 404 on both URLs",
  "observed": "2026-08-27",
  "claim": "Eliminate implied comparisons by transitivity, drastically cutting explicit questions.",
  "limitations": "SOURCE UNVERIFIED. Also assumes transitivity, which human taste violates (Ailon 2010).",
  "disposition": "VERIFY FIRST, then test"
}
Non-transitivity-tolerant active rankingrecord 46
{
  "id": "I46",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Non-transitivity-tolerant active ranking",
  "evidence_class": "observed",
  "source": "Ailon 2010 (arXiv 1011.0108)",
  "observed": "2026-08-27",
  "claim": "Explicitly handles 'non-transitivity paradoxes which may arise naturally due to human mistakes or irrationality'.",
  "limitations": "No closed-form bound extracted from the abstract.",
  "disposition": "study"
}
Copeland counting as a simplicity checkrecord 47
{
  "id": "I47",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Copeland counting as a simplicity check",
  "evidence_class": "observed",
  "source": "Shah & Wainwright 2015",
  "observed": "2026-08-27",
  "claim": "Rank by comparisons won: optimal up to CONSTANT factors, no conditions on the probability matrix. If this matches the GP model's output, the GP is not earning its complexity.",
  "limitations": "Top-k recovery framing, not utility-vector recovery.",
  "disposition": "ADOPT as a baseline check"
}
Dueling-bandit formulationrecord 48
{
  "id": "I48",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Dueling-bandit formulation",
  "evidence_class": "observed",
  "source": "Yue & Joachims; Zoghi RUCB; Sui et al. survey",
  "observed": "2026-08-27",
  "claim": "Relative-feedback online optimisation.",
  "limitations": "Regret-minimisation over continuous operation, not fixed-budget identification — a different objective from ours.",
  "disposition": "reject as primary framing"
}
Per-knob independent BT models vs joint GPrecord 49
{
  "id": "I49",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Per-knob independent BT models vs joint GP",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Independent per-knob models are simpler and interpretable; a joint GP captures correlation.",
  "limitations": "Independence is empirically false (§4 of first-principles).",
  "disposition": "test — joint expected to win"
}
Bootstrap confidence intervals over pairwise judgmentsrecord 50
{
  "id": "I50",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Bootstrap confidence intervals over pairwise judgments",
  "evidence_class": "observed",
  "source": "lmarena/arena-hard-auto",
  "observed": "2026-08-27",
  "claim": "Battle-tested approach to reporting uncertainty on BT fits from human votes.",
  "limitations": "—",
  "disposition": "ADOPT"
}
TrueSkill-style Gaussian belief per packrecord 51
{
  "id": "I51",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "TrueSkill-style Gaussian belief per pack",
  "evidence_class": "observed",
  "source": "Herbrich et al. 2007; use openskill.py (MIT), NOT sublee/trueskill",
  "observed": "2026-08-27",
  "claim": "Calibrated per-item uncertainty that Elo does not give.",
  "limitations": "LICENCE TRAP: sublee/trueskill's LICENSE body bars commercial use despite a BSD badge.",
  "disposition": "study — openskill.py only"
}
Stop when expected information gain of the best next question < thresholdrecord 52
{
  "id": "I52",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Stop when expected information gain of the best next question < threshold",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "A wide posterior no available question can narrow is a reason to stop, not to continue.",
  "limitations": "Threshold (~0.25 bits) needs calibration.",
  "disposition": "ADOPT"
}
Three-condition stop: confidence OR ceiling OR user-satisfiedrecord 53
{
  "id": "I53",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Three-condition stop: confidence OR ceiling OR user-satisfied",
  "evidence_class": "inferred",
  "source": "derived §6; Brochu 2010 20-iteration abandonment; Midjourney 'past round 15... small and subtle'",
  "observed": "2026-08-27",
  "claim": "Primary rule is confidence; 15 rounds is a ceiling; user satisfaction always terminates.",
  "limitations": "Ceiling is grounded in two sources, not a derivation.",
  "disposition": "ADOPT"
}
Hard floor of 5 roundsrecord 54
{
  "id": "I54",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Hard floor of 5 rounds",
  "evidence_class": "inferred",
  "source": "derived; Sequential Gallery mean 5.36",
  "observed": "2026-08-27",
  "claim": "Early apparent convergence is usually the prior, not the data.",
  "limitations": "—",
  "disposition": "ADOPT"
}
Per-knob confidence reporting, never a single scalarrecord 55
{
  "id": "I55",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Per-knob confidence reporting, never a single scalar",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Some knobs resolve in two rounds; others may never resolve, and 'this client does not care about shadow' is a finding not a failure. A single '87% confident' number would mislead.",
  "limitations": "More complex to present to a non-technical client.",
  "disposition": "ADOPT"
}
'No preference' as a first-class per-knob outcomerecord 56
{
  "id": "I56",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "'No preference' as a first-class per-knob outcome",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Explicitly represent indifference rather than forcing a point estimate.",
  "limitations": "Downstream pack selection must handle a free knob.",
  "disposition": "ADOPT"
}
Two-week test-retest stability as the primary validation metricrecord 57
{
  "id": "I57",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Two-week test-retest stability as the primary validation metric",
  "evidence_class": "inferred",
  "source": "Slovic 1995 construction of preference",
  "observed": "2026-08-27",
  "claim": "A model that fits the clicks but changes answer next Tuesday is worthless. Measures whether the thing we claim to measure exists.",
  "limitations": "Requires client time two weeks apart; hard to get.",
  "disposition": "ADOPT — this is falsifier F7"
}
Held-out pick prediction as the fit metricrecord 58
{
  "id": "I58",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Held-out pick prediction as the fit metric",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Reserve 2 rounds, predict them from the model fitted on the rest.",
  "limitations": "Small n makes the estimate noisy.",
  "disposition": "ADOPT"
}
Neighbour-discrimination testrecord 59
{
  "id": "I59",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Neighbour-discrimination test",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Show the client their pack plus the 2nd and 3rd nearest, unlabelled. If they cannot pick their own above chance, the catalogue is denser than perception.",
  "limitations": "—",
  "disposition": "ADOPT — falsifier F5"
}
Perturbation-tolerance testrecord 60
{
  "id": "I60",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Perturbation-tolerance test",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Show the converged pack and a +/-1-level perturbation. If clients reliably distinguish and reject, the tolerance argument in §2 collapses and round counts triple.",
  "limitations": "—",
  "disposition": "ADOPT — falsifier F1, run FIRST"
}
Separation-aware stoppingrecord 61
{
  "id": "I61",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Separation-aware stopping",
  "evidence_class": "observed",
  "source": "Chen & Suh 2015 — complexity scales inversely with separation",
  "observed": "2026-08-27",
  "claim": "Detect when remaining candidates are genuinely near-tied and stop rather than burning rounds distinguishing the indistinguishable.",
  "limitations": "—",
  "disposition": "ADOPT"
}
Comparison-graph connectivity checkrecord 62
{
  "id": "I62",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Comparison-graph connectivity check",
  "evidence_class": "observed",
  "source": "Hunter 2004 MM convergence condition",
  "observed": "2026-08-27",
  "claim": "BT/MM fitting requires a strongly connected comparison graph; verify before fitting.",
  "limitations": "Constrains which stimulus sets are legal.",
  "disposition": "ADOPT as a gate"
}
Abandonment-rate monitoring as a UX stop signalrecord 63
{
  "id": "I63",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Abandonment-rate monitoring as a UX stop signal",
  "evidence_class": "observed",
  "source": "Brochu et al. 2010: 20 iterations is 'roughly the point at which users start to quit'",
  "observed": "2026-08-27",
  "claim": "Instrument drop-off per round; if it rises before the ceiling, lower the ceiling.",
  "limitations": "—",
  "disposition": "ADOPT"
}
Time-per-round budgetrecord 64
{
  "id": "I64",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Time-per-round budget",
  "evidence_class": "observed",
  "source": "Sequential Gallery: 14.8s per plane-search subtask",
  "observed": "2026-08-27",
  "claim": "Budget total elicitation time, not just round count. 10 rounds x ~15s is ~2.5 minutes.",
  "limitations": "Whole-screen comparison likely slower than the cited subtask.",
  "disposition": "ADOPT"
}
Confidence decay over calendar timerecord 65
{
  "id": "I65",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Confidence decay over calendar time",
  "evidence_class": "inferred",
  "source": "Netflix supersession rule; Hinge 24h window",
  "observed": "2026-08-27",
  "claim": "Treat the profile as decaying, prompting re-elicitation rather than assuming permanence.",
  "limitations": "Decay rate unknown; needs I57 data first.",
  "disposition": "test"
}
Re-elicitation triggered by client rejection of built outputrecord 66
{
  "id": "I66",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Re-elicitation triggered by client rejection of built output",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "If a client rejects the delivered design, that is strong evidence to re-run rather than patch.",
  "limitations": "Expensive and reads as failure.",
  "disposition": "test"
}
Continuous preference vector internally, closed pack at the boundaryrecord 67
{
  "id": "I67",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Continuous preference vector internally, closed pack at the boundary",
  "evidence_class": "inferred",
  "source": "derived §8; token-pack-science gates A-J; lane synthesis 'learn a vector not a pack ID'",
  "observed": "2026-08-27",
  "claim": "Catalogue can grow without re-eliciting; shipped artefact is always gate-passing.",
  "limitations": "Nearest-neighbour metric weighting is an open parameter.",
  "disposition": "ADOPT — primary recommendation"
}
Reject continuous token interpolationrecord 68
{
  "id": "I68",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Reject continuous token interpolation",
  "evidence_class": "observed",
  "source": "token-pack-science §5 Gate C — WCAG luminance is non-linear in channel values",
  "observed": "2026-08-27",
  "claim": "The midpoint of two AA-passing palettes can fail AA; an interpolated pack has passed none of gates A-J.",
  "limitations": "Loses expressive range between packs.",
  "disposition": "ADOPT the rejection"
}
Closed-pack snapping AS the anti-overfitting regularizerrecord 69
{
  "id": "I69",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Closed-pack snapping AS the anti-overfitting regularizer",
  "evidence_class": "hypothesis",
  "source": "derived §8",
  "observed": "2026-08-27",
  "claim": "Output space of ~20-30 packs rather than 36,000 means the model cannot overfit into a bespoke corner on ten noisy clicks.",
  "limitations": "If the catalogue grows large this protection weakens.",
  "disposition": "ADOPT"
}
Explain the pack choice in knob termsrecord 70
{
  "id": "I70",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Explain the pack choice in knob terms",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "'We chose this because you are here in knob space' is worth real money in the client conversation.",
  "limitations": "Requires knob names clients understand.",
  "disposition": "ADOPT"
}
Weighted knob-space distance metricrecord 71
{
  "id": "I71",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Weighted knob-space distance metric",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Typography and chroma likely dominate perception; equal weighting is probably wrong.",
  "limitations": "Weights must be fitted, adding parameters.",
  "disposition": "test"
}
Second- and third-choice packs offered alongside the firstrecord 72
{
  "id": "I72",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Second- and third-choice packs offered alongside the first",
  "evidence_class": "inferred",
  "source": "derived; supports I59",
  "observed": "2026-08-27",
  "claim": "Offering the top 3 both hedges model error and generates validation data for free.",
  "limitations": "Reintroduces choice at the moment we claimed to have decided.",
  "disposition": "ADOPT"
}
TasteProfile schema: per-knob posterior mean + interval + n_observationsrecord 73
{
  "id": "I73",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "TasteProfile schema: per-knob posterior mean + interval + n_observations",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Makes uncertainty first-class and auditable in the output contract.",
  "limitations": "—",
  "disposition": "ADOPT"
}
PreferenceConfidence as a per-knob vector, not a scalarrecord 74
{
  "id": "I74",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "PreferenceConfidence as a per-knob vector, not a scalar",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Mirrors I55 in the output contract.",
  "limitations": "—",
  "disposition": "ADOPT"
}
DesignDNA as a versioned, reproducible artefactrecord 75
{
  "id": "I75",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "DesignDNA as a versioned, reproducible artefact",
  "evidence_class": "inferred",
  "source": "token-pack-science pack.version discipline",
  "observed": "2026-08-27",
  "claim": "Pin the model version, catalogue version, and elicitation transcript so a profile is reproducible.",
  "limitations": "—",
  "disposition": "ADOPT"
}
Preference vector pre-filters P05 component suggestionsrecord 76
{
  "id": "I76",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Preference vector pre-filters P05 component suggestions",
  "evidence_class": "observed",
  "source": "live demo states this intent",
  "observed": "2026-08-27",
  "claim": "The same vector that picks a pack can rank components, so elicitation pays for itself twice.",
  "limitations": "Component preference may not be the same construct as pack preference.",
  "disposition": "test"
}
Preference vector seeds image-generation roundsrecord 77
{
  "id": "I77",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Preference vector seeds image-generation rounds",
  "evidence_class": "observed",
  "source": "live demo states this intent",
  "observed": "2026-08-27",
  "claim": "Reuse for generated imagery.",
  "limitations": "Unverified that the vector transfers to a different modality.",
  "disposition": "test"
}
Per-context confidence (marketing page vs dense dashboard)record 78
{
  "id": "I78",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Per-context confidence (marketing page vs dense dashboard)",
  "evidence_class": "hypothesis",
  "source": "derived; parts.json open question",
  "observed": "2026-08-27",
  "claim": "A client's density preference on a landing page may differ from a dashboard. Model context as a conditioning variable rather than assuming one profile fits all surfaces.",
  "limitations": "Multiplies the parameter count by the number of contexts, directly worsening the sample-size problem.",
  "disposition": "test — this is the open question with the worst cost/benefit"
}
Single profile with per-context OFFSETS rather than separate profilesrecord 79
{
  "id": "I79",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Single profile with per-context OFFSETS rather than separate profiles",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "A cheaper resolution of I78: one taste vector plus small learned context deltas.",
  "limitations": "Offsets still need data to fit.",
  "disposition": "test — preferred form of I78"
}
Stated vs revealed preference agreement metricrecord 80
{
  "id": "I80",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Stated vs revealed preference agreement metric",
  "evidence_class": "observed",
  "source": "Pinterest deprecated onboarding interests in favour of engagement clusters",
  "observed": "2026-08-27",
  "claim": "Measure whether what the client PICKS matches what they later APPROVE in delivered work. The single most commercially important validation.",
  "limitations": "Requires delivered projects, so it is a slow metric.",
  "disposition": "ADOPT — long-horizon"
}
Fall back to a safe default pack on low confidencerecord 81
{
  "id": "I81",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Fall back to a safe default pack on low confidence",
  "evidence_class": "inferred",
  "source": "Shopify SDUI unknown-layout fallback; Netflix popular-set fallback",
  "observed": "2026-08-27",
  "claim": "If confidence is too low, ship an industry-appropriate default rather than a badly-inferred pack.",
  "limitations": "—",
  "disposition": "ADOPT"
}
Log spec/pack collisions as a coverage signalrecord 82
{
  "id": "I82",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Log spec/pack collisions as a coverage signal",
  "evidence_class": "observed",
  "source": "ui-pick-to-spec §6 failure mode 4",
  "observed": "2026-08-27",
  "claim": "If two different clients converge to identical packs, that tells you which knob to add.",
  "limitations": "—",
  "disposition": "ADOPT"
}
Human review gate before pack is committed to a buildrecord 83
{
  "id": "I83",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Human review gate before pack is committed to a build",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "A designer confirms the selected pack before it drives a client deliverable.",
  "limitations": "Costs human time per client; undermines the automation claim.",
  "disposition": "ADOPT for v1, revisit"
}
Profile portability across a client's multiple projectsrecord 84
{
  "id": "I84",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Profile portability across a client's multiple projects",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Reuse the DesignDNA for the same client's later work.",
  "limitations": "Preference may be project-specific, not client-specific.",
  "disposition": "test"
}
Explicit re-elicitation rather than silent drift trackingrecord 85
{
  "id": "I85",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Explicit re-elicitation rather than silent drift tracking",
  "evidence_class": "inferred",
  "source": "Netflix supersession; Pinterest UIC replacement",
  "observed": "2026-08-27",
  "claim": "Both major platforms found stated onboarding preference decays against behaviour. Prefer a cheap explicit re-run over an invisible decay model.",
  "limitations": "Costs client time.",
  "disposition": "ADOPT"
}
Revealed-preference signal from delivered-work approvalsrecord 86
{
  "id": "I86",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Revealed-preference signal from delivered-work approvals",
  "evidence_class": "inferred",
  "source": "Hinge infers rankings from like/pass; Stitch Fix updates in real time",
  "observed": "2026-08-27",
  "claim": "Client approvals/rejections of delivered designs are the highest-quality preference signal available and cost nothing to collect.",
  "limitations": "Slow, sparse, and confounded with non-aesthetic factors (deadline, budget).",
  "disposition": "ADOPT — long-horizon"
}
Regulated industries constrain the stimulus space a priorirecord 87
{
  "id": "I87",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Regulated industries constrain the stimulus space a priori",
  "evidence_class": "inferred",
  "source": "b2b-template-shelf-report.md: healthcare PHI, law firm privilege, insurance document authority",
  "observed": "2026-08-27",
  "claim": "Healthcare, legal, insurance, mortgage clients should not be shown high-chroma, high-motion, low-contrast packs at all — prune before eliciting.",
  "limitations": "Risks the pruning being wrong and the client wanting distinctiveness.",
  "disposition": "ADOPT"
}
Accessibility floor is not a preference axisrecord 88
{
  "id": "I88",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Accessibility floor is not a preference axis",
  "evidence_class": "observed",
  "source": "token-pack-science Gates C and D",
  "observed": "2026-08-27",
  "claim": "Contrast below AA is never offered regardless of stated preference. The client cannot choose to fail WCAG.",
  "limitations": "Removes a genuine aesthetic direction (low-contrast minimalism).",
  "disposition": "ADOPT"
}
End-user vs operator preference divergencerecord 89
{
  "id": "I89",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "End-user vs operator preference divergence",
  "evidence_class": "inferred",
  "source": "b2b-template-shelf-report.md: property_management tenant portal is 'a distinct untrusted identity'; it_services_msps 'per-client tenancy is the product'",
  "observed": "2026-08-27",
  "claim": "For the 6/17 industries where `portal` is secondary, the person choosing (the operator) is not the person using (the end customer). Elicit from the operator but weight toward end-user norms.",
  "limitations": "No mechanism yet for representing the absent end user.",
  "disposition": "ADOPT — open design problem"
}
Density preference conditioned on archetyperecord 90
{
  "id": "I90",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Density preference conditioned on archetype",
  "evidence_class": "inferred",
  "source": "b2b-template-shelf: case_workflow primary for 6/17, portal secondary for 6/17",
  "observed": "2026-08-27",
  "claim": "Two archetypes carry the majority of catalogue demand; density expectations differ sharply between a dense case-workflow console and a sparse client portal.",
  "limitations": "Adds a conditioning variable.",
  "disposition": "ADOPT — cheapest useful form of I78"
}
Conservative-industry prior on chroma and motionrecord 91
{
  "id": "I91",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Conservative-industry prior on chroma and motion",
  "evidence_class": "inferred",
  "source": "accounting_firms, law_firms, insurance_agencies, mortgage_brokers variant deltas",
  "observed": "2026-08-27",
  "claim": "Seed these industries from a low-chroma, low-motion region of knob space.",
  "limitations": "Stereotyping risk; must be a prior the client can override in 2-3 rounds.",
  "disposition": "ADOPT"
}
Expressive-industry priorrecord 92
{
  "id": "I92",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Expressive-industry prior",
  "evidence_class": "inferred",
  "source": "course_creators, marketing_social_media_agencies, real_estate variant deltas",
  "observed": "2026-08-27",
  "claim": "Seed from a higher-chroma, higher-expressiveness region.",
  "limitations": "Same stereotyping risk.",
  "disposition": "ADOPT"
}
Industry prior must be overridable within 3 roundsrecord 93
{
  "id": "I93",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Industry prior must be overridable within 3 rounds",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Guarantee the prior is a starting point, not a cage: a client must be able to move out of their industry's region quickly.",
  "limitations": "Requires measuring prior strength.",
  "disposition": "ADOPT"
}
Bespoke motion and illustration explicitly out of scoperecord 94
{
  "id": "I94",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Bespoke motion and illustration explicitly out of scope",
  "evidence_class": "observed",
  "source": "ui-pick-to-spec §6 failure modes 1 and 2",
  "observed": "2026-08-27",
  "claim": "The spec bottleneck loses hand-tuned parallax and custom artwork. Do not let a knob imply a capability that does not exist.",
  "limitations": "Removes a real source of client delight.",
  "disposition": "ADOPT"
}
The preference learner must itself be A/B tested against a static defaultrecord 95
{
  "id": "I95",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "The preference learner must itself be A/B tested against a static default",
  "evidence_class": "observed",
  "source": "Spotify Engineering: a bandit is 'a feature you've built, not an experimental method'",
  "observed": "2026-08-27",
  "claim": "Without a holdout comparing elicited packs to an industry-default pack, we cannot claim the elicitation works at all.",
  "limitations": "Needs enough clients for a powered comparison — likely the binding constraint.",
  "disposition": "ADOPT — the single most important methodological gate"
}
Measure the knob-correlation matrix before assuming low effective dimensionrecord 96
{
  "id": "I96",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Measure the knob-correlation matrix before assuming low effective dimension",
  "evidence_class": "hypothesis",
  "source": "derived §4; Jamieson & Nowak's d log n is contingent on the embedding assumption holding",
  "observed": "2026-08-27",
  "claim": "The whole few-rounds argument rests on effective d < 7. Measure it, do not assume it.",
  "limitations": "—",
  "disposition": "ADOPT — falsifier F3"
}
Do not cite Miller 7+/-2 to justify gallery sizerecord 97
{
  "id": "I97",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Do not cite Miller 7+/-2 to justify gallery size",
  "evidence_class": "observed",
  "source": "Miller 1956 is about absolute judgment and memory span, not interface option counts",
  "observed": "2026-08-27",
  "claim": "Citing it would not survive a technical client's scrutiny.",
  "limitations": "—",
  "disposition": "ADOPT as a writing rule"
}
Do not cite the jam study as a design principlerecord 98
{
  "id": "I98",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Do not cite the jam study as a design principle",
  "evidence_class": "observed",
  "source": "Scheibehenne et al. 2010: D = 0.02, CI95 [-0.09, 0.12], 63 conditions, N = 5,036",
  "observed": "2026-08-27",
  "claim": "Choice overload does not robustly replicate; size galleries on information grounds.",
  "limitations": "—",
  "disposition": "ADOPT as a writing rule"
}
Do not promise a specific accuracy number before measuringrecord 99
{
  "id": "I99",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Do not promise a specific accuracy number before measuring",
  "evidence_class": "observed",
  "source": "ui-pick-to-spec §4.2 no-figure-found; the 728-vs-113 connector error class",
  "observed": "2026-08-27",
  "claim": "No published benchmark measures our task; any quoted figure would be extrapolation.",
  "limitations": "—",
  "disposition": "ADOPT as a writing rule"
}
Treat the live demo's convergence claims as unmeasuredrecord 100
{
  "id": "I100",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "Treat the live demo's convergence claims as unmeasured",
  "evidence_class": "observed",
  "source": "actionist-taste.pages.dev asserts '~10 picks instead of ~50' with no measurement",
  "observed": "2026-08-27",
  "claim": "Our own prototype's marketing claims must not be recycled as evidence into a client deliverable.",
  "limitations": "—",
  "disposition": "ADOPT as a writing rule"
}
Continuous preference vector internally, closed pack at the boundaryrecord 101
{
  "id": "I67-TOP10",
  "kind": "innovation",
  "group": "output-compilation",
  "name": "Continuous preference vector internally, closed pack at the boundary",
  "evidence_class": "inferred",
  "source": "derived §8; token-pack-science gates A-J; lane synthesis 'learn a vector not a pack ID'",
  "observed": "2026-08-27",
  "claim": "Catalogue can grow without re-eliciting; shipped artefact is always gate-passing.",
  "limitations": "Nearest-neighbour metric weighting is an open parameter.",
  "disposition": "ADOPT — primary recommendation",
  "refers_to": "I67",
  "rank": 1,
  "rationale": "Resolves the central architectural question (d): continuous vector internally, closed gate-passing pack at the boundary. Everything downstream — catalogue growth, accessibility guarantees, overfitting control — depends on this split being right."
}
The preference learner must itself be A/B tested against a static defaultrecord 102
{
  "id": "I95-TOP10",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "The preference learner must itself be A/B tested against a static default",
  "evidence_class": "observed",
  "source": "Spotify Engineering: a bandit is 'a feature you've built, not an experimental method'",
  "observed": "2026-08-27",
  "claim": "Without a holdout comparing elicited packs to an industry-default pack, we cannot claim the elicitation works at all.",
  "limitations": "Needs enough clients for a powered comparison — likely the binding constraint.",
  "disposition": "ADOPT — the single most important methodological gate",
  "refers_to": "I95",
  "rank": 2,
  "rationale": "Without a holdout against a static industry-default pack we cannot claim the learner works at all. Spotify's framing is decisive: a personalisation system is a feature, not an experimental method. This is the gate that turns the lane from plausible into measured."
}
Warm-start prior from a population aesthetic modelrecord 103
{
  "id": "I40-TOP10",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Warm-start prior from a population aesthetic model",
  "evidence_class": "observed",
  "source": "LAION aesthetic predictor; Brochu et al. 2010 cut iterations 11.25 -> 6.5 with a learned prior",
  "observed": "2026-08-27",
  "claim": "Learn the client's DEVIATION from a population baseline rather than their taste from scratch. The largest single efficiency lever reported anywhere in the applied literature.",
  "limitations": "Population aesthetic is not per-client taste and may encode a generic bias.",
  "disposition": "ADOPT",
  "refers_to": "I40",
  "rank": 3,
  "rationale": "A warm-start population prior is the largest single efficiency lever reported anywhere in the applied literature: Brochu et al. 2010 cut iterations from 11.25 to 6.5. It is also cheap, since the LAION aesthetic predictor already exists."
}
Perturbation-tolerance testrecord 104
{
  "id": "I60-TOP10",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Perturbation-tolerance test",
  "evidence_class": "hypothesis",
  "source": "derived",
  "observed": "2026-08-27",
  "claim": "Show the converged pack and a +/-1-level perturbation. If clients reliably distinguish and reject, the tolerance argument in §2 collapses and round counts triple.",
  "limitations": "—",
  "disposition": "ADOPT — falsifier F1, run FIRST",
  "refers_to": "I60",
  "rank": 4,
  "rationale": "The perturbation-tolerance test is falsifier F1 and should run FIRST, because the entire round-count argument rests on +/-1-level tolerance being acceptable. If it is not, the requirement roughly triples and the product changes shape."
}
Brand-constraint hard cutoffs before elicitationrecord 105
{
  "id": "I08-TOP10",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "Brand-constraint hard cutoffs before elicitation",
  "evidence_class": "observed",
  "source": "Sawtooth ACBC must-have/unacceptable mechanism (vendor manual)",
  "observed": "2026-08-27",
  "claim": "Existing brand colours/fonts become hard prunes of the stimulus space, not soft preferences; all later stimuli satisfy them.",
  "limitations": "Over-pruning can empty the space; needs a fallback.",
  "disposition": "ADOPT",
  "refers_to": "I08",
  "rank": 5,
  "rationale": "Sawtooth's must-have/unacceptable cutoff mechanism is the missing piece for brand constraints, and it is proven commercial practice. A hard prune is categorically better than learning a client's non-negotiables as a soft preference."
}
Simulate-to-size: Monte Carlo power analysis for round countrecord 106
{
  "id": "I42-TOP10",
  "kind": "innovation",
  "group": "adaptive-policy",
  "name": "Simulate-to-size: Monte Carlo power analysis for round count",
  "evidence_class": "observed",
  "source": "skpr (R)",
  "observed": "2026-08-27",
  "claim": "Answer 'how many rounds' EMPIRICALLY by simulating synthetic clients under a BT likelihood rather than arguing from bounds.",
  "limitations": "Power analysis is built for GLM responses; needs adapting.",
  "disposition": "ADOPT — this is how the N question should actually be settled",
  "refers_to": "I42",
  "rank": 6,
  "rationale": "Simulate-to-size settles the 'how many choices' question empirically rather than by argument. Given that no paper gives a sufficiency guarantee for our setting, a Monte Carlo power analysis over synthetic clients is the only route to a defensible number."
}
Two-week test-retest stability as the primary validation metricrecord 107
{
  "id": "I57-TOP10",
  "kind": "innovation",
  "group": "stopping-and-confidence",
  "name": "Two-week test-retest stability as the primary validation metric",
  "evidence_class": "inferred",
  "source": "Slovic 1995 construction of preference",
  "observed": "2026-08-27",
  "claim": "A model that fits the clicks but changes answer next Tuesday is worthless. Measures whether the thing we claim to measure exists.",
  "limitations": "Requires client time two weeks apart; hard to get.",
  "disposition": "ADOPT — this is falsifier F7",
  "refers_to": "I57",
  "rank": 7,
  "rationale": "Two-week test-retest is falsifier F7 — the one that would kill the part rather than adjust it. If preference is not stable, no amount of elicitation precision matters. Test early and cheaply."
}
P05 corpus components as re-themable controlled stimulirecord 108
{
  "id": "I03-TOP10",
  "kind": "innovation",
  "group": "stimulus-design",
  "name": "P05 corpus components as re-themable controlled stimuli",
  "evidence_class": "observed",
  "source": "research/21st-corpus-audit-2026-08-27.md — 86.7% of colour-bearing CSS rules already resolve through var(--token); 87.3% carry a shadcn oklch :root block",
  "observed": "2026-08-27",
  "claim": "Use the 8,515-identity corpus as the stimulus pool: identical structure, varying token pack, re-theming is a find-and-replace on ~30 custom properties.",
  "limitations": "271 components lack previews; component SOURCE is absent from the corpus.",
  "disposition": "ADOPT",
  "refers_to": "I03",
  "rank": 8,
  "rationale": "The P05 corpus turns stimulus generation from a cost into an asset: 8,515 identities, 86.7% of colour rules already indirected through tokens, re-theming as a ~30-property find-and-replace. Identical structure with varying treatment is the ideal experimental article."
}
4-up gallery with explicit outside optionrecord 109
{
  "id": "I19-TOP10",
  "kind": "innovation",
  "group": "question-format",
  "name": "4-up gallery with explicit outside option",
  "evidence_class": "inferred",
  "source": "derived §3 + Sawtooth 4-5 per screen + live demo",
  "observed": "2026-08-27",
  "claim": "2.32 noiseless bits/round; ~0.91 effective at p=0.15; balances information against deliberation cost and avoids the hostility of 'worst'.",
  "limitations": "Best-worst is strictly more informative per question.",
  "disposition": "ADOPT — primary recommendation",
  "refers_to": "I19",
  "rank": 9,
  "rationale": "The primary format recommendation — 4-up with an explicit outside option — supported independently by the information derivation, Sawtooth's 4-5-per-screen guidance, and our own working prototype."
}
End-user vs operator preference divergencerecord 110
{
  "id": "I89-TOP10",
  "kind": "innovation",
  "group": "drift-industry-risk",
  "name": "End-user vs operator preference divergence",
  "evidence_class": "inferred",
  "source": "b2b-template-shelf-report.md: property_management tenant portal is 'a distinct untrusted identity'; it_services_msps 'per-client tenancy is the product'",
  "observed": "2026-08-27",
  "claim": "For the 6/17 industries where `portal` is secondary, the person choosing (the operator) is not the person using (the end customer). Elicit from the operator but weight toward end-user norms.",
  "limitations": "No mechanism yet for representing the absent end user.",
  "disposition": "ADOPT — open design problem",
  "refers_to": "I89",
  "rank": 10,
  "rationale": "The operator-vs-end-user divergence is a genuine unsolved design problem that affects 6 of 17 industries, and nobody in the commercial denominator has solved it. Naming it now prevents shipping a system that optimises for the wrong person."
}