Whole-screen stimuli with only targeted knobs variedrecord 1
{
"id": "I01",
"kind": "innovation",
"group": "stimulus-design",
"name": "Whole-screen stimuli with only targeted knobs varied",
"evidence_class": "inferred",
"source": "D-optimal / Bayesian adaptive DCE design (idefix); live demo already does this",
"observed": "2026-08-27",
"claim": "Render complete realistic screens but vary only the knobs the round targets, holding others at the posterior mean.",
"limitations": "Assumes posterior mean is a sensible hold value early on.",
"disposition": "ADOPT"
}
Content held constant across a comparisonrecord 2
{
"id": "I02",
"kind": "innovation",
"group": "stimulus-design",
"name": "Content held constant across a comparison",
"evidence_class": "inferred",
"source": "basic experimental control",
"observed": "2026-08-27",
"claim": "Use the same headline/copy/imagery in all cards of a round so content preference cannot confound the token-pack signal.",
"limitations": "Some knobs (density) interact with content length.",
"disposition": "ADOPT"
}
P05 corpus components as re-themable controlled stimulirecord 3
{
"id": "I03",
"kind": "innovation",
"group": "stimulus-design",
"name": "P05 corpus components as re-themable controlled stimuli",
"evidence_class": "observed",
"source": "research/21st-corpus-audit-2026-08-27.md — 86.7% of colour-bearing CSS rules already resolve through var(--token); 87.3% carry a shadcn oklch :root block",
"observed": "2026-08-27",
"claim": "Use the 8,515-identity corpus as the stimulus pool: identical structure, varying token pack, re-theming is a find-and-replace on ~30 custom properties.",
"limitations": "271 components lack previews; component SOURCE is absent from the corpus.",
"disposition": "ADOPT"
}
Stimulus realism gradient: abstract swatches -> component -> full screenrecord 4
{
"id": "I04",
"kind": "innovation",
"group": "stimulus-design",
"name": "Stimulus realism gradient: abstract swatches -> component -> full screen",
"evidence_class": "hypothesis",
"source": "none",
"observed": "2026-08-27",
"claim": "Early rounds use cheap abstract stimuli, later rounds full screens.",
"limitations": "Risks measuring preference for the abstraction, not the design.",
"disposition": "test"
}
Never re-show an identical stimulusrecord 5
{
"id": "I05",
"kind": "innovation",
"group": "stimulus-design",
"name": "Never re-show an identical stimulus",
"evidence_class": "inferred",
"source": "Zajonc 1968 mere exposure",
"observed": "2026-08-27",
"claim": "Prevents familiarity inflating preference for repeated designs.",
"limitations": "Constrains the generator's search near convergence.",
"disposition": "ADOPT"
}
Log stimulus-repetition rate as a measured confoundrecord 6
{
"id": "I06",
"kind": "innovation",
"group": "stimulus-design",
"name": "Log stimulus-repetition rate as a measured confound",
"evidence_class": "inferred",
"source": "Zajonc 1968",
"observed": "2026-08-27",
"claim": "Even if repeats occur, make the exposure effect measurable rather than invisible.",
"limitations": "—",
"disposition": "ADOPT"
}
Industry-conditioned stimulus priorsrecord 7
{
"id": "I07",
"kind": "innovation",
"group": "stimulus-design",
"name": "Industry-conditioned stimulus priors",
"evidence_class": "inferred",
"source": "b2b-template-shelf-report.md 17-industry variant deltas",
"observed": "2026-08-27",
"claim": "Seed the stimulus space from an industry prior (a law firm starts in a more conservative region than a course creator).",
"limitations": "Risks stereotyping; must remain a prior, not a constraint.",
"disposition": "ADOPT"
}
Brand-constraint hard cutoffs before elicitationrecord 8
{
"id": "I08",
"kind": "innovation",
"group": "stimulus-design",
"name": "Brand-constraint hard cutoffs before elicitation",
"evidence_class": "observed",
"source": "Sawtooth ACBC must-have/unacceptable mechanism (vendor manual)",
"observed": "2026-08-27",
"claim": "Existing brand colours/fonts become hard prunes of the stimulus space, not soft preferences; all later stimuli satisfy them.",
"limitations": "Over-pruning can empty the space; needs a fallback.",
"disposition": "ADOPT"
}
Dark-mode variant shown as part of the stimulusrecord 9
{
"id": "I09",
"kind": "innovation",
"group": "stimulus-design",
"name": "Dark-mode variant shown as part of the stimulus",
"evidence_class": "inferred",
"source": "token-pack-science §3.9 — dark is not inversion",
"observed": "2026-08-27",
"claim": "Show each candidate in both modes since packs must ship both.",
"limitations": "Doubles visual load per card.",
"disposition": "test"
}
Density stress-test stimuli using realistic long copyrecord 10
{
"id": "I10",
"kind": "innovation",
"group": "stimulus-design",
"name": "Density stress-test stimuli using realistic long copy",
"evidence_class": "inferred",
"source": "ui-pick-to-spec §6 failure mode 6",
"observed": "2026-08-27",
"claim": "Include a card rendered with realistically long client copy so density preference is measured under real conditions.",
"limitations": "—",
"disposition": "test"
}
Fragment rounds reserved for fine disambiguation onlyrecord 11
{
"id": "I11",
"kind": "innovation",
"group": "stimulus-design",
"name": "Fragment rounds reserved for fine disambiguation only",
"evidence_class": "hypothesis",
"source": "none",
"observed": "2026-08-27",
"claim": "Use component fragments only when two knobs remain confounded at whole-screen scale.",
"limitations": "Unproven that fragments resolve better.",
"disposition": "test — this is falsifier F4"
}
Adaptive grid size by viewportrecord 12
{
"id": "I12",
"kind": "innovation",
"group": "stimulus-design",
"name": "Adaptive grid size by viewport",
"evidence_class": "observed",
"source": "Sequential Gallery reduced 5x5 to 3x3 on a 13-inch display",
"observed": "2026-08-27",
"claim": "Gallery size should adapt to display, not be a fixed product constant.",
"limitations": "—",
"disposition": "ADOPT"
}
Stimuli drawn from the closed pack catalogue onlyrecord 13
{
"id": "I13",
"kind": "innovation",
"group": "stimulus-design",
"name": "Stimuli drawn from the closed pack catalogue only",
"evidence_class": "inferred",
"source": "token-pack-science gates A-J",
"observed": "2026-08-27",
"claim": "Never show a card that is not a real, gate-passing pack — so the preview IS the deliverable.",
"limitations": "Restricts early exploration to catalogue coverage.",
"disposition": "ADOPT"
}
Preview must be re-rendered from spec, never the original thumbnailrecord 14
{
"id": "I14",
"kind": "innovation",
"group": "stimulus-design",
"name": "Preview must be re-rendered from spec, never the original thumbnail",
"evidence_class": "observed",
"source": "ui-pick-to-spec §6 operational consequence",
"observed": "2026-08-27",
"claim": "Closes the expectation gap by construction and makes previews a free end-to-end pipeline test.",
"limitations": "Requires the render pipeline before elicitation ships.",
"disposition": "ADOPT"
}
Anti-prototypicality proberecord 15
{
"id": "I15",
"kind": "innovation",
"group": "stimulus-design",
"name": "Anti-prototypicality probe",
"evidence_class": "inferred",
"source": "Tuch et al. 2012; Reber et al. 2004",
"observed": "2026-08-27",
"claim": "Deliberately include one low-prototypicality card per round to detect clients who want distinctive over safe.",
"limitations": "May be systematically rejected, wasting a slot.",
"disposition": "test"
}
Seed from client's existing website via extractionrecord 16
{
"id": "I16",
"kind": "innovation",
"group": "stimulus-design",
"name": "Seed from client's existing website via extraction",
"evidence_class": "observed",
"source": "Canva AI extracts brand context from a public URL",
"observed": "2026-08-27",
"claim": "Warm-start the prior from the client's current site rather than a uniform prior.",
"limitations": "Their current site may be exactly what they want to escape.",
"disposition": "test"
}
Competitor-anchored stimulirecord 17
{
"id": "I17",
"kind": "innovation",
"group": "stimulus-design",
"name": "Competitor-anchored stimuli",
"evidence_class": "hypothesis",
"source": "none",
"observed": "2026-08-27",
"claim": "Show packs near and far from named competitors to elicit differentiation preference.",
"limitations": "Conflates taste with positioning.",
"disposition": "test"
}
Stimulus set diversity floorrecord 18
{
"id": "I18",
"kind": "innovation",
"group": "stimulus-design",
"name": "Stimulus set diversity floor",
"evidence_class": "inferred",
"source": "Negahban et al. spectral gap dependence",
"observed": "2026-08-27",
"claim": "Enforce a minimum pairwise distance among the 4 cards so the comparison graph stays well-connected.",
"limitations": "Competes with least-resolved-knob targeting, which wants NEAR pairs.",
"disposition": "ADOPT with tension noted"
}
4-up gallery with explicit outside optionrecord 19
{
"id": "I19",
"kind": "innovation",
"group": "question-format",
"name": "4-up gallery with explicit outside option",
"evidence_class": "inferred",
"source": "derived §3 + Sawtooth 4-5 per screen + live demo",
"observed": "2026-08-27",
"claim": "2.32 noiseless bits/round; ~0.91 effective at p=0.15; balances information against deliberation cost and avoids the hostility of 'worst'.",
"limitations": "Best-worst is strictly more informative per question.",
"disposition": "ADOPT — primary recommendation"
}
Best-worst of 4 (MaxDiff Case 1)record 20
{
"id": "I20",
"kind": "innovation",
"group": "question-format",
"name": "Best-worst of 4 (MaxDiff Case 1)",
"evidence_class": "observed",
"source": "Marley & Louviere 2005; Sawtooth manual",
"observed": "2026-08-27",
"claim": "3.58 noiseless bits/round, ~2x the 4-up rate; would cut rounds to ~3.",
"limitations": "Roughly doubles per-screen deliberation; 'pick the worst' is hostile in a client sales context; no verified numeric information-gain multiplier exists.",
"disposition": "test as an A/B arm — falsifier F6"
}
Attribute-level best-worst (MaxDiff Case 2)record 21
{
"id": "I21",
"kind": "innovation",
"group": "question-format",
"name": "Attribute-level best-worst (MaxDiff Case 2)",
"evidence_class": "inferred",
"source": "Marley, Flynn & Louviere 2008; support.BWS2",
"observed": "2026-08-27",
"claim": "'Which knob is most and least right on THIS design' — maps almost exactly onto per-knob credit assignment, solving the attribution problem a whole-card pick has.",
"limitations": "Requires the client to reason about knobs explicitly, which is exactly what we said they cannot do.",
"disposition": "test — high upside, high risk"
}
Binary pairwiserecord 22
{
"id": "I22",
"kind": "innovation",
"group": "question-format",
"name": "Binary pairwise",
"evidence_class": "inferred",
"source": "derived §3",
"observed": "2026-08-27",
"claim": "Simplest, but ~11 rounds at p=0.15 and most sensitive to near-ties.",
"limitations": "Nearly double the rounds of 4-up.",
"disposition": "reject as primary"
}
Three-way pairwise with a neutral middlerecord 23
{
"id": "I23",
"kind": "innovation",
"group": "question-format",
"name": "Three-way pairwise with a neutral middle",
"evidence_class": "observed",
"source": "Midjourney legacy Style Tuner: 'leave the middle box selected to skip the pair'",
"observed": "2026-08-27",
"claim": "A shipped outside-option affordance inside a pairwise format.",
"limitations": "Fewer bits than 4-up.",
"disposition": "reference — the affordance, not the format"
}
Full ranking of 4record 24
{
"id": "I24",
"kind": "innovation",
"group": "question-format",
"name": "Full ranking of 4",
"evidence_class": "inferred",
"source": "Plackett 1975; Hajek, Oh & Xu 2014",
"observed": "2026-08-27",
"claim": "4.58 noiseless bits, the densest format tested.",
"limitations": "Ranking four whole screens is a heavy cognitive task; inherits IIA.",
"disposition": "reject"
}
9-up galleryrecord 25
{
"id": "I25",
"kind": "innovation",
"group": "question-format",
"name": "9-up gallery",
"evidence_class": "inferred",
"source": "derived §3; Sequential Gallery used 5x5 and 3x3",
"observed": "2026-08-27",
"claim": "3.17 noiseless bits.",
"limitations": "Nine whole-screen renders exceed comfortable simultaneous comparison; selection noise rises in a way the bit count does not capture.",
"disposition": "reject for whole screens"
}
Mixed format by phase: 4-up early, pairwise laterecord 26
{
"id": "I26",
"kind": "innovation",
"group": "question-format",
"name": "Mixed format by phase: 4-up early, pairwise late",
"evidence_class": "hypothesis",
"source": "none",
"observed": "2026-08-27",
"claim": "Broad exploration wants breadth; fine disambiguation is naturally pairwise.",
"limitations": "Format switching may confuse users and complicates the likelihood.",
"disposition": "test"
}
Slider direct manipulation as a fallbackrecord 27
{
"id": "I27",
"kind": "innovation",
"group": "question-format",
"name": "Slider direct manipulation as a fallback",
"evidence_class": "observed",
"source": "Figma First Draft radius/spacing sliders",
"observed": "2026-08-27",
"claim": "Offer sliders to clients who DO know what they want, skipping elicitation entirely.",
"limitations": "Most clients cannot name a radius; the slider is the competitor's approach.",
"disposition": "ADOPT as escape hatch"
}
Sequential line search (1D slider through knob space)record 28
{
"id": "I28",
"kind": "innovation",
"group": "question-format",
"name": "Sequential line search (1D slider through knob space)",
"evidence_class": "observed",
"source": "Koyama et al. SIGGRAPH 2017 (MIT-licensed implementation exists)",
"observed": "2026-08-27",
"claim": "User picks a point on a 1D slice; reported 15-iteration budget, good by iteration 4-5 at 6D/7D.",
"limitations": "A slider over a design continuum is harder to render than 4 discrete packs, and our output space is discrete anyway.",
"disposition": "study"
}
Outside option as weak negative evidence on all shown cardsrecord 29
{
"id": "I29",
"kind": "innovation",
"group": "question-format",
"name": "Outside option as weak negative evidence on all shown cards",
"evidence_class": "hypothesis",
"source": "derived; diverges from Midjourney which makes skipping non-informative",
"observed": "2026-08-27",
"claim": "Preserves information from declines while keeping the re-roll honest.",
"limitations": "Weight is a free parameter needing calibration.",
"disposition": "ADOPT — calibrate the weight"
}
Cap consecutive outside-option selections at 3record 30
{
"id": "I30",
"kind": "innovation",
"group": "question-format",
"name": "Cap consecutive outside-option selections at 3",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Three declines in a row means the generator is in the wrong region; trigger a re-seed rather than another round.",
"limitations": "Threshold is arbitrary pending data.",
"disposition": "ADOPT"
}
'This is right' always-available terminal affordancerecord 31
{
"id": "I31",
"kind": "innovation",
"group": "question-format",
"name": "'This is right' always-available terminal affordance",
"evidence_class": "observed",
"source": "Sequential Gallery satisfaction button; mean 5.36 rounds to satisfaction",
"observed": "2026-08-27",
"claim": "User's own judgement is the criterion the whole exercise proxies for.",
"limitations": "Users may satisfice early.",
"disposition": "ADOPT"
}
Preview-before-commit on the selected packrecord 32
{
"id": "I32",
"kind": "innovation",
"group": "question-format",
"name": "Preview-before-commit on the selected pack",
"evidence_class": "observed",
"source": "Webflow 'Preview in Designer'",
"observed": "2026-08-27",
"claim": "Let the client see the pack applied to their real content before committing.",
"limitations": "Requires the render pipeline.",
"disposition": "ADOPT"
}
Paired-comparison tie option distinct from 'none'record 33
{
"id": "I33",
"kind": "innovation",
"group": "question-format",
"name": "Paired-comparison tie option distinct from 'none'",
"evidence_class": "inferred",
"source": "prefmod handles undecided/ties",
"observed": "2026-08-27",
"claim": "'These two are equally good' is different information from 'both are wrong'.",
"limitations": "Adds a third response category to model.",
"disposition": "test"
}
Confidence-weighted responses (client marks how sure they are)record 34
{
"id": "I34",
"kind": "innovation",
"group": "question-format",
"name": "Confidence-weighted responses (client marks how sure they are)",
"evidence_class": "hypothesis",
"source": "none",
"observed": "2026-08-27",
"claim": "Let the client flag a low-confidence pick so it is down-weighted.",
"limitations": "Self-reported confidence is poorly calibrated; adds friction.",
"disposition": "reject"
}
BALD acquisition on a GP preference modelrecord 35
{
"id": "I35",
"kind": "innovation",
"group": "adaptive-policy",
"name": "BALD acquisition on a GP preference model",
"evidence_class": "observed",
"source": "Houlsby et al. 2011 (arXiv 1112.5745)",
"observed": "2026-08-27",
"claim": "Information gain expressed via predictive entropies, explicitly extended to GP preference learning; tractable.",
"limitations": "Myopic one-step criterion.",
"disposition": "ADOPT"
}
qEUBO acquisitionrecord 36
{
"id": "I36",
"kind": "innovation",
"group": "adaptive-policy",
"name": "qEUBO acquisition",
"evidence_class": "observed",
"source": "Astudillo et al. AISTATS 2023; shipped in BoTorch",
"observed": "2026-08-27",
"claim": "One-step Bayes optimal under noise-free responses; simple regret converges at o(1/n); qEI can FAIL to converge for PBO.",
"limitations": "Asymptotic guarantee gives no finite-sample count.",
"disposition": "ADOPT"
}
Target the least-resolved knob each roundrecord 37
{
"id": "I37",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Target the least-resolved knob each round",
"evidence_class": "observed",
"source": "live demo already does this",
"observed": "2026-08-27",
"claim": "Constructs cards differing on the widest-posterior axis while holding others steady.",
"limitations": "Greedy per-knob targeting can miss interaction effects.",
"disposition": "ADOPT"
}
Explicit exploration/exploitation schedule ('candy vs medicine')record 38
{
"id": "I38",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Explicit exploration/exploitation schedule ('candy vs medicine')",
"evidence_class": "observed",
"source": "Stitch Fix Style Shuffle practice (press-reported)",
"observed": "2026-08-27",
"claim": "Deliberately mix near-certain-hit cards (engagement) with high-information cards (learning).",
"limitations": "Ratio unpublished; costs information per round.",
"disposition": "ADOPT — calibrate ratio"
}
Hierarchical Bayes shrinkage toward a population priorrecord 39
{
"id": "I39",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Hierarchical Bayes shrinkage toward a population prior",
"evidence_class": "observed",
"source": "ChoiceModelR; standard conjoint practice",
"observed": "2026-08-27",
"claim": "Designed exactly for the few-observations-per-individual regime; the primary anti-overfitting mechanism alongside closed packs.",
"limitations": "Needs a population of prior clients to shrink toward.",
"disposition": "ADOPT — but cold-start it from I40"
}
Warm-start prior from a population aesthetic modelrecord 40
{
"id": "I40",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Warm-start prior from a population aesthetic model",
"evidence_class": "observed",
"source": "LAION aesthetic predictor; Brochu et al. 2010 cut iterations 11.25 -> 6.5 with a learned prior",
"observed": "2026-08-27",
"claim": "Learn the client's DEVIATION from a population baseline rather than their taste from scratch. The largest single efficiency lever reported anywhere in the applied literature.",
"limitations": "Population aesthetic is not per-client taste and may encode a generic bias.",
"disposition": "ADOPT"
}
D-optimal / Bayesian-efficient stimulus set selectionrecord 41
{
"id": "I41",
"kind": "innovation",
"group": "adaptive-policy",
"name": "D-optimal / Bayesian-efficient stimulus set selection",
"evidence_class": "observed",
"source": "idefix (R, GPL-3)",
"observed": "2026-08-27",
"claim": "Choose the 4 cards to maximise design efficiency under the current posterior.",
"limitations": "GPL-3 licence; requires priors.",
"disposition": "ADOPT-METHOD, reimplement"
}
Simulate-to-size: Monte Carlo power analysis for round countrecord 42
{
"id": "I42",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Simulate-to-size: Monte Carlo power analysis for round count",
"evidence_class": "observed",
"source": "skpr (R)",
"observed": "2026-08-27",
"claim": "Answer 'how many rounds' EMPIRICALLY by simulating synthetic clients under a BT likelihood rather than arguing from bounds.",
"limitations": "Power analysis is built for GLM responses; needs adapting.",
"disposition": "ADOPT — this is how the N question should actually be settled"
}
Orthogonal-array fixed design as the passive baselinerecord 43
{
"id": "I43",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Orthogonal-array fixed design as the passive baseline",
"evidence_class": "observed",
"source": "DoE.base",
"observed": "2026-08-27",
"claim": "The non-adaptive control our adaptive policy must beat.",
"limitations": "—",
"disposition": "ADOPT as control arm"
}
Projective preferential BO for high-dim knob spacesrecord 44
{
"id": "I44",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Projective preferential BO for high-dim knob spaces",
"evidence_class": "observed",
"source": "AaltoPML/PPBO (MIT)",
"observed": "2026-08-27",
"claim": "Queries along projections, designed for human-in-the-loop high-dimensional elicitation.",
"limitations": "Small research repo (20 stars).",
"disposition": "study"
}
Transitivity-based pair elimination (PAPRIKA-style)record 45
{
"id": "I45",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Transitivity-based pair elimination (PAPRIKA-style)",
"evidence_class": "hypothesis",
"source": "1000minds method; COULD NOT VERIFY — 404 on both URLs",
"observed": "2026-08-27",
"claim": "Eliminate implied comparisons by transitivity, drastically cutting explicit questions.",
"limitations": "SOURCE UNVERIFIED. Also assumes transitivity, which human taste violates (Ailon 2010).",
"disposition": "VERIFY FIRST, then test"
}
Non-transitivity-tolerant active rankingrecord 46
{
"id": "I46",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Non-transitivity-tolerant active ranking",
"evidence_class": "observed",
"source": "Ailon 2010 (arXiv 1011.0108)",
"observed": "2026-08-27",
"claim": "Explicitly handles 'non-transitivity paradoxes which may arise naturally due to human mistakes or irrationality'.",
"limitations": "No closed-form bound extracted from the abstract.",
"disposition": "study"
}
Copeland counting as a simplicity checkrecord 47
{
"id": "I47",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Copeland counting as a simplicity check",
"evidence_class": "observed",
"source": "Shah & Wainwright 2015",
"observed": "2026-08-27",
"claim": "Rank by comparisons won: optimal up to CONSTANT factors, no conditions on the probability matrix. If this matches the GP model's output, the GP is not earning its complexity.",
"limitations": "Top-k recovery framing, not utility-vector recovery.",
"disposition": "ADOPT as a baseline check"
}
Dueling-bandit formulationrecord 48
{
"id": "I48",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Dueling-bandit formulation",
"evidence_class": "observed",
"source": "Yue & Joachims; Zoghi RUCB; Sui et al. survey",
"observed": "2026-08-27",
"claim": "Relative-feedback online optimisation.",
"limitations": "Regret-minimisation over continuous operation, not fixed-budget identification — a different objective from ours.",
"disposition": "reject as primary framing"
}
Per-knob independent BT models vs joint GPrecord 49
{
"id": "I49",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Per-knob independent BT models vs joint GP",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Independent per-knob models are simpler and interpretable; a joint GP captures correlation.",
"limitations": "Independence is empirically false (§4 of first-principles).",
"disposition": "test — joint expected to win"
}
Bootstrap confidence intervals over pairwise judgmentsrecord 50
{
"id": "I50",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Bootstrap confidence intervals over pairwise judgments",
"evidence_class": "observed",
"source": "lmarena/arena-hard-auto",
"observed": "2026-08-27",
"claim": "Battle-tested approach to reporting uncertainty on BT fits from human votes.",
"limitations": "—",
"disposition": "ADOPT"
}
TrueSkill-style Gaussian belief per packrecord 51
{
"id": "I51",
"kind": "innovation",
"group": "adaptive-policy",
"name": "TrueSkill-style Gaussian belief per pack",
"evidence_class": "observed",
"source": "Herbrich et al. 2007; use openskill.py (MIT), NOT sublee/trueskill",
"observed": "2026-08-27",
"claim": "Calibrated per-item uncertainty that Elo does not give.",
"limitations": "LICENCE TRAP: sublee/trueskill's LICENSE body bars commercial use despite a BSD badge.",
"disposition": "study — openskill.py only"
}
Stop when expected information gain of the best next question < thresholdrecord 52
{
"id": "I52",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Stop when expected information gain of the best next question < threshold",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "A wide posterior no available question can narrow is a reason to stop, not to continue.",
"limitations": "Threshold (~0.25 bits) needs calibration.",
"disposition": "ADOPT"
}
Three-condition stop: confidence OR ceiling OR user-satisfiedrecord 53
{
"id": "I53",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Three-condition stop: confidence OR ceiling OR user-satisfied",
"evidence_class": "inferred",
"source": "derived §6; Brochu 2010 20-iteration abandonment; Midjourney 'past round 15... small and subtle'",
"observed": "2026-08-27",
"claim": "Primary rule is confidence; 15 rounds is a ceiling; user satisfaction always terminates.",
"limitations": "Ceiling is grounded in two sources, not a derivation.",
"disposition": "ADOPT"
}
Hard floor of 5 roundsrecord 54
{
"id": "I54",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Hard floor of 5 rounds",
"evidence_class": "inferred",
"source": "derived; Sequential Gallery mean 5.36",
"observed": "2026-08-27",
"claim": "Early apparent convergence is usually the prior, not the data.",
"limitations": "—",
"disposition": "ADOPT"
}
Per-knob confidence reporting, never a single scalarrecord 55
{
"id": "I55",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Per-knob confidence reporting, never a single scalar",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Some knobs resolve in two rounds; others may never resolve, and 'this client does not care about shadow' is a finding not a failure. A single '87% confident' number would mislead.",
"limitations": "More complex to present to a non-technical client.",
"disposition": "ADOPT"
}
'No preference' as a first-class per-knob outcomerecord 56
{
"id": "I56",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "'No preference' as a first-class per-knob outcome",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Explicitly represent indifference rather than forcing a point estimate.",
"limitations": "Downstream pack selection must handle a free knob.",
"disposition": "ADOPT"
}
Two-week test-retest stability as the primary validation metricrecord 57
{
"id": "I57",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Two-week test-retest stability as the primary validation metric",
"evidence_class": "inferred",
"source": "Slovic 1995 construction of preference",
"observed": "2026-08-27",
"claim": "A model that fits the clicks but changes answer next Tuesday is worthless. Measures whether the thing we claim to measure exists.",
"limitations": "Requires client time two weeks apart; hard to get.",
"disposition": "ADOPT — this is falsifier F7"
}
Held-out pick prediction as the fit metricrecord 58
{
"id": "I58",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Held-out pick prediction as the fit metric",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Reserve 2 rounds, predict them from the model fitted on the rest.",
"limitations": "Small n makes the estimate noisy.",
"disposition": "ADOPT"
}
Neighbour-discrimination testrecord 59
{
"id": "I59",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Neighbour-discrimination test",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Show the client their pack plus the 2nd and 3rd nearest, unlabelled. If they cannot pick their own above chance, the catalogue is denser than perception.",
"limitations": "—",
"disposition": "ADOPT — falsifier F5"
}
Perturbation-tolerance testrecord 60
{
"id": "I60",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Perturbation-tolerance test",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Show the converged pack and a +/-1-level perturbation. If clients reliably distinguish and reject, the tolerance argument in §2 collapses and round counts triple.",
"limitations": "—",
"disposition": "ADOPT — falsifier F1, run FIRST"
}
Separation-aware stoppingrecord 61
{
"id": "I61",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Separation-aware stopping",
"evidence_class": "observed",
"source": "Chen & Suh 2015 — complexity scales inversely with separation",
"observed": "2026-08-27",
"claim": "Detect when remaining candidates are genuinely near-tied and stop rather than burning rounds distinguishing the indistinguishable.",
"limitations": "—",
"disposition": "ADOPT"
}
Comparison-graph connectivity checkrecord 62
{
"id": "I62",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Comparison-graph connectivity check",
"evidence_class": "observed",
"source": "Hunter 2004 MM convergence condition",
"observed": "2026-08-27",
"claim": "BT/MM fitting requires a strongly connected comparison graph; verify before fitting.",
"limitations": "Constrains which stimulus sets are legal.",
"disposition": "ADOPT as a gate"
}
Abandonment-rate monitoring as a UX stop signalrecord 63
{
"id": "I63",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Abandonment-rate monitoring as a UX stop signal",
"evidence_class": "observed",
"source": "Brochu et al. 2010: 20 iterations is 'roughly the point at which users start to quit'",
"observed": "2026-08-27",
"claim": "Instrument drop-off per round; if it rises before the ceiling, lower the ceiling.",
"limitations": "—",
"disposition": "ADOPT"
}
Time-per-round budgetrecord 64
{
"id": "I64",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Time-per-round budget",
"evidence_class": "observed",
"source": "Sequential Gallery: 14.8s per plane-search subtask",
"observed": "2026-08-27",
"claim": "Budget total elicitation time, not just round count. 10 rounds x ~15s is ~2.5 minutes.",
"limitations": "Whole-screen comparison likely slower than the cited subtask.",
"disposition": "ADOPT"
}
Confidence decay over calendar timerecord 65
{
"id": "I65",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Confidence decay over calendar time",
"evidence_class": "inferred",
"source": "Netflix supersession rule; Hinge 24h window",
"observed": "2026-08-27",
"claim": "Treat the profile as decaying, prompting re-elicitation rather than assuming permanence.",
"limitations": "Decay rate unknown; needs I57 data first.",
"disposition": "test"
}
Re-elicitation triggered by client rejection of built outputrecord 66
{
"id": "I66",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Re-elicitation triggered by client rejection of built output",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "If a client rejects the delivered design, that is strong evidence to re-run rather than patch.",
"limitations": "Expensive and reads as failure.",
"disposition": "test"
}
Continuous preference vector internally, closed pack at the boundaryrecord 67
{
"id": "I67",
"kind": "innovation",
"group": "output-compilation",
"name": "Continuous preference vector internally, closed pack at the boundary",
"evidence_class": "inferred",
"source": "derived §8; token-pack-science gates A-J; lane synthesis 'learn a vector not a pack ID'",
"observed": "2026-08-27",
"claim": "Catalogue can grow without re-eliciting; shipped artefact is always gate-passing.",
"limitations": "Nearest-neighbour metric weighting is an open parameter.",
"disposition": "ADOPT — primary recommendation"
}
Reject continuous token interpolationrecord 68
{
"id": "I68",
"kind": "innovation",
"group": "output-compilation",
"name": "Reject continuous token interpolation",
"evidence_class": "observed",
"source": "token-pack-science §5 Gate C — WCAG luminance is non-linear in channel values",
"observed": "2026-08-27",
"claim": "The midpoint of two AA-passing palettes can fail AA; an interpolated pack has passed none of gates A-J.",
"limitations": "Loses expressive range between packs.",
"disposition": "ADOPT the rejection"
}
Closed-pack snapping AS the anti-overfitting regularizerrecord 69
{
"id": "I69",
"kind": "innovation",
"group": "output-compilation",
"name": "Closed-pack snapping AS the anti-overfitting regularizer",
"evidence_class": "hypothesis",
"source": "derived §8",
"observed": "2026-08-27",
"claim": "Output space of ~20-30 packs rather than 36,000 means the model cannot overfit into a bespoke corner on ten noisy clicks.",
"limitations": "If the catalogue grows large this protection weakens.",
"disposition": "ADOPT"
}
Explain the pack choice in knob termsrecord 70
{
"id": "I70",
"kind": "innovation",
"group": "output-compilation",
"name": "Explain the pack choice in knob terms",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "'We chose this because you are here in knob space' is worth real money in the client conversation.",
"limitations": "Requires knob names clients understand.",
"disposition": "ADOPT"
}
Weighted knob-space distance metricrecord 71
{
"id": "I71",
"kind": "innovation",
"group": "output-compilation",
"name": "Weighted knob-space distance metric",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Typography and chroma likely dominate perception; equal weighting is probably wrong.",
"limitations": "Weights must be fitted, adding parameters.",
"disposition": "test"
}
Second- and third-choice packs offered alongside the firstrecord 72
{
"id": "I72",
"kind": "innovation",
"group": "output-compilation",
"name": "Second- and third-choice packs offered alongside the first",
"evidence_class": "inferred",
"source": "derived; supports I59",
"observed": "2026-08-27",
"claim": "Offering the top 3 both hedges model error and generates validation data for free.",
"limitations": "Reintroduces choice at the moment we claimed to have decided.",
"disposition": "ADOPT"
}
TasteProfile schema: per-knob posterior mean + interval + n_observationsrecord 73
{
"id": "I73",
"kind": "innovation",
"group": "output-compilation",
"name": "TasteProfile schema: per-knob posterior mean + interval + n_observations",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Makes uncertainty first-class and auditable in the output contract.",
"limitations": "—",
"disposition": "ADOPT"
}
PreferenceConfidence as a per-knob vector, not a scalarrecord 74
{
"id": "I74",
"kind": "innovation",
"group": "output-compilation",
"name": "PreferenceConfidence as a per-knob vector, not a scalar",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Mirrors I55 in the output contract.",
"limitations": "—",
"disposition": "ADOPT"
}
DesignDNA as a versioned, reproducible artefactrecord 75
{
"id": "I75",
"kind": "innovation",
"group": "output-compilation",
"name": "DesignDNA as a versioned, reproducible artefact",
"evidence_class": "inferred",
"source": "token-pack-science pack.version discipline",
"observed": "2026-08-27",
"claim": "Pin the model version, catalogue version, and elicitation transcript so a profile is reproducible.",
"limitations": "—",
"disposition": "ADOPT"
}
Preference vector pre-filters P05 component suggestionsrecord 76
{
"id": "I76",
"kind": "innovation",
"group": "output-compilation",
"name": "Preference vector pre-filters P05 component suggestions",
"evidence_class": "observed",
"source": "live demo states this intent",
"observed": "2026-08-27",
"claim": "The same vector that picks a pack can rank components, so elicitation pays for itself twice.",
"limitations": "Component preference may not be the same construct as pack preference.",
"disposition": "test"
}
Preference vector seeds image-generation roundsrecord 77
{
"id": "I77",
"kind": "innovation",
"group": "output-compilation",
"name": "Preference vector seeds image-generation rounds",
"evidence_class": "observed",
"source": "live demo states this intent",
"observed": "2026-08-27",
"claim": "Reuse for generated imagery.",
"limitations": "Unverified that the vector transfers to a different modality.",
"disposition": "test"
}
Per-context confidence (marketing page vs dense dashboard)record 78
{
"id": "I78",
"kind": "innovation",
"group": "output-compilation",
"name": "Per-context confidence (marketing page vs dense dashboard)",
"evidence_class": "hypothesis",
"source": "derived; parts.json open question",
"observed": "2026-08-27",
"claim": "A client's density preference on a landing page may differ from a dashboard. Model context as a conditioning variable rather than assuming one profile fits all surfaces.",
"limitations": "Multiplies the parameter count by the number of contexts, directly worsening the sample-size problem.",
"disposition": "test — this is the open question with the worst cost/benefit"
}
Single profile with per-context OFFSETS rather than separate profilesrecord 79
{
"id": "I79",
"kind": "innovation",
"group": "output-compilation",
"name": "Single profile with per-context OFFSETS rather than separate profiles",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "A cheaper resolution of I78: one taste vector plus small learned context deltas.",
"limitations": "Offsets still need data to fit.",
"disposition": "test — preferred form of I78"
}
Stated vs revealed preference agreement metricrecord 80
{
"id": "I80",
"kind": "innovation",
"group": "output-compilation",
"name": "Stated vs revealed preference agreement metric",
"evidence_class": "observed",
"source": "Pinterest deprecated onboarding interests in favour of engagement clusters",
"observed": "2026-08-27",
"claim": "Measure whether what the client PICKS matches what they later APPROVE in delivered work. The single most commercially important validation.",
"limitations": "Requires delivered projects, so it is a slow metric.",
"disposition": "ADOPT — long-horizon"
}
Fall back to a safe default pack on low confidencerecord 81
{
"id": "I81",
"kind": "innovation",
"group": "output-compilation",
"name": "Fall back to a safe default pack on low confidence",
"evidence_class": "inferred",
"source": "Shopify SDUI unknown-layout fallback; Netflix popular-set fallback",
"observed": "2026-08-27",
"claim": "If confidence is too low, ship an industry-appropriate default rather than a badly-inferred pack.",
"limitations": "—",
"disposition": "ADOPT"
}
Log spec/pack collisions as a coverage signalrecord 82
{
"id": "I82",
"kind": "innovation",
"group": "output-compilation",
"name": "Log spec/pack collisions as a coverage signal",
"evidence_class": "observed",
"source": "ui-pick-to-spec §6 failure mode 4",
"observed": "2026-08-27",
"claim": "If two different clients converge to identical packs, that tells you which knob to add.",
"limitations": "—",
"disposition": "ADOPT"
}
Human review gate before pack is committed to a buildrecord 83
{
"id": "I83",
"kind": "innovation",
"group": "output-compilation",
"name": "Human review gate before pack is committed to a build",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "A designer confirms the selected pack before it drives a client deliverable.",
"limitations": "Costs human time per client; undermines the automation claim.",
"disposition": "ADOPT for v1, revisit"
}
Profile portability across a client's multiple projectsrecord 84
{
"id": "I84",
"kind": "innovation",
"group": "output-compilation",
"name": "Profile portability across a client's multiple projects",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Reuse the DesignDNA for the same client's later work.",
"limitations": "Preference may be project-specific, not client-specific.",
"disposition": "test"
}
Explicit re-elicitation rather than silent drift trackingrecord 85
{
"id": "I85",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Explicit re-elicitation rather than silent drift tracking",
"evidence_class": "inferred",
"source": "Netflix supersession; Pinterest UIC replacement",
"observed": "2026-08-27",
"claim": "Both major platforms found stated onboarding preference decays against behaviour. Prefer a cheap explicit re-run over an invisible decay model.",
"limitations": "Costs client time.",
"disposition": "ADOPT"
}
Revealed-preference signal from delivered-work approvalsrecord 86
{
"id": "I86",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Revealed-preference signal from delivered-work approvals",
"evidence_class": "inferred",
"source": "Hinge infers rankings from like/pass; Stitch Fix updates in real time",
"observed": "2026-08-27",
"claim": "Client approvals/rejections of delivered designs are the highest-quality preference signal available and cost nothing to collect.",
"limitations": "Slow, sparse, and confounded with non-aesthetic factors (deadline, budget).",
"disposition": "ADOPT — long-horizon"
}
Regulated industries constrain the stimulus space a priorirecord 87
{
"id": "I87",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Regulated industries constrain the stimulus space a priori",
"evidence_class": "inferred",
"source": "b2b-template-shelf-report.md: healthcare PHI, law firm privilege, insurance document authority",
"observed": "2026-08-27",
"claim": "Healthcare, legal, insurance, mortgage clients should not be shown high-chroma, high-motion, low-contrast packs at all — prune before eliciting.",
"limitations": "Risks the pruning being wrong and the client wanting distinctiveness.",
"disposition": "ADOPT"
}
Accessibility floor is not a preference axisrecord 88
{
"id": "I88",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Accessibility floor is not a preference axis",
"evidence_class": "observed",
"source": "token-pack-science Gates C and D",
"observed": "2026-08-27",
"claim": "Contrast below AA is never offered regardless of stated preference. The client cannot choose to fail WCAG.",
"limitations": "Removes a genuine aesthetic direction (low-contrast minimalism).",
"disposition": "ADOPT"
}
End-user vs operator preference divergencerecord 89
{
"id": "I89",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "End-user vs operator preference divergence",
"evidence_class": "inferred",
"source": "b2b-template-shelf-report.md: property_management tenant portal is 'a distinct untrusted identity'; it_services_msps 'per-client tenancy is the product'",
"observed": "2026-08-27",
"claim": "For the 6/17 industries where `portal` is secondary, the person choosing (the operator) is not the person using (the end customer). Elicit from the operator but weight toward end-user norms.",
"limitations": "No mechanism yet for representing the absent end user.",
"disposition": "ADOPT — open design problem"
}
Density preference conditioned on archetyperecord 90
{
"id": "I90",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Density preference conditioned on archetype",
"evidence_class": "inferred",
"source": "b2b-template-shelf: case_workflow primary for 6/17, portal secondary for 6/17",
"observed": "2026-08-27",
"claim": "Two archetypes carry the majority of catalogue demand; density expectations differ sharply between a dense case-workflow console and a sparse client portal.",
"limitations": "Adds a conditioning variable.",
"disposition": "ADOPT — cheapest useful form of I78"
}
Conservative-industry prior on chroma and motionrecord 91
{
"id": "I91",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Conservative-industry prior on chroma and motion",
"evidence_class": "inferred",
"source": "accounting_firms, law_firms, insurance_agencies, mortgage_brokers variant deltas",
"observed": "2026-08-27",
"claim": "Seed these industries from a low-chroma, low-motion region of knob space.",
"limitations": "Stereotyping risk; must be a prior the client can override in 2-3 rounds.",
"disposition": "ADOPT"
}
Expressive-industry priorrecord 92
{
"id": "I92",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Expressive-industry prior",
"evidence_class": "inferred",
"source": "course_creators, marketing_social_media_agencies, real_estate variant deltas",
"observed": "2026-08-27",
"claim": "Seed from a higher-chroma, higher-expressiveness region.",
"limitations": "Same stereotyping risk.",
"disposition": "ADOPT"
}
Industry prior must be overridable within 3 roundsrecord 93
{
"id": "I93",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Industry prior must be overridable within 3 rounds",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Guarantee the prior is a starting point, not a cage: a client must be able to move out of their industry's region quickly.",
"limitations": "Requires measuring prior strength.",
"disposition": "ADOPT"
}
Bespoke motion and illustration explicitly out of scoperecord 94
{
"id": "I94",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Bespoke motion and illustration explicitly out of scope",
"evidence_class": "observed",
"source": "ui-pick-to-spec §6 failure modes 1 and 2",
"observed": "2026-08-27",
"claim": "The spec bottleneck loses hand-tuned parallax and custom artwork. Do not let a knob imply a capability that does not exist.",
"limitations": "Removes a real source of client delight.",
"disposition": "ADOPT"
}
The preference learner must itself be A/B tested against a static defaultrecord 95
{
"id": "I95",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "The preference learner must itself be A/B tested against a static default",
"evidence_class": "observed",
"source": "Spotify Engineering: a bandit is 'a feature you've built, not an experimental method'",
"observed": "2026-08-27",
"claim": "Without a holdout comparing elicited packs to an industry-default pack, we cannot claim the elicitation works at all.",
"limitations": "Needs enough clients for a powered comparison — likely the binding constraint.",
"disposition": "ADOPT — the single most important methodological gate"
}
Measure the knob-correlation matrix before assuming low effective dimensionrecord 96
{
"id": "I96",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Measure the knob-correlation matrix before assuming low effective dimension",
"evidence_class": "hypothesis",
"source": "derived §4; Jamieson & Nowak's d log n is contingent on the embedding assumption holding",
"observed": "2026-08-27",
"claim": "The whole few-rounds argument rests on effective d < 7. Measure it, do not assume it.",
"limitations": "—",
"disposition": "ADOPT — falsifier F3"
}
Do not cite Miller 7+/-2 to justify gallery sizerecord 97
{
"id": "I97",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Do not cite Miller 7+/-2 to justify gallery size",
"evidence_class": "observed",
"source": "Miller 1956 is about absolute judgment and memory span, not interface option counts",
"observed": "2026-08-27",
"claim": "Citing it would not survive a technical client's scrutiny.",
"limitations": "—",
"disposition": "ADOPT as a writing rule"
}
Do not cite the jam study as a design principlerecord 98
{
"id": "I98",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Do not cite the jam study as a design principle",
"evidence_class": "observed",
"source": "Scheibehenne et al. 2010: D = 0.02, CI95 [-0.09, 0.12], 63 conditions, N = 5,036",
"observed": "2026-08-27",
"claim": "Choice overload does not robustly replicate; size galleries on information grounds.",
"limitations": "—",
"disposition": "ADOPT as a writing rule"
}
Do not promise a specific accuracy number before measuringrecord 99
{
"id": "I99",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Do not promise a specific accuracy number before measuring",
"evidence_class": "observed",
"source": "ui-pick-to-spec §4.2 no-figure-found; the 728-vs-113 connector error class",
"observed": "2026-08-27",
"claim": "No published benchmark measures our task; any quoted figure would be extrapolation.",
"limitations": "—",
"disposition": "ADOPT as a writing rule"
}
Treat the live demo's convergence claims as unmeasuredrecord 100
{
"id": "I100",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "Treat the live demo's convergence claims as unmeasured",
"evidence_class": "observed",
"source": "actionist-taste.pages.dev asserts '~10 picks instead of ~50' with no measurement",
"observed": "2026-08-27",
"claim": "Our own prototype's marketing claims must not be recycled as evidence into a client deliverable.",
"limitations": "—",
"disposition": "ADOPT as a writing rule"
}
Continuous preference vector internally, closed pack at the boundaryrecord 101
{
"id": "I67-TOP10",
"kind": "innovation",
"group": "output-compilation",
"name": "Continuous preference vector internally, closed pack at the boundary",
"evidence_class": "inferred",
"source": "derived §8; token-pack-science gates A-J; lane synthesis 'learn a vector not a pack ID'",
"observed": "2026-08-27",
"claim": "Catalogue can grow without re-eliciting; shipped artefact is always gate-passing.",
"limitations": "Nearest-neighbour metric weighting is an open parameter.",
"disposition": "ADOPT — primary recommendation",
"refers_to": "I67",
"rank": 1,
"rationale": "Resolves the central architectural question (d): continuous vector internally, closed gate-passing pack at the boundary. Everything downstream — catalogue growth, accessibility guarantees, overfitting control — depends on this split being right."
}
The preference learner must itself be A/B tested against a static defaultrecord 102
{
"id": "I95-TOP10",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "The preference learner must itself be A/B tested against a static default",
"evidence_class": "observed",
"source": "Spotify Engineering: a bandit is 'a feature you've built, not an experimental method'",
"observed": "2026-08-27",
"claim": "Without a holdout comparing elicited packs to an industry-default pack, we cannot claim the elicitation works at all.",
"limitations": "Needs enough clients for a powered comparison — likely the binding constraint.",
"disposition": "ADOPT — the single most important methodological gate",
"refers_to": "I95",
"rank": 2,
"rationale": "Without a holdout against a static industry-default pack we cannot claim the learner works at all. Spotify's framing is decisive: a personalisation system is a feature, not an experimental method. This is the gate that turns the lane from plausible into measured."
}
Warm-start prior from a population aesthetic modelrecord 103
{
"id": "I40-TOP10",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Warm-start prior from a population aesthetic model",
"evidence_class": "observed",
"source": "LAION aesthetic predictor; Brochu et al. 2010 cut iterations 11.25 -> 6.5 with a learned prior",
"observed": "2026-08-27",
"claim": "Learn the client's DEVIATION from a population baseline rather than their taste from scratch. The largest single efficiency lever reported anywhere in the applied literature.",
"limitations": "Population aesthetic is not per-client taste and may encode a generic bias.",
"disposition": "ADOPT",
"refers_to": "I40",
"rank": 3,
"rationale": "A warm-start population prior is the largest single efficiency lever reported anywhere in the applied literature: Brochu et al. 2010 cut iterations from 11.25 to 6.5. It is also cheap, since the LAION aesthetic predictor already exists."
}
Perturbation-tolerance testrecord 104
{
"id": "I60-TOP10",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Perturbation-tolerance test",
"evidence_class": "hypothesis",
"source": "derived",
"observed": "2026-08-27",
"claim": "Show the converged pack and a +/-1-level perturbation. If clients reliably distinguish and reject, the tolerance argument in §2 collapses and round counts triple.",
"limitations": "—",
"disposition": "ADOPT — falsifier F1, run FIRST",
"refers_to": "I60",
"rank": 4,
"rationale": "The perturbation-tolerance test is falsifier F1 and should run FIRST, because the entire round-count argument rests on +/-1-level tolerance being acceptable. If it is not, the requirement roughly triples and the product changes shape."
}
Brand-constraint hard cutoffs before elicitationrecord 105
{
"id": "I08-TOP10",
"kind": "innovation",
"group": "stimulus-design",
"name": "Brand-constraint hard cutoffs before elicitation",
"evidence_class": "observed",
"source": "Sawtooth ACBC must-have/unacceptable mechanism (vendor manual)",
"observed": "2026-08-27",
"claim": "Existing brand colours/fonts become hard prunes of the stimulus space, not soft preferences; all later stimuli satisfy them.",
"limitations": "Over-pruning can empty the space; needs a fallback.",
"disposition": "ADOPT",
"refers_to": "I08",
"rank": 5,
"rationale": "Sawtooth's must-have/unacceptable cutoff mechanism is the missing piece for brand constraints, and it is proven commercial practice. A hard prune is categorically better than learning a client's non-negotiables as a soft preference."
}
Simulate-to-size: Monte Carlo power analysis for round countrecord 106
{
"id": "I42-TOP10",
"kind": "innovation",
"group": "adaptive-policy",
"name": "Simulate-to-size: Monte Carlo power analysis for round count",
"evidence_class": "observed",
"source": "skpr (R)",
"observed": "2026-08-27",
"claim": "Answer 'how many rounds' EMPIRICALLY by simulating synthetic clients under a BT likelihood rather than arguing from bounds.",
"limitations": "Power analysis is built for GLM responses; needs adapting.",
"disposition": "ADOPT — this is how the N question should actually be settled",
"refers_to": "I42",
"rank": 6,
"rationale": "Simulate-to-size settles the 'how many choices' question empirically rather than by argument. Given that no paper gives a sufficiency guarantee for our setting, a Monte Carlo power analysis over synthetic clients is the only route to a defensible number."
}
Two-week test-retest stability as the primary validation metricrecord 107
{
"id": "I57-TOP10",
"kind": "innovation",
"group": "stopping-and-confidence",
"name": "Two-week test-retest stability as the primary validation metric",
"evidence_class": "inferred",
"source": "Slovic 1995 construction of preference",
"observed": "2026-08-27",
"claim": "A model that fits the clicks but changes answer next Tuesday is worthless. Measures whether the thing we claim to measure exists.",
"limitations": "Requires client time two weeks apart; hard to get.",
"disposition": "ADOPT — this is falsifier F7",
"refers_to": "I57",
"rank": 7,
"rationale": "Two-week test-retest is falsifier F7 — the one that would kill the part rather than adjust it. If preference is not stable, no amount of elicitation precision matters. Test early and cheaply."
}
P05 corpus components as re-themable controlled stimulirecord 108
{
"id": "I03-TOP10",
"kind": "innovation",
"group": "stimulus-design",
"name": "P05 corpus components as re-themable controlled stimuli",
"evidence_class": "observed",
"source": "research/21st-corpus-audit-2026-08-27.md — 86.7% of colour-bearing CSS rules already resolve through var(--token); 87.3% carry a shadcn oklch :root block",
"observed": "2026-08-27",
"claim": "Use the 8,515-identity corpus as the stimulus pool: identical structure, varying token pack, re-theming is a find-and-replace on ~30 custom properties.",
"limitations": "271 components lack previews; component SOURCE is absent from the corpus.",
"disposition": "ADOPT",
"refers_to": "I03",
"rank": 8,
"rationale": "The P05 corpus turns stimulus generation from a cost into an asset: 8,515 identities, 86.7% of colour rules already indirected through tokens, re-theming as a ~30-property find-and-replace. Identical structure with varying treatment is the ideal experimental article."
}
4-up gallery with explicit outside optionrecord 109
{
"id": "I19-TOP10",
"kind": "innovation",
"group": "question-format",
"name": "4-up gallery with explicit outside option",
"evidence_class": "inferred",
"source": "derived §3 + Sawtooth 4-5 per screen + live demo",
"observed": "2026-08-27",
"claim": "2.32 noiseless bits/round; ~0.91 effective at p=0.15; balances information against deliberation cost and avoids the hostility of 'worst'.",
"limitations": "Best-worst is strictly more informative per question.",
"disposition": "ADOPT — primary recommendation",
"refers_to": "I19",
"rank": 9,
"rationale": "The primary format recommendation — 4-up with an explicit outside option — supported independently by the information derivation, Sawtooth's 4-5-per-screen guidance, and our own working prototype."
}
End-user vs operator preference divergencerecord 110
{
"id": "I89-TOP10",
"kind": "innovation",
"group": "drift-industry-risk",
"name": "End-user vs operator preference divergence",
"evidence_class": "inferred",
"source": "b2b-template-shelf-report.md: property_management tenant portal is 'a distinct untrusted identity'; it_services_msps 'per-client tenancy is the product'",
"observed": "2026-08-27",
"claim": "For the 6/17 industries where `portal` is secondary, the person choosing (the operator) is not the person using (the end customer). Elicit from the operator but weight toward end-user norms.",
"limitations": "No mechanism yet for representing the absent end user.",
"disposition": "ADOPT — open design problem",
"refers_to": "I89",
"rank": 10,
"rationale": "The operator-vs-end-user divergence is a genuine unsolved design problem that affects 6 of 17 industries, and nobody in the commercial denominator has solved it. Naming it now prevents shipping a system that optimises for the wrong person."
}