P06 · Experience · Rendered from source

source register

Design taste and preference learner

27 lines19,288 bytessha256 9adbfde5b776
P06-SR-001record 1
{
  "id": "P06-SR-001",
  "evidence_class": "observed",
  "source": "research/ui-pick-to-spec-2026-08-27.md",
  "observed": "2026-08-27",
  "claim": "Local prior establishing pick-as-query-not-asset, a closed semantic section vocabulary of ~6-12 enum axes per section type, SDUI prior art (Airbnb/Lyft/Spotify/Shopify/DoorDash), a four-layer validation gate, and the requirement that previews be re-rendered from spec rather than shown as original thumbnails.",
  "limitations": "Report is internal synthesis; its own enum completeness is marked CONVENTION and unvalidated against a labelled sample.",
  "disposition": "foundational_input"
}
P06-SR-002record 2
{
  "id": "P06-SR-002",
  "evidence_class": "observed",
  "source": "research/token-pack-science-2026-08-27.md",
  "observed": "2026-08-27",
  "claim": "Local prior establishing the closed token-pack gallery thesis, DTCG 2025.10 as interchange format, Radix 12-step role scales with APCA targets, M3 HCT seed-to-role determinism, a complete pack schema and machine gates A-J.",
  "limitations": "States explicitly that a 20-30 pack catalogue size is a product decision, not an evidence-backed number.",
  "disposition": "foundational_input"
}
P06-SR-003record 3
{
  "id": "P06-SR-003",
  "evidence_class": "observed",
  "source": "research/actionmodel-builder-research-2026-08-26/phase-8/external-opus-inputs/CLAUDE-LANES-SYNTHESIS.md",
  "observed": "2026-08-27",
  "claim": "Names the seven mechanical preference knobs (contrast, radius, shadow, border, typography, density, chroma), identifies Bradley-Terry/Luce as relevant models, and records that forced choice without a 'none' option pollutes the preference model.",
  "limitations": "Conversation-derived synthesis; the knob set is asserted, not measured for independence.",
  "disposition": "foundational_input"
}
P06-SR-004record 4
{
  "id": "P06-SR-004",
  "evidence_class": "observed",
  "source": "site/system-map/data/parts.json (P06 entry)",
  "observed": "2026-08-27",
  "claim": "Defines P06 ownership: preference axes and stimulus generation, experiment policy, outside option and re-roll behaviour, client design DNA and confidence, preference update over time. Open questions include whether the minimum is 7, 14 or another number of choices.",
  "limitations": "Part definition, not evidence.",
  "disposition": "scope_definition"
}
P06-SR-005record 5
{
  "id": "P06-SR-005",
  "evidence_class": "observed",
  "source": "https://en.wikipedia.org/wiki/Bradley%E2%80%93Terry_model",
  "observed": "2026-08-27",
  "claim": "Bradley-Terry gives Pr(i>j) = p_i/(p_i+p_j) with positive real scores; logit form logit Pr(i>j) = beta_i - beta_j. Original citation Bradley & Terry (1952) 'Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons', Biometrika 39(3/4):324-345, doi:10.2307/2334029. Previously studied by Zermelo (1929). Plackett-Luce generalizes to full rankings and satisfies Luce's choice axiom.",
  "limitations": "Encyclopedia summary of the model; does not itself cover tie models (Davidson 1970 not present on this page). Primary sources not read in this run.",
  "disposition": "model_foundation"
}
P06-SR-006record 6
{
  "id": "P06-SR-006",
  "evidence_class": "observed",
  "source": "https://en.wikipedia.org/wiki/Bradley%E2%80%93Terry_model",
  "observed": "2026-08-27",
  "claim": "Crowd-BT (Chen, Bennett, Collins-Thompson, Horvitz, WSDM 2013, doi:10.1145/2433396.2433420) reduces the number of comparisons required by modelling per-judge reliability and screening unreliable raters; outperformed plain Bradley-Terry and TrueSkill in a 624-judge document-difficulty task.",
  "limitations": "Reported via encyclopedia summary; the original paper was not read in this run. The task domain is document difficulty, not visual taste.",
  "disposition": "directly_relevant_extension"
}
P06-SR-007record 7
{
  "id": "P06-SR-007",
  "evidence_class": "observed",
  "source": "https://en.wikipedia.org/wiki/Best-worst_scaling",
  "observed": "2026-08-27",
  "claim": "Best-worst scaling: with four items, naming best and worst reveals 5 of the 6 implied pairwise relations (only the middle pair stays unknown); with five items, 7 of 10. Invented by Jordan Louviere in 1987; concept traces to A.A.J. Marley; definitive text Louviere, Flynn & Marley (2015, Cambridge University Press).",
  "limitations": "Information-recovery figures are combinatorial facts about implied relations, not a measurement of estimator efficiency under response noise.",
  "disposition": "format_evidence"
}
P06-SR-008record 8
{
  "id": "P06-SR-008",
  "evidence_class": "observed",
  "source": "https://en.wikipedia.org/wiki/Best-worst_scaling",
  "observed": "2026-08-27",
  "claim": "MaxDiff is a strict subset of BWS assuming respondents evaluate all n(n-1) ordered pairs and pick the maximum-difference pair; the source reports that across 14 years of presentations the authors virtually never found a practitioner who admitted using that strategy, with most describing sequential best-then-worst strategies.",
  "limitations": "Anecdotal evidence about respondent strategy, reported by the method's own authors.",
  "disposition": "caution_on_model_assumptions"
}
P06-SR-009record 9
{
  "id": "P06-SR-009",
  "evidence_class": "observed",
  "source": "https://en.wikipedia.org/wiki/Best-worst_scaling",
  "observed": "2026-08-27",
  "claim": "BWS Case 2 (profile case) has proven usable by populations that struggle with conventional discrete choice experiments; one cited study found nearly all older respondents produced usable BWS data while only about half did for the DCE.",
  "limitations": "Single cited study, health domain, no effect size read in this run.",
  "disposition": "accessibility_evidence"
}
P06-SR-010record 10
{
  "id": "P06-SR-010",
  "evidence_class": "observed",
  "source": "https://ui.shadcn.com/docs/theming",
  "observed": "2026-08-27",
  "claim": "shadcn theming exposes semantic token pairs (background/foreground, card, popover, primary, secondary, muted, accent, destructive), structural tokens (border, input, ring), chart-1..5, a sidebar token group, and a radius scale. Surface tokens pair with a matching -foreground token for text and icons.",
  "limitations": "Vendor documentation; describes convention, not enforcement.",
  "disposition": "confirms_local_corpus_finding"
}
P06-SR-011record 11
{
  "id": "P06-SR-011",
  "evidence_class": "observed",
  "source": "research/21st-corpus-audit-2026-08-27.md",
  "observed": "2026-08-27",
  "claim": "The 21st bundle corpus carries a standard shadcn :root oklch token block in 87.3% of sampled bundles (n=200) and 86.7% of colour-bearing CSS rules already resolve through var(--token) (n=25), making bundles re-themable as controlled stimuli by swapping ~30 custom properties.",
  "limitations": "Sampled figures, not corpus-wide; WebGL components (~16% of n=25) and Tailwind v3 builds (8.3% of n=60) are excluded cases.",
  "disposition": "stimulus_supply_evidence"
}
P06-SR-012record 12
{
  "id": "P06-SR-012",
  "evidence_class": "inferred",
  "source": "lanes/S1-L3/owner-finding-p06-choice-budget.md",
  "observed": "2026-08-27",
  "claim": "Information-theoretic derivation: 7 knobs x 4 discriminable levels = 16,384 profiles = 14 bits for exact identification; selecting among a closed 30-pack gallery is only ~4.9 bits, cutting the elicitation budget by roughly two thirds.",
  "limitations": "Derivation over assumed parameters (4 levels/knob, 0.4-0.6 bits/comparison efficiency). The efficiency band was NOT verified against a specific paper in this run and must not be quoted as a measured constant. SUPERSEDED by the channel-capacity derivation in first-principles.md §3, which replaces the assumed 0.4-0.6 bits/comparison band with an explicit binary-symmetric-channel noise model and a tolerance-adjusted target of 4.04 bits.",
  "disposition": "superseded_by_SR_015"
}
P06-SR-013record 13
{
  "id": "P06-SR-013",
  "evidence_class": "observed",
  "source": "https://actionist-taste.pages.dev/",
  "observed": "2026-08-27",
  "claim": "Live taste-picker prototype INSPECTED. Four rendered interface cards per round ('every card is a bundle of five knobs'), explicit outside option 'none of these' plus 'start over', ~10 picks total, later rounds target the 'least-resolved knob' holding other inferred preferences steady. Five of seven knobs varied (palette, radius, type, density, shadow). Live belief panel 'What the system believes about you'. Estimator is win-rate counting with Laplace smoothing, a zeroth-order Bradley-Terry; the page names choix or a GP preference model as production intent.",
  "limitations": "A demo, not a measured system. Its '~10 picks instead of ~50' convergence claim has NO measurement behind it and must not be recycled into a client deliverable. UI inconsistency: labels itself 'pick 1 of 10' while showing four cards.",
  "disposition": "RESOLVED — adopt the interaction model, replace the estimator, add a real stopping rule"
}
P06-SR-014record 14
{
  "id": "P06-SR-014",
  "evidence_class": "observed",
  "source": "Commercial preference-elicitation denominator (Stitch Fix, Netflix cold-start, Spotify onboarding, Pinterest interest picker, Sawtooth, Conjointly, Qualtrics, Canva Brand Kit, Figma First Draft)",
  "observed": "2026-08-27",
  "claim": "Commercial denominator BUILT: 103 surfaces across 10 categories in top-companies.jsonl, 10 ranked.",
  "limitations": "WEAK LEG. Only 13 of 103 are 'observed'; 57 are 'hypothesis'. Consumer style quizzes are client-side JS behind bot protection — 403 from Stitch Fix, Warby Parker, Looka, Behr, Hinge; JS shell from Function of Beauty; 404 from 1000minds and Conjointly. Question counts and adaptivity for that whole category are NOT fetchable and need a browser session.",
  "disposition": "PARTIALLY RESOLVED — enumerated, thinly verified"
}
P06-SR-015record 15
{
  "id": "P06-SR-015",
  "evidence_class": "observed",
  "source": "first-principles.md §1-§3 (derivation, this run)",
  "observed": "2026-08-27",
  "claim": "Channel-capacity derivation: 7 knobs at plausible level counts give 36,000 packs = 15.14 bits; resolving each knob to +/-1 level leaves 11.09 bits deliberately unresolved, so only 4.04 bits must be acquired. Through a binary symmetric channel at p=0.15 a 4-up-plus-outside-option screen delivers ~0.91 effective bits, giving ~5 rounds; recommend 8-12 for myopic acquisition, exploration overhead and knob correlation.",
  "limitations": "Level counts per knob are plausible assumptions, not measured. The +/-1-level tolerance is the load-bearing assumption and is falsifier F1. Sensitivity across four knob-space scenarios is reported in §3.4 and the recommendation survives all four.",
  "disposition": "primary_derivation"
}
P06-SR-016record 16
{
  "id": "P06-SR-016",
  "evidence_class": "observed",
  "source": "https://arxiv.org/abs/2005.04107 — Sequential Gallery, Koyama, Sato & Goto, SIGGRAPH 2020 (full text)",
  "observed": "2026-08-27",
  "claim": "Closest published prior art. Verbatim: 'The mean iteration count necessary for finding satisfactory results was 5.36 with SD = 2.69'; 'One plane-search subtask took 14.8 seconds on average'. Participants continued to 15 iterations regardless; 5 of 6 pressed satisfaction within that. Synthetic functions at 5D, 15D, 10D, 20D. Grid was 5x5 for photo enhancement but reduced to 3x3 on a 13-inch display for body-shape design.",
  "limitations": "n=6 participants (5 students + 1 researcher), a preliminary study. 'Satisfactory' is self-reported, not distance to a ground-truth optimum. Grid size is application- and display-dependent; quoting one number would repeat the 728-vs-113 error class.",
  "disposition": "primary_empirical_anchor"
}
P06-SR-017record 17
{
  "id": "P06-SR-017",
  "evidence_class": "observed",
  "source": "https://arxiv.org/abs/1109.3701 — Jamieson & Nowak, Active Ranking using Pairwise Comparisons, NIPS 2011",
  "observed": "2026-08-27",
  "claim": "Theoretical licence for few comparisons over 7 knobs. Verbatim: objects in d-dimensional Euclidean space ranked by distance from a reference, 'the number of possible rankings grows like n^{2d}', and an algorithm identifies a randomly selected ranking using 'just slightly more than d log n adaptively selected pairwise comparisons, on average'; if comparisons are chosen at random 'almost all pairwise comparisons must be made'.",
  "limitations": "AVERAGE-CASE over a randomly selected ranking, NOT worst case — for d>=2 there exist placements requiring at least n-1 queries. Entirely contingent on adaptivity AND on the low-dimensional embedding assumption actually holding, which is unmeasured for our knobs (Gate 1).",
  "disposition": "theoretical_licence"
}
P06-SR-018record 18
{
  "id": "P06-SR-018",
  "evidence_class": "observed",
  "source": "https://arxiv.org/abs/1606.08842 — Heckel, Shah, Ramchandran & Wainwright; Annals of Statistics 2019",
  "observed": "2026-08-27",
  "claim": "The counterweight that prevents over-promising. An adaptive counting algorithm with confidence-interval stopping recovers the ranking 'optimal up to logarithmic factors' with no structural assumption on the pairwise probability matrix; and a lower bound proves parametric choices (Bradley-Terry, Thurstone) 'offer at most logarithmic gains for stochastic comparisons'.",
  "limitations": "Requires pairwise probabilities bounded away from zero. Ranks by probability of beating a random item, not by recovering a utility vector.",
  "disposition": "constrains_optimism"
}
P06-SR-019record 19
{
  "id": "P06-SR-019",
  "evidence_class": "observed",
  "source": "https://sawtoothsoftware.com/help/lighthouse-studio/manual/maxdiff-designing-study.html (fetched)",
  "observed": "2026-08-27",
  "claim": "The only hard question-count formula found in commercial practice, quoted verbatim: 'we recommend displaying either four or five items at a time (per set or question)'; never more than half the study's item count per set; ask enough sets 'such that each item has the opportunity to appear from three to five times per respondent'; formula '3K/k' where K = total items, k = items per set. Beyond ~5 items per set 'The gains in precision of the estimates are minimal.' Assumes HB estimation.",
  "limitations": "This is a design guideline for ITEM SCALING (ranking K items), not for locating a point in a continuous knob space — it maps to our problem by analogy, not identity.",
  "disposition": "format_corroboration"
}
P06-SR-020record 20
{
  "id": "P06-SR-020",
  "evidence_class": "observed",
  "source": "https://docs.midjourney.com/hc/en-us/articles/41308374558221-Style-Creator",
  "observed": "2026-08-27",
  "claim": "A SHIPPED adaptive visual-preference elicitor. Uses 'the styles you pick (and the ones you don't!)' to build a reusable --sref code. Each regenerated preview set is one refinement round. Documented convergence: 'Most styles stabilize after 5-10 rounds', rounds 10-15 add detail, 'Past round 15, changes are small and subtle'. Skipping is available but explicitly NON-informative: 'skipping does not affect your style development'. User-terminated via End Session.",
  "limitations": "Direct fetch returned 403; content read from the search index of the official docs page. Grid size per round not stated. No published accuracy or user study. We deliberately DIVERGE on the skip semantics.",
  "disposition": "closest_shipped_analogue"
}
P06-SR-021record 21
{
  "id": "P06-SR-021",
  "evidence_class": "observed",
  "source": "https://medium.com/pinterest-engineering/pinner-progression-better-use-case-representation-driving-weekly-active-user-growth-at-pinterest-bd2131ab238a",
  "observed": "2026-08-27",
  "claim": "The most important negative result in the commercial sweep. Pinterest REPLACED onboarding followed-interests as a retrieval condition with user interest clusters derived from actual engagement, because the onboarding signal 'skewed heavily toward dominant interests' and was 'static, not evolving with behavior'.",
  "limitations": "Engineering blog, not a product doc; timing and scope of the replacement not fully specified. Does not imply elicitation is useless at cold start, which is precisely the case we face — it implies the profile is perishable.",
  "disposition": "counter_evidence_on_drift"
}
P06-SR-022record 22
{
  "id": "P06-SR-022",
  "evidence_class": "observed",
  "source": "https://help.netflix.com/en/node/100639",
  "observed": "2026-08-27",
  "claim": "Vendor-documented supersession rule. Title picker at profile creation is explicitly OPTIONAL; skipping falls back to 'a diverse and popular set of titles'. Once engagement begins it 'supersedes' the initial preferences, with recent activity outweighing older.",
  "limitations": "Help page states behaviour but gives no title count and no weighting detail.",
  "disposition": "drift_precedent"
}
P06-SR-023record 23
{
  "id": "P06-SR-023",
  "evidence_class": "observed",
  "source": "https://doi.org/10.1086/651235 — Scheibehenne, Greifeneder & Todd 2010 (full text)",
  "observed": "2026-08-27",
  "claim": "Settles a question that would otherwise silently shape the UI. Meta-analysis of 63 conditions from 50 experiments (N=5,036): mean effect size D = 0.02, CI95 [-0.09, 0.12]; trimmed D = 0.001. 'No sufficient conditions could be identified' and 'adverse consequences due to having too much choice are not a robust phenomenon'.",
  "limitations": "Heterogeneity I2 = 68% untrimmed suggests moderators may exist. Meta-analysis to 2010. Slight publication bias noted.",
  "disposition": "frees_gallery_sizing"
}
P06-SR-024record 24
{
  "id": "P06-SR-024",
  "evidence_class": "observed",
  "source": "https://doi.org/10.1037/h0043158 — Miller 1956; and https://doi.org/10.1037/0003-066x.50.5.364 — Slovic 1995",
  "observed": "2026-08-27",
  "claim": "Two writing rules. Miller's 7+/-2 concerns absolute judgment of unidimensional stimuli and immediate memory span for chunks, NOT interface option counts; Miller himself treats the recurrence of ~7 with irony. Slovic establishes that preferences are frequently CONSTRUCTED during elicitation rather than retrieved, which bounds how much precision is even meaningful and argues for test-retest stability as the validation metric.",
  "limitations": "Both verified via Crossref metadata; the Miller interpretive caution is standard in the literature but the 1956 text was not re-read this run.",
  "disposition": "epistemic_guardrails"
}
P06-SR-025record 25
{
  "id": "P06-SR-025",
  "evidence_class": "observed",
  "source": "https://github.com/sublee/trueskill (LICENSE body read via gh api)",
  "observed": "2026-08-27",
  "claim": "Licence trap of exactly the class this project was warned about. The repo (803 stars) shows a BSD-3-Clause badge, but the LICENSE body reads: 'Microsoft permits only Xbox Live games or non-commercial projects to use TrueSkill(TM). If your project is commercial, you should find another rating system.' GitHub's API reports only NOASSERTION. Conversely huawei-noah/HEBO shows a bare NONE badge but carries a real MIT licence one directory down at HEBO/LICENSE.",
  "limitations": "Badges were wrong in BOTH directions. Licences must be read from the body, and CRAN licences from DESCRIPTION rather than GitHub mirrors.",
  "disposition": "DISQUALIFIES sublee/trueskill for paid client work; openskill.py (MIT) is the substitute"
}
P06-SR-026record 26
{
  "id": "P06-SR-026",
  "evidence_class": "observed",
  "source": "https://engineering.atspotify.com/2026/1/why-we-use-separate-tech-stacks-for-personalization-and-experimentation",
  "observed": "2026-08-27",
  "claim": "Methodological gate. With a contextual bandit there is no single 'best button', only a 'best system', so it still requires an experiment against the static original — 'the bandit is a feature you've built, not an experimental method'. Optimizely's docs corroborate the mechanism: no statistical significance is calculated for contextual-bandit optimisations.",
  "limitations": "Engineering blog plus vendor docs read via search summary.",
  "disposition": "makes_Gate_6_non_negotiable"
}