# A/B Test Design: AI Music Filter **Status: PRE-REGISTERED before the simulation was run.** Decision rules are written down here first so results cannot be rationalized afterwards. **All data in this experiment is SYNTHETIC.** Simulated users, simulated outcomes. A real version requires platform data; this demonstrates the method and the decision discipline. --- ## Hypothesis Giving listeners a control to hide fully AI-generated music increases 30-day retention, because catalogue trust is a retention driver for the segment that resents AI content (Deezer/Ipsos: 45% of streaming users say they want this filter). ## Variants | Arm | Description | |---|---| | **Control** | Status quo, Sept 2026: AI Personas excluded from recommendations by default. No user control. | | **B** | Filter toggle available in settings, **off by default**. Filters fully AI-generated content from search, radio, and recommendations. | | **C** | Same toggle, **on by default**. Users can turn it off. | The toggle keys on "AI Generated" only, never "AI Assisted", in both arms. ## Unit and assignment User-level randomization, one-third per arm, assignment fixed for the full window. Simulated population: 30,000 users. ## Primary metric and decision rule **30-day retention**, intent-to-treat (measured on everyone in the arm, whether or not they touched the toggle). - Ship threshold, pre-registered in the decision memo: **+2.0 percentage points** over control, significant at p < 0.05 (two-proportion z-test). - Below threshold but significant: do not ship as-is; investigate targeting. - Not significant: the trust-retention link is unproven at this scale. ## Secondary metrics - **Adoption**: share of arm B users who enable the toggle. Pre-registered in the memo: below 5% means the demand was loud but narrow, and Reddit was over-read. - **Opt-out**: share of arm C users who turn the filter off. High opt-out means default-on is hostile, not helpful. ## Guardrails - Streams per user per week must not drop more than 2% in any arm (filtering content risks thinning sessions). - Complaint volume about "missing music" (proxied in simulation; in production, support contacts per 1,000 users). ## Power Baseline retention 60%. To detect +2.0pp at alpha 0.05, power 0.80, a two-proportion test needs roughly **9,400 users per arm**. 10,000 per arm clears it with margin. This is why the simulation uses 30,000 users, not 600: underpowered tests produce confident noise. ## The synthetic effect model (the honest part) The simulation does not assume the feature works uniformly. Users are drawn from three segments, and every parameter below is visible and tunable in `config/experiment.yaml`: | Segment | Share | Effect of an active filter on their retention | |---|---|---| | AI-averse | 30% | +6pp (the trust effect the hypothesis claims) | | Indifferent | 60% | 0 | | AI-positive | 10% | −2pp (they wanted that content) | Adoption behaviour, also assumed and visible: in arm B, 25% of AI-averse users find and enable the toggle, 2% of indifferent, 0% of AI-positive. In arm C, 95% of AI-averse keep it on, 90% of indifferent never touch it, 70% of AI-positive turn it off. These numbers are guesses. That is the point: the simulation shows what the *readout discipline* looks like given assumptions you can argue with, one number at a time. ## What feeds back (stage 8) Whatever the outcome, the result is written back against cluster C003 in the registry, so the opportunity that triggered this experiment carries its experimental verdict in its history.