# PRD: AI Music Filter | | | |---|---| | **Author** | Rashi Raut | | **Date** | August 18, 2026 | | **Status** | Complete. Written last, on purpose (see Process note) | | **Artifacts** | Discovery pipeline (`run.py`), decision memo (Aug 5), landscape brief (Aug 18), prototype (`prototype.html`), A/B design + simulation (`experiment.py`) | **Process note.** This PRD was written after the prototype and the simulated experiment, not before. The build order was: collect feedback, cluster it, pick a problem and defend it, prototype, test, then spec. Prototyping first surfaced decisions a blind spec would have missed (the reveal interaction, the default question); writing the spec last means it records what was learned, not what was guessed. To keep that honest, the decision memo and experiment design were dated and locked before the prototype and simulation existed, so the reasoning cannot have been retrofitted. --- ## 1. TL;DR Give listeners a setting that hides fully AI-generated music from their discovery surfaces. Off by default at launch, with the experiment plan built to test whether the default should flip. Filters, never blocks. Keys on the industry "AI Generated" tag only; AI-assisted music by human artists is labeled, never hidden. ## 2. The problem, and how it was found **Found by pipeline, not intuition.** A discovery agent ingesting 628 pieces of public Spotify feedback (App Store, Google Play, r/truespotify) clustered complaints by meaning and tracked them across runs. AI-content backlash emerged as the only genuinely new problem in the set: 5 of the top 25 r/truespotify posts in a month (~1,070 combined upvotes), the cleanest complaint-to-praise ratio of any cluster, and users building their own AI filters, which is unmet demand made visible. **The problem statement.** Listeners are losing trust in the catalogue. AI-generated tracks now appear in search, radio, and playlists without the listener having any say. On Deezer, 44% of daily uploads are AI-generated; listeners cannot reliably tell (97% fail unaided, per Deezer/Ipsos), and 45% of streaming users say they want the ability to filter fully AI-generated music out. **What Spotify shipped instead.** On August 11, Spotify announced AI Persona badges: AI-generated artist *identities* get labeled and excluded from recommendations by default. This is disclosure plus a platform-chosen default. It leaves three gaps this PRD addresses: 1. No listener control anywhere 2. Identity-based, not output-based: a fully AI-generated track under a human-looking artist name is untouched 3. Recommendations only: search and radio are untouched Deezer and Tidal both already ship user-facing filters. This is a parity feature with a design argument, not an invention. ## 3. Goals and non-goals **Goals** - Give listeners control over whether fully AI-generated music appears in their discovery surfaces (search, radio and mixes, recommendations) - Improve 30-day retention among AI-averse listeners, the segment driving the complaint cluster - Preserve trust while filtering: the user always sees that filtering happened and can reverse it in one tap **Non-goals** - Filtering or hiding AI-*assisted* music. Most working musicians touch AI somewhere in production; a coarse filter buries real artists - Blocking playback. Followed AI artists, direct links, and shared playlists still play (this is where we deliberately diverge from Tidal) - Demonetization, detection, or labeling itself. The filter consumes existing labels (AI Persona badges, DDEX/AI Credits disclosures); it does not create them - User-applied AI tags. Crowdsourced labeling is a brigading vector where false positives cost real artists income ## 4. Users From the experiment's segment model, which is an explicit assumption, not a measurement: - **AI-averse (~30%)**: resent AI content in their surfaces, the source of the complaint cluster. The filter is for them - **Indifferent (~60%)**: will mostly never open the setting. Must not be harmed or confused by it - **AI-positive (~10%)**: actively want AI content. Must retain access, which is why filtering beats blocking and why the toggle must be reversible ## 5. The feature **Setting: "Hide fully AI-generated music"** in Content preferences. - Applies to search results, radio and mixes, and recommendations. Per-surface sub-toggles, with recommendations noted as already default-on for AI Personas since September - Keys on: tracks marked AI Generated under the RIAA/IFPI two-tier standard, plus tracks by badged AI Persona artists - Never touches: AI Assisted tracks, which show their existing AI Credits label instead - **Transparency chip**: when results are hidden, the surface says "N AI-generated tracks hidden" with a one-tap Show that reveals them dimmed. A filter that silently disappears music creates a second trust problem while solving the first - Exceptions that always play: artists the user follows, direct links, tracks already in the user's library and playlists Interactive prototype: `prototype.html` (toggle, filtered search, reveal state, with design notes overlay). ## 6. Decision log The core of this document. Each row is a real fork; the losing option is recorded with the reason it lost. | # | Decision | Chosen | Rejected | Why | |---|---|---|---|---| | 1 | Which problem to work on | AI-content trust (C003) | Free-tier friction (C001), the scoring model's pick with 3x the mentions | Free-tier friction is the conversion mechanism working as designed, not a defect. The model cannot distinguish a bug from a business model. Recorded as a known limitation of the tool | | 2 | Control vs disclosure | User toggle | Labels only (Spotify's shipped approach) | Labels inform, they don't protect the listener's time. 45% of users say they want filtering, and two competitors already ship it | | 3 | Filter vs block | Hide from discovery | Tidal's unplayable-when-off approach | Blocking breaks shared playlists and external links, punishing users who never opted into the fight. Filtering achieves the goal with less collateral damage | | 4 | What gets filtered | Fully AI-generated only | Everything with any AI disclosure | AI-assisted is most working musicians. The two-tier industry standard (AI Generated / AI Assisted) exists precisely to draw this line. Single most important call in the spec | | 5 | Who applies labels | Platform + standards (AI Persona review, DDEX) | User-applied tags | Crowdsourced tagging is brigadable, and a false positive costs a real artist income with no appeals process. User reports may feed a review queue, never a live label | | 6 | Default state | Off at launch, experiment on the flip | On by default from day one | See experiment results below. On-by-default wins the metric but spends goodwill with the AI-positive segment; that trade should be made with real data, not simulated data | | 7 | Silent vs transparent filtering | Transparency chip with reveal | Silently cleaner results | A user who learns music was hidden from them without notice stops trusting every list the product shows them | | 8 | Spec timing | PRD written last | Spec-first waterfall | The prototype surfaced the reveal interaction; the experiment reframed the default as the real decision. Both would have been guesses in a spec-first order. Pre-registered memo and experiment design keep it honest | ## 7. Experiment: design and results Full design in `2026-08-18-ab-test-design.md` (pre-registered), simulation in `experiment.py`. **All data synthetic**: 30,000 simulated users, three arms, segment and adoption assumptions visible in `config/experiment.yaml`. The simulation demonstrates decision discipline; a real rollout reruns this on platform data. Pre-registered rules: ship at +2.0pp 30-day retention (ITT), p < 0.05. Arm B adoption below 5% means demand was narrow. Guardrail: streams per user must not drop more than 2%. | Arm | Retention lift | p | Verdict against pre-registered bar | |---|---|---|---| | B: toggle, off by default | +1.07pp | 0.12 | Not significant. Fails | | C: toggle, on by default | +2.34pp | 0.0007 | Clears the bar | Arm B adoption: 8.7%, clearing the 5% floor, so demand is real, not just loud. Guardrails held in all arms. **The finding: the default, not the feature, is the product decision.** The per-user effect is identical in both arms; what differs is how many users ever experience it. Off-by-default satisfies the vocal minority without moving the platform metric. On-by-default moves the metric and takes goodwill from the segment that wanted AI content. ## 8. Rollout 1. **Phase 1, launch**: toggle ships off by default (arm B), because the simulated case for default-on rests on assumed segment sizes. Instrument adoption, retention, streams, and support contacts 2. **Phase 2, targeted prompting**: offer the filter contextually to users whose behaviour suggests AI-aversion (repeated skips of AI-tagged tracks). Hypothesis from the experiment: this captures most of arm C's lift without flipping anyone's default 3. **Phase 3, default experiment on real data**: rerun B vs C as a true A/B with the pre-registered rules. Flip the default only if the real data clears the bar 4. **Feedback loop**: the discovery pipeline keeps running weekly. Cluster C003's trend line is the production follow-up: if the filter works, its complaint volume should shrink. The experiment verdict is recorded against C003 in the cluster registry ## 9. Risks - **Label coverage is the ceiling.** The filter only hides what's tagged. AI Persona review currently starts at audience thresholds; untagged AI tracks pass through. Mitigation: DDEX disclosures widen coverage over time; communicate the filter as "reduces," never "eliminates" - **Catalogue economics.** AI tracks are cheap catalogue; filtering reduces margin per stream. Counterweights: AI content is only 1-3% of streams, and 85% of those are fraudulent on Deezer's numbers, so the legitimate revenue at risk is small - **Adversarial labeling.** Demonetization and filtering create incentive to evade tagging. Detection is an arms race owned by the labeling systems, not this feature, but this feature inherits their misses - **Segment assumptions.** The 30/60/10 split and effect sizes are guesses. Phase 1 instrumentation exists to replace them with measurements before the default decision ## 10. Open questions Each carries how it would get resolved, because an open question without a resolution path is just a worry. - **Linear surfaces can't hide, they substitute.** Search can drop a row; a radio queue has to play *something* in slot 4. Does the transparency chip even belong in a queue, or does substitution happen silently there? Resolve by prototype: the next screen to build after launch scope is locked. - **Whose filter wins in shared contexts?** Blend and collaborative playlists mix users with different settings. Strictest-member-wins protects the averse user but lets one person's setting reshape a shared space; per-viewer rendering means the "same" playlist differs by viewer. Resolve by prototype plus a small user test: this is a felt-experience question, not a data question. - **Does the filter extend to podcasts and audiobooks** as AI generation spreads there? Resolve by re-running the discovery pipeline scoped to podcast feedback: build only if the complaint cluster exists there too. - **When listener reporting of suspected AI Personas ships**, does report volume feed this filter's coverage, and with what appeals safeguard? Resolve by policy review with the labeling team: reports as review-queue signal, never as a live label (decision #5). - **Can targeted prompting replace the default flip entirely?** The experiment suggests contextual offers to AI-skipping users could capture most of arm C's lift without changing anyone's default. Resolve by phase 2 experiment: prompt-triggered adoption vs. arm B's organic 8.7%. --- *Part of an end-to-end build: opportunity agent → decision memo → landscape brief → prototype → simulated A/B → this PRD. Repo README tells the full story.*