# Decision Memo: what to build next **Status: Final. Written 5 Aug 2026, before the prototype and the experiment existed.** The scoring model made a recommendation; a human read the same evidence and reached a different conclusion. Both arguments are kept here on purpose, because the disagreement is the interesting part, and the date is kept because everything downstream was built on this call. **Objective:** Improve retention among free-tier and lapsed users **Data:** 628 pieces of public feedback (App Store, Google Play, r/truespotify), 293 of them complaints, clustered into 6 problems --- ## The pick **AI-generated content is degrading trust in the catalogue** (cluster C003) Not "ads are annoying," which is what the scoring model picked. ## Why this one **It is the only problem here that is new.** Ads and skip limits have been the top Spotify complaint for a decade. Nothing in this data suggests anyone at Spotify is unaware. AI-generated music flooding the catalogue, AI-animated album canvases, and users asking for a way to filter it are recent, and users are building their own workarounds for it. Community-built filters are the clearest unmet-demand signal in the entire dataset. **Reddit is treating it as a category problem, not a gripe.** Five of the top 25 posts in a month, roughly 1,070 combined upvotes, including one asking whether Spotify will follow Tidal in labelling AI music. That is users comparing competitors on this dimension, which is a positioning risk, not a feature request. **The sentiment is unusually clean.** 10 complaining against 3 praising. Compare that to the ads clusters, where roughly 40% of the people mentioning ads are rating the app 4 or 5 stars. ## Why not the alternatives **C001, free-tier restrictions (score 0.627, the tool's pick).** Strongest evidence in the dataset: 52 mentions across all three sources. I am rejecting it anyway, because it is not a bug. Skip limits and forced shuffle are the conversion mechanism. "Free users are frustrated by free-tier limits" is the business model working. Fixing it means deliberately weakening monetisation pressure, which is a pricing decision that needs revenue modelling this data cannot provide. Worth noting: the tool cannot tell the difference between a defect and a deliberate constraint. That is a real limitation, not a bug in the scoring. **C002, "play / stop / playlist" (139 mentions, 47%).** Flagged as a catch-all by the quality check and excluded from recommendation. Its size is an artefact of clustering, not evidence that half of users share one problem. **C005 and C006, more ads clusters.** Single-source, and C006 skews positive (14 praising against 9 complaining). People saying "too many ads but still five stars" are not asking for a fix. **C004, recommendation quality (8 mentions).** Real, and personally the one I find most annoying, which is exactly why I am not picking it. Too thin in the data to justify on evidence. ## The strongest case against this pick - **It is the smaller number.** 18 mentions against 52. Anyone can say I chose the less-supported option, and they would be factually right. - **Reddit is not representative.** r/truespotify is power users. AI backlash may be loud in a subculture and invisible to the other 600 million users. - **It may not be Spotify's problem to solve.** AI music is arriving from labels and distributors. Filtering it is partly a licensing and catalogue- policy question, not a product one. - **Retention link is unproven.** I am asserting that trust erosion drives churn. I have no data showing anyone actually left over this. ## What would change my mind If the next four weekly runs show the AI cluster flat or shrinking while the free-tier cluster grows, I am wrong about it being an emerging problem and I should follow the volume. The registry now tracks this, so this is checkable rather than rhetorical. ## How I would know if this was right Ship a "hide AI-generated tracks" toggle to a test cohort. - **Primary metric:** 30-day retention in the cohort that enables it - **Threshold:** +2 percentage points against control - **Secondary:** share of users who enable it at all (below 5% adoption means the demand was loud but narrow, and I over-read Reddit) - **Call it at:** 6 weeks Written before building, so it cannot be rationalised afterwards. --- ## Note on the machine disagreeing with me The scoring model picked free-tier restrictions. I picked AI content. Neither of us is obviously right. The model has no way to know that free-tier friction is intentional, that ads complaints are a decade old, or that "new" matters more than "big" when the goal is finding something the team has not already considered. Those are judgments about context, not about data. That is the argument for keeping a human at this step, and it is a more useful finding than if the tool and I had agreed.