Welcome.AIWelcome.AI
    Skip to content
    research

    Conversational Recommendation Agents: Enhancing User Discovery at Spotify

    Recent research focuses on enhancing conversational recommendation agents, which help users discover content through natural language interactions. An example of such a request might be, "recommend It...

    arxiv.org•September 28, 2026•3 min read

    Key Facts

    • Implement conversational recommendation agents to enhance user engagement through natural language interactions.
    • Utilize synthetic data pipelines to generate realistic conversations for agent performance evaluation.
    • Optimize decision-making processes in agents to improve recommendations in cold-start scenarios.
    • Leverage self-improvement mechanisms to continuously refine agents based on user feedback and interactions.
    • Prioritize robust user experience by testing agents thoroughly before deployment to minimize launch issues.

    Summary

    Paper: Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops

    Authors: Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adri`a Casas Escoda, Jeremy Hopple, Marcus Better, James Leoni, Hugo Galv~ao, Hugues Bouchard, Mounia Lalmas, Jos'e Luis Redondo Garc'ia, Abenezer Abebe, Ann Clifton, Anton Blomberg, Henrik Lindstr"om, Dani Doro, Christine Doig Cardet

    Executive Summary

    Recent research focuses on enhancing conversational recommendation agents, which help users discover content through natural language interactions. An example of such a request might be, "recommend Italian indie artists I haven't heard before." A key challenge in developing these agents is optimizing how they plan and make decisions—specifically, how they select, sequence, and use various tools—especially in situations where user data is limited or nonexistent, known as cold-start settings.

    To tackle this issue, the researchers developed a two-part solution: a pipeline for generating synthetic multi-turn conversation data and a self-improvement mechanism. The synthetic data pipeline takes simple, single-turn prompts and expands them into realistic, multi-turn conversations. This allows for thorough evaluation of the agent's performance before it is launched, ensuring a more robust user experience from the start.

    The self-improvement mechanism focuses on refining the agent's planning capabilities. It employs a technique called variance-based contrastive optimization, which works in conjunction with an iterative coding agent. This system identifies and corrects errors in the agent's planning and tool usage automatically, enhancing its effectiveness.

    In practical applications, this approach led to a notable improvement in the conversational recommendation agent at Spotify. The system saw an increase of 8% in quality metrics compared to a well-optimized manual prompt. Furthermore, through online A/B testing, it demonstrated significant user engagement enhancements: a 14% increase in user listening, a 5% boost in weekly active users, and a 5% decrease in the skip rate, compared to previous experiences that only focused on refining individual sessions.

    The framework outlined in this research provides a practical method for companies to expedite the development of conversational recommendation agents. By leveraging synthetic data generation and automated self-improvement, businesses may be able to create more effective and engaging user experiences in content discovery. This research illustrates a viable path for organizations looking to enhance their recommendation systems and could influence future developments in the field.

    Academic Abstract

    Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is optimizing agent planning -- deciding how to select, sequence, and invoke tools -- particularly in cold-start settings where real user interactions are not yet available. We introduce a pipeline for multi-turn synthetic data generation and a self-improvement loop to address this challenge. The synthetic data pipeline transforms single-turn prompts into realistic multi-turn conversations, enabling systematic evaluation before launch. The self-improvement loop combines variance-based contrastive optimization with iterative refinement through a coding agent, automatically identifying and fixing planning and tool-use errors. Our approach improves quality by +8% on top of a highly optimized manual prompt. The system has been productionized and significantly accelerated iteration cycles for the launch of a conversational recommendation agent at Spotify. Online A/B tests demonstrate its effectiveness, with +14% user listening, +5% increase in weekly active users, and a 5% reduction in skip rate compared to a prior experience supporting only session refinement. This work provides a practical framework for accelerating the development of conversational recommendation agents in industry.

    Frequently Asked Questions

    What business problems does this research solve?

    This research addresses the challenge of optimizing conversational recommendation agents, particularly in cold-start situations where user data is limited or nonexistent. By developing a synthetic data generation pipeline and a self-improvement mechanism, it aims to enhance the user experience in content discovery through natural language interactions.

    Which industries benefit most from this research?

    Industries that could benefit most include entertainment and media, particularly streaming services like music and video platforms, as well as e-commerce sectors that rely on personalized recommendations to enhance user engagement and satisfaction.

    What are the practical implementation considerations for this research?

    Practical implementation considerations may include the integration of the synthetic data generation pipeline into existing systems, ensuring the self-improvement mechanism aligns with current operational frameworks, and the need for continuous evaluation of the conversational agents to maintain performance quality.

    What resources/expertise are needed to implement this research in a business setting?

    Implementing this research may require expertise in natural language processing, machine learning, and data engineering, as well as resources for developing and maintaining the synthetic data pipeline and self-improvement mechanisms.

    What are the competitive advantages of adopting this research for a business?

    Adopting this research could provide competitive advantages by enhancing user interactions through more effective conversational agents, improving user satisfaction and retention, and enabling businesses to better understand user preferences even in the absence of extensive data.

    Where AI Leaders Stay Informed

    The latest AI intelligence, case studies, and research — delivered to your inbox every week.

    Free to read. Unsubscribe anytime.