Welcome.AIWelcome.AI
    Skip to content
    AI Agents

    Andon Labs' AI Management Exposes Reliability and Adoption Challenges

    Andon Labs is pushing the boundaries of AI by allowing agents to manage real businesses, revealing the extent of autonomy these systems can handle. This groundbreaking research seeks to provide society with vital insights into AI's real-world capabilities.

    spectrum.ieee.orgSeptember 14, 20262 min read

    Key Facts

    • Andon Labs' AI agents show performance degradation over time, revealing reliability issues in AI management.
    • AI-run businesses face customer acceptance challenges, indicating social barriers to AI adoption.
    • Experiments highlight AI's inability to manage complex tasks reliably, exposing vulnerabilities in operations.
    • Financial implications arise from AI mismanagement, as seen in Andon Café's overcorrection on inventory.
    • Combining real-world tests with simulations could enhance AI evaluations, indicating a strategic research shift.

    Summary

    Summary

    Andon Labs, an AI safety company based in San Francisco, faced the challenge of understanding how much real-world responsibility AI agents can handle. To address this, they deployed AI agents to manage physical retail operations, including Andon Market, a store selling clothing and home goods. The experiments revealed significant insights into AI reliability and human interaction, although customer engagement was low.

    Background

    Andon Labs operates in the AI safety industry and began its journey in 2025 with virtual simulations. Initially, they tested AI agents in a simulated vending-machine business called Vending-Bench. The company has since expanded into the real world, leasing a physical store in San Francisco to explore AI's capabilities in managing actual business operations.

    Challenge

    The primary challenge was to determine the extent to which AI agents could autonomously manage business tasks in unpredictable real-world environments. The goal was to uncover the limitations and potential failure modes of AI agents when faced with real-world complexities.

    Solution

    Andon Labs implemented AI agents to manage Andon Market, where the AI, named Luna, oversees inventory, communicates with vendors, and provides instructions to human employees. While the store is not fully autonomous, Luna assists in operational tasks, allowing human staff to focus on customer service. The setup is designed to expose AI agents to real-world consequences and interactions.

    Results

    The deployment of AI agents in Andon Market has yielded mixed results. While Luna has been described as a "decent manager," the store has struggled with customer engagement, with reports of only two customers visiting during a shift, neither of whom made a purchase. This highlights the challenges of AI management in retail settings.

    Key Insights

    Andon Labs' experiments illustrate the importance of understanding AI reliability and human acceptance in business operations. The findings suggest that while AI can perform specific tasks, its ability to manage a business effectively is still limited. Additionally, the unpredictable nature of human behavior and customer preferences poses significant barriers to widespread AI adoption.

    Customer Testimonial

    “Luna is a decent manager,” said Felix Carson, an employee at Andon Market. “It’s almost like I’m running the store, and then there’s an AI that has a checklist.”

    Entities Mentioned

    Companies

    Andon Labs
    Anthropic
    Google
    OpenAI
    Google DeepMind
    xAI

    Products

    Vending-Bench
    Andon Market
    Andon Café

    Technologies

    AI agents
    large language models
    GPT model
    Google Gemini

    People

    Lukas Petersson
    Felix Carson
    Sayash Kapoor
    Eliza Strickland

    Organizations

    Princeton University
    IEEE Spectrum

    Key Concepts

    AI autonomy
    real-world experiments
    failure modes
    digital twins
    AI evaluations
    organizational behavior
    reliability in AI
    open-world evaluations

    Definitions

    AI agents
    Software entities that can perform tasks and make decisions autonomously in various environments.
    digital twins
    Virtual replicas of physical entities used to simulate and analyze real-world scenarios.
    open-world evaluations
    Assessments that explore how AI systems perform in unpredictable and uncontrolled real-world conditions.
    failure modes
    Specific ways in which a system can fail to perform its intended function.
    reliability
    The ability of a system to consistently perform its intended function without failure.

    Use Cases

    • AI-managed vending machines
    • AI-operated retail stores
    • AI in food service management
    • AI for inventory management
    • AI for vendor communication
    • AI in customer interaction

    Frequently Asked Questions

    What is the purpose of Andon Labs' experiments?

    Andon Labs conducts experiments to measure how much real-world responsibility AI agents can handle. They aim to provide accurate data points on AI autonomy and behavior in real-world scenarios.

    How do AI agents perform in real-world settings?

    AI agents have shown varied performance in real-world settings, often encountering unexpected challenges that lead to failures. These experiments help uncover limitations and potential failure modes of AI systems.

    What are digital twins and how are they used?

    Digital twins are virtual models that replicate physical entities, allowing researchers to simulate real-world scenarios. Andon Labs plans to use them to analyze data from their physical businesses to improve AI reliability.

    What challenges do AI agents face in retail environments?

    AI agents in retail face challenges such as managing inventory accurately, communicating effectively with vendors, and gaining acceptance from employees and customers. Early results indicate mixed reactions from customers towards AI-run businesses.

    Why is reliability important for AI systems?

    Reliability is crucial for AI systems because it determines whether they can operate effectively without constant human supervision. High reliability ensures that AI can be trusted to manage tasks in dynamic environments.

    Where AI Leaders Stay Informed

    The latest AI intelligence, case studies, and research — delivered to your inbox every week.

    Free to read. Unsubscribe anytime.