24/7 AI Salesperson: What Avatarin Learned from 30,000 Users

Find out what avatarin learned from 30,000 users with a 24/7 AI sales assistant, and how e-commerce integrates voice, RAG, products, human escalation, and KPIs.

Short answer: A 24/7 AI sales representative is useful when it isn't limited to pre-written responses. It must listen to the customer’s actual needs, retrieve accurate information from the catalog, and ask the right clarifying questions before recommending a product. This is what avatarin and Yamada Holdings tested with the Kurashi-Marugoto AI Agent.

Approximately 30,000 users participated in the two-week public trial at the Yamada Denki online store, and 92% of the responses to the post-use survey were positive. This finding is a strong indication of the experience’s acceptance, not proof of increased sales: the official presentation does not disclose the conversion rate, average order value, or a comparison with human salespeople.

The most useful lesson for an e-commerce business is about architecture. A natural voice requires product grounding, commercial knowledge requires conversation design, and every proposal requires boundaries, up-to-date data, and a scaling path. This is the same foundation required by a Private chatbot with RAG and curated corporate knowledge, but applied to guided product selection.

What to keep: The AI salesperson is neither the catalog nor the ultimate decision-maker for the purchase. It serves as the conversational layer that connects the customer’s needs with verified product data, states its limitations, and defines a path for escalation to a human when the decision calls for it.

Contents

What avatarin and Yamada Denki Tested

Avatarin is an AI customer service company that emerged from ANA Holdings. In collaboration with Yamada Holdings, it developed the Kurashi-Marugoto AI Agent, a multilingual shopping agent designed for natural voice and text conversations. The goal was to bring part of the experience provided by home appliance salespeople into the online environment and make it available beyond regular store hours.

The need is easy to identify. A customer isn’t always looking for a «500-liter refrigerator.» They might say they have a family of four, a small kitchen, a limited budget, and uncertainty about capacity. The right recommendation requires combining these details with actual dimensions, features, availability, and store policies.

The public campaign served as a laboratory experiment in product selection. Avatarin’s official announcement clarified that it was not linked to Yamada’s member app or related customer data. This is important: the future vision of a unified experience across the web, mobile, and in-store should not be presented as if it were already a fully realized integration.

From Keyword Matching to Guided Counseling

A conventional chatbot typically waits to recognize an intent and then returns a pre-prepared response. Avatarin’s shopping agent is designed to understand context: uncertainty, changing preferences, practical constraints, and the details that turn a general question into a useful suggestion.

The difference isn't just a more natural copy. It's a different service flow. The system must recognize which details are missing, ask questions without feeling like a tedious form, and return to a safe scope when the conversation veers off track. When choosing a platform and features, it helps to distinguish guided shopping from general Conversational support solutions for e-commerce.

Four Levels of a Useful AI Salesperson Experience

Simple FAQ Bot

It returns pre-written responses regarding hours of operation, shipping, or policies when it recognizes the correct intent.

StaticSmall scope

Guided Discovery

They ask about space, family, intended use, and budget before narrowing down the suitable product options.

Follow-upContext

Product Grounding

It bases every proposal on up-to-date specifications, availability, and policies from the actual catalog.

RAGProvenance

Safe escalation

It indicates uncertainty and hands the context over to a human when data is missing or the decision involves high risk.

Human handoffLimits

Why GPT-Realtime Has Transformed the Voice Experience

Avatarin had already been using OpenAI technologies for speech recognition, query analysis, and employee training. For the Yamada Denki project, it chose GPT-Realtime to combine voice, text, and visual understanding into a low-latency conversational experience.

Latency is not just a technical metric. In voice communication, it affects interruptions, pauses, and the sense that the agent is listening to the user. Slow or rigid turn-taking can turn even correct answers into a poor experience. Conversely, an immediate response is worthless if it is based on an incorrect product or an outdated policy.

That’s why the real-time model is only one part of the system. We need prompts for tone and behavior, retrieval tools, monitoring, cost per conversation, and a fallback when the voice or network fails. Speed must be measured alongside accuracy, task completion, and handoff quality.

The natural flow of the conversation would be compromised if the responses strayed from the actual product catalog. Avatarin used retrieval-augmented generation so that the agent could retrieve relevant product information, while GPT-Realtime kept the conversation flowing smoothly.

The practical principle is clear: the model is not the catalog. The catalog, PIM, ERP, or authorized knowledge base are the sources of truth for features, inventory, prices, compatibility, shipping, and returns. The model is the layer that transforms a natural language query into a controlled retrieval and explains the result.

Every entry must have an owner, an update date, and a clear source. A discontinued product or a policy that has changed should not continue to appear simply because it remained in an old index. In e-commerce, the retrieval pipeline must be updated with the same level of professionalism required by Creating and operating an e-shop, from strategy to launch.

The salesperson's knowledge became conversation design

The information a salesperson needs varies by product category. For a refrigerator, the external dimensions, capacity, door opening, and the family’s needs are important. For a washing machine, the available space, frequency of use, noise level, and wash cycles may be the most important factors.

Avatarin incorporated Yamada Denki’s customer service expertise into the conversation flows and prompts. So the technology did not replace business knowledge; rather, it required that knowledge to be translated into sequences of questions, rules, tone, and clear points at which the agent stops making assumptions.

This is also important for the brand. An agent must sound like the specific business without hiding the fact that it is AI. The tone, persistence, budget management, and the way it requests additional information are all design decisions. The same principle is evident in more mature examples where the AI agents are changing the customer service model, not just the text of a reply.

An agent who asks questions—not just answers them

Avatarin took a proactive approach. The agent asked follow-up questions to uncover the user’s true need and shift the conversation from a simple Q&A to guided product discovery. If the user changed their requirements or temporarily strayed from the topic, the flow could be adjusted, while guardrails helped keep them within the shopping experience.

The key point is query efficiency. The agent should not collect every possible field. It should request only the information that substantially changes the ranking of the options. For a small refrigerator, the width of the niche can eliminate dozens of models before secondary features are even considered.

They also need to explain why they’re asking. A brief phrase such as «the width determines which models will fit without obstructing the door from opening» builds trust and reduces the feeling of being interrogated. The explanation is part of the UX, not a superfluous technical comment.

What Do 30,000 Users and 92% Really Mean?

The public trial lasted two weeks, attracted approximately 30,000 users, and offered multilingual voice and text support around the clock. 92% of the responses in the post-use survey were positive. These are the verified results published by OpenAI for this case.

The published metrics of the avatarin experience

The figures refer specifically to this public campaign and do not serve as a general sales benchmark for all e-commerce businesses.

2weeks

The duration of the public trial period for the Yamada Denki online store.

30.000approximately [number] users

Shoppers who interacted with the agent during the campaign.

92%positive responses

The percentage of positive responses in the post-use survey.

24/7availability

Multilingual support via voice and text outside of store hours.

The 92% is a useful acceptance indicator, but it does not reveal which users completed the survey, how many made a purchase, how much each successful conversation cost, or whether the experience surpassed that of a human salesperson. It should not be interpreted as a claim regarding 92% accuracy, 92% conversion rates, or 92% satisfaction across the entire population.

The Correct Way to Read 92%

Measure acceptance, accuracy, and business impact separately.

The survey indicates whether participants enjoyed the experience. Grounded-answer tests show whether the suggestions were correct. Conversion, assisted revenue, cost per resolution, and quality handoffs indicate whether the system generates sustainable value.

The conversation became a source of customer insight

For Avatarin, the key insight wasn’t just the scale. The conversations revealed what interests buyers, why they hesitate, and what information can help them make a decision. In conventional online shopping, these nuances often don’t show up in click data.

Each conversation ended with a brief voice survey. The official report states that customers appreciated the ability to repeat questions without feeling anxious and to discuss budgets or uncertainties without sales pressure. These are qualitative findings specific to this case and do not represent behavior that should automatically be generalized to every audience.

Analyzing such conversations requires clear objectives. The frequency of a question is one metric, while a person’s purchase intent is another. The team should anonymize data wherever possible, restrict access, and link these insights to responsible changes in product descriptions, filters, content, and sales training.

Limits, Data Security, and Human Scaling

Avatarin’s announcement warned that AI responses were provided for reference only and might not be accurate, complete, or up to date. This warning is not sufficient on its own, but it points in the right direction: a natural-sounding voice should not mask the system’s uncertainty.

A production rollout requires rules governing when the agent refuses to make a recommendation, when it requests a second confirmation, and when it initiates a handoff. Examples include the inability to verify availability, incompatibility between sources, a request for a binding guarantee, or a situation where the customer requests handling of payment and personal data.

Conversations, transcripts, and survey responses must have a specific purpose, adhere to a minimum data collection principle, follow a retention policy, and have access controls in place. The practical link between support, marketing, and data protection is also analyzed in the guide for CRM compliance without slowing down marketing. In the area of AI, NIST recommends continuous risk management using the "govern," "map," "measure," and "manage" functions throughout the entire lifecycle.

Six Steps to Becoming an AI Salesperson in E-commerce

The Avatarin case study is not a ready-made technical blueprint to be copied. However, it offers a practical set of priorities: start with a category where advice actually influences the choice, organize product sources, turn salespeople’s knowledge into questions, and measure the experience along with accuracy.

From a use case to a controlled shopping experience

  1. Step 1Choose a category that presents a real dilemma

    Start with products where space, intended use, compatibility, or budget significantly affect the right choice, rather than looking at the entire catalog.

  2. Step 2Map Out the Questions Asked by Top Salespeople

    Note which information eliminates options, which sequence of questions reduces friction, and which wording aligns with the company’s voice.

  3. Step 3Link only approved product data

    Here are reliable sources of information on features, availability, price, delivery, and policies, including the owner and update frequency for each feed.

  4. Step 4Design boundaries and human handoffs

    Specify what the agent is not allowed to assume, what uncertainties it encounters, and when it forwards the summary of the conversation to a human.

  5. Step 5Evaluate me using real-life shopping scenarios

    Test for changes in preferences, missing information, conflicting products, background noise, and outdated recordings before making the stream available to the entire audience.

  6. Step 6Evaluate experience, accuracy, and cost-effectiveness separately

    Track positive feedback, grounded-answer accuracy, task completion, conversion, assisted revenue, cost per resolution, and handoff quality without combining them into a single score.

The rollout should begin with a limited scope and a clear baseline. The agent can initially suggest a few categories, display the sources it used, and request human confirmation before taking any critical action. Autonomy is expanded only when logs and outcomes demonstrate consistent quality.

The key lesson from the 30,000 users isn’t that every e-shop needs a voice avatar. It’s that the conversational experience creates value when it brings together three things: natural communication, real product data, and business knowledge translated into proactive questions. Without these, an AI salesperson is just a fun demo.

Business automation and AI from TWO DOTS

Design an AI salesperson who knows the product catalog and its limitations.

TWO DOTS maps out products, data sources, conversation flows, integrations, human handoffs, and KPIs so that the conversational agent can drive actual online sales without hiding the uncertainty.

Frequently Asked Questions (FAQs)

What is the Kurashi-Marugoto AI Agent?;

It is the multilingual shopping agent developed by avatarin and Yamada Holdings for guided product selection through natural conversation, voice, and text.

What technology did Avatarin use?;

The agent was based on GPT-Realtime for low-latency multimodal conversation and on retrieval-augmented generation to ground responses in relevant product information.

What did the 30,000 users and the 92% reveal?;

Approximately 30,000 shoppers used the agent during a two-week public trial, and 92% of the responses to the post-use survey were positive.

Does the 92% show that sales have increased?;

No. The official report does not publish conversion rates, revenue, average order value, or comparisons with human salespeople. The 92% is a measure of positive survey responses.

Why did the agent ask follow-up questions?;

To identify the customer’s actual constraints—such as available space, family needs, or budget—and to transform a general request into a guided product selection.

How was the accuracy of the information maintained?;

Avatarin used retrieval-augmented generation so that the responses would be based on relevant product information, along with guardrails that kept the conversation within the scope of shopping.

Did it use Yamada's public customer data?;

Avatarin's announcement clarified that the lab experience was not linked to Yamada's member app or related customer data.

What is the first step toward a smaller e-shop?;

Select a category where advice is changing the market, organize reliable product data, and document the questions and handoffs of experienced salespeople.

Newsletter

Enter your email address below to subscribe to our newsletter