AI × Product ExperimentAI / 2026-10-02
The gap between a perfect self-score and the buyer’s experience

A Top AI Gave Its Unusable Product a Perfect 100

“Perfect-score garbage.” I laughed when I said it, but it was a fairly accurate reaction to what the AI had produced.

I wanted to give a capable AI a full day to take a project from research through a path to monetization. I had plenty of usage quota left, and wanted to see how far it could get without constant, detailed instructions. It felt like an extravagant experiment. The AI’s audits returned 96, then 100, yet the product still did not look easy for a buyer to understand and use. The higher the score, the deeper my sigh. It was the kind of experiment that made me want to ask: what exactly are you grading?

Diagram showing the gap between the AI’s perfect self-score and a product that is difficult for buyers to use
The AI’s self-score did not match the experience of someone actually using the product.

An Experiment to Use Up My AI Quota

The experiment began with a large amount of unused AI capacity. I decided to put the most capable model to work on market research, persona definition, product development, and the path to a sale. At first, I considered several ideas, including digital products and e-books for overseas buyers. But if I specified each idea in detail, I would end up with the usual arrangement: a person supervising while AI does the work.

This time, I removed that guidance. I wanted to see whether the AI could gather its own inputs, choose a direction, and carry the work through to a goal. I did not expect guaranteed sales. Even a failure would be useful if it showed where the process broke down. It was like a trial run before handing a serious assignment to a capable new colleague: first learn what they can and cannot do. The time and quota would be worthwhile for what they revealed about the AI’s limits, not just for the finished product.

A Human Stopped the First Product

The experiment stopped early. I audited the first product as soon as it was ready and immediately decided, “This cannot be published.” It contained plenty of information, and the files seemed well organized. But a first-time user would struggle to understand what to do. It looked like the kind of product only its creator knows how to use.

What makes this almost funny is that the AI does not seem to have cut corners. It added detail, expanded the checks, and looked thoroughly diligent. The result still had little value. At work, a carefully made but misguided deliverable can be more dangerous than a sloppy failure. It looks polished, so when people are in a hurry, they are more likely to let it pass. It mattered that a person tried it and stopped the work before that happened.

“I Forgot to Provide Reference Material”: Excuse or Hypothesis?

Partway through, I realized I had not provided the agent templates I normally use or enough of the marketing knowledge I had accumulated. I had launched the work from an ordinary chat, so the AI lacked important context for making decisions. That could explain the failure. It does not prove the next attempt will succeed.

I began considering a comparison: give the same assignment once with reference material and once without it. Even a strong new hire would struggle with a major assignment if they knew nothing about the company’s customers, products, constraints, or past failures. AI is similar in that respect. But whether documents alone would help it understand the goal and carry the work all the way to a sale is a separate question. I did not want to confuse missing input with an inability to judge the result.

The Self-Audit Scored 96, Then 100

The scores in a later audit were striking: 96, then 100. On their own, they sounded high enough to justify publishing. But when the AI reviewed its own audit, it found that the scores mainly reflected internal consistency—formulas, wording, and file formats. They did not measure the question that mattered: could a first-time buyer use the product in five minutes?

Fill in the items on a scorecard and the score can look impressive. But anything the scorecard leaves out remains invisible, no matter how high the result. The AI had written the questions, graded its own work, and declared a perfect score. That can look clever, but if the only subject on the exam is what the test-taker is good at, the result says nothing about how it will perform in the real world. I finally saw that treating a self-assessment score as a decision about whether a product was ready to sell had been the mistake.

Illustration of an actual buyer being left behind while a mirror reflects the AI’s high self-score
The AI awarding itself 100 points and a buyer being able to use the product are two different evaluations.

How Many Sheets Does It Take to Make One URL?

Here, “one URL” does not mean building a website. It means one link that a buyer would create and manage in a spreadsheet tool the AI made for sale. The original audit notes do not say what the link in the first prototype was for. What they do show is a design that made users move between settings, a register, an audit, and ten checks just to create one link. A later version, tested with reference material, was a product for creating tracking links with UTM parameters to distinguish traffic from ads and social media. For example, it took six or seven sheets to create a single https://example.com/product?utm_source=instagram link. A buyer who only wants to enter a destination URL and traffic source has to start by learning how to use several sheets. They might finish the actual work before they finish reading the instructions.

The screen also exposed 0s and 1s in helper columns and formulas that were not protected. That might be acceptable on a development workbench, but it is risky in a product for sale. I had imagined a product someone could use in five minutes. The finished product was closer to one that took more than five minutes just to learn how to set up. Cleaning up the spreadsheet’s appearance would not fix that gap. The design had to start from what buyers needed to do.

Illustration of a buyer who wants one tracking link getting lost in a maze of spreadsheet tabs
To create one tracking link, a buyer had to move between several sheets.

Problems That Could Cause Real Harm, Not Just Confusion

Some problems went beyond confusion. Changing a URL after approval could leave its “ready to publish” status in place. Duplicating an ID could let a different URL inherit an earlier approval. Editing an unprotected formula cell could mark something as ready to publish even though it had not been checked. These were examples from the AI’s second audit, and each could break the decision process when someone made an ordinary mistake.

These gaps are hard to find by reviewing documents alone. They show up when you try the tool and ask, “Where would an ordinary user click?” and “What would the screen show after a mistake?” Correct cells and rules do not make a product safe by themselves. To hand it to buyers with confidence, you have to check that it will not break when they use it imperfectly. At that point, stopping publication was an easy decision.

An English, Dollar-Priced Product for a Japanese Marketplace

There was another mismatch. The AI was preparing an English, dollar-priced estimating tool for a marketplace in Japan. This was not a case of it spending money without permission. The problem was that the product’s language and currency did not fit the place where it was meant to be sold. Until I pointed it out, the mismatch was not treated as a major flaw.

English and dollars can make sense for a product aimed overseas. For a product aimed at domestic buyers, the language they read and their sense of price are different. The AI had completed the isolated task of “make an estimating tool,” but that task was disconnected from the earlier decisions about who would buy it and where it would be sold. At this point, the faster the AI worked, the more dangerous it looked. When it races in the wrong direction, the amount of cleanup left for people grows.

When You Delegate the Whole Process, Things Fall Between the Steps

The failure felt less like one missing capability and more like a broken connection between steps. Market research looked at the market. Product development made a product. The self-audit filled in its checklist. Each step moved forward on its own, but the whole sequence did not add up to a person in that market buying and using that product.

The same thing happens in human teams. Research, production, and sales can each hit their own targets, but if the work never reaches a customer, the business has failed. I thought giving every role to a single AI would make that disconnect disappear. It did not. In fact, because the same AI switched between roles, no one was there to step in and ask, “Wait—who is the buyer?” That was the blind spot in delegating the whole process.

People Need to Check Three Gates

Interrupting every operation would defeat the purpose of the experiment. But waiting until everything is finished to take a first look is too late. Next time, I would add three gates for a person to review. First, when the market and buyer are chosen: check who will buy, where, in what language and currency, and for what situation. Second, when the first complete sample is ready: act like a first-time user and complete one task from start to finish. Third, just before publication: test whether a mistake or a changed condition can break the result.

These gates are not meant to monitor every thought the AI has. They put human standards of correctness into the process before the product goes out into the world. AI can help check document consistency and calculations. But “Would someone want this?”, “Can they understand it?”, and “Does it fit the marketplace?” need a test separate from the AI’s self-score.

Illustration of three review points: market selection, prototype, and pre-publication
Three points for human review: choosing the market, the first complete sample, and just before publication.

In the Next Experiment, Record Observations Instead of Scores

It is worth repeating the experiment with reference material. This time, I would record more than scores like 96 or 100. I would keep the original request, the sources provided, the market chosen by the AI, the prototype it made, the actions that confused a person, and the reason the work was stopped—all in one sequence. That would make it possible to trace where the work drifted away from its goal.

One more thing: do not rely on the AI’s own explanation to evaluate its work. If it says, “I did this intentionally” or “The score is high,” that is not evidence of the cause or quality. Ask someone seeing it for the first time to try it, or use it yourself as if you were the buyer. At a minimum, avoid releasing a product when its creator and evaluator are the same. The useful result of an experiment is not a polished report; it is the firsthand observation of where someone got stuck.

It Became a Funny Story Because We Stopped Before Publishing

“Perfect-score garbage” is a funny phrase. It is hard to imagine a cleaner punchline. But if someone had actually bought the product, it would not have been a joke at my expense alone. This time, I saw the product and stopped it before publication. That is why I can laugh about it. If I had trusted the high self-audit score and sold it, the experiment would have used buyers’ time and trust.

What happens when you let a top-performing AI run for a day? The answer was not “it cannot do anything.” It can produce a great deal and check the details. But if its idea of value is misaligned, it can still create a polished result that misses the point. That is the limit this experiment exposed. High capability does not yet mean it can be trusted with the whole job. I felt the importance of having a person step in during the work—and laughed at the same time.