AI Outcomes in the AI Era Depend on More Than
model performanceA Five-Factor Framework for Generative AI
Generative AI has become remarkably fast and capable. Does that mean people using the same model will eventually get the same results?
For more than a year, I have regularly built reference knowledge, brainstormed with AI, automated work, and used AI agents. Even with the same model, the results still vary widely.
I found it useful to explain the difference by treating AI outcomes not as one performance score but as the product of five factors.
AI outcomes differ not because of
who uses the smartest model, but because of
What you combine with the same AI.
Will everyone get the same results as models get smarter?
Improvements in model capability have rapidly expanded what generative AI can do: reason, understand long documents, research, code, combine disciplines, and work autonomously. Workflows that draw on multiple reference sources have also become more practical.
But as AI gets stronger, another question arises: will everyone approach the same quality if they use the same model?
In practice, they will not. Even with the same model, results differ based on what it reads, what you ask, what feels wrong in the output, and how you carry lessons from failure into the next attempt.
Think of AI outcomes as the product of five factors
This is not an academically validated formula. It is a conceptual model for organizing my experience with AI and deciding what to improve.
A conceptual model for reviewing AI work
Model capability
Baseline ability in reasoning, instruction-following, long-context processing, research, coding, and autonomous execution.
Unique context
Company information, customer understanding, past decisions, successes and failures, and domain-specific reference knowledge.
Question quality
The ability to define what needs thought and expand the discussion by connecting disciplines.
Review skill
The ability to catch and correct contradictions, leaps in logic, missing assumptions, infeasible steps, and actions beyond authority.
Learning speed
The speed of cycling through outputs, identifying issues, correcting them, reassessing, and preserving the lesson as a decision criterion.
Model capability is a strong foundation shared by everyone
As models improve, their ceiling rises. They can follow complex instructions, retain long context, combine knowledge, and carry out implementation or research.
As models improved, I became able to work across a large body of reference knowledge that I could not use effectively before. That was a major change for me.
At the same time, new models are generally available to many people. They may offer a short-term advantage, but access to a model alone is unlikely to differentiate you over the long term. Model capability is the foundation of an outcome, not the whole outcome.
The same model becomes a different AI with different background data
When you provide information AI does not have by default, its responses move from generic advice toward your reality. Company information, industry knowledge, customer traits, past decisions, successes and failures, and your own criteria give its answers definition.
I started by using Google UX Design course materials as reference knowledge. Over time, I moved beyond summarizing them: I turned them into UX criteria, used them for UX audits, and connected them with marketing, psychology, behavioral economics, and analytics.
Broad knowledge expands the questions you can ask
Even as AI gets smarter, deciding what to think about remains important. Knowing only marketing can deepen questions in that field. Some understanding of UX, customers, systems, data, and operations design can broaden the questions across disciplines.
“Does it work as an experience, not just as a conversion?”
“How does it affect retention, not just short-term response?”
“Even if it is good for customers, can operations handle the workload?”
“Given these customer traits, what differs from the general case?”
“How can we test that hypothesis with data?”
In the AI era, you do not need to be an expert in every field to become more effective. There is also potential in understanding the structures of several disciplines and connecting them to form better questions.
Review skill matters more than generation ability
No matter how smart AI is, you may accept a wrong answer if you cannot spot the error. Skill in using AI shows not only in eliciting good answers but also in stopping bad ones.
Is the classification really appropriate for the goal?
Does this logic contradict an earlier claim?
Does it work as a business, as well as from a UX perspective?
Does this sample support such a strong conclusion?
Do we actually have the authority to carry out the recommendation?
Even if it reads well, are facts and judgments being mixed?
Without this ability to spot issues, fluent mistakes can be mistaken for high-quality work. As AI takes on more generation, human value shifts toward review, selection, and accountability.
Turn one failure into better quality next time
Over time, a gap opens between someone who asks a question once and stops and someone who identifies issues, revises, reassesses, and saves the reasoning as a decision criterion.
Cycles of hypothesis, production, review, and revision that once took days can now happen in minutes. If you also feed the reasons for corrections back into your knowledge base, one failure improves future quality. That is how AI use compounds.
A weak factor limits the whole, even if another is strong
The point of using multiplication is to show that one strong factor cannot drive results if the others remain low. The numbers below are illustrative, not measured; they demonstrate how a bottleneck works.
Only the model is strong
- Model capability95
- Unique context90
- Question quality20
- Review skill20
- Learning speed20
All five factors working together
- Model capability95
- Unique context90
- Question quality80
- Review skill85
- Learning speed90
The goal is not only to improve factors that are already strong. Find the weakest factor holding back the overall result and raise it.
Some differences will disappear, but framing and selection will remain
As models improve, the value of prompt tricks, format instructions, basic research, information organization, and simple routing will decline. AI will infer more intent and handle intermediate work.
But people will still decide which problems to recognize, what to give AI, which disciplines to connect, what is wrong in the output, and what to accept or discard.
As model capability improves for everyone, differences in the other factors may become more visible. The gap in AI performance may not disappear; differences in human workflow design may stand out more.
The ability to bridge disciplines matters more than the amount you know
Before AI, broad but shallow knowledge could be seen as unfocused. But if generative AI can provide deeper expertise when needed, people can understand the basics across disciplines and use AI to explore the necessary depth.
Marketing, UX, pets, data, systems, and AI may each look like a separate specialty. But connecting customer understanding to UX, UX to business metrics, data to operations design, and making those connections testable with AI can create value beyond a simple average of skills.
The value of breadth is not the number of fields you know. It is the ability to pose questions between disciplines and connect them in a way AI can use.
Review the five factors before switching models
You do not need to switch models first to improve your AI workflow. Start by checking which factor is holding back the result.
Does the model meet the reasoning, long-context, implementation, and research needs of the task?
Are you giving it your customers, constraints, past decisions, successes, and failures instead of only general information?
Before asking for an answer, have you defined the issues and effects across other disciplines?
Can a person verify acceptance criteria, prohibitions, evidence, and authority to act?
Are the reasons for corrections carried into future criteria instead of ending in a single conversation?
Improve the weakest factor you find. The multiplicative model is not a score for ranking AI users; it helps determine where to work next.