Don’t Ask AI for an Answer. Teach It How to Think.
Sep 22, 2026
Most organizations are teaching employees how to prompt AI.
Far fewer are teaching them how to think with it.
That distinction matters most when AI is being used for work that requires significant critical or executive thinking: diagnosing a business problem, synthesizing a large body of information, evaluating strategic options, making a recommendation, or generating a new solution.
For these tasks, the challenge isn't simply getting AI to produce a polished answer. It's making sure the thinking behind that answer is good.
Research suggests that won't happen automatically.
A Microsoft and Carnegie Mellon study asked 319 knowledge workers about 936 real-world uses of generative AI. Workers who had more confidence in AI reported engaging in less critical thinking, while those with greater confidence in their own ability reported engaging in more.
That's not an argument against using AI for important work.
It's an argument for giving people a better process for doing it.
For work that requires significant critical thinking, leaders should teach employees to do three things:
- Teach AI how to perform the thinking by giving it mental models to use.
- Evaluate its thinking with questions matched to the type of thinking it performed.
- Coach AI by sharing the final product and asking it to compare to its last version.
Call it Teach → Evaluate → Coach.
The research does not test this exact three-step method. Teach → Evaluate → Coach is Zarvana's synthesis of what the evidence suggests about structured reasoning, active evaluation, and feedback.
The important shift is that this goes well beyond prompt engineering.
Prompt engineering tells AI what you want. Critical thinking systems tell it how to get there.
Much of the advice about prompt engineering focuses on improving how we describe the task, context, role, format, and desired output. More sophisticated prompting can also prescribe reasoning steps.
Those practices matter.
But for difficult work, organizations need something broader: a repeatable system for both structuring the thinking before AI responds and evaluating the particular kind of thinking it produces afterward.
Research provides some support for this distinction.
In a 2024 study on metacognitive prompting, researchers tested a technique called metacognitive prompting across four large language models and ten natural-language-understanding benchmarks. Rather than simply asking the models for answers, the technique walked them through a structured process of interpretation, preliminary judgment, critical evaluation, and final decision-making.
The structured approach generally outperformed simpler prompting methods. On several of the more demanding medical and legal benchmarks highlighted by the researchers, it improved performance over Plan-and-Solve prompting by roughly 4 to 12 percentage points.
That does not mean every critical thinking framework will improve every AI response by 4 to 12 points. These were specific benchmark tasks, not business strategy decisions.
How you ask AI to think can matter, not just what you ask it to produce.
Consider an executive evaluating whether a B2B company should expand beyond its core mid-market customers into larger enterprise accounts.
The executive has customer interviews, sales data, churn and retention metrics, competitor research, market-size estimates, implementation costs, and feedback from the sales team.
The conventional prompt might be:
Review this information and recommend whether we should expand into enterprise customers.
A better prompt might provide additional background, specify the desired format, and ask AI to explain its reasoning.
But you can go much further.
You can teach AI an actual critical-thinking process.
Step 1: Teach AI how to think through the problem
The first mistake would be to accept the proposed decision—Should we expand into enterprise accounts?—as the starting point.
The proposal itself contains an assumption: that enterprise expansion is a solution to an important problem.
So start by working backward.
Use POCS to understand the problem before evaluating the solution
POCS stands for Problem, Outcome, Cause, Solution.
In this case, you can reverse-engineer the model.
Enterprise expansion is the proposed solution.
What desired outcome is it intended to produce?
What problem is preventing that outcome today?
And what is causing that problem?
Perhaps management believes growth has slowed because the mid-market segment is reaching saturation.
But that is only one possibility.
The real problem could instead be increasing churn, deteriorating sales productivity, an underdeveloped product offering, weak pricing, or insufficient penetration within the company's existing market.
This is where the prompt itself can encode the mental model:
Use the POCS framework—Problem, Outcome, Cause, Solution—to reverse-engineer the proposal to expand into enterprise accounts. Do not assume expansion is the right solution. Start with “expand into enterprise accounts” as the proposed Solution and work backward.
1. Outcome: What business outcome is enterprise expansion intended to produce?
2. Problem: What current condition is preventing us from achieving that outcome? State the problem without embedding a solution in it.
3. Cause: What are the leading hypotheses for why that problem exists? Distinguish what the available evidence establishes from what is still an assumption.
4. Solution: Only after defining the problem and likely causes, assess whether enterprise expansion logically addresses them. Identify circumstances under which it would not.
If the evidence suggests a different underlying problem than the one implied by the enterprise-expansion proposal, reframe the POCS accordingly rather than forcing the evidence to support the original proposal.
Notice what this does. The prompt doesn't merely ask AI to explain its reasoning. It gives AI a specific structure for reasoning and explicitly tells it not to let the proposed solution determine the problem.
Use hypothesis-driven synthesis to identify the problem the evidence actually supports
The next prompt should build directly on that POCS analysis rather than inventing a hypothesis in advance.
The executive could say:
Something is happening that is causing us to consider expanding beyond our core mid-market customers. Review all of the data I've provided—including customer interviews, sales data, churn and retention metrics, competitor research, market-size estimates, implementation costs, and sales-team feedback—and identify the 1–3 most plausible options for what the underlying Problem in our POCS could be.
For each potential problem:
• State it precisely without embedding a solution.
• Identify the strongest evidence supporting it.
• Identify evidence that weakens or contradicts it.
• Distinguish facts from assumptions or missing information.
• Explain why this problem could logically lead us to consider enterprise expansion.
Then compare the 1–3 options and identify which problem is best supported by the evidence currently available. Do not assume that enterprise expansion is the appropriate solution.
Now the prompts ladder.
The first establishes how to think about the proposal using POCS.
The second asks AI to synthesize the available evidence to determine what should fill the Problem box.
Once AI identifies the strongest problem hypothesis, the leader can dig further into its causes, test the logic, or ask for additional analysis before moving toward a recommendation.
The result is not simply a shorter version of the source material. It is synthesis directed toward improving the problem definition created in the previous step.
Once the problem is clearer, AI can move from synthesis toward recommendation.
Use HAE to build the recommendation
A strong recommendation can be broken into three elements:
Hypothesis: What should we do?
Assertions: What must be true for that recommendation to be right?
Evidence: What demonstrates that those assertions are true?
Suppose AI recommends expanding into enterprise accounts.
One assertion might be:
Enterprise customers offer superior economics to the company's current segment.
But what counts as sufficient evidence?
A larger average contract value alone isn't enough.
You might establish a minimum evidence standard requiring analysis of customer acquisition cost, implementation cost, retention, expansion revenue, gross margin, and sales-cycle length.
AI shouldn't merely find some evidence supporting an assertion. It should know what evidence would constitute an adequate case.
Use the Rule of 3 to prevent the proposal from becoming the default
Finally, require AI to develop at least three meaningful alternatives.
- Expand into enterprise customers.
- Double down on the existing mid-market segment.
- Expand into a different adjacent segment, product, or offering.
Now AI isn't simply answering:
Can I make a reasonable case for enterprise expansion?
It must help answer the harder question:
Is enterprise expansion better than the realistic alternatives?
This is what it means to teach AI to think.
You aren't asking for more explanation after it generates an answer.
You're specifying the thinking process before the answer is produced.
Step 2: Evaluate the particular kind of thinking AI performed
Even a strong thinking process doesn't eliminate the need for human judgment.
But “review the AI's work” isn't much of a process.
Generic evaluation criteria—Is it accurate? Is it complete? Does it make sense?—are better than blindly accepting the answer, but they miss something important.
Different types of thinking fail in different ways.
A synthesis can fail because it identifies the wrong insights or weights them incorrectly.
A recommendation can fail even if every fact in it is technically accurate.
A creative solution can be factually sound yet completely conventional.
So the questions used to evaluate AI should depend on what kind of thinking AI just performed.
That's one of the central ideas behind Zarvana's Critical Thinking Roadmap, which separates critical thinking into different phases such as Execute, Synthesize, Recommend, and Generate.
Return to the enterprise-expansion example.
Test the problem understanding with a logic model
AI has concluded that slowing growth stems from saturation in the mid-market segment.
Don't just ask whether that sounds reasonable.
Reverse-engineer the logic:
We should expand into enterprise accounts
because
our current market cannot support sufficient future growth
because
the mid-market segment is approaching saturation
because
specific customer, market, sales, and competitive evidence indicates that saturation is occurring.
Then work back through the chain in the opposite direction.
Does the evidence actually prove saturation?
If the evidence is true, does saturation necessarily explain the slowdown?
Could the same evidence support a different conclusion—for example, that the company hasn't executed effectively within its existing market?
The evaluation process matches the synthesis task AI performed.
Apply the 3 Ins to the recommendation
Then examine the assertions supporting AI's recommendation using three failure modes:
Inaccurate: Is an assertion actually false, or does its evidence fail to establish it?
Insufficient: Even if the existing assertions are true, is another assertion required for the recommendation to hold?
Incomplete: Could the recommendation be correct on its own terms but create secondary consequences significant enough to change the decision?
For example:
Enterprise customers generate more revenue per account.
That could be perfectly accurate.
But it may be insufficient if the recommendation also requires enterprise customers to produce stronger unit economics.
And the analysis may be incomplete if enterprise expansion creates far greater implementation complexity, longer sales cycles, and organizational demands that outweigh the additional revenue.
Finally, return to the alternatives:
Even if enterprise expansion could work, is one of the other options better?
Now “evaluate the AI output” has become a repeatable process rather than an exhortation to be careful.
That's important because research increasingly distinguishes evaluation from prompt construction itself. A 2026 study developing a Strategic Prompting Scale found three separate dimensions of sophisticated human-AI interaction: planning, adaptation, and evaluation. In two samples totaling nearly 600 participants, these behaviors were positively associated with metacognition and critical thinking and negatively associated with disengagement during AI use.
The problem with AI isn't just whether employees check its answers.
It's whether they have a repeatable process for checking the particular kind of thinking AI performed.
Step 3: Coach AI from the difference between its work and yours
There is still one step left.
Suppose AI ultimately recommends a broad expansion into enterprise accounts.
After reviewing the logic, evidence, risks, and alternatives, the executive reaches a different conclusion: run a six-month enterprise pilot first.
Why?
AI underweighted implementation complexity.
It treated total addressable market as more important than customer economics.
It failed to recognize that two large customers accounted for much of the apparent enterprise demand.
And it did not give enough weight to the longer sales cycle.
The executive should make those changes.
But the process shouldn't end there.
Give AI its version and the final human-edited version. Then explain the meaningful differences:
You recommended a full enterprise expansion. I ultimately recommended a limited pilot. Compare your version with mine. Notice that I placed substantially more weight on implementation cost, demand concentration, and sales-cycle length. In future analyses like this, distinguish total market opportunity from economically attractive opportunity and test whether apparent demand is disproportionately driven by a small number of observations.
There is a useful analogy here to deliberate practice.
The deliberate practice research has sometimes been overstated—practice alone does not explain all differences in expertise. But one enduring principle is that improvement requires more than repetition.
That principle offers a useful analogy for how we can work with AI: performance improves when feedback is specific about the gap between the initial output and a stronger standard.
That is the logic of the coaching step.
Don't simply correct AI's work and move on.
Show it the gap between its performance and your standard.
The analogy has limits: an AI model does not learn from feedback in the same way a human develops expertise, and what persists will depend on the system, context, and memory available.
But from the user's perspective, the principle is useful:
Your edits contain information about your judgment. Don't throw that information away.
From prompt engineering to thinking system design
Prompt engineering is useful.
But for consequential work, it isn't enough.
A good prompt can tell AI the task, provide context, and specify the desired output.
A critical thinking system goes further.
It establishes:
How should AI perform the thinking?
How should the human evaluate the particular kind of thinking AI performed?
How can the gap between AI's work and the final human judgment improve the next interaction?
Research on AI-assisted decision-making points in a similar direction. In a CHI 2025 study of a complex investment task, researchers compared a conventional recommendation-oriented AI with a system designed to build upon and extend users' own reasoning. The latter integrated more naturally with participants' thinking and produced slightly better outcomes, while direct AI recommendations required less cognitive effort.
That tension matters.
The easiest AI interaction is often:
Give me the answer.
For routine production tasks, that may be perfectly appropriate.
But when the work requires significant critical or executive thinking, organizations should aim higher.
Don't just teach employees how to write better prompts.
Give them repeatable critical-thinking processes they can teach AI.
Give them evaluation tools matched to the type of thinking AI performs.
And when their judgment improves the result, teach AI what they saw that it didn't.
Don't just ask AI for an answer. Teach it how to think.