promptdojo_

Challenge the recommendation, then rerun the memo — step 1 of 7

The recommendation arrives fully dressed

Somewhere past the charts, the AI stops describing and starts advising: "Recommendation: sunset the trial tier. It converts at 4.1%, the worst of the three tiers, and removing it would simplify pricing." Clean sentence. Number attached. Sounds like the analyst you wish you were at 6pm.

Here's what that sentence doesn't know. Trial isn't supposed to convert like the paid tiers — it's the top of the funnel for them. If trial feeds 60% of your team-plan upgrades, "worst-converting tier" is true and "kill it" is a $2M mistake, and both fit in the same chart. The recommendation is where AI's confident-narrative failure mode does the most damage, because a recommendation travels: it leaves your notebook, enters a deck, and starts making decisions while you're at lunch. And unless someone builds a review step with a name on it, nothing stands in its way — the advice ships with a company logo on it and no one on record having examined it.

Analysts have exactly the right tool for this, borrowed from the most paranoid industry there is. Since 2011, the Federal Reserve and OCC's SR 11-7 guidance has required banks to validate every model against three things, and regulators now apply the same triad to AI and gen-AI models:

  1. Conceptual soundness — does the method actually fit the question? (Is conversion rate even the metric a kill-the-tier decision needs?)
  2. Ongoing monitoring — once it ships, what tells us it's drifting or breaking? (Define the canary before the change, not during the postmortem.)
  3. Outcomes analysis — did its past answers match what actually happened? (The last recommendation from this same pipeline predicted +8% MRR. Reality delivered +1.9%. That calibration is data.)

One more rule rides along with the triad, and it's the one people resist: the validator is not the author. Banking learned this the hard way; self-review isn't review, and that applies with extra force when the "author" is a model that agrees with whoever pushed back last. Someone who didn't write the query — and is allowed to say no — runs the challenge.

If this sounds like a burden, reframe it: it's the job description now. The drafting has moved to the machine; what's left of the role — the part someone gets paid for — is standing behind the output when it's questioned. The challenge record you'll build in this lesson is that skill with a filename. And because a challenge is only as good as your ability to rerun the analysis it challenged, the second half of the lesson attacks the quietest failure of all: the memo that changes between two runs on the same data.