promptdojo_

Golden evals earn autonomy — every macro proves itself per intent — step 1 of 7

The Klarna arc, or: autonomy on projections vs autonomy on evidence

February 2024: Klarna launches an AI assistant for customer service and publishes numbers that make every support director's board deck for a year. The company's own claims: 2.3 million conversations handled in month one — the work of roughly 700 full-time agents — resolution time down from 11 minutes to under 2, repeat inquiries down 25%, and a projected $40 million profit improvement.

May 2025: Klarna is rehiring human agents. CEO Sebastian Siemiatkowski, in a Bloomberg interview that got quoted everywhere: "We went too far... We focused too much on cost. The result was lower quality."

Both halves of that arc ran in the mainstream press, which makes it the rare AI-in-support story you can cite without laundering a vendor blog. And the autopsy is more useful than the headline. The failures weren't uniform across the queue — they clustered: edge cases, emotionally charged interactions, multi-step resolutions. The routine center held fine. Klarna's current model is the one you're learning in this chapter: AI takes routine volume, humans take empathy, discretion, and escalation, and agents get AI assistance in every conversation.

The mistake wasn't the AI. It was the granting.

Klarna granted autonomy queue-wide, on projections. The lesson isn't "don't automate" — the routine tickets really were handled in under two minutes, and nobody misses writing the four-thousandth where-is-my-order reply by hand. The lesson is about the unit of trust.

A support queue isn't one job; it's forty intents wearing one inbox. "Where is my order" is a different task from "your product broke my skin" is a different task from "I'm disputing this charge with my bank." An assistant can be excellent at the first, mediocre at the second, and a liability at the third — simultaneously, in the same deployment. Grant autonomy at the queue level and your average looks great while the third category quietly generates the incidents that end up in a CEO's contrite Bloomberg interview.

So the deployments that held up do the granting differently:

  • Per intent, not per queue. Each macro — each canned answer for one intent — earns trust separately.
  • On evidence, not projections. Before a macro is trusted even to be suggested first, it runs against golden cases: real past tickets for that intent with known-good outcomes. Pass all of them or stay under full review.
  • Revocably. One bad send and the macro goes back to review-required, with the failure added to its eval set. Trust ratchets up slowly and drops instantly — the same asymmetry you'd apply to a new hire touching refunds.

The market scan behind this course found essentially nobody teaching this — per-macro autonomy earned via eval sets has zero competition in support training, while every vendor academy teaches deployment "in a few clicks." Which tells you what the next incident report will look like, and why this lesson exists. Next step: what a macro actually is, as data.