promptdojo_promptdojo_promptdojo_promptdojo_promptdojo_promptdojo_promptdojo_promptdojo_promptdojo_promptdojo_promptdojo_

Golden evals earn autonomy — every macro proves itself per intent — step 1 of 7

The Klarna arc, or: autonomy on projections vs autonomy on evidence

You're in the support studio. What you're doing this lesson: decide which canned answers the AI may send on its own, and which ones still need a human — based on past tickets, not a hopeful projection. You do not need the policy-pack lesson to start.

A few terms, once:

  • Autonomy here means "the AI may send this reply without a human clicking send."
  • A macro is a canned answer for one kind of ticket ("where is my order," "how do I return this").
  • A golden case is a real past ticket plus the reply that was actually right.

February 2024: Klarna launches an AI assistant for customer service and publishes numbers that make every support director's board deck for a year. The company's own claims: 2.3 million conversations handled in month one — the work of roughly 700 full-time agents — resolution time down from 11 minutes to under 2, repeat inquiries down 25%, and a projected $40 million profit improvement.

May 2025: Klarna is rehiring human agents. CEO Sebastian Siemiatkowski, in a Bloomberg interview that got quoted everywhere: "We went too far... We focused too much on cost. The result was lower quality."

Both halves of that arc ran in the mainstream press, which makes it the rare AI-in-support story you can cite without laundering a vendor blog. And the autopsy is more useful than the headline. The failures weren't uniform across the queue — they clustered: edge cases, emotionally charged interactions, multi-step resolutions. The routine center held fine. Klarna's current model is the one you're learning in this chapter: AI takes routine volume, humans take empathy, discretion, and escalation, and agents get AI assistance in every conversation.

The mistake wasn't the AI. It was the granting.

Klarna granted autonomy queue-wide, on projections. The lesson isn't "don't automate" — the routine tickets really were handled in under two minutes, and nobody misses writing the four-thousandth where-is-my-order reply by hand. The lesson is about the unit of trust.

A support queue isn't one job; it's forty intents wearing one inbox. "Where is my order" is a different task from "your product broke my skin" is a different task from "I'm disputing this charge with my bank." An assistant can be excellent at the first, mediocre at the second, and a liability at the third — simultaneously, in the same deployment. Grant autonomy at the queue level and your average looks great while the third category quietly generates the incidents that end up in a CEO's contrite Bloomberg interview.

So the deployments that held up do the granting differently:

  • Per intent, not per queue. Each macro earns trust separately.
  • On evidence, not projections. Before a macro is trusted even to be suggested first, it runs against golden cases. Pass all of them or stay under full review.
  • Revocably. One bad send and the macro goes back to review-required, with the failure added to its eval set (the list of golden cases it must still pass). Trust ratchets up slowly and drops instantly — the same asymmetry you'd apply to a new hire touching refunds.

Vendor academies teach deployment "in a few clicks." This lesson teaches the slower thing: per-macro autonomy earned against an eval set, revoked the first time it fails. That's why the next incident report looks the way it does. Next step: what a macro actually is, as data.