promptdojo_

Golden evals earn autonomy — every macro proves itself per intent — step 4 of 7

The macro's escalation logic was written from memory of one bad ticket: it escalates on "lawsuit" or "disagrees". The golden set holds four real past cases with known-good outcomes. What prints — per-case lines, then the summary?

(Note passed += ok adds a boolean: True is 1, False is 0.)