Evals — Smoke, Regression, Golden

Evals or Hope, Pick One

EvalSkillCronsilent failure

It’s 11:42 PM Thursday. The friday-wrapup that has run flawlessly for six weeks just shipped a leadership canvas with $0 in pipeline because someone renamed a HubSpot stage and the skill silently filtered everything out. My COO read it first. I found out from her Slack DM at 7:14 AM Friday — “Vlad, did we have a bad week or is the skill broken?” Both answers are bad. The skill was broken for nine days. I had no .

The $0 canvas, in production
The $0 canvas, in production screenshot the leadership Slack canvas with the empty pipeline section, the timestamp, and the COO's DM underneath — readers need to feel the receipt, not just hear about it.

The 9-day silent failure#

Here’s what nine days of silence costs. Nine canvases shipped to the leadership channel. Each one wrong in the same way — pipeline section empty, deal-motion section thin, executive summary contradicting itself. None of my reports flagged it because three of them had stopped reading the canvas closely two weeks earlier (it had become wallpaper) and the fourth assumed the empty pipeline meant a quiet stretch. The skill didn’t crash. It didn’t error. It returned a beautifully formatted canvas with a ghost inside.

The trigger was a HubSpot stage rename — “Qualified” became “Qualified — Round 1” because a new VP of Sales wanted to track a sub-stage. The skill’s filter was hard-coded against dealstage = 'Qualified'. Zero matches, zero drama. The model generated graceful prose around the empty result set. “A measured week with a focus on top-of-funnel motion” — that was the lede. There was no top-of-funnel motion. There was no funnel motion at all because the query returned nothing.

If you’re keeping score: a skill that failed silently for 216 hours, in front of every leader at the company, written by me, owned by me, with no instrumentation between the model and my COO’s screen. That’s not a model failure. That’s an operator failure. I shipped a worker into production and forgot to ship the supervisor.

What an eval actually is#

Strip the jargon. An eval is a function that runs against your skill’s output and answers one question: did this output meet a minimum bar? It returns a boolean and a reason. That’s it.

You don’t need an eval framework. You don’t need a benchmark suite. You don’t need Promptfoo, Braintrust, LangSmith, or anything with a logo. You need three lines:

def eval_friday_wrapup(canvas_text: str) -> tuple[bool, str]:
    if "$0" in canvas_text or "no pipeline" in canvas_text.lower():
        return False, "pipeline section is empty — likely stage filter drift"
    return True, "ok"

That heuristic would flag the described $0 canvas, but it is not a correctness test. A genuinely empty pipeline can be valid; a broken connector can also produce convincing nonzero prose. Check source status and structured counts before trusting the text. The runnable example below makes that distinction.

You are evaluating the workflow, not only the model. Start with source health and artifact checks; add labeled scenarios and factuality checks when the job needs them. Internal reports need correct numbers too. The model can be perfectly fine and the workflow can still be broken because something upstream changed shape: a stage rename, an API rate limit, or an auth error quietly summarized as “no recent activity.”

An eval is a smoke detector for the artifact. Not for the model.

The four eval types every operator needs#

Four shapes cover roughly 90% of what shipped skills actually need. Building one of each, even badly, beats building a perfect framework.

Smoke evals ask “did the artifact arrive and contain the obvious things?” Length over 200 chars. All required sections present. Headers in the right order. Money figures parse as numbers. These catch the dumb failures — empty output, truncated output, malformed JSON. Run them on every output, every time.

Regression evals compare today’s artifact to yesterday’s. Did the canvas length drop 80%? Did the deal count go from 47 to zero? Did the executive summary section disappear? You don’t need ML for this. You need a stored snapshot and a delta function. If today’s pipeline value is less than 10% of last week’s pipeline value, raise a flag — the skill might be right (a genuinely terrible week) or wrong (broken filter), and either way a human should look.

Golden-set evals are the smallest deliberate test data you can write. Three or four hand-built input scenarios with known correct outputs. You ship a skill change, you run it against the golden set, you check that the four answers still look right. This is the eval most operators skip because it feels like overhead. It is overhead. It’s also the cheapest insurance against a CLAUDE.md edit silently changing your pipeline math.

Adversarial evals assume the upstream world is hostile. Stage names change. APIs return 503. Connectors decide to require new scopes. Empty arrays appear. The adversarial eval feeds your skill the worst plausible inputs — empty result sets, malformed dates, surprise null fields — and confirms it fails loudly instead of producing graceful nonsense. Most silent failures live in the gap between “API returned nothing” and “model wrote graceful prose around the nothing.”

You don’t need all four on day one. Build the smoke eval first, then add fixtures for the failures your workflow actually faces. There is no measured coverage percentage promised by this starter.

Running evals on cron#

Here’s the second job nobody talks about — the one that runs the eval, not the workflow.

The earlier setup ran a dry run at 4:30 PM ET before a 5:00 PM Friday delivery. That is useful early warning: a failed check gives the owner time to fix an upstream issue. But a passing dry run does not validate another model response or a changed source thirty minutes later.

The required delivery sequence is fetch complete inputs, generate once, validate that exact artifact, then publish those same bytes. Failed reads, malformed output, or failed checks block publication and notify the owner. If you regenerate, validate again. Keep the evaluator and approved baseline outside the producing agent’s write permissions. Chapter 7 covers cursor recovery and duplicate delivery; an eval does not replace those controls.

A second model run has a cost; measure it on your workload. The local checks below require no model calls. The dry run is optional advance notice, while the delivery gate is mandatory.

The eval failure budget#

Evals fire false positives. If you treat every fired eval as a fire drill, you’ll mute the eval inside three weeks and be back to nine-day silent failures. The fix is a failure budget — how often the eval is allowed to be wrong before you change the eval, not the skill.

My rule: an eval that pages me more than once every two weeks gets refined. An eval that pages me less than once a quarter gets dropped or hardened — either it’s not catching anything real, or it’s so loose it’s not actually watching. The two evals I run for friday-wrapup have fired four times in the last six months. One real failure (stage rename), one near-real (HubSpot rate limit cascading into thin output), two false positives (genuinely quiet weeks where pipeline did drop hard). The 50% true-positive rate is on the low end of what I’d accept; if it drops below 25% I’ll tighten the threshold.

The eval is also a skill. It’s not divine. It can drift. It can be wrong. The thing you’re protecting against is silent failure, not all failure — accept the false positives as the cost of catching the silent ones.

Three evidence types, not one dataset#

These sources answer different questions. They support taking evaluation seriously, but they are not three equivalent measurements of silent production failures.

Operator account: the friday-wrapup incident described above is my account of workflow drift, not an independently measured failure rate.

Qualitative research: Anthropic’s study interviewed 80,508 people across 159 countries and 70 languages about their experiences, hopes, and concerns around AI. It was not an incident database or a study of 81,000 reported agent failures. Reports of unreliability matter, but they do not establish the frequency or mechanism of my pipeline bug.

Benchmark security research: Berkeley RDI’s April 2026 report describes an automated scanning agent finding exploitable evaluation infrastructure in eight benchmarks. That demonstrates weaknesses in the tested evaluators, not that every published score is fraudulent or that models were trained into greater underlying capability.

Primary sources checked on 2026-09-05. My operator inference is to test your own stage filters, source failures, and held-out scenarios, and to protect the evaluator from the agent it evaluates. The studies do not supply a production failure percentage for your workflow.

The 30-minute starter eval#

This is an executable, standard-library artifact checker, not a complete scheduler or connector. Put it in eval_yourskill.py. Your trusted connector adapter must supply source.status and source.rows after fetching every page; do not ask the model to invent its own source-health evidence. The model supplies only the canvas text.

import argparse
import json
from pathlib import Path
import sys

def smoke(output: dict) -> tuple[bool, str]:
    source = output.get("source")
    if not isinstance(source, dict) or source.get("status") != "ok":
        return False, "source read failed or completeness is unverified"
    rows = source.get("rows")
    if type(rows) is not int or rows < 0:
        return False, "source row count must be a non-negative integer"
    text = output.get("canvas", "")
    if not isinstance(text, str):
        return False, "canvas must be text"
    if len(text) < 200:
        return False, f"canvas too short: {len(text)} chars"
    required = ["Pipeline", "Deal Motion", "Executive Summary"]
    missing = [s for s in required if s not in text]
    if missing:
        return False, f"missing sections: {missing}"
    return True, "ok"

def regression(output: dict, baseline: dict) -> tuple[bool, str]:
    if len(output["canvas"]) < 0.4 * len(baseline["canvas"]):
        return False, "canvas length dropped more than 60%; review required"
    if output["source"]["rows"] < 0.1 * baseline["source"]["rows"]:
        return False, "source row count dropped more than 90%; review required"
    return True, "ok"

def read_output(path: str) -> dict:
    output = json.loads(Path(path).read_text(encoding="utf-8"))
    if not isinstance(output, dict):
        raise ValueError("artifact must be a JSON object")
    ok, reason = smoke(output)
    if not ok:
        raise ValueError(reason)
    return output

def main() -> int:
    parser = argparse.ArgumentParser()
    parser.add_argument("artifact")
    parser.add_argument("baseline", help="human-reviewed JSON artifact")
    args = parser.parse_args()
    try:
        output = read_output(args.artifact)
        baseline = read_output(args.baseline)
        ok, reason = regression(output, baseline)
        if not ok:
            raise ValueError(reason)
    except (OSError, ValueError) as error:
        print(f"BLOCKED: {error}", file=sys.stderr)
        return 1
    print("PASS: artifact checks passed; no delivery attempted")
    return 0

if __name__ == "__main__":
    sys.exit(main())

Start with this synthetic fixture, not a production receipt, in candidate.json. Review it and create a separate approved-baseline.json with the same structure for the first local test:

{
  "source": { "status": "ok", "rows": 0 },
  "canvas": "Pipeline: $0 in this synthetic empty-week fixture. Deal Motion: no changes were returned after a complete source read. Executive Summary: a quiet test window, not a connector failure. The operator reviewed this example before accepting it as a baseline."
}
python3 eval_yourskill.py candidate.json approved-baseline.json

A passing fixture prints PASS and exits 0. A missing baseline, failed source, malformed artifact, or regression exits nonzero and must block delivery. A valid zero-row result can pass against a reviewed quiet baseline; a sudden drop from a busy baseline requires review. The checker never creates or replaces its baseline, contacts a service, or publishes anything.

These thresholds are illustrative, not measured guarantees. Section presence does not prove that the prose matches the source totals: add numeric reconciliation and held-out domain cases before production. Your scheduler must route nonzero exits to its failure channel and publish only the exact checked artifact after success, with the cursor and delivery rules from Chapter 7. Do not regenerate between checking and sending.

If your skill writes to a -driven workflow or uses an connector that talks to HubSpot, Stripe, or Slack — anything covered in Chapter 12 — the same eval shape applies. Smoke test the artifact, regression test against yesterday, page yourself when the world drifts.

The closer#

The operator reads the canvas, not the eval. Keep the early warning if it buys useful repair time, but put the blocking check on the artifact that actually ships. A quiet result, a failed source, and a missing baseline must not all look like success.

Spotted something wrong, missing, or sharper? Email Vlad with feedback on this chapter →
Stay close

The next edition lands when this list says it does.

No course. No paywall. Operator playbooks weekly. 10K+ subscribers.