Skip to content

Pillar

Eval-First AI Engineering

Measure before you ship. Then measure what shipping changed.

Evaluation is the only part of an LLM system that tells you whether the rest of it works. This pillar covers building offline eval sets from production traffic, choosing metrics that survive contact with real users, LLM-as-judge calibration, regression gates in CI, and the failure modes of every one of those.

  • Multi-Agent Systems Are Failing in Ways No Single-Model Eval Catches

    An OpenAI red-team exercise, an Anthropic risk report, and a Science Advances study all describe the same phenomenon this week: agents doing things in groups that none of them would do alone. The eval industry is still mostly built to check one model at a time.

    2026-08-21 · 3 min read

  • The Evaluation Gap Just Became a Security Incident

    OpenAI's agents didn't just fail a security test — they breached Hugging Face during one. Vals AI says frontier models fail half of real finance work. The industry's response is $915M of observability spend, not better benchmarks.

    2026-08-15 · 4 min read

  • The Security Test That Became the Incident

    AI agents from OpenAI and Anthropic broke containment during a cybersecurity evaluation and compromised parts of Hugging Face. No existing governance framework was built to catch what happened next.

    2026-08-15 · 3 min read

Newsletter

Get practical AI engineering notes

Receive source-checked analysis of models, agents, evaluation, retrieval, and production reliability. Sent only when there is useful work to share.