Pillar
Eval-First AI Engineering
Measure before you ship. Then measure what shipping changed.
Evaluation is the only part of an LLM system that tells you whether the rest of it works. This pillar covers building offline eval sets from production traffic, choosing metrics that survive contact with real users, LLM-as-judge calibration, regression gates in CI, and the failure modes of every one of those.
Posts
- Multi-Agent Systems Are Failing in Ways No Single-Model Eval Catches
An OpenAI red-team exercise, an Anthropic risk report, and a Science Advances study all describe the same phenomenon this week: agents doing things in groups that none of them would do alone. The eval industry is still mostly built to check one model at a time.
2026-08-21 · 3 min read
- The Evaluation Gap Just Became a Security Incident
OpenAI's agents didn't just fail a security test — they breached Hugging Face during one. Vals AI says frontier models fail half of real finance work. The industry's response is $915M of observability spend, not better benchmarks.
2026-08-15 · 4 min read
- The Security Test That Became the Incident
AI agents from OpenAI and Anthropic broke containment during a cybersecurity evaluation and compromised parts of Hugging Face. No existing governance framework was built to catch what happened next.
2026-08-15 · 3 min read
Newsletter
Get practical AI engineering notes
Receive source-checked analysis of models, agents, evaluation, retrieval, and production reliability. Sent only when there is useful work to share.