
First in a series on probabilistic forecasting for physical signals. Next: what happens when you roll the forecast forward more than one step.

A viral debate over loops versus graphs points to a bigger shift in how we build AI systems. Here’s what graph engineering actually means, how it differs from prompt, context, and loop engineering, and why it matters.

Why the skill-inflation panic is aimed at the wrong thing, and what it costs to make agent knowledge a build artifact instead of a file.

Building a FastAPI endpoint for churn prediction, and everything that broke between "it runs" and "it's live

A practitioner's guide to estimating what an opt-in AI feature actually did, when nobody randomized it.

A real Weave project that regression-tests three OpenAI models against the exact reply format your app depends on.

What large operators can teach us about turning alert fatigue into faster, safer service assurance

I built a system that automatically discovers, verifies, and applies relevant requirements from earlier interactions without asking the user where they came from.

Why AI coding makes software design more important

Frequentist confidence intervals and Bayesian credible intervals answer different questions, and confusing them can distort product decisions

A viral debate over loops versus graphs points to a bigger shift in how we build AI systems. Here’s what graph engineering actually means, how it differs from prompt, context, and loop engineering, and why it matters.

A practitioner's guide to estimating what an opt-in AI feature actually did, when nobody randomized it.

Why AI coding makes software design more important

Frequentist confidence intervals and Bayesian credible intervals answer different questions, and confusing them can distort product decisions

Why autonomous agents expose a new explainability problem in fraud detection

A model is only as reliable as the assumptions behind it

First in a series on probabilistic forecasting for physical signals. Next: what happens when you roll the forecast forward more than one step.

When Codex is the right shape for the problem, when Claude Code is, and how I split 5 specialist agents between them on dense AI capacity work.

Why the skill-inflation panic is aimed at the wrong thing, and what it costs to make agent knowledge a build artifact instead of a file.

A real Weave project that regression-tests three OpenAI models against the exact reply format your app depends on.

I built a system that automatically discovers, verifies, and applies relevant requirements from earlier interactions without asking the user where they came from.

Statistical thinking beyond formulas