Every engagement starts as one of these. Most end up touching two.
01
Does the model actually work?
Evaluation suites built against your real distribution, not a public
benchmark that stopped being representative two years ago. Held-out
sets with documented provenance, slice-level reporting, and failure
taxonomies you can act on.
02
Did the change cause the outcome?
Correlation is cheap and abundant. We design experiments and
quasi-experiments that isolate effect from coincidence — and say
plainly when the data cannot support the claim being asked of it.
03
What happens next, and how sure are we?
Forecasts are only useful when they carry honest uncertainty. We
build time-series and demand models that report calibrated intervals,
and we backtest them against the decisions they are meant to inform.
04
Will we notice when it breaks?
Instrumentation, drift detection and regression harnesses that run
on every deploy. The goal is unglamorous: the day performance moves,
somebody finds out from a test rather than from a customer.