Asymptote

Independent analytics laboratory

Measurement for systems
that learn.

Machine-learning systems fail quietly. Accuracy drifts, benchmarks stop resembling production, and a dashboard keeps showing green. We build the measurement layer that makes the failure visible — and the decision obvious.

The limit is never reached. The distance to it is the thing worth measuring.

Practice

Four questions we are built to answer

Every engagement starts as one of these. Most end up touching two.

01

Does the model actually work?

Evaluation suites built against your real distribution, not a public benchmark that stopped being representative two years ago. Held-out sets with documented provenance, slice-level reporting, and failure taxonomies you can act on.

  • Benchmark design
  • Slice analysis
  • Human evaluation protocols
02

Did the change cause the outcome?

Correlation is cheap and abundant. We design experiments and quasi-experiments that isolate effect from coincidence — and say plainly when the data cannot support the claim being asked of it.

  • Experiment design
  • Causal inference
  • Uplift modelling
03

What happens next, and how sure are we?

Forecasts are only useful when they carry honest uncertainty. We build time-series and demand models that report calibrated intervals, and we backtest them against the decisions they are meant to inform.

  • Time-series modelling
  • Calibration
  • Backtesting
04

Will we notice when it breaks?

Instrumentation, drift detection and regression harnesses that run on every deploy. The goal is unglamorous: the day performance moves, somebody finds out from a test rather than from a customer.

  • Drift detection
  • Evaluation harnesses
  • Monitoring design

Method

How an engagement runs

Short cycles, written artefacts at every step, and no dependency on us at the end.

  1. 1

    Scope

    We write down the decision the analysis is meant to support, and the evidence that would change it. If no realistic result would change the decision, we say so and the engagement stops here.

    Output A one-page decision brief

  2. 2

    Baseline

    Before any modelling, we establish what current performance genuinely is — reproducibly, from raw data, with the pipeline committed to version control. Baselines are where most surprises surface.

    Output Reproducible baseline and data audit

  3. 3

    Experiment

    Hypotheses registered in advance, so results cannot be quietly reinterpreted afterwards. Negative results are reported with the same weight as positive ones.

    Output Analysis notebook, code, and findings memo

  4. 4

    Handover

    Everything we build ships as code your team owns and can run without us: documented, tested, and walked through until somebody on your side can reproduce the result unaided.

    Output Running pipeline and a working session with your team

Principles

What we hold to

Reproducible or it did not happen

Every number we report can be regenerated from raw data by a command your team can run.

Uncertainty is part of the answer

A point estimate without an interval hides exactly the information a decision needs.

Negative results get reported

"The effect is not there" is a finding worth paying for. We deliver it with the same care.

Handover beats retainer

Success is your team running the work without us. We would rather be re-hired than depended on.

Contact

Bring us a question you cannot answer internally

Describe the decision at stake and the data you hold. We will reply with an honest read on whether the question is answerable, what it would take, and whether you need us at all.

hello@asmtg.site