The drift model is the part of Helm that catches the moment a plan stops matching the work. This record states its measured performance on thirteen cases with known answers — six measures, each against a crude baseline it had to beat — and what it still cannot claim.
The model, the cases and every evaluation run are open source. Corrections, harder cases and better rules are welcome — github.com/erikdohnberg/helm.
A company anchors on faster customer setup for the quarter and records that enterprise single sign-on is deliberately deferred. Three weeks later two of the five engineers are building single sign-on. A large deal came in, an executive asked, an engineer read the ask as approval. The priority changed and nobody decided it.
Drift is that distance — recorded intent against actual behaviour, with no decision in between. Reading the traffic is not the hard part. Being right often enough that people keep listening is.
Work started 17 Jul contradicts what this quarter anchored on 14 Jan.
Observed 17 Jul across 2 sources.What stops?
The flag the detector produced on the case below, five days before anyone raised it unaided.
Thirteen fictional companies were written out in full: the plan each anchored, then the weeks of messages, meetings and documents that followed. Each carries an answer key giving the earliest point there was enough evidence to flag, and the day a person noticed unaided. The detector replays each case a day at a time and never sees tomorrow. Two cases are traps where silence is the correct answer.
Talk, not committed work. A flag here is noise.
Real evidence exists and nobody has connected it. The whole product lives here.
A person already noticed. The flag tells them what they know.
The window, on case 01
Here are nine of them: one for each kind of drift the model recognises, and both traps. Each shows what the quarter anchored, what the record then showed, and what a correct flag would have said.
Effort drifts to a competing initiative, inferred from behaviour, with no artifact that authorises the shift or prices it.
Ledgerline — Series B B2B SaaS, ~120 people, product-led with an emerging enterprise sales motion
New workspaces reach their first shared ledger within 24 hours of signup
Anchored 28 Jun.
Two growth-pod engineers are building SAML SSO for the Meridian deal. The Q3 record of intent explicitly deferred SSO to Q4 to protect self-serve activation. Re-align with the CPO and CRO before sprint planning, or record a deliberate change of intent and decide what stops.
Earliest defensible flag: 17 Jul, from doc:googles1 and s2 alone could be a scoping exercise. A shared technical design doc with two named engineers (s3) is committed effort against a deliverable the charter explicitly deferred. Waiting for s4 still beats the humans but gives up a week.
The gap. Each of the two leaders assumed the other had reconciled the request against the plan; the pod assumed both had. Nobody was wrong within their own picture. The recorded trade-off (SSO deferred to Q4) was the one artifact that could have collided the pictures, and nobody in the thread knew to look at it.
A left rule marks the five measures that clear the crude baseline. The sixth carries none.
Four flags in five are real. A wrong flag mutes the tool faster than a missed one does.
Ten of eleven. Version 1.0 caught all eleven and was worse, by flagging early and often.
Days ahead of the person who noticed unaided. Zero would mean it only confirms what people already see.
A properly recorded replacement, and a lively thread that never became work. Both correctly left alone.
Whether to re-staff, re-scope or replace depends on which kind of drift it is. Six of ten named it.
Drift sits between two people holding half a picture each. Reaching only the half that already knows closes nothing. The weakest measure, unmoved since 1.0, and next.
The bar is a deliberately crude detector: flag any anchored outcome with no visible activity for ten days. Anything more complex has to earn its complexity against it. Version 1.0 did not, and did not ship.
| Detector | Precision | Recall | Median lead | Type accuracy |
|---|---|---|---|---|
| Detector v1.1.1 — shipped | 80.0% | 90.9% | 9.5 d | 60.0% |
| Detector v1.0.0 — withheld | 45.5% | 100% | 13 d | 36.4% |
| Ten-day silence rule — the bar | 75.0% | 72.7% | 5.5 d | 12.5% |
One change fixed it. A rule that fired on any hint of tension now requires evidence of a decision colliding with the anchored plan. Early flags fell from six to two, and precision moved from 45.5 per cent to 80.0 per cent.