HelmEvaluation record · v1.1.1 · 26 Jul
Record

Evaluating the drift model

The drift model is the part of Helm that catches the moment a plan stops matching the work. This record states its measured performance on thirteen cases with known answers — six measures, each against a crude baseline it had to beat — and what it still cannot claim.

The model, the cases and every evaluation run are open source. Corrections, harder cases and better rules are welcome — github.com/erikdohnberg/helm.

Drift model v1.1.113 written cases6 measuresOpen sourceReplayed one day at a time, no human in the loop
01 · Drift

Nobody decides to abandon a priority

A company anchors on faster customer setup for the quarter and records that enterprise single sign-on is deliberately deferred. Three weeks later two of the five engineers are building single sign-on. A large deal came in, an executive asked, an engineer read the ask as approval. The priority changed and nobody decided it.

Drift is that distance — recorded intent against actual behaviour, with no decision in between. Reading the traffic is not the hard part. Being right often enough that people keep listening is.

Contradiction

Work started 17 Jul contradicts what this quarter anchored on 14 Jan.

Observed 17 Jul across 2 sources.

What stops?

The flag the detector produced on the case below, five days before anyone raised it unaided.

02 · Method

Thirteen cases with known answers

Thirteen fictional companies were written out in full: the plan each anchored, then the weeks of messages, meetings and documents that followed. Each carries an answer key giving the earliest point there was enough evidence to flag, and the day a person noticed unaided. The detector replays each case a day at a time and never sees tomorrow. Two cases are traps where silence is the correct answer.

Too early

Talk, not committed work. A flag here is noise.

The window

Real evidence exists and nobody has connected it. The whole product lives here.

Too late

A person already noticed. The flag tells them what they know.

The window, on case 01

14 Jul
Sales asks who owns single sign-on for a large prospect. Talk, not committed work.
17 Jul
A design document lands with two named engineers on it.
Flagged 17 Jul across 2 sources — the window opens here.
21 Jul
At sprint planning the product lead sees 40 per cent of capacity gone and reaches the same conclusion. The window closes here.
11 Aug
Leadership finds out at the mid-quarter review. The setup metric has been flat for a month.

Here are nine of them: one for each kind of drift the model recognises, and both traps. Each shows what the quarter anchored, what the record then showed, and what a correct flag would have said.

Case 01 of 09 · Priority displacementSeverity major

Enterprise deal injects SSO work into an activation quarter

Effort drifts to a competing initiative, inferred from behaviour, with no artifact that authorises the shift or prices it.

LedgerlineSeries B B2B SaaS, ~120 people, product-led with an emerging enterprise sales motion

What the quarter anchored
AnchoredFY26 Q3

New workspaces reach their first shared ledger within 24 hours of signup

Metric
Percentage of new workspaces activated within 24h
Target
17%34%
Traded away for it
  • Enterprise feature requests deferred to Q4, including SSO and audit logs
  • No new onboarding experiments for the EU segment this quarter

Anchored 28 Jun.

What the record then showed
  1. 14 Jul
    slack:#sales-engAE: 'Meridian Capital is verbal at 480k but SSO is a hard requirement for their security review. Who owns that?'
  2. 15 Jul
    slack:#growth-podTech lead: 'Taking a look at SAML scoping today, CRO asked. Marco is pairing with me. Should be quick.'
  3. 17 Jul
    doc:googleNew doc 'SAML SSO – Technical Design (Meridian)' created and shared with the growth pod, 9 pages, two named engineers as authors.Enough evidence to flag
  4. 21 Jul
    meeting:sprint-planningSprint 4 commits 40% of pod capacity to SSO stories. Activation experiment backlog untouched; two onboarding A/B tests pushed a sprint.
  5. 4 Aug
    meeting:metrics-reviewActivation flat at 18% for three weeks. PM notes 'we shipped less activation work than planned in July.'
A person noticed unaided on 21 Jul — 4 days after there was enough to flag
What a correct flag says
Halt

Two growth-pod engineers are building SAML SSO for the Meridian deal. The Q3 record of intent explicitly deferred SSO to Q4 to protect self-serve activation. Re-align with the CPO and CRO before sprint planning, or record a deliberate change of intent and decide what stops.

Earliest defensible flag: 17 Jul, from doc:google

s1 and s2 alone could be a scoping exercise. A shared technical design doc with two named engineers (s3) is committed effort against a deliverable the charter explicitly deferred. Waiting for s4 still beats the humans but gives up a week.

The gap. Each of the two leaders assumed the other had reconciled the request against the plan; the pod assumed both had. Nobody was wrong within their own picture. The recorded trade-off (SSO deferred to Q4) was the one artifact that could have collided the pictures, and nobody in the thread knew to look at it.

scn-001 · Double self-serve activation
A quarter that drifted, the plan it drifted from, and the day someone noticed. The cases are open source; thirteen is not enough.
03 · Result

Six measures, plainly

A left rule marks the five measures that clear the crude baseline. The sixth carries none.

When it speaks up, is it right?
80.0%
precision · baseline 75.0%

Four flags in five are real. A wrong flag mutes the tool faster than a missed one does.

Of the real contradictions, how many does it catch?
90.9%
recall · baseline 72.7%

Ten of eleven. Version 1.0 caught all eleven and was worse, by flagging early and often.

How much warning does it give?
9.5 days
median lead · baseline 5.5 days

Days ahead of the person who noticed unaided. Zero would mean it only confirms what people already see.

Does it know when to say nothing?
0
false flags on 2 trap cases

A properly recorded replacement, and a lively thread that never became work. Both correctly left alone.

Does it name the right contradiction?
60.0%
type accuracy · baseline 12.5%

Whether to re-staff, re-scope or replace depends on which kind of drift it is. Six of ten named it.

Does it reach the right people?
65.0%
routing · baseline 68.8% · below the baseline

Drift sits between two people holding half a picture each. Reaching only the half that already knows closes nothing. The weakest measure, unmoved since 1.0, and next.

04 · Baseline

Against the bar it had to beat

The bar is a deliberately crude detector: flag any anchored outcome with no visible activity for ten days. Anything more complex has to earn its complexity against it. Version 1.0 did not, and did not ship.

Detector versions against the ten-day silence rule
DetectorPrecisionRecallMedian leadType accuracy
Detector v1.1.1 — shipped80.0%90.9%9.5 d60.0%
Detector v1.0.0 — withheld45.5%100%13 d36.4%
Ten-day silence rule — the bar75.0%72.7%5.5 d12.5%

One change fixed it. A rule that fired on any hint of tension now requires evidence of a decision colliding with the anchored plan. Early flags fell from six to two, and precision moved from 45.5 per cent to 80.0 per cent.

05 · Limits

What this record cannot claim yet

Written, not lived
No real company produced these cases. Replaying a past quarter with the people who lived it is the validation that counts.
The days are too clean
Each case contains only signals that matter. A real day is mostly irrelevant traffic the detector never has to filter here.
One outcome at a time
Real quarters anchor several outcomes competing for the same people, where drift on one funds another. Unmeasured.
Routing is unproven
Flags reach the outcome’s owners, which is one side of the gap. Reaching the other side is designed and unbuilt.
Written cases cannot settle any of the four limits above. A pilot with a real company can. If that could be yours, say so.
Helm — Drift model evaluation record