A ranked list that changes how reps spend their day is exactly the kind of feature that can feel great and do nothing, or worse. This is how I'd measure it: what winning looks like, what I'd refuse to break to get there, and the single number that decides whether it ships.
This plan measures the feature specified in the Next Best Visit PRD. Read that first for what the feature is; read this for how we'd know it worked.
If reps are given a ranked, explained list of who to visit each day, they will shift a meaningful share of their visits toward high-value and at-risk accounts, and that shift will show up as more opportunities advanced per week, without reducing how many visits they make in total.
Stated as a testable claim: the treatment group's share of visits to priority accounts rises versus control, and their opportunity-progression rate rises with it, while total visits logged holds flat.
Good experiment design starts by separating the one number that defines success from the numbers that keep success honest.
| Role | Metric | Why it's here |
|---|---|---|
| North Star | Priority-visit rate: share of a rep's weekly visits that go to high-value or at-risk accounts | This is the behavior the feature is designed to change. It captures "visiting better," which is the actual goal, rather than "visiting more," which is not. |
| Supporting | Opportunities advanced per rep per week | Confirms the behavior change produces business value, not just a different-looking calendar. This is the outcome the North Star is a proxy for. |
| Supporting | List engagement: share of days a rep acts on at least one listed account | Tells us whether the feature is used at all, which distinguishes "didn't work" from "wasn't adopted." |
| Guardrail | Total visits logged per rep per week | The feature must not make reps visit fewer accounts overall. If priority-visit rate rises only because total visits fell, that's a loss disguised as a win. |
| Guardrail | Rep retention / active usage of the app | A list that feels like surveillance could drive reps off the product entirely. We watch for that directly. |
| Guardrail | Override rate | Very high overrides mean the ranking is wrong or distrusted. A useful diagnostic even if the North Star moves. |
The tempting North Star is "visits logged," because it's easy to measure and always goes up when a feature nudges activity. But the feature's entire thesis is that which accounts get visited matters more than how many. Choosing a rate over a count is what stops the experiment from rewarding busywork. It's also the choice most likely to be argued about, which is exactly why it belongs in writing before the test starts.
| Decision | Choice | Reasoning |
|---|---|---|
| Type | Randomized controlled A/B, feature-flagged | Clean causal read on a feature that changes daily behavior. |
| Unit of randomization | The rep, not the account or the visit | The feature changes a rep's behavior across their whole book. Randomizing accounts would let treatment and control leak into the same rep's day. |
| Split | 50 / 50 treatment and control | Maximizes statistical power for a first read where we have no strong prior on effect size. |
| Exposure | Reps with an active book of accounts above a minimum size | The feature is meaningless for a rep with three accounts. We test where it can plausibly matter. |
Illustrative, to show the shape of the reasoning rather than to assert real numbers. Suppose priority-visit rate sits around 40% in control and we care about detecting a 5-percentage-point absolute lift, at 80% power and 95% confidence. That implies a few hundred reps per arm. Field sales behavior is also weekly and lumpy, so the test runs a minimum of four full weeks to average over week-to-week noise and to let reps move past the novelty of a new screen.
The headline result can hide opposite effects in different populations, so a few cuts are planned in advance rather than fished for afterward.
Written before results exist, so the read is honest.
| Outcome | Decision |
|---|---|
| Priority-visit rate up by the target lift, supporting metrics up, all guardrails held | Ship. Roll out to all reps, keep weights tunable. |
| North Star up but a guardrail breached (fewer total visits, or retention dip) | Iterate. The idea works but the execution costs too much. Fix the guardrail before shipping. |
| Engagement high but North Star flat | Iterate on the model. Reps use it but it isn't changing which accounts they pick. The ranking, not the feature, is wrong. |
| Engagement low and North Star flat | Reconsider. The feature wasn't adopted. Investigate why before investing further; a kill candidate. |
A win doesn't end the measurement. On rollout, the North Star and guardrails move into ongoing monitoring so a model that decays, or weights that stop fitting a changing book of accounts, get caught. The configurable weights from the PRD exist precisely so that tuning is a dial we can turn on live data, not a rebuild.
Not that I can run statistics, but that I can decide what "better" means before the data arrives, protect the things that shouldn't break to get there, and commit to a decision rule in writing so the result can't be rationalized after the fact. Choosing priority-visit rate over visit count is the whole argument in one decision. All figures here are illustrative.