Data mining · sequence and retention

retail-sequence-mining

A large lift can be mathematically real in the history and still be the wrong reason to act. Counts, intervals, and causal limits stay beside every association.

Public · synthetic demo
The problem

Purchase order matters, but an association is not a cause.

Snapshot baskets miss sequence. Bare percentages hide denominator risk. This pipeline mines ordered patterns, measures retention, and keeps correlation separate from explanation.

Input

A deterministic synthetic transaction history.

Customers: 600
Transactions: 4992
Observation cutoff: day 400
Inactivity window: 90 days
Minimum support: 30 customers
Churned customers: 206
The money shot

The planted A → B pattern appears with its denominator and uncertainty.

Frequent patterns
155
Churn associations
156
A → B support
427
A → B lift
1.73×

Customers with A → B churned at 39.1% with a 95% Wilson interval of 34.6%–43.8%. Customers without it churned at 22.5%, interval 17.0%–29.3%.

How it's verified

Thin groups are omitted and surviving rates keep confidence bands.

TraitSupportChurn with traitLift
frequency bucket = low27563.6%6.67×
monetary bucket = low20070.5%4.34×
sequence A → B42739.1%1.73×
Honest limitations

The strongest-looking rows are not causal findings.

  • The data and A → B signal are synthetic.
  • Low frequency and low monetary value partly overlap mechanically with a 90-day inactivity definition.
  • The miner has no timing-gap constraint and is capped at short patterns.
  • Wilson intervals show sampling uncertainty, not unmeasured confounding. A controlled experiment is still required.
The report recommends an experiment before anyone treats a historical pattern as a lever. jigonyoo.com · Back to hub