Market Basket Analysis for FMCG: How to Understand Shopper Behavior

TL;DR: Market basket analysis for FMCG identifies products purchased together in recorded transactions. Support measures how common a combination is, confidence measures a directional conditional purchase rate, and lift compares that rate with the second product’s baseline frequency. Together, these metrics can inform assortment and promotion hypotheses, but they do not establish causation, loyalty or guaranteed sales uplift. Reliable analysis starts with clean transaction IDs, consistent product mapping and a defined retail universe. Findings should then be checked across periods and stores, with promotions, seasonality and availability considered before any commercial action.

A receipt shows which products appeared in the same recorded purchase. It does not, by itself, explain why the shopper chose them. That distinction makes market basket analysis useful when its findings are treated as evidence for investigation rather than a shortcut to claims about consumer motivation.

This guide explains the calculations, a practical data preparation workflow and the limits of interpretation. All numerical examples are hypothetical, not TrendBox customer results. For the broader framework, read about FMCG data analytics.

What is market basket analysis?

Market basket analysis identifies patterns of co-purchase within transactions and expresses them as associations between products or product groups. It asks which combinations occur and how those combinations compare with baseline purchase frequencies.

An association rule such as A → B means that baskets containing A are examined for the presence of B. The arrow is a direction of calculation, not a statement that buying A causes someone to buy B. It does not establish the order in which items were selected either.

IBM’s explanation of the Apriori algorithm describes how frequent itemsets can be used to derive association rules. The algorithm is one possible method; the business value still depends on the quality of the transactions and the interpretation of the result.

Which transaction data is required?

The essential data is a reliable transaction identifier linked to the products recorded in that transaction. Store, time and product attributes make the findings easier to interpret and validate.

Useful fields for an FMCG basket dataset
Field Purpose
Transaction ID Groups product lines into a single basket without merging unrelated purchases.
Product or SKU ID Identifies the item consistently across records.
Store ID Enables store-level checks and prevents ID collisions between stores.
Transaction time Supports period, weekday and seasonal comparisons.
Quantity and value Allow separately defined quantity or commercial analyses.
Category and pack attributes Support meaningful grouping and comparisons between variants.
Promotion information Helps distinguish recurring associations from campaign-period patterns.

If receipt numbers repeat across stores or days, transaction ID alone may not be unique. Construct and document a reliable basket key. Otherwise, purchases that never happened together may be combined into one artificial basket.

Basic product association analysis does not require identifying a person. If customer-level data is used for a separate purpose, privacy, access and permitted use must be addressed independently. An anonymous basket pattern should not be presented as a finding about an identifiable household.

How should basket data be prepared?

Prepare basket data by defining the analysis universe, resolving product identities and applying consistent rules to invalid or repeated records. Every transformation should preserve what a transaction means.

  1. Specify the stores, channel, period and eligible transaction types.
  2. Confirm that the basket key is unique across the dataset.
  3. Remove duplicate records according to a documented rule.
  4. Handle voids, returns and cancelled transactions explicitly.
  5. Map product IDs and categories consistently over time.
  6. For binary basket metrics, count a product once per basket even when several units are purchased.
  7. Review reporting gaps and store participation before calculating comparisons.

Product grouping changes the question. An SKU-level analysis can examine specific pack combinations; a category-level analysis can examine broader co-purchase. Combining every variant into a category may hide a useful pattern, while keeping every minor variant separate may leave too few observations. Choose and record the level before reviewing attractive rules.

What is support?

Support is the proportion of eligible baskets containing an itemset. For a pair A and B, it is the number of baskets containing both products divided by the total number of eligible baskets.

Support describes prevalence, not the strength or cause of a relationship. A combination appearing frequently may simply involve two popular products. A rare combination may show a strong association but have limited commercial scale or too few observations for a stable conclusion.

A single product’s basket penetration is its share of eligible baskets; it is not the share of shoppers who bought it.

Some tools report a support count as well as a rate. Label them clearly: “40 baskets” and “4% of eligible baskets” are different representations. Also define the denominator. All-store baskets and baskets restricted to a particular mission or category do not produce directly comparable support values.

What is confidence?

Confidence for A → B is the proportion of baskets containing A that also contain B. It uses A baskets as the denominator, so reversing the rule can produce a different value.

A high confidence value may reflect B’s general popularity rather than a distinctive affinity. If B appears in most baskets, many A baskets will contain B even without an unusually strong association. Confidence should therefore be read alongside baseline frequency and lift.

It is also a historical descriptive rate for the selected dataset. It is not automatically a probability that a future shopper will buy B after seeing A, nor proof that displaying the products together will change purchasing.

What is lift?

Lift compares the observed co-purchase rate with the rate expected under independence. For A → B, it is confidence divided by the overall basket frequency of B.

IBM’s lift documentation explains this comparison. A lift above 1 indicates co-occurrence above the independence benchmark in the measured data; a lift below 1 indicates co-occurrence below it. A value of 1 is the benchmark itself, not proof that every possible relationship is absent.

High lift is not sufficient for action. A rule based on very few baskets can be unstable. Store mix, promotions or restricted availability can also shape the result. Read lift with observation counts and check whether the pattern persists in relevant segments.

How are the basket metrics calculated?

Support, confidence and lift can be calculated from the same transaction counts, provided the universe and product definitions stay consistent. The table uses binary product presence within each basket.

Association rule formulas for A → B
Metric Formula
Support of A and B Baskets containing both A and B / all eligible baskets.
Confidence of A → B Baskets containing both A and B / baskets containing A.
Baseline frequency of B Baskets containing B / all eligible baskets.
Lift of A → B Confidence of A → B / baseline frequency of B.

Support and confidence can be shown as percentages. Lift is a ratio, not a percentage-point gain. If software reports additional metrics, confirm their definitions rather than assuming similarly named fields use the same calculation.

What does a worked basket example look like?

A worked example shows why directional confidence and baseline-adjusted lift answer different questions. The following counts are educational figures, not observed brand performance.

  • Total eligible baskets: 1,000.
  • Baskets containing product A: 200.
  • Baskets containing product B: 100.
  • Baskets containing both A and B: 40.

Pair support is 40 / 1,000 = 4%. Confidence for A → B is 40 / 200 = 20%. B’s baseline frequency is 100 / 1,000 = 10%, so lift is 20% / 10% = 2.

In this dataset, B occurs in A baskets at twice its overall basket frequency. That does not mean a promotion will double sales. When the rule is reversed, confidence for B → A is 40 / 100 = 40%. A’s baseline is 20%, so reverse lift is also 2.

The example illustrates a useful distinction: confidence is directional, while the pair’s lift is symmetric under these definitions. Neither result explains whether the combination reflects a shopping mission, a promotion, store assortment or another influence.

How can basket analysis inform assortment decisions?

Basket analysis can identify combinations worth investigating when reviewing assortment. It provides evidence about observed co-purchase, not a complete instruction to add or remove a product.

For example, a recurring combination may justify examining whether relevant stores stock both categories. A weak combination may instead reflect that the products were rarely available together. Comparing stores with different assortments without recognizing that constraint can create a misleading conclusion about shopper preference.

Connect the association with sales value, margin information where available, replenishment considerations and category role. An attractive statistical rule may have little operational relevance. Equally, a commercially important product might not appear in a strong pair rule because it serves varied missions.

How can basket analysis inform promotion hypotheses?

Basket analysis can suggest products to test in a promotion or cross-selling experiment. It does not establish that a bundle, display or recommendation will cause incremental purchases.

Write the hypothesis explicitly: which combination, which stores, which period and which outcome? Keep association metrics separate from the experiment’s success measures. A promotion may increase pair purchases while reducing margin or substituting purchases that would otherwise have occurred.

Compare the proposed action against a relevant control or comparison design where practical. Define the evaluation before the campaign begins. Avoid calling a post-promotion increase “incremental uplift” without a method that can reasonably distinguish the promotion from other changes.

How should findings be validated?

Validate findings by checking data quality, observation counts and stability across relevant periods or store segments. A rule that appears only under one unusual condition should not be generalized to the entire market.

  • Review whether the pattern exists outside promotion periods.
  • Compare equivalent seasonal windows rather than unrelated calendar periods.
  • Check store-level availability and assortment constraints.
  • Inspect whether a small number of stores drives the result.
  • Use a separate period to see whether the association persists.
  • Document rule-selection thresholds and avoid reporting only the most attractive findings.

A separate validation period reduces reliance on patterns selected from the same data used to discover them. Many candidate combinations can produce apparently impressive rules; the commercially useful question is whether a relevant and sufficiently supported pattern survives further checks.

What are the limits of interpreting shopper behavior?

Transaction associations describe recorded purchases, not the shopper’s intentions, reasons or long-term relationship with a brand. Those questions need additional evidence and an appropriately designed analysis.

One basket may cover several people, occasions or shopper missions. Repeated co-purchase does not automatically indicate loyalty, and the absence of an item does not prove rejection when it was unavailable. Survey or qualitative research can explore motivations, but its results should not be silently inferred from receipts.

TrendBox’s public analytics features and measurement services describe basket-related analysis in a broader commercial context. Combine those questions with channel conditions, including the limits of traditional trade measurement. Get in Touch to discuss the question and required data.

Frequently asked questions about FMCG basket analysis

These distinctions help keep a useful association finding from becoming an unsupported claim about causation or commercial impact.

Does high confidence mean a strong product affinity?

Not by itself. Confidence can be high because the second product is common overall. Compare it with baseline frequency, lift and observation counts.

What does a lift of 2 mean?

It means the observed pair frequency is twice the independence benchmark in the defined dataset. It does not mean a campaign will double sales.

Should repeated units count as repeated basket purchases?

Not for the binary presence metrics used here. A product is counted once per basket; quantity analysis is a separate, explicitly defined calculation.

Can basket analysis prove shopper loyalty?

No. Co-purchase in transactions alone does not establish loyalty or repeated customer behavior. Those questions require suitable longitudinal evidence and permitted data use.

Should every high-lift combination become a promotion?

No. Check support, stability, commercial relevance, availability and the planned evaluation. Treat the combination as a hypothesis to test, not a guaranteed outcome.