I'm currently seeking new Product Management opportunities.

How One CAT DILR Set Led Me to a Billion Dollar Retail Analytics Problem

Two Years Ago, CAT Gave Me the Table. Today, I’m Building the System That Creates It.

Illustrative retail analytics dashboard showing zone engagement, dwell time and visitor metrics
Retail analytics solution glimpse

It Started With This CAT Question

Two years ago, while preparing for CAT, I came across the DILR set below. I solved it, moved on, and did not think there was anything particularly special about it.

One question from the set became strangely relevant to a real problem I worked on almost two years later.

“The number of people who left the temple within 2 hours of entering is at least?”

CAT DILR entry and exit question used as the starting point for the retail analytics article
Figure 1 Original CAT DILR entry and exit set. Source credit is visible in the image.

The Only CAT Question That Matters for This Story

See Q2 in the image above. The table gives aggregated entries and exits, but it does not tell us which person entering in one interval corresponds to which person exiting later. The task is to infer the minimum number whose stay was within two hours.

My handwritten working arrived at 449 people as the minimum short dwell cohort.

Handwritten working for the CAT DILR minimum short dwell cohort question
Figure 2 My handwritten working for Q2 based on entry, time spent and exit.
The connection
The number 449 is not the part that stayed with me. The structure did. ENTRY → TIME SPENT → EXIT. The CAT problem was asking me to reason about a dwell time cohort, even though I would never have called it that at the time.

Almost Two Years Later the Same Structure Reappeared

This time there was no temple. There was a physical retail store. The questions were no longer exam questions.

  • How many customers entered the store?
  • How long did each visitor stay?
  • Where inside the store did they spend that time?
  • Which visitors were barely passing through and which were meaningfully engaged?
  • Which journeys showed stronger purchase intent?
  • Which visits ultimately converted?

The structure was familiar. ENTRY → BEHAVIOUR + TIME SPENT → EXIT

The key difference
CAT already gave me the table. Retail did not. In the real problem we first had to build the system that could generate the table from camera footage.

The Real Retail Problem

Digital commerce is naturally instrumented. A website can log page views, searches, clicks, add to cart events, checkout starts and purchases. A physical store usually sees only fragments of that journey.

Imagine 1,000 people enter a store and 120 purchases are recorded. The billing system tells us about the 120 transactions. A footfall counter may tell us that 1,000 people entered. But the middle of the journey is mostly invisible.

  • Did a visitor leave after 40 seconds or stay for 25 minutes?
  • Which section did they explore?
  • Did they return to the same category?
  • Which zones attracted traffic but failed to hold attention?
  • Which visitors showed engagement without converting?

The product opportunity is to turn physical movement into structured behavioural events.

What We Actually Built

The solution was not one model and it was not one dashboard. It was a pipeline that converted raw camera footage into anonymous visitor journeys and then into retail metrics.

The working pipeline
Camera Frames → Person Detection → Tracking and ReID → Anonymous Visitor Session → Zone Events → Entry and Exit Timestamps → Dwell Metrics → Engagement Signals → Retail Analytics

1. Turn Video Into Observations

Suppose a visitor is inside the store for 18 minutes. At 30 frames per second that period contains:

30 × 60 × 18 = 32,400 frames

One visitor can appear in thousands of frames. The first job is therefore not to count frames. It is to detect people inside those frames.

2. Detect People With YOLO

A person detection model such as YOLO can detect where people are present in a frame. It converts pixels into machine readable observations such as person at position A and person at position B.

In short, YOLO tells us there is a person here. It does not tell us whether that person is the same visitor we saw earlier.

Detection is not identity
Thousands of detections do not mean thousands of customers.

3. Preserve Anonymous Visitor Identity

Now imagine a visitor walks behind a shelf and disappears for a few seconds. A track may end and a new track may begin when the person becomes visible again. If the system treats both tracks as different visitors, footfall inflates and dwell time breaks.

ReID helps determine whether a person seen now is the same anonymous person seen earlier, even after a temporary disappearance.

Track 47 10:02:14 to 10:08:31

↓

Temporary loss

↓

Track 83 10:08:39 to 10:19:47

↓

Both tracks are associated with Visitor V0183

The product does not need to know the visitor’s real world name for this use case. It only needs enough temporary identity to understand that multiple observations belong to the same anonymous visit.

4. Create a Visitor Session

Once detections can be associated with the same anonymous visitor, the system can create a session object instead of treating every frame independently.

visitor_session_id  V0183
first_seen          10:02:14
last_seen           10:19:47
status              completed

This session becomes the container for everything that happens during the visit. Entry, exit, zone movement and dwell can all be attached to the same anonymous visitor.

5. Map the Store Into Business Zones

Tracking coordinates alone is not useful to a retailer. A retailer cares that a visitor moved from the entrance to the eyewear section, not that a bounding box moved from one pixel coordinate to another.

Zone Meaning
AEntrance
BEyeglasses
CSunglasses
DAccessories
ECheckout
FExit

Once these areas are configured, movement can be translated into events the product understands.

6. Turn Movement Into Events

For Visitor V0183 the session might generate this event history.

Event Timestamp
Entered store10:02:14
Entered Eyeglasses10:04:11
Left Eyeglasses10:11:36
Entered Sunglasses10:12:05
Left Sunglasses10:16:22
Reached Checkout10:17:10
Exited store10:19:47

Underneath the table, the product can store the same information as structured events.

visitor_id  V0183
event       ZONE_ENTER
zone        EYEGLASSES
timestamp   10:04:11

visitor_id  V0183
event       ZONE_EXIT
zone        EYEGLASSES
timestamp   10:11:36

This is the important transition. Video has now become event data.

7. Calculate Dwell Time

Visitor V0183 enters at 10:02:14 and exits at 10:19:47.

Dwell Time = Exit Timestamp − Entry Timestamp = 17 min 33 sec

The same logic works inside each zone.

Metric Value
Total store dwell17m 33s
Eyeglasses dwell7m 25s
Sunglasses dwell4m 17s
Eyeglasses share of visit≈ 42.3%

That changes the interpretation from the customer stayed for 17 minutes to the customer stayed for 17 minutes and spent roughly 42 percent of the visit in one category.

8. Handle Revisits Instead of Hiding Them

A visitor may enter the same zone more than once. Treating the whole visit as one continuous block would lose useful behaviour.

For example:

Eyeglasses visit 1   7m 25s
Eyeglasses visit 2   4m 10s

Total zone dwell     11m 35s
Zone visit count     2

This lets the system distinguish someone who casually passed through a zone from someone who returned to it repeatedly.

9. Build Engagement From Multiple Signals

Dwell time alone should not decide whether a visitor is valuable. A stronger view combines several behavioural signals.

Signal Example
Total store dwell17m 33s
Relevant zone dwell11m 42s
Zones explored2
Zone revisits1
Checkout reachedYes

These signals can feed retailer defined engagement or prospect logic.

Behaviour is a signal
Dwell Time + Zone Dwell + Revisits + Journey Pattern → Engagement and Intent Signals

Long dwell does not equal conversion. Someone can stay 40 minutes and buy nothing. Another visitor can buy in 12 minutes. Conversion needs an actual business outcome such as a linked checkout or transaction.

From Individual Journeys to Retail Analytics

One visitor gives us a journey. Thousands of visitor sessions create the metrics a retailer can use across a store.

  • Footfall
  • Average dwell time
  • Zone dwell time
  • Zone concentration
  • Peak occupancy
  • Engagement cohorts
  • High intent signals
  • Conversion rate when purchase data is connected

Illustrative in store engagement dashboard with zone analytics, footfall, dwell time and high intent visitor metrics
Figure 3 Illustrative in store engagement dashboard showing how anonymous visitor journeys become retail metrics.

What Decisions Can the Retailer Make

The useful question is not whether the product can draw a heatmap. The useful question is whether the data changes a decision.

Signal Possible action
High traffic and high dwell Protect a strong engagement zone and test strategic product placement
High traffic and low dwell Test assortment, merchandising, signage or whether the area is only a transit path
Low traffic in an important category Test discoverability, layout and relocation
Peak occupancy by hour Adjust staffing, queues and customer to staff ratios
Layout change followed by a dwell shift Measure whether the redesign changed behaviour

This is where physical retail begins to look more like digital product analytics. Instrument behaviour, test a hypothesis, change the experience and measure whether behaviour moves.

The Dashboard Looks Simple but the System Is Not

The physical world creates edge cases that a clean dashboard hides.

  • Two people overlap and one temporarily blocks the other
  • A visitor leaves the camera view and later returns
  • A visitor exits the store and re enters
  • Employees remain in the store for hours and should not inflate customer metrics
  • Families and groups may need different logic from individuals
  • A person stands on the boundary between two zones
  • Camera position or lighting changes affect detection quality
  • Crowding creates more occlusion and harder identity association

Errors also compound.

Detection Error

↓

Tracking and Identity Error

↓

Journey Error

↓

Metric Error

↓

Bad Business Insight

Privacy Is Part of the Product

The system does not need a real world name to measure an anonymous visit. It only needs enough temporary identity to connect observations within a session. Data minimisation, retention, access control and consent requirements therefore belong in the product design.

The PM Layer

A computer vision team may think about detection precision, embeddings, tracking, inference speed and occlusion handling. A retailer thinks about conversion, merchandising, staffing, store layout and customer engagement. Product work connects those two languages.

The product chain
Camera Frames → Person Detection → Tracking and ReID → Anonymous Visitor Session → Zone Events → Dwell Metrics → Engagement Signals → Conversion Outcome → Retail Decision

Back to the CAT Table

The CAT table is still useful because entry and exit counts can produce aggregate metrics such as footfall and occupancy.

Current Occupancy = Previous Occupancy + Entries − Exits

But the CAT table cannot tell us which exit belongs to which earlier entry. That missing identity is exactly why the exam asks for minimum and maximum possibilities.

Entry ⇄ Person ⇄ Exit

In CAT that relationship has been removed, so we infer what might have happened. In the retail system, tracking and ReID try to preserve that relationship before aggregation.

Why the CAT reference matters
CAT gave me aggregated entry and exit counts and asked me to reconstruct possible journeys. The retail system starts earlier with camera footage and tries to preserve the anonymous journey first. Camera → Detection → ReID → Visitor Session → Entry and Exit → Dwell → Aggregates.

Two years ago, the table was the input. This time, building the system that creates the table was the actual problem.