← All playbooks·Cross-media measurement
Playbooks · Cross-media measurement
Match audience identifiers across platforms and vendors
This job produces a validated cross-vendor ID match rate and a corrected unduplicated reach number for the plan. It feeds the decision on whether the naive summed reach across platforms can be used or must be replaced by the overlap-adjusted figure.
What you need first
- Raw ID export per vendor/platform for the campaign period, tagged by identifier type (STB, MAID, cookie, hashed email, RampID) — from each vendor's delivery or audience report
- Identity resolution vendor's accepted match-key documentation (e.g. LiveRamp, ID5, UID2) stating which ID types it resolves and how
- Clean room access (Habu, InfoSum, AWS Clean Rooms) or equivalent match report from a prior campaign for benchmarking
- Total addressable universe size for the target demo, from the vendor's or panel's reach documentation
- Stated overlap methodology from the identity vendor — deterministic, probabilistic, or hybrid
The procedure
- Pull raw ID counts per vendor/platform for the campaign period, tagging each set with its identifier type
- Normalize formats, strip nulls and test/QA IDs, and dedupe within each vendor set, producing a clean count per vendor
- Load the cleaned sets into the clean room and run resolution against the common identity graph, producing a matched-ID count
- Calculate match rate as matched IDs divided by the smaller of the two sets, producing a match rate for the pair
- Compare the match rate against the identifier-type benchmark range and flag any pair that falls outside it
- Compute unduplicated reach as set A plus set B minus matched IDs, divide by universe size, producing the corrected reach percentage
- Document the ID types, methodology, and match rate in the plan appendix
Worked through with numbers
Netherlands, adults 18-54, universe 8,000,000. TV set (linear STB + BVOD, deduped) = 3,600,000 IDs, claimed reach 45%. Digital set (CTV + mobile programmatic, MAID/cookie, deduped) = 2,800,000 IDs, claimed reach 35%. Naive summed reach: 45% + 35% = 80%, which double-counts anyone reached on both. Both sets go into the clean room against the RampID graph: matched IDs = 840,000. Match rate = 840,000 / 2,800,000 (smaller set) = 30%, which sits inside the expected 20-40% range for a MAID/cookie-to-RampID match. Unduplicated reach = 3,600,000 + 2,800,000 - 840,000 = 5,560,000 IDs, divided by 8,000,000 universe = 69.5%. Read it as: true reach lands well below the naive 80% sum but above either platform alone, and the 30% match rate tells you the correction is trustworthy rather than an artifact of a thin match.
Where it goes wrong
- Dividing the matched count by the larger set instead of the smaller one inflates the apparent match rate; use the reference set the clean room specifies
- Benchmarking a cookie-based match against deterministic email or RampID thresholds; benchmark by identifier type, since a 25% cookie match and a 25% email match mean different things
- Summing platform-level reach percentages without subtracting the matched overlap, which double-counts the shared audience
- Accepting a matched-ID count without checking for expired or rotated device IDs on the vendor side, which can inflate matches on paper without inflating real reach
How to know it is right
The corrected unduplicated reach should land between the higher of the two individual platform reach figures and their naive sum; if it falls outside that range, recheck the match rate before sending the number on.
Terms used