Buyer's guide

Media planning tools: what to compare before you buy

Demos show the interface. The arithmetic underneath decides whether your plan survives a sales house negotiation. Ask about the arithmetic.

By the Tilsim team · · 7 min read

Media Planning Tools: What to Compare Before You Buy

Media planning tools convert a rating target into a budget a buyer can defend, then back again. A sales house argues with that conversion, never with the dashboard around it. So compare tools on three things: reproducibility (same inputs, same numbers, every time), transparency (every figure decomposes into named steps), and local market mechanics (whether the tool encodes how deals are priced where you buy). Everything below fits in a 45-minute call.

Will the tool give you the same answer twice?

Reopen a saved plan and the numbers should be identical. If they drift, what you have is an opinion with a timestamp. Three questions settle it.

Is the calculation core frozen, and how is that enforced? Frozen means the math cannot move unless someone changes it deliberately and re-proves parity. Ask for the test count, then ask what those tests assert. Ours sits behind roughly 750 backend and 2,100 frontend tests — close to 2,850 — checked against an agency's own Excel model to under 0.01% deviation. The reference matters more than the count. A tool tested only against itself proves nothing.

What happens when an iterative calculation lands between two steps? Budget and discount depend on each other: a bigger declared budget unlocks a deeper discount tier, which changes the budget. Iteration solves the loop, and sometimes it oscillates between two tiers. No answer here is correct, only documented. We take the larger budget, the conservative side for the buyer — over-reserving beats finding a shortfall at reconciliation. Ask what the vendor does. "It converges" is not an answer.

Does the tool ever silently substitute a default? The most underrated line on a procurement checklist. When a required input goes missing and the system quietly assumes zero, or last month's value, or an industry average, nobody finds out until an audit. We went the other way: a missing input halts the calculation and names the field. What you see is what was computed.

Can you decompose any number on the screen?

Pick one line of the output and ask the vendor to walk it back to inputs, out loud, without opening the code. A cost-per-point should come apart into a chain you can restate in a sentence.

CPP = base_cpp30 × (1 + seasonality) × (1 − combined discount) × (1 + uplift)

Four factors, in that order, and the order matters. Push the same way on ratings: weighted GRPs are weighted TRPs divided by affinity over 100. An affinity of 200 means your target watches twice as densely as the average person, so a target rating point buys half the general-population equivalent. Spot length too — the 30-second spot is the unit at coefficient 1.0, shorter formats weigh less. A vendor who cannot produce these chains in conversation is handing you an export nobody can audit, and the first skeptical CFO wins that argument.

One caution on transparency claims, ours included: a CPP inside a planning tool is a proxy until you feed it your own negotiated rate card. Right structure for comparing scenarios; no statement about what the campaign will cost or what it will return. Any vendor promising ROI out of a planning model is selling a number they cannot source.

Where do generic platforms quietly break?

International platforms tend to be strong on inventory and workflow, thin on the arithmetic of the market you actually buy in. Four checks catch most of it.

  1. Do sales house discounts multiply or add? In many markets they multiply: 20% and 10% combine as 1 − (0.8 × 0.9) = 28%, not 30%. Two points on a large annual deal is no rounding error, and a tool that adds them flatters the buyer on paper and the seller at signature.
  2. Is off-prime a share of money or a share of ratings? It is normally negotiated as a share of budget, then converted into a share of ratings using affinity and the off-prime discount. A tool that asks for off-prime as a rating share is answering a different question from the one your contract asks.
  3. What does the prime-time uplift field actually touch? In the deal structures we work with it applies to the whole deal, not just prime inventory. Applied narrowly, every plan the tool produces comes out under-costed.
  4. Are declared and executed budgets tracked separately? The Shop List — the budget you declare to the seller, which sets your discount tier — and the executed budget are two money layers. One budget field cannot describe the deal you signed.

What file is the tool actually reading?

For anything touching audience measurement, ask what the tool ingests. A real gap sits between the raw delivery files from the measurement provider and the text export most tools read. That export carries age as a band, 1 through 5. No 18–49 or 25–54 break can be assembled from bands, because the boundaries fall inside them. Blame the export format, not the measurement: the raw delivery carries exact ages, 4 through 65. That is why TV Planner reads the raw binary deliveries (EVS/RDS/RSP).

Then ask for proof of accuracy, and define proof before they answer. Not "matches the industry standard" — a cell-level reconciliation. Ours: an official audience table reproduced cell by cell, 60,858 cells, all 21,548 non-zero cells matching within 0.5 of a person (the source rounds to whole people), across 17 of 17 demographic breaks. Nielsen is the measurement authority here and our partner. We reproduce Nielsen's numbers — that is a parity claim about our arithmetic, not an endorsement or accreditation. A vendor who blurs that line is showing you how it handles its other claims.

Two more data questions, worth thirty seconds each. How is multi-day aggregation computed? Averaging daily percentages breaks when universes differ; it has to be weighted by universe. What happens when a corrected day arrives? With SHA-256-based idempotent ingestion, a repeated delivery is skipped and a corrected one replaces the old rows outright — no silent duplicates in the quarterly report.

Is the post-buy module real analysis or a spot log?

Ask what statuses a delivered spot can receive. "Aired / not aired" means you are buying a log. Each spot should land in exactly one mutually exclusive bucket; ours uses seven: matched, out of flight, unplanned channel, out of week, out of daypart, wrong length, no lines. That taxonomy turns reconciliation with the seller into a conversation instead of an argument.

Ask one more thing: are client-facing reports frozen snapshots on revocable links? A report that changes after you send it is a liability.

Which of this ships today?

Close every demo bluntly: which of what you just showed me ships, and which is roadmap? For us, forecasting and autonomous plan generation are roadmap — not shipped features, and we say so before anyone asks. Flight shaping ships: 17 flight templates (Flat, Front-Loaded, Mid-Peak and Crescendo among them) and 5 intensity levels from 0.5 to 1.7, with the logic grounded in published work — Broadbent's adstock decay of roughly 2.5 weeks for FMCG, Jones's STAS, Binet and Field. Ask any vendor to name its sources for anything shaped automatically.

Where we fit. TV Budgeting runs the two-way TRP-to-budget and budget-to-TRP calculation on the frozen core above; TV Planner reads raw measurement deliveries and handles the post-buy. Our audience data covers one market today, Moldova, and expands from there. If this checklist is the conversation you want, we would rather have it against your own reference file than our demo data.

FAQ

What is the fastest way to test a media planning tool?

Bring one finished plan you already trust and ask the vendor to reproduce it, then compare line by line. Deviations under 0.1% are usually rounding; anything larger is a modeling difference you need explained before you sign.

Does a bigger test suite mean a better tool?

Only if the tests assert against an external reference. Thousands of tests confirming that a tool agrees with itself prove consistency, not correctness. Ask what the tests compare to.

Why can't I get an 18–49 break from a tool that reads standard exports?

Because the export stores age as bands 1–5, and industry breaks cut through the middle of them. It is a property of the export format, not the measurement — the raw delivery carries exact ages, 4 through 65.

How should roadmap features be scored during procurement?

At zero. Buy on what ships today and read the roadmap as an indicator of direction — and of how candid the vendor is willing to be with you.

Bring your own reference plan

Run this checklist against TV Budgeting with a plan you already trust — frozen core, multiplied discounts, tier iteration resolved and shown.

Request a demo →