OVO field guide
How do I measure influencer marketing ROI?
OVO Measurement Contract
Version 1.0 · Updated 2026-08-21
Define what success means, what can be observed, and what can be claimed before creator work begins, so the final report answers the question the campaign was built to answer.
- Campaign objective
- Primary business action
- Purchase cycle
- Available event sources
- Launch and readout dates
- Name one decision
Write the decision this campaign must support and choose one primary result that would change that decision.
- Separate credit from cause
State whether the campaign needs attribution, causal lift, or both, then label each number by the question it can answer.
- Lock the instruments
Choose the event source, tags, codes, survey, comparison rule, and attribution window before any creator content goes live.
- Record the baseline
Capture a prelaunch period long enough to include the normal purchase cycle and document material changes during that period.
- Set the readout rule
Name the reporting date, calculation, known blind spots, and result owner before the team has seen campaign performance.
OutputA one-page measurement agreement that names the primary result, evidence source, comparison, window, limitations, reporting date, and decision owner.
Attribution and incrementality answer different questions
Attribution tells you which sales touched a creator. Incrementality tells you which sales the creator caused. The gap between those two numbers is not small, and in the evidence below it runs in the same direction every time: the observational estimate comes in larger than the experimental one.
The cleanest demonstration comes from eBay's economists. Blake, Nosko and Tadelis ran a large field experiment on paid search and reported that standard methods produced "a ROI of over 4,100% without time and geographic controls, and a ROI of over 1,400% with such controls," while their experimental method on the same spend found "a ROI of -63%, with a 95% confidence interval of [-124%, -3%], rejecting the hypothesis that the channel yields positive returns at all." The spend and the underlying data were the same in both cases. Only the measurement method changed.
The same pattern shows up at scale. Gordon, Moakler and Zettelmeyer analyzed 663 large-scale experiments at Facebook and found median true lift of 5% for lower-funnel outcomes, against 24% and 64% from the two standard observational methods applied to the same campaigns. Gordon and colleagues had more than 5,000 user-level features available and still concluded: "despite having access to large-scale experiments and rich user-level data, we are unable to reliably estimate an ad campaign's causal effect."
Influencer content is particularly exposed to this bias, because creators reach audiences already inclined toward the category. A fitness creator's followers were going to buy protein powder from somebody.
What each measurement method can and cannot tell you
No single instrument answers the whole question. The useful move is knowing which direction each one is wrong in, so you can read the numbers with the bias built in rather than treating any one of them as truth.
| Method | What it tells you | What it cannot tell you | Direction of error |
|---|---|---|---|
| Unique discount codes | Which creator a buyer heard from, and roughly when | Whether that buyer would have purchased anyway, or bought by a different route | Over-credits when redeemers were existing buyers; under-credits when the buyer forgot the code |
| Tracking links and UTMs | Session-level path from a creator's link to a purchase | Anything that happened on a different device, inside an app, or after cookie expiry | Under-credits, sometimes heavily, on mobile and long purchase cycles |
| Platform-reported conversions | What the platform credits to the inventory it sells | Whether the platform's model agrees with yours; it will not | Over-credits, since view-through windows count exposure as influence |
| Multi-touch attribution | How credit distributes across a known path | Anything unobserved, including offline sales and word of mouth | Inherits the bias of whatever data it sees; cannot separate correlation from cause |
| Post-purchase survey | What the buyer says prompted the purchase, in their words | What actually drove it, since recall is unreliable and self-report is noisy | Mixed; useful mainly as a directional cross-check on other instruments |
| Holdout lift test | Causal effect against a randomized control group | Anything at all, if you did not define the control group before launch | Unbiased in expectation, though often too wide to be decisive at small scale |
| Geo or matched-market test | Causal effect at market level without user-level tracking | Effects smaller than between-market noise, or anything in a single market | Unbiased if markets are well matched; sensitive to spillover between neighbors |
| Marketing mix modeling | Channel-level contribution across long time horizons | Anything about a single campaign or a single creator | Depends entirely on how much spend variation exists in the historical data |
Codes and links are attribution instruments. Holdouts, geo tests and mix models are incrementality instruments. Reporting an attribution number as ROI is the most common error in this category.
Instrument before launch, because most of this cannot be added later
The measurement decisions that matter are made before any content goes live. Meta's lift study documentation states the rule plainly: "Study start time must be in the future." No system lets you build a control group after a campaign has run, so a campaign that launches without one is stuck with correlational tools for its whole life.
Six things to have in place before the first post:
- A defined holdout or control geography, agreed and locked before launch. Meta's documentation also notes that once a study starts you cannot update its start time or treatment percentage.
- A clean pre-period baseline covering at least one full purchase cycle, captured before any creator content publishes.
- The post-purchase survey question live on your confirmation page before launch, so you have baseline answers to compare against. A survey that starts on launch day has nothing to compare to.
- One unique discount code per creator, at a value that exists nowhere else on your site or in email.
- An Amazon Attribution tag per creator if any traffic lands on Amazon. Amazon Ads describes it as "a free measurement solution" for professional sellers in Brand Registry, vendors, KDP authors, and agencies with clients selling on Amazon, and its reports cover a 14-day attribution window with new to brand among the conversion metrics.
- One named primary metric and a written measurement window, agreed before results start arriving rather than after.
How to run an incrementality test without buying one
Meta notes that "Conversion Lift Measurement is currently limited" and directs advertisers to contact a Meta representative for access, so most brands cannot simply switch it on. The practical alternative is a geographic test, and the tooling for it is free and open source.
Meta publishes GeoLift under an MIT license, described as an "end-to-end geo-experimental methodology based on Synthetic Control Methods used to measure the true incremental effect (Lift) of ad campaign." Google publishes CausalImpact under Apache 2.0 as "an R package for causal inference in time series." Both repositories were updated in 2026, and neither requires a vendor contract or user-level tracking.
A workable geo test needs enough markets to build a synthetic control, a pre-period long enough for the model to learn the baseline relationship between those markets, and treatment markets that do not spill heavily into control markets. Suppressing creator content geographically is the hard part, which is why paid amplification of creator posts makes clean holdouts easier than organic-only programs. Whitelisted content runs through ad delivery, and ad delivery can be geo-fenced.
Why your platform numbers will never match your store
Platform reporting and store data are built to disagree. Google Analytics 4 uses data-driven attribution by default, with a session conversion window where, per Google's documentation, "by default, it's 90 days." Meta applies different click and view-through windows, and a different credit model, to the inventory it sells. Two systems using different windows and different credit rules will never produce the same number from the same sales.
Browser behavior widens the gap. WebKit's engineering blog states that Safari's tracking prevention has capped the expiry of client-side cookies at seven days since February 2019, and that a matching seven-day limit was later extended to LocalStorage, IndexedDB and Service Worker registrations under the heading "7-Day Cap on All Script-Writeable Storage," clearing that data after seven days of Safari use without user interaction on the site. If your average purchase cycle runs longer than a week, link-based attribution built on JavaScript-set cookies is losing Safari buyers before they convert, and no amount of tag auditing recovers them.
Last-click reporting cuts against creator work specifically. Amazon's measurement researchers describe the problem this way: last-touch attribution "credits the full value of a conversion to one ad, usually the last one on the path to purchase, overlooking the impact of earlier touchpoints like awareness campaigns that build interest before shoppers are ready to buy." Influencer content sits mostly upper and mid funnel, so last-click undercounts it while discount codes over-count it.
Reconcile platform reports against store data once, early, to catch broken tags. Then pick one system as the record and stop relitigating the difference. Google's documentation notes that "more than 95% of key events get attributed within the first 14 days," which is a reasonable default window when the purchase cycle is short.
Engagement is not a proxy for sales
Engagement gets used as an ROI stand-in because it is easy to collect. The peer-reviewed evidence says it should not be.
Leung and colleagues, writing in the Journal of Marketing in 2022, estimated that "a 1% increase in influencer marketing spend increases engagement by .457%," and that the firms in their data set "could increase consumer engagement by 16.6% if they allocated their budgets proportional to these elasticities and base engagement levels, rather than their current allocations." The outcome variable in that study is engagement, not sales, so it does not tell you what influencer spend does to revenue.
A meta-analysis in the Journal of the Academy of Marketing Science, synthesizing 1,531 effect sizes from 251 papers, splits results into non-transactional outcomes such as attitude and purchase intention and transactional outcomes such as purchase behavior and sales. Different things drive each. For non-transactional outcomes it reports that "follower characteristics (social identity) have the strongest effects on consumer attitudes and behavioral engagement." For transactional outcomes it reports that "influencer characteristics (influencer communication) have the strongest effects on purchase behavior." A creator who maximizes one is not automatically maximizing the other. That is why an engagement number cannot stand in for a sales number.
The honest ceiling on what any measurement can resolve
Below a certain spend and sample size, single-campaign ROI is not measurable at useful precision, regardless of method or vendor.
The Gordon, Moakler and Zettelmeyer work is the clearest evidence of that ceiling. They had 663 randomized experiments and more than 5,000 user-level features to work with, which is more measurement infrastructure than most brands will ever assemble, and their conclusion was that they were "unable to reliably estimate an ad campaign's causal effect." A single creator flight measured with codes and links has a tiny fraction of that data behind it, so it will not resolve a precise return figure no matter how the report is formatted.
The practical response is to change the unit of measurement. Judge a portfolio of creator activity across quarters instead of demanding a verdict on each post. Codes and links are then operational tools, useful for deciding which creators to renew and which content to amplify. Incrementality tests belong to the bigger question of whether the channel earns its budget, which is worth answering a few times a year rather than continuously. Anyone promising a precise ROI figure on a small single-flight campaign is reporting attribution and calling it something else.
Numbers you will see quoted that do not hold up
Two figures dominate this topic, an "11x return" and "$6.50 for every $1 spent," and neither arrives with the things that would make it checkable. Neither is published with a sample size, a stated method, or a control group, and neither traces to a peer-reviewed study. They circulate as industry constants anyway.
These figures persist because the demand for them is real. The CMO Survey's 35th edition, fielded January 7 to 29, 2026 with 308 responding marketing leaders at U.S. for-profit companies, found that 58.8% of those answering the question reported increased CEO pressure to prove the value of marketing, and 55.5% reported the same from the CFO. A single memorable multiple travels well under that kind of pressure.
Ask three questions of any return multiple a vendor quotes you: what was the control group, what was the sample size, and who paid for the measurement. A number that cannot answer all three is a marketing asset rather than a result.
Frequently asked questions
How long should my influencer campaign measurement window be?
Fourteen days is a reasonable default for short purchase cycles. Google's Analytics 4 documentation states that "more than 95% of key events get attributed within the first 14 days," while the default conversion window for sessions in GA4 runs 90 days. If your product has a considered purchase cycle of several weeks, extend the window to cover at least one full cycle and hold it constant across campaigns so results stay comparable. Measuring only the posting day is the most common version of this mistake, and it makes every campaign look weaker than it was.
Do unique discount codes over-count or under-count creator sales?
Both, in different directions, which is why the net number is unreadable without a control group. Codes over-count when redeemers were existing customers who would have purchased anyway, and when a code leaks to coupon aggregators and gets used by people who never saw the creator. Codes under-count when buyers watch the content, buy later without the code, or purchase through a marketplace where the code does not apply. Two checks are worth running: search your codes on aggregator sites mid-flight, and compare redeemers against your new-customer rate. If most redeemers are existing customers, you funded a discount rather than an acquisition.
Can I run a holdout test on organic influencer posts?
Not cleanly at the user level, because you cannot control who sees an organic post. The practical route is a geographic test: hold creator activity out of a set of matched markets and compare sales against markets that received it. Meta's GeoLift package, published under an MIT license, is built for exactly this using synthetic control methods, and Google's CausalImpact does the same work in R under Apache 2.0. Both are free. The binding constraint is market count. With too few comparable markets the synthetic control has almost nothing to fit against, and the confidence interval ends up wider than the effect you are trying to detect.
What is a good ROI for influencer marketing?
There is no defensible benchmark figure. The multiples that circulate are single-campaign case studies published without a sample size, a stated method, or a control group. What actually determines your result is margin, purchase cycle length, category, whether the audience already knew your brand, and how much the creator's followers overlap with buyers you would have reached anyway. Precision is hard even with far better data than a brand will ever hold: Gordon, Moakler and Zettelmeyer had 663 randomized experiments and more than 5,000 user-level features available and still reported being "unable to reliably estimate an ad campaign's causal effect."
Do I need to buy a vendor tool to run an incrementality test?
No. The two most credible geo-experiment packages are free and open source: Meta publishes GeoLift under an MIT license for measuring incremental campaign effect via synthetic control methods, and Google publishes CausalImpact under Apache 2.0 for causal inference in time series. Platform-native lift studies exist as well, though Meta's documentation states that "Conversion Lift Measurement is currently limited" and requires contacting a Meta representative for access. What you genuinely need is analyst time and a clean pre-period, not a software contract.
Why do Meta and Google Analytics report different numbers for the same campaign?
Meta and Google Analytics disagree by design and will never reconcile. Google Analytics 4 uses data-driven attribution by default, with a session conversion window that Google's documentation sets at 90 days by default. Meta applies different click and view-through windows, and a different credit model, to the inventory it sells. Different windows and different credit rules produce different totals from identical sales. Reconcile them once early to catch genuinely broken tags, then designate one system as the record and stop spending analyst hours on the difference.
What should I instrument before a campaign launches rather than after?
Six things, and three of them cannot be added afterwards: the control group, the pre-period baseline, and the post-purchase survey. Meta's lift study documentation is explicit that "Study start time must be in the future," so a holdout cannot be constructed retroactively. Capture a baseline covering at least one full purchase cycle before content publishes, and put the post-purchase survey question live before launch so you have answers to compare against. The other three are easiest at the start but assignable later: one unique code per creator, Amazon Attribution tags if any traffic lands on Amazon, and a single primary metric plus measurement window agreed in writing before the first numbers arrive.
What mistakes make campaign results unreadable?
The worst one is to launch creators, paid social and an email promotion in the same week, because no model separates them afterwards. If you change the offer or the code value mid-flight, nothing before the change is comparable to anything after it. Judge results against last month instead of a control group and the baseline is moving under you, which is how eBay's economists got a reported return above 4,100% from data that a randomized test scored at negative 63%. The fourth is engagement used as a sales proxy: the Journal of Marketing estimate of 0.457% engagement growth per 1% spend increase measures engagement, not sales.
- NBER: Consumer Heterogeneity and Paid Search Effectiveness, A Large Scale Field Experiment (2014)
- arXiv: Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement (2022)
- Journal of Marketing: Influencer Marketing Effectiveness (2022)
- Journal of the Academy of Marketing Science: Influencer marketing effectiveness, a meta-analytic review (2024)
- Meta for Developers: Lift Studies, Marketing API (2026)
- Amazon Ads: Amazon Attribution (2026)
- Google Analytics Help: Comparing metrics, Google Analytics vs. Universal Analytics (2026)
- WebKit: Full Third-Party Cookie Blocking and More (2020)
- The CMO Survey: 35th Edition Topline Report (2026)
- arXiv: Amazon Ads Multi-Touch Attribution (2025)
- GitHub: facebookincubator/GeoLift (2026)
- GitHub: google/CausalImpact (2026)
Planning a creator campaign? See how the OVO team scopes and delivers it.
Plan a campaign with OVOReady now? Start a campaign inquiry
Inquiries are reviewed personally by the founding team, and every one gets a reply within 48 hours.