This framework defines how GEO outcomes are attributed and how return on investment is established without self-deception.
Attribution problem
GEO outcomes move for three reasons: your work, engine changes, and market drift. The framework’s job is to separate these. Citation-rate movement after a model update is not your achievement, and holding a baseline flat through a hostile update can be a success.
Control structure
- Fixed instruments: the prompt portfolio and annotation rules are frozen; changes are re-baselined.
- Version capture: every run records model versions; movements coinciding with version changes are flagged as model-driven.
- Competitor control: track competitors on the same instrument; if the whole market moves, the movement is market-driven.
- Staggered deployment: where possible, deploy changes to segments in sequence so effects are observable against the unchanged segment.
Metric hierarchy
Primary: citation rate across the prompt portfolio, per engine. Secondary: retrievability of key pages, entity-consistency score, cited-source quality, share of voice vs competitors in answers. Business: visits from AI answer channels, pipeline influenced by AI mentions — measured at the channel level, never claimed as full attribution.
ROI calculation
Apply the ROI calculator model with measured, not assumed, inputs: actual citation movements from the control structure, actual visit and conversion data from analytics. Report both the model and its sensitivity — an honest ROI statement shows what would happen if each assumption moved.
Reporting rhythm
Monthly measurement, quarterly attribution review, annual ROI statement. The attribution review explicitly re-baselines after model releases and portfolio changes. This rhythm is what GEO reporting implements, and the attribution service exists for teams that need deeper causal analysis.
What this framework refuses
It refuses to credit programmes for model-update luck, to average across engines with different behaviours, or to present unverifiable “visibility scores” as results. If an agency cannot produce raw runs for its reported numbers, it is not running this framework.