This methodology defines how prompt portfolios are designed and citation tests are executed so the numbers produced are trustworthy, comparable and repeatable.

Why methodology matters

Prompt testing sounds simple — ask engines questions, see if you are named. In practice, results swing with prompt phrasing, model versions, temperature, session state and run order. Without methodology, the numbers are noise.

Step 1 — Portfolio design

Define 20-50 prompts from real buyer behaviour: sales and support transcripts, search data, and market knowledge. Classify each by intent (category research, comparison, solution, brand). Write prompts in the natural language buyers use, not in optimised query syntax. Freeze the portfolio.

Step 2 — Run conditions

Fix the conditions that affect answers: the engine and model version, response length, and the sequence in which prompts run. Record temperature and session behaviour where the engine exposes them. Run each prompt in fresh sessions to avoid conversation-context contamination.

Step 3 — Annotation rules

Define before measuring what counts as a citation for each engine: a visible source link, a named brand mention, or both. Decide how paraphrased mentions are classified. Two annotators should independently score a sample to check consistency, and disagreements must be resolved by the written rules, not by intuition.

Step 4 — Baseline and cadence

Capture a full baseline run, then re-run monthly under identical conditions. Record model versions in every run — a model update between runs must be logged and reported rather than hidden.

Step 5 — Analysis

Report citation rate per prompt family per engine, with the raw runs attached. Investigate movements: a change after a content or technical deployment is evidence; a change between model versions is information, not credit. Attribution requires the discipline of GEO attribution methodology, not a dashboard.

Step 6 — Evolution

Review the portfolio quarterly: retire prompts that no longer reflect buying behaviour, add new ones, and re-baseline when the portfolio changes. A living instrument beats a frozen one — as long as changes are controlled and documented.

Limitations

Prompt testing measures a sample of behaviour under fixed conditions. It cannot see everything an engine does, and engines change. What it provides is a comparable, auditable series — which is more than most reporting in this industry provides. The prompt optimisation service runs this methodology for clients; GEO reporting turns the series into management reporting.