The short answer
A credible AEO study predefines its question, target set, prompts, engines, settings, dates, repetitions, coding rules, exclusions, analysis, and limitations before collecting results.
Business outcome
You produce evidence worth citing and discussing because a skeptical reader can inspect how the result was made and where it stops being true.
The process
Build it in five passes
Choose a decision-shaped question
Ask something bounded that matters to a buyer or operator, such as which sources appear in local provider recommendations or how often engines agree on a top option. Define the target population or selected cases and explain why that scope is useful.
Freeze the collection protocol
Write exact prompts, engine products, model labels when visible, search settings, account state, location, collection dates, repetition count, and stopping rule before collecting answers. Record any paid-request count and approved budget separately when applicable.
Predefine coding
Create rules for mentions, recommendation levels, positions, citations, claim accuracy, duplicates, refusals, partial failures, and ambiguous answers. Pilot a small sample to clarify the rubric, then freeze it before the main analysis.
Preserve observations and calculate
Store raw answers or approved references, normalized rows, reviewer decisions, and analysis formulas. Keep excluded rows with reasons. Use simple descriptive statistics unless the design supports stronger inference, and avoid causal language for observational snapshots.
Publish method, data, and limits
Lead with the finding and its practical meaning, then give dates, sample or target selection, protocol, uncertainty, failures, and limitations. Publish a useful summary dataset and data dictionary when rights and privacy allow. Invite correction and make version changes visible.
Before it ships
Quality checklist
- The question, scope, and selection method are written before collection.
- Prompts, products, settings, dates, repetitions, and stopping rules are frozen.
- Coding rules cover ambiguous, failed, and duplicate observations.
- Raw evidence, normalized data, exclusions, and formulas are preserved.
- The headline avoids causal or population claims the design cannot support.
- Method, limitations, correction path, and reusable evidence are published.
Copyable artifact
Study protocol preregistration
Complete and timestamp this before collection. Any later change belongs in a deviations log.
STUDY TITLE: [working title] PRIMARY QUESTION: [bounded question] DECISION VALUE: [who can use the answer and how] TARGET SET / POPULATION: [definition] SELECTION METHOD: [how cases are included] COLLECTION WINDOW: [dates and time zone] PROMPTS: [stable IDs and exact text] ENGINES / SURFACES: [exact products] MODEL LABELS: [as displayed] WEB SEARCH: [on/off/automatic] ACCOUNT / LOCATION STATE: [details] REPETITIONS: [count] STOPPING RULE: [rule] PRIMARY METRICS: [definitions] CODING RUBRIC: [mention, recommendation, citation, accuracy] FAILURE RULES: [timeout, refusal, partial answer] EXCLUSIONS: [predefined conditions] ANALYSIS: [formulas or code plan] DATA RELEASE: [raw, normalized, summary, none + reason] KNOWN LIMITATIONS: [list] CORRECTION CONTACT: [email] DEVIATIONS LOG: [initially none]
Validation
How you know it is ready
- 01Another researcher can reproduce the collection settings from the protocol.
- 02Every headline table reconciles to preserved observation rows and formulas.
- 03A skeptical reviewer can identify both the useful finding and its boundary.
Do not overclaim
AI products are dynamic and often personalized. A dated experiment can reveal a real pattern without representing every user, future model, or market. Treat volatility as part of the result.
Questions
What teams usually ask
How large should an AEO study be?
Large enough to answer one bounded question under a consistent protocol. A small transparent study can be more useful than a large opaque scrape.
Do repeated model runs make the result scientific?
No. Repetitions can describe variability, but selection, coding, settings, and claim scope still determine what the study supports.
Should every raw answer be published?
Not automatically. Consider provider terms, privacy, licensing, sensitive content, and reader utility. Publish enough structured evidence to support the result and explain any release limits.
Sources reviewed