The short answer
An AI visibility baseline is a dated set of buyer questions, engines, answers, mentions, citations, competitors, and accuracy notes collected with the same method every time.
Business outcome
You leave with a defensible before-state and a clean way to prove whether visibility, citation share, or answer quality improved.
The process
Build it in five passes
Choose one buying decision
Name the category, buyer, market, and decision you want to measure. A useful scope sounds like 'US operations leads choosing an employee scheduling tool,' not 'all prompts about our company.' Keep branded reputation questions in a separate set because they behave differently from category discovery.
Lock the prompt and engine set
Write 12 to 20 natural questions across discovery, comparison, objection, and purchase intent. Record the exact wording, engine, model or product surface, search mode, account state, and location. Do not edit prompts after seeing an answer. If a prompt changes, start a new version.
Capture more than mentions
For every answer, record whether the brand appeared, its position or context, the exact claim made, linked or named sources, named alternatives, and any material error. A mention without recommendation context is not the same outcome as being a top option with a citation.
Calculate a small scorecard
Report prompt coverage, recommendation share, citation share, accurate-answer rate, and competitor lead. Preserve the raw answer beside each coded field. A single blended score is convenient, but the component metrics tell you what to fix.
Schedule the next matched snapshot
Repeat with the same prompt set and engine settings after a meaningful publishing or distribution cycle. Note model changes and dates. Compare directional movement, not tiny differences that may be ordinary answer volatility.
Before it ships
Quality checklist
- One category, audience, market, and buying decision are written at the top.
- Every question has a stable ID and intent stage.
- Engine, surface, date, location, and search setting are recorded.
- Raw answer text or an approved archival reference is retained.
- Mentions, recommendations, citations, competitors, and factual errors use separate columns.
- The next measurement date and change window are named before publishing starts.
Copyable artifact
Baseline worksheet header
Paste this header into a spreadsheet. Use one row per prompt and engine combination.
snapshot_date,prompt_id,intent_stage,prompt_text,engine,product_surface,model_label,web_search_on,location,brand_mentioned,recommendation_position,recommendation_context,cited_url,cited_domain,answer_accurate,material_error,top_competitor,raw_answer_reference,reviewer_notes 2026-08-13,D01,discovery,"[exact buyer question]",[engine],[surface],[model label],yes,[market],yes,[top/middle/list/no],[why the brand appeared],[full URL],[domain],yes,[none],[competitor],[file or record ID],[notes]
Validation
How you know it is ready
- 01A second reviewer can code five sample answers and reach the same result.
- 02Every summary number can be traced to raw rows, not copied from a slide.
- 03The rerun can be executed without guessing any prompt, setting, or scoring rule.
Do not overclaim
AI answers are volatile and may vary by account, location, model, source index, and wording. A baseline is a controlled snapshot, not a census of everything every buyer will see.
Questions
What teams usually ask
How many prompts belong in a baseline?
Start with 12 to 20 high-value questions for one buying decision. Depth and repeatability matter more than a large synthetic prompt list.
Should I run the same prompt several times?
Repetition can reveal volatility, but an early pilot should first establish one consistent run per prompt and engine. If you add repetitions later, label them and keep the method stable.
What is the best headline metric?
Recommendation share is useful for executives, but pair it with citation share and answer accuracy. That prevents a favorable but incorrect mention from looking like a win.
Sources reviewed