How to build a prompt bank for AI visibility measurement
“We want to measure whether AI recommends us. Where do the questions come from?” The prompt bank is the instrument. If the questions are wrong, every number downstream is wrong with perfect execution behind it. This is how a bank gets built so the numbers mean something.
Start from themes, not from the brand
Nobody asks an assistant for a category. They ask about a situation inside it. So the first step is listing the themes a real buyer brings, and the themes come from documented demand, not from a whiteboard:
- Search Console queries, if you have access
- Google autocomplete and People Also Ask for the category
- Forums and communities where the category gets discussed
- Review platforms, for the language of satisfaction and complaint
- Call center FAQs and internal site search
A sunscreen category, worked example, yields themes like: protecting children, sensitive skin, daily facial use, beach and sport, ingredient safety, price and value, where to buy. Seven themes, each from questions people actually type or ask.
Give every theme several real angles
One question per theme is a sample of one, and assistants answer different angles differently. Four angle families, worked on the same theme:
- Who is asking: “bloqueador para bebé de menos de seis meses”
- The situation: “cómo evito que mis hijos se quemen en la playa”
- How far along the decision is: “cuál bloqueador para niños es el mejor”
- The comparison being weighed: “bloqueador mineral o químico para niños”
An assistant answers those four in four different ways. They are not one question reworded, and the bank needs the spread.
Three layers, counted separately
The bank holds three layers, and each answers a different question:
- Unbranded. Questions that never name you. The only honest way to know whether you appear on your own, because “para saber si tu marca aparece por sí sola, la pregunta no puede mencionarla”.
- Brand-named. What happens when someone names you: what the model believes, which critique arrives attached, where it sends buyers, whether your subbrands get credited.
- Competitive. Who owns the conversation: head-to-heads, and the comparisons happening without you in them.
Each layer gets its own numbers, and they are never averaged. Mix them and every figure moves for reasons nobody can explain. A bank of 50 to 100 questions usually gives every layer cuts big enough to report.
The three failure modes of homemade banks
1. The questions name the brand. Name the brand and the model will happily discuss it; you measured your own question, not your presence. This contaminates exactly the number the unbranded layer exists to produce.
2. One question per theme. A sample of one per topic measures noise. Assistants answer angles differently, so the bank needs angle variety per theme.
3. The wrong competitor set. Models routinely treat a premium brand as interchangeable with the supermarket private label next to it on the shelf. If the bank only pits you against the rivals your sales team respects, the share numbers come out wrong in a way nobody notices.
Each failure corrupts a different number, silently. That is what makes banks worth building carefully.
The frozen-bank rule
Once the bank produces its first baseline, it freezes. New questions added later start their own baseline, and banks are never spliced into a trend. Run two, months later, asks the same questions, so the runs compare and “it worked” or “nothing moved” means something. Document the bank so any future edit is deliberate, not accidental.
DIY or done with you
Everything above is buildable in-house with research tools and discipline. If you want it built or built together, that is the service: themes with documented demand, angles validated against real phrasing, per-question origin annotations, and a coverage report: /banco-de-prompts/.
Next step: The prompt bank service