Methodology · prompt_fit_rubric@0.2.0
How the Prompt Fit Score works
The Prompt Fit Score grades an AI visibility prompt set out of 100. It weights seven dimensions: buyer relevance (25%), format diversity, branded contamination and topic coverage (15% each), and market fit, long-tail depth and redundancy (10% each). It is deterministic: the same set and brief always get the same score, and no AI model is involved in scoring.
- Scale
- 0 to 100
- Dimensions
- 7, weighted
- Set size
- 3 to 2,000 prompts
- AI in the score
- None
The question it answers
Does this prompt set measure how your buyers actually ask AI assistants about your category? AI visibility trackers such as Peec, Profound and Semrush run a fixed list of prompts against ChatGPT, Perplexity, Gemini and others, and report how often your brand appears. Every number they report is an average over that list. RateMyPrompts grades the list itself, before anyone trusts the numbers it produces.
It does not measure how visible your brand is. That is the tracker's job. It measures whether the tracker is being asked the right questions.
The seven dimensions
Each scores 0 to 10. Several are measured over the measurable core: unbranded prompts that are about your category rather than a neighbouring one, because variety about the wrong subject is not variety.
| Dimension | Weight | Asks | Scores highest when | Scores lowest when |
|---|---|---|---|---|
| D1 ICP / commercial relevance | 25% | Are these the questions your buyers ask about your category? | 45% or more of prompts are unbranded, on-category questions. | Capped at 2 when a quarter of prompts use SEO or GEO trade language, and at 3 when 40% or more name your brand. |
| D2 Format diversity | 15% | Do the prompts ask in the different ways buyers ask? | Five or more effective formats among unbranded prompts: rankings, comparisons, conversational questions, personas and so on. | A set that is nearly all one shape counts as one format, however many stragglers it has. |
| D3 Branded contamination | 15% | How much of the set names your own brand? | Under 5% of prompts name the brand. | 40% or more branded scores 2 out of 10. |
| D4 Taxonomy / node coverage | 15% | Does the set cover several topics, each with enough prompts to read? | Five or more topics with at least two prompts each, and none above 40% of the set. | One topic holding 70% or more scores 2. |
| D5 Geographic / language fit | 10% | Do the prompts fit the markets you track? | At least 10% of prompts carry a place or currency from your declared markets. | 20% or more pointing at undeclared markets scores 4. |
| D6 Stress / long-tail | 10% | Are there specific, constraint-rich prompts? | Half or more of on-category prompts run to eight words or more and carry a constraint such as a budget, team size or integration. | Under 15% scores 2. |
| D7 Dedup / redundancy | 10% | How much of the set is paraphrase, or one template repeated? | Under 15% of prompts are near-duplicates or share the most common opening. | A third or more scores 2. |
From dimensions to the headline
The overall score is ten times the weighted sum of the seven dimensions, rounded: round(10 × Σ weight × dimension).
Worked example. Dimensions of 6, 7, 8, 5, 7, 4 and 8 give 10 × (0.25×6 + 0.15×7 + 0.15×8 + 0.15×5 + 0.10×7 + 0.10×4 + 0.10×8) = 10 × 6.4 = 64, Mixed, grade C.
Confidence is shown beside the score and does not change it: low under 20 prompts, medium without a brand and category in the brief, high otherwise.
| Band | Score | Grade | What it means |
|---|---|---|---|
| Excellent | 85 to 100 | A | Measures buyers well. Keep it stable and track it. |
| Strong | 70 to 84 | B | Trustworthy, with a few fixes worth making. |
| Mixed | 50 to 69 | C | Usable, but some numbers it produces will mislead. |
| Weak | 30 to 49 | D | Unlikely to reflect buyers. Rebuild before reporting. |
| Critical | 0 to 29 | F | Measures something other than discovery. |
Beside the score, never in it
Red flags call out problems a person should see whatever the score: personal data in prompts, sets too small to detect a change (under 20 prompts), leading prompts that assume the answer, year-stamped prompts and heavy competitor lock-in. Where a flag's cause is already a dimension, it says so rather than counting twice.
Coverage shows how much of the category's buying decision the set reaches, on a grid of topics and six journey stages: discovery, problem, evaluation, comparison, validation and post-purchase. It is scored out of 100 and banded complete (80+), broad (60+), partial (35+) or thin. It sits beside the headline because a coverage number is partly a score of the grid, and the grid is generated, so it is printed in full and can be edited.
| Coverage part | Points | Measures |
|---|---|---|
| Topics reached | 30 | Share of the category's topics with at least one prompt |
| Depth on what matters | 25 | Share of high-weight topics with five or more prompts |
| Stages occupied | 20 | Share of the six journey stages reached |
| Balance | 15 | Spread across topics and stages, cut when one stage dominates |
| Staying on the grid | 10 | Less the share of prompts that fit no topic |
Recommendations rank the changes worth making. Those the tool can make itself, such as trimming branded prompts or dropping near-duplicates, show the points they are worth, worked out by re-running the scorer on the changed set rather than estimated. Re-scoring applies the ones you tick and grades the result as a new, linked scorecard.
What it does not measure
- How visible your brand is. That is what the tracker reports.
- How many people ask each question. No auditable volume figure exists for AI prompts, and the score does not invent one.
- Where your prompts came from. The brief lets you declare it, and it is shown, but it is never scored.
- Meaning, beyond what rules can detect. Two prompts sharing words can still ask different things; redundancy is measured on wording.
Questions about the method
- Does RateMyPrompts use AI to score prompt sets?
- No. The Prompt Fit Score is deterministic: the same prompt set and brief always get the same score, and no AI model is involved in scoring. A model may help build the reference grid for coverage, which sits beside the score and is labelled when it does.
- Why is buyer relevance weighted highest?
- Because it decides what the visibility number means. A set that asks category jargon or names your brand returns more mentions than buyers actually see, so buyer relevance carries 25% and caps apply when jargon or branded prompts dominate.
- Why doesn't the score reward prompts that contain the category name?
- Real buyers usually describe their situation rather than naming the category. In one live 156-prompt mortgage set, only 13 of 94 unbranded prompts contained the word mortgages. Scoring the keyword would reward keyword-shaped sets, which are the ones that inflate visibility.
- How is the rubric checked?
- Against prompt sets scored by hand. Tests assert that no set a person rated Critical or Weak scores 50 or more, and no set rated Strong or Excellent scores below 70. Every scorecard records the rubric version it was graded against.
See it on your own set
Free, in under a minute. Terms are defined in the glossary.
Published · Rubric prompt_fit_rubric@0.2.0