GEO Maturity Model
Quick facts
- Difficulty
- Intermediate
- Time
- About half a day for a first team assessment and about 1 hour for a reassessment
- Prerequisites
- Full GEO Audit, GEO Metrics
- Purpose
- Assess five dimensions across five levels, identify the next exit test, and reassess after addressing it
- The scoring rule
- Your overall level is the lowest of the five dimension scores, never the average
- Industry standard?
- No. Webflow, Semrush, Superlines, and others publish maturity models, but they do not use a common number of levels or shared names. This 5 × 5 grid and its minimum-score rule are GEO Wiki's synthesis
- Current publishers
- Platform vendors and agencies publish these models. As of August 2026, Gartner, Forrester, and McKinsey do not publish a GEO or AEO maturity model
- Time required
- About half a day for a first team assessment and about 1 hour for a reassessment
1. What a GEO maturity assessment answers
The useful question is not “How good is our GEO?” It is “Which capability should we build next, and what evidence will show that we have built it?” A maturity assessment answers that question by pairing every level with a testable exit condition.
This framework uses levels instead of a 0–100 readiness score. A single score combines several distinct capabilities and can hide the one that is limiting progress. GEO Metrics treats vendor composite scores with the same caution, as does the full GEO audit when it reports its findings.
A GEO audit provides a severity-ranked view of the site’s current condition. AI Citation Tracking provides recurring evidence of what engines actually do. A maturity assessment uses both to identify the next exit condition and the actions required to meet it.
About the term. A maturity model is an established type of framework descended from CMMI. The official CMMI model lists levels 0 through 5: Incomplete, Initial, Managed, Defined, Quantitatively Managed, and Optimizing (CMMI Institute). Maturity models for GEO and AEO are also in commercial use. Webflow publishes an AEO Maturity Model with four categories and five levels. Semrush launched a four-stage “People and Process Maturity Matrix” at Adobe Summit (announcement). Superlines publishes a four-stage GEO Maturity Framework.
These commercial models are not standardized. Published versions have four or five levels and do not share level names. Webflow uses Keywords → Answers → Structure → Pillar → Authority, while Superlines uses Ad Hoc → Emerging → Integrated → AI-First. Their publishers are platform vendors and agencies rather than analyst firms. Forrester offers an AEO tactics guide but no maturity model, and pure-play AI visibility tools publish no maturity model. Webflow’s category-by-level grid is the closest published model to this one. The ownership-and-cadence dimension and the minimum-score rule are GEO Wiki’s working synthesis, not a ratified standard.
2. The five dimensions and the minimum-score rule
A one-dimensional model cannot show which independent capability is holding an organization back. This assessment uses five dimensions. Four correspond to layers in the GEO audit, while the fifth evaluates whether the work operates as a recurring process.
| # | Dimension | Corresponding audit layer | Capability assessed |
|---|---|---|---|
| D1 | Access & delivery | Layers 1–2 | Engines can fetch the site, and its content remains available after rendering. |
| D2 | Structure & content | Layers 3–4 | The fetched content is both parseable and citable. |
| D3 | Off-site authority | Layer 5 | Sources outside the organization’s domain corroborate the entity. |
| D4 | Measurement | Layer 6 | The organization can measure engine behavior reproducibly. |
| D5 | Ownership & cadence | None | The work has an owner, a regular cycle, and a budget. |
D5 captures the transition from a one-time project to a recurring process, which no audit layer can measure. Earlier SEO maturity frameworks reached a similar conclusion. Heather Physioc’s organizational search maturity stages progress from “initial and ad hoc” to “efficient and optimizing,” applying CMMI concepts to search. Martijn Scheijbeler’s SEO Maturity Curve uses multiple dimensions instead of a single progression, assessing technical work, content, authority, support, and strategy.
Your overall level is the lowest of the five dimension scores, not the average. This rule follows the audit’s dependency logic. A D2 score of 4 paired with a D1 score of 1 describes excellent content that no engine has fetched. An average would obscure that dependency, while the lowest score identifies the condition currently limiting citation performance.
Work completed above the limiting dimension is not wasted. It can affect results once the lower-level condition is resolved. A team with strong content and a blocked crawler has not wasted a quarter; the content can begin contributing after the robots.txt issue is fixed.
3. The five levels and their exit tests
| L | Name | Current state | Testable exit condition |
|---|---|---|---|
| L1 | Unmanaged | There is no owner or baseline, and the site is often unreachable to AI user-agents. | AI user-agents can fetch the site, primary content survives rendering, and one owner is named. |
| L2 | Instrumented | The organization has a declared prompt set, a baseline, and a repeatable method. | A second measurement round reproduces the first under a written method. |
| L3 | Systematic | Content and structure work follows a regular cadence rather than operating as a one-time project. | A change in a re-audit metric can be attributed to a specific action. |
| L4 | Competitive | The organization has a declared competitor set, works by engine, and manages off-site authority as an ongoing program. | Citation Share remains stable or grows against the declared competitor set for at least two periods. |
| L5 | Reference | Engines use the organization as a default source, and it defines the category’s vocabulary. | The organization sustains the first-cite position on core topics across at least three engines and publishes its method. |
It must be possible to fail each exit condition. If a level has no condition that can be tested and failed, it cannot resolve disagreements about the organization’s current capability.
4. How to complete the 5 × 5 rubric
| Dimension | L1 Unmanaged | L2 Instrumented | L3 Systematic | L4 Competitive | L5 Reference |
|---|---|---|---|---|---|
| D1 Access & delivery | AI user-agents are blocked or have never been audited, and primary content is rendered on the client. | User-agents have been verified, and primary content appears in the server HTML. | Access and rendering are checked again with every deployment. | Fetching behavior is understood and optimized for each engine. | Infrastructure changes must pass AI fetch tests before release. |
| D2 Structure & content | There is no structured data, and the prose cannot be extracted in self-contained chunks. | Key templates have schema, and one content-chunk audit has been completed. | Citability standards are part of the authoring workflow, and content follows a freshness schedule. | Coverage is adjusted to query intent, and content gaps are addressed relative to competitors. | Engines quote the organization as the source of the topic’s vocabulary. |
| D3 Off-site authority | The entity is unresolved, and brand mentions are not monitored. | Mentions are counted, and the state of the entity in knowledge graphs is known. | Earning brand mentions is an established workstream. | Share of Voice is tracked against a declared competitor set. | Third parties cite the organization as the definition. |
| D4 Measurement | There is no prompt set or baseline. | The organization has a declared prompt set, a baseline, and a written method. | Changes in metrics can be attributed to specific actions. | Results are segmented by competitor and by engine. | The method is published and reproducible by others. |
| D5 Ownership & cadence | No one owns the work. | One owner is named, and work occurs on an ad hoc schedule. | A budget supports quarterly audits and monthly tracking. | Engineering, content, and communications collaborate, and GEO work appears on the roadmap. | GEO requirements influence the roadmap before implementation begins. |
Support every cell with a dated artifact rather than an opinion. Saying that the site’s schema is probably correct does not satisfy a D2 condition. A dated result from the Rich Results Test does. Dated evidence makes a reassessment comparable with the initial assessment, which is the reason to repeat the exercise.
If a GEO audit is available, map its severity-ranked findings to D1 through D4. A Blocker at Layer 1 keeps D1 at L1 regardless of the other dimension scores.
5. Assessment by level
Each level uses the same five-part format: the current condition, one exit test, priority actions, the KPIs appropriate to the level, and a common pitfall.
5.1 L1: Unmanaged
| Current condition | The robots.txt rules for AI user-agents are unknown, primary content is rendered on the client, and no one has formal responsibility for the work. |
| Exit test | AI user-agents can fetch the site, primary content appears in the server HTML, and one owner is named. |
| Priority actions | 1. Audit crawler access with the AI Crawler Access Audit. 2. Make primary content server-rendered by following SSR for AI Crawlers. 3. Name one owner and document the assignment. |
| KPIs | Use Source Diversity Score and Mention Frequency only as existence checks. Do not use a competitor set yet. |
Google documents user-agent tokens and robots.txt behavior in its common crawlers list. AI Crawlers provides additional reference detail.
Common pitfall: buying a monitoring subscription before the site is fetchable. The organization pays to observe zero results while the dashboard creates the appearance of progress.
5.2 L2: Instrumented
| Current condition | You can report a Citation Rate and document how it was calculated, and the prompt set is written down rather than remembered. |
| Exit test | A second measurement round reproduces the first under a written method. |
| Priority actions | 1. Establish recurring measurement with AI Citation Tracking. 2. Document the engine set and time window. 3. Run the first full GEO audit. 4. Establish a structured-data baseline with Schema Implementation. |
| KPIs | Add Citation Rate and Average Position using definition A, which is citation order (GEO Metrics). |
The definition of Average Position is part of the exit test, not a footnote. The same brand receives different results when position means citation order, mention order, or list position. Early measurements have two additional sources of distortion. AI-referred sessions are difficult to identify through standard analytics (AI Search Attribution), and counting mentions instead of citations can change a figure by two to three times (Citation vs Mention). Vendor formulas also vary. Otterly publishes its KPI definitions, providing a useful comparison when another tool does not disclose its formulas.
Common pitfall: expanding measurement instead of acting on it. Adding engines and prompts may look like progress, but it does not satisfy an exit test.
5.3 L3: Systematic
| Current condition | Citability standards are part of the authoring workflow rather than a one-time cleanup, and someone follows an established reassessment schedule. |
| Exit test | A change in at least one re-audit metric can be attributed to a specific action. |
| Priority actions | 1. Make extraction quality part of the workflow with Citability and Writing for AI Citation. 2. Set a freshness schedule with Content Freshness. 3. Publish and maintain llms.txt. 4. Fund quarterly audits and monthly tracking. |
| KPIs | Add Answer Inclusion Rate to measure coverage across the declared prompt set. Do not introduce a competitor set yet. |
This level has the strongest published causal evidence. In Aggarwal et al. 2024 (arXiv:2311.09735), content-level changes that added citations, statistics, and quotations increased answer visibility by up to roughly 40%. The study measured this increase for a single participant optimizing in isolation, so it should not be treated as a guaranteed result.
The order of the changes matters. Wan et al. 2024 (arXiv:2402.11782) supports placing relevance before stylistic refinement. When the researchers tested which of two conflicting pages a model preferred, relevance to the query was the dominant factor. Signals that human readers associate with credibility, including scientific references, a neutral tone, and formal citation style, had little effect on model preference. The paper reports the direction of the effects rather than point estimates, so use it to order actions rather than forecast a numeric improvement.
Common pitfall: claiming improvement without a re-audit change or attributing an engine-side change to your own work. The before-and-after evidence does not support either claim.
5.4 L4: Competitive
| Current condition | The organization has a declared competitor set, has evidence of differences among engines, and has assigned responsibility for off-site work. |
| Exit test | Citation Share remains stable or grows against the declared competitor set for at least two periods. |
| Priority actions | 1. Manage off-site authority as a recurring program with Brand Mention Tracking, informed by Brand Mentions and Entity Recognition. 2. Address measured differences among Perplexity AI, ChatGPT Search, and Google AI Overviews. 3. Apply the relevant industry approach for SaaS / B2B, E-commerce, or Media. |
| KPIs | Add Citation Share, Share of Voice, First-Cite Rate, and Brand Sentiment. |
Competitor-relative metrics depend on how they are calculated. Ahrefs weights Share of Voice by search volume to approximate impressions (Brand Radar methodology), while most vendors use raw mention counts. As a result, two metrics labeled “SOV” rarely measure the same construct. First-Cite Rate also depends on whether an engine exposes citation order. Perplexity’s API returns citations as an ordered array, while several engines do not.
Common pitfall: optimizing for one engine in ways that do not transfer to others, then reporting competitor-relative and absolute findings in the same column. These findings represent different types of measurement.
5.5 L5: Reference
| Current condition | Third parties cite the organization as the definition, and GEO requirements influence the product and content roadmap before implementation rather than appearing later as corrective work. |
| Exit test | The organization sustains the first-cite position on core topics across at least three engines and publishes its method. |
| Priority actions | 1. Publish the measurement method so others can reproduce it. 2. Maintain the position through retrieval and model resets. 3. Preserve the E-E-A-T signals that make the entity a default source. |
| KPIs | Use all ten KPIs and a custom composite. A composite is appropriate only at this level, once every input is tracked independently and its weighting is public. |
Google states that its AI features require no special structured data (AI features and your site). At this level, structured data should therefore be evaluated for entity clarity rather than as a requirement for eligibility.
Common pitfall: treating the position as permanent. One retrieval or model update can reset it. An organization must continue to meet the L5 conditions rather than assuming it will retain the level.
6. Time and cost for each transition
These durations describe the process that must occur, not guaranteed completion dates. The table pairs each level transition with the corresponding elapsed time and category in GEO ROI Models.
| Transition | Typical elapsed time | What must occur | Related GEO ROI category |
|---|---|---|---|
| L1 → L2 | Days to weeks | The crawler fetches the site again after the access or rendering fix, followed by one baseline measurement round. | Technical foundation (0–30 days) |
| L2 → L3 | 1–2 quarters | The organization completes two comparable measurement rounds and makes one change whose effect can be attributed. | Citation Value (30–120 days) |
| L3 → L4 | 2–4 quarters | Content passes through repeated regrounding cycles, and the competitor set remains stable enough for comparison. | Substituted Traffic Value (60–180 days) |
| L4 → L5 | 4 or more quarters, and often never | Off-site mentions accumulate and influence the model’s prior understanding of the entity. | Brand Authority Value (180–540 days) |
GEO ROI Models covers three cost categories: one-time expenses, ongoing expenses, and internal time that usually does not appear as a line item. It also provides a tool-pricing range. Before presenting any price to leadership, verify it on the vendor’s current page because market prices change faster than published ranges.
Maturity levels can also decline. Ahrefs’ analysis of 17 million citations found that AI-cited content was 25.7% fresher on average than organic Google results. Maintaining a regular content-update schedule is therefore an ongoing L3 condition rather than a one-time achievement. Content Freshness explains what that freshness difference means and why it is smaller than the headline may suggest.
7. Actions to defer at each level
Four sequencing rules apply to specific levels:
- Do not optimize for individual engines before L4. Until the organization can measure changes by engine reliably, the same core work generally serves every engine: access, rendering, structure, citable content, and entity corroboration. Engine-specific analysis becomes useful after that exit test.
- Do not use competitor-relative KPIs before L4. Share of Voice cannot support a meaningful comparison without a stable competitor set and a declared engine set.
- Do not use a composite score before L5. Even at L5, use one only when its weighting is public. GEO Metrics and the GEO audit apply the same requirement.
- Do not use an industry-specific playbook before L3. The approaches for SaaS, e-commerce, and media begin to diverge after a content system exists. Before then, the work is nearly identical across those industries.
All four rules address the same failure: teams report the KPIs they can already produce instead of the KPIs that would satisfy the next exit test. A Level 2 team that reports Share of Voice is measuring a result it cannot yet interpret.
8. Choosing where to stop
Most organizations should not target L5. The appropriate target is a business decision. GEO ROI Models identifies five cases in which the expected return does not support moving beyond L1’s technical foundation: niches with very low volume, commodity products dominated by marketplaces, organizations driven entirely by paid or product-led acquisition, relationship-sales teams already limited by capacity, and companies that have not reached product-market fit.
For most mid-market B2B organizations, L3 is a realistic target. Pursuing L4 should be a deliberate, funded decision. Choosing not to pursue it is not a failure.
A maturity assessment cannot determine three important factors:
- The effect of engine-side changes. A retrieval update can change the metrics more than a quarter’s worth of work by the organization.
- The category’s structural limit. Some categories direct users toward marketplaces or aggregators, and a higher maturity level does not remove that limit.
- The cause of a competitor’s gain. Relative movement may result from the competitor’s work or from an engine change, and the available data rarely distinguishes between them.
Treat vendor “AI maturity scores” in the same way as vendor metric composites: identify them without endorsing them, and do not interpret them without a public formula (GEO Metrics). Most published maturity models, including those cited in §1, are marketing assets tied to a self-assessment funnel. This does not make their conclusions incorrect, but the assigned level tends to fall immediately below the product tier designed to address it.
9. What to document and when to reassess
Produce the following items for every assessment:
- The completed 5 × 5 rubric, supported by a dated artifact for every cell.
- The level statement, which includes the lowest score and the score for each dimension. Format it as
L1: D1:3 D2:2 D3:1 D4:2 D5:2. Reporting the full vector prevents the overall level from obscuring the dimension scores. - The next exit test, quoted exactly from §3.
- The prioritized action list for meeting that test, ranked by impact × confidence × ease.
- The reassessment date.
Reassess quarterly and after any event that requires another audit (GEO Audit): a site migration, a rendering change, a robots.txt edit, a major content launch, or a known engine update.
A reassessment that changes no cell is still informative. It shows that the quarter’s work did not affect the limiting dimension, which is more useful than reporting a number that changed for an unknown reason.
10. Further reading
- Current-state diagnosis: Full GEO Audit provides severity-ranked findings that map to D1 through D4.
- Recurring measurement: AI Citation Tracking explains how to measure the D4 outcomes over time.
- Metric definitions: GEO Metrics defines every KPI used here, including formulas and vendor variations.
- Business case: GEO ROI Models covers cost categories, payback curves, and the cases for stopping described in §8.
- Core concept: Generative Engine Optimization explains the broader discipline.
- Academic evidence: Aggarwal et al. 2024 examines the effects of content changes, while Wan et al. 2024 examines which signals affect a model’s preference.
Frequently asked questions
Is 'GEO maturity model' a real industry framework, or an invention?
We have great content but AI crawlers are blocked. What level are we?
Which level should we actually be targeting?
When should we start optimizing for individual engines?
How is this different from running a GEO audit?
Can we skip a level?
Related playbooks & wiki
- Full GEO Audit
- AI Citation Tracking
- Citability Audit
- AI Crawler Access Audit
- Schema Implementation
- Writing for AI Citation
- Brand Mention Tracking
- Deploying llms.txt
- geo-for-saas
- geo-for-ecommerce
- geo-for-media
- Generative Engine Optimization
- GEO ROI Models
- GEO Metrics
- Citability
- E-E-A-T
- Brand Mentions
- Entity Recognition
- Citation vs Mention vs Link
- Content Freshness
- AI Crawlers
- ssr-for-ai-crawlers
- ai-search-attribution
Sources
Primary
- CMMI Levels of Capability and Performance · CMMI Institute / ISACA
- GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024) · arXiv / KDD '24 · 2024-08-25
- GEO: Generative Engine Optimization (KDD '24 Proceedings) · ACM SIGKDD · 2024-08-25
- What Evidence Do Language Models Find Convincing? (Wan et al., ACL 2024) · arXiv / ACL 2024 · 2024-02-19
- What Evidence Do Language Models Find Convincing? (ACL Anthology) · Association for Computational Linguistics · 2024-08-11
- Otterly.ai — Brand Report KPI Definitions · Otterly.ai
- Ahrefs Brand Radar Methodology · Ahrefs
- Perplexity API — Chat Completions Reference · Perplexity
- List of Google's common crawlers (Googlebot, Google-Extended) · Google Search Central · 2026-04-23
- Google Search Central — AI features and your site · Google · 2025-12-10
Secondary
- Announcing Webflow's AEO Maturity Model · Webflow
- Semrush Unveils Brand Visibility Framework at Adobe Summit · Semrush
- The 4-Stage GEO Maturity Framework: How Ready Is Your Brand For AI Search? · Superlines
- Preparing Your Brand for AI-Driven Search with the AEO Maturity Model · SEO Hacker
- Win Visibility In AI Search With Answer Engine Optimization · Forrester
- SEO Maturity: How to Grow Search at Your Company · Page One Power (Heather Physioc)
- The SEO Maturity Curve: Where Does Your Strategy Stand? · Martijn Scheijbeler
- New Study: AI Assistants Prefer to Cite 'Fresher' Content · Ahrefs