Methodology
How the score is calculated, and what it does not claim
See what the score measures, how missing evidence is handled, and where the analyzer deliberately stops. The short explanations come first; exact rules, limits and sources remain available when you need them.
Criteria set 2026.09.25.1 · updated · individual source verification dates are listed below
Read the result in this order
- 1. Eligibility: can this URL be crawled and indexed?
- 2. Coverage: how much usable evidence was available?
- 3. Score: how much supported optimization credit was earned?
The score is an auditable engineering model. It is not a Google ranking prediction.
Category weights
Eight category measurements retain 100 legacy design points for reproducible exports. Only Crawlability, Technical, On-Page and measured Performance evidence contribute to the central readiness verdict; the other categories are supporting evidence.
Crawlability & Indexing
Whether search engines can reach, fetch, and index this URL at all — HTTP status, robots directives, robots.txt and sitemap discovery.
Why it matters and sources
HTTP errors can prevent content indexing. A robots.txt disallow prevents crawling, but the URL can still appear in search results through links elsewhere. A noindex directive must be accessible to the crawler to prevent indexing; this audit does not query Google’s index.
Technical Foundations
HTTPS, redirect behaviour, canonical correctness, language declarations, hreflang, viewport, character encoding, mixed content and DOM complexity.
Why it matters and sources
These checks combine HTML/HTTP requirements, documented search guidance and explicitly labelled heuristics. Findings can reveal canonical conflicts, unsupported language annotations or mobile rendering problems; each check explains the evidence and its limits.
On-Page Signals
Title link, meta description snippet, heading outline, image alt semantics, and internal/external link quality.
Why it matters and sources
Google uses the title element as its primary source for title links and page content plus the meta description for snippets. Headings and descriptive link text help Google understand page structure and destination topics.
Content Signals
Observable content signals: text structure, readability, repetition, freshness markers and contact links. These are limited proxies; they do not establish usefulness, expertise, accuracy or whether a page satisfies its purpose.
Why it matters and sources
Google’s guidance on helpful, reliable, people-first content is the most emphasised part of its documentation. There is no required word count — what matters is whether the page satisfies the search intent it targets.
Performance
Server-measured transfer characteristics, Core Web Vitals from real-user field data, and a separately identified Lighthouse lab score when available.
Why it matters and sources
Core Web Vitals (LCP, INP, CLS) are part of Google’s page-experience signals and are evaluated at the 75th percentile of real user visits. Browser-measured metrics cannot be derived from server-side HTML, so they are only scored when real measurements are obtained.
Structured Data
JSON-LD, Microdata and RDFa discovery, syntax validation, Schema.org type recognition, and required/recommended property checks for Google-supported features.
Why it matters and sources
Google states that structured data by itself is not a generic ranking factor; it helps Google understand a page and can make it eligible for rich results, and eligibility is never guaranteed. Only types in Google’s current Search Gallery can produce a rich result.
Static accessibility markup
Static, HTML-verifiable accessibility signals: document language, alternative text, form labels, accessible names for controls and links, landmarks and iframe titles.
Why it matters and sources
Accessibility is not a documented Google ranking factor, but the same semantics that assistive technology depends on are what search engines parse. These checks are HTML-level only and cannot establish WCAG conformance.
Social Metadata
Open Graph and X (Twitter) card metadata, favicon, site identity and discoverable social profiles.
Why it matters and sources
Social metadata has no effect on Google ranking. It controls how links look when shared, which influences click-through and referral traffic — so it is scored, but with the lowest weight in the model.
Why Social Metadata is a supporting measurementSharing previews matter to users, but they do not determine the central SEO verdict.
Open Graph and X card tags have no effect on Google ranking whatsoever. They control how a link looks when shared, which is a real click-through concern but not a ranking claim. Social Metadata, Accessibility, Structured Data and heuristic Content observations therefore remain separate supporting measurements and do not change search implementation readiness. Their legacy design weights remain only so older exported reports can reproduce their original arithmetic.
The arithmetic
Deterministic and fully re-derivable from the check list in any report.
- 1. Each criterion earns points. Every criterion declares a point value within its category. Binary checks normally award full credit for a pass, partial credit for a warning and none for a failure. Graduated checks use a measured ratio: a passing check can still earn partial points. The status describes its threshold; the displayed earned and available points determine its contribution. Reports label such a pass Partial credit and list every lost point, by category, under “Why this is not 100”.
- 2. Categories are ratios, not sums. A category score is points earned divided by points available, over the criteria that could actually be measured. Criteria marked informational, not checked or unable to verify carry zero weight, so an unmeasurable criterion never costs the page points.
- 3. Sparse categories are withheld. A normal category score is not shown when no weighted criterion was measured, fewer than three weighted criteria were measured, or fewer than half of that category's listed criteria contributed points. It is also withheld when one heuristic supplies at least half of the measured weight with fewer than five measured criteria, or heuristics supply at least 75%. These are conservative SEOLens confidence rules, not Google ranking rules. Performance is also withheld when its evidence is transfer-only, because response-transfer observations cannot substitute for a Field or Lab rendering measurement. A legacy raw score across every category with at least one scored check remains in exports for compatibility and auditability; it is not a second readiness verdict.
- 4. Primary categories are re-weighted. Search implementation readiness uses Crawlability, Technical, On-Page and Performance when they have sufficient evidence. Content heuristics, Structured Data, Accessibility and Social Metadata remain visible supporting measurements with zero effective weight. Included primary categories are re-normalised proportionally. Confidence is low below 65% primary-model coverage, moderate below 90% or when any included category has moderate evidence, and high otherwise.
- 5. Eligibility is a separate gate. The implementation value does not override HTTP errors, robots blocking, noindex or a conflicting canonical. A URL with one of those blockers is labelled not scorable for search implementation readiness. If a search exclusion has no observed conflict, the roadmap is withheld until the user confirms whether the exclusion is intentional. When the URL answers 4xx or 5xx, advice about the page itself — its canonical, structured data or social tags — is withheld from the roadmap entirely: Google states an error response is not indexed, so there is nothing yet to optimise. Those findings stay in the criteria list with their evidence.
- 6. A grade needs most of the weight. Re-normalising makes the score speak only for the categories that had evidence, so the same number means different things at 100% and at 68% of the criterion weight. A readiness score therefore carries a quality band only when at least 90% of the primary design weight was measurable and evidence confidence is high. Between 65% and 90% the number is published as a measured score with its quality grade withheld, without a band or verdict colour; below 65% no readiness number is published at all. Every report states the percentage it rests on beside the gauge. Weighted model coverage and real-user field-data availability are shown separately because one cannot stand in for the other. These limits are SEOLens confidence rules, not Google rules.
Missing evidence changes the basis of comparison. Read category coverage and confidence alongside the technical value, which answers: how much of the sufficiently supported scoring credit did the observed checks earn?
Worked example · criteria 2026.09.25.1
Illustrative observations calculated with the report scoring engine:
- Binary check: pass · 4/4 points.
- Graduated check: pass · 3/4 points.
- Missing measurement: unavailable · 0/0 points.
Category score: 7 ÷ 8 × 100 = 88 after rounding. The unavailable measurement contributes no points or denominator. Two passes therefore need not mean a score of 100.
Score bands
Excellent
90–100
Good
75–89
Needs work
50–74
Poor
0–49
What each status means
Distinguishing “we checked and it failed” from “we could not check” is the difference between a report you can trust and one you cannot.
- Pass
- The page meets this criterion.
- Partial credit
- The page clears this criterion’s threshold but earns only part of its points. The missing points are listed under “Why this is not 100”.
- Warning
- The page partially meets this criterion.
- Fail
- The page does not meet this criterion.
- Info
- An observation with no pass or fail judgement. Not scored.
- Not checked
- Out of scope for this analysis. Excluded from the score.
- Unable to verify
- We tried but could not obtain trustworthy data. Excluded from the score rather than estimated.
Every threshold used
Split into values that come from published documentation and values that are industry convention. The distinction is shown on each finding too.
Documented by the source
These come directly from Google or a standards body. They are requirements or officially published thresholds.
- LCP
- good ≤ 2.5 s · poor > 4.0 sweb.dev (Google)
- INP
- good ≤ 200 ms · poor > 500 msweb.dev (Google)
- CLS
- good ≤ 0.1 · poor > 0.25web.dev (Google)
- TTFB target
- under 800 msweb.dev (Google)
- robots.txt size processed by Google
- 500 KiBGoogle Search Central
- Sitemap limits
- 50,000 URLs and 50 MB uncompressed per fileGoogle Search Central
- Core Web Vitals evaluation point
- 75th percentile of real user visitsweb.dev (Google)
Industry heuristics12 review conventions and proxies — never search-engine requirements.
These are review conventions and limited proxies, not search-engine requirements or verified predictions. Any finding that uses one carries a Heuristic marker. Review its relevance to the page before acting.
- Title length (informational review cue)
- 30–60 characters
Google publishes no title length limit. Character count is unscored and does not establish descriptive quality or an exact display cut-off.
- Meta description length
- 70–160 characters
Google publishes no description limit and may generate its own snippet regardless.
- Word count
- informational, unscored
There is no required word count. Text length alone cannot establish quality or whether a page answers its intended question; a short answer or utility can be sufficient.
- Text-to-HTML ratio
- 15% passes the inventory check; lower ratios need corroborating evidence
A low ratio is informational and unscored when readable text is already present in the served HTML. It becomes a problem only when the response also contains no readable body text.
- Readability analyzer outcome
- pass ≥ 60 · warning 30–59.9 · fail < 30
These are analyzer verdict bands, separate from Flesch labels such as “difficult”. Flesch Reading Ease is an English-only writing heuristic, not a ranking factor. Technical terms can lower it even when the copy suits its readers.
- Keyword stuffing (single term)
- 4.5% of all words, min. 8 occurrences
Google’s spam policies describe the behaviour but publish no density figure.
- Subheadings expected above
- 350 words
A readability convention, not a documented requirement.
- Multiple H1 warning
- 4 or more
Multiple H1s are valid HTML and Google handles them. We only warn when no primary topic is identifiable.
- Excessive link count
- 300 links
Google withdrew its old “under 100 links” guidance. This is a display threshold only.
- Alt-text coverage to pass
- 95% of content images
Decorative images with alt="" are excluded from the denominator, since that is correct markup.
- URL length
- under 120 characters
A readability convention. Google recommends simple URLs but sets no limit.
- External script references
- pass ≤ 25 · review 26–50 · high inventory > 50
The count is an HTML inventory heuristic. Inline bytes remain visible evidence but do not prove transfer size, execution cost, unused code or interaction delay.
Structured data: what is and is not claimed
Schema.org vocabulary and search-engine presentation rules are separate layers. The analyzer reports both without treating one search engine's feature rules as universal.
One vocabulary, platform-specific features
The generator creates standard Schema.org JSON-LD, which can be read by any consumer that supports that vocabulary. Google, Bing, Yandex and other crawlers decide independently which types, properties and search experiences they support. A valid Schema.org graph is therefore not a promise of a Google rich result, a Bing enhancement, a Yandex appearance, indexing or ranking.
SEOLens applies documented Google feature rules only where Google publishes a verifiable requirement set. It does not invent parallel eligibility rules for engines whose public documentation does not define an equivalent test. Validate the vocabulary first, then inspect the published URL with each target engine. Bing URL Inspection requires the URL to belong to a site added to Bing Webmaster Tools.
Every declared type and property is compared against the published Schema.org vocabulary — 939 types and 1,538 properties from the release retrieved 2026-09-23. A term that does not exist in it, such as nmae for name, is reported as a fault, because a consumer ignores what it does not recognise rather than reporting it. A real property used on a type it is not defined for is reported without being failed: Schema.org publishes its domains as guidance and states the vocabulary is extensible. This is a vocabulary and documented-requirement check, not a replica of Google's Rich Results Test — confirm eligibility there.
Retired Google featuresValid Schema.org vocabulary can outlive the Search feature that once used it.
A Schema.org type can remain valid after Google retires the Search feature that used it. The retired feature no longer produces its former Search appearance. Check the feature-specific retirement notice before deciding whether to keep markup that no current consumer needs.
A plain WebSite entity is current for Google site names and is not classified as retired. Only the observed WebSite potentialAction → SearchAction sitelinks search-box pattern is reported as a retired Google feature.
- FAQPage
- FAQ rich results stopped appearing in Google Search on 7 May 2026. FAQPage remains a Schema.org type, but it no longer produces that Google Search feature.
- HowTo
- HowTo rich results were retired in 2023 and the documentation was removed. The markup is inert in Google Search.
- ClaimReview
- The fact-check rich result was phased out as part of Google’s 2025 search-results simplification.
- SpecialAnnouncement
- The special announcement feature was deprecated on 31 July 2025.
- JobTraining
- The estimated-salary and related job-training features were retired in 2025.
- LearningResource
- The learning-video feature was retired in 2025.
- Vehicle
- The vehicle-listing feature was retired in 2025.
- Practice
- The practice-problem type was removed from Google Search in January 2026.
- Dataset
- Google Search no longer uses Dataset markup: since November 2025 it is used only by Google Dataset Search, so it produces no Google Search feature. Keep it if datasets should be discoverable in Dataset Search.
Recognised Google feature contexts (35)A documented precondition is not a guarantee that Google will show a special presentation.
Rich-result types are mapped from Google’s Search Gallery. WebSite is mapped separately from Google’s current site-name guidance; site names are not supported in the Rich Results Test. Presence here is a documented precondition, never a guarantee that Google will use a particular presentation.
- ArticleArticle
- NewsArticleArticle
- BlogPostingArticle
- BreadcrumbListBreadcrumb
- CourseCourse list
- CourseInstanceCourse list
- DiscussionForumPostingDiscussion forum
- QuizEducation Q&A
- EmployerAggregateRatingEmployer aggregate rating
- EventEvent
- ImageObjectImage metadata
- JobPostingJob posting
- LocalBusinessLocal business
- RestaurantLocal business
- StoreLocal business
- OrganizationOrganization
- OnlineStoreOrganization
- ProductProduct
- ProductGroupProduct
- ProfilePageProfile page
- QAPageQ&A
- RecipeRecipe
- ReviewReview snippet
- AggregateRatingReview snippet
- SoftwareApplicationSoftware app
- MobileApplicationSoftware app
- WebApplicationSoftware app
- VacationRentalVacation rental
- VideoObjectVideo
- MathSolverMath solver
- MemberProgramLoyalty program
- MerchantReturnPolicyMerchant return policy
- ShippingServiceMerchant shipping policy
- SpeakableSpecificationSpeakable (beta)
- WebSiteSite name
Google Search Central — Structured data markup that Google Search supportsGoogle Search Central — Site names
Two carousel rule familiesGeneral host carousels and the separate limited beta have different rules.
Google’s general host carousel supports Course, Movie, Recipe, Restaurant. It requires at least two same-type items; Course lists require at least three. Summary-page URLs must be unique and on the same domain, while all-in-one item URLs must point to anchors on the current page.
A separate beta documents nested Product, Event, LocalBusiness, Restaurant, Store, Hotel, VacationRental patterns with at least 3 items. Its required common fields are name, image and URL. It is available only for specified regions and query categories, and some markets require an interest form or CSS program. SEOLens therefore labels a matching beta graph as review-only and never promises eligibility. Ratings and reviews are never suggested unless genuine, visible data already exists.
Rule sources verified 11 September 2026.
What this tool cannot determine
Stated plainly, because a limitation you know about is manageable and one you do not is a trap.
JavaScript is not executedThe primary audit reads the server response; rendered-page evidence remains a separate comparison.
The HTML audit reads the server response without running the page’s JavaScript. Content added by client-side rendering is outside that audit. Every current report can compare complete rendered HTML pasted by the reader without fetching or executing it. A deployment may also offer an isolated Chromium worker. A separate Lighthouse run provides performance diagnostics, but neither source replaces the served-HTML evidence.
Core Web Vitals need a field-data providerMissing real-user data stays unavailable and never becomes an estimated measurement.
Core Web Vitals checks use LCP, INP and CLS at the 75th percentile of real user visits. Each field metric that is missing is “unable to verify” and excluded from the score, even when Lighthouse returns a lab result. The report identifies URL-level or origin-level coverage. Lab measurements remain diagnostic, and the separately labelled Lighthouse performance criterion can still contribute points. When neither Field nor Lab evidence is usable, transfer checks such as TTFB and response size remain visible, but their diagnostic score is withheld from readiness arithmetic. New or low-traffic pages may lack field data even when the provider is configured.
Colour contrast is not measuredReliable contrast checking requires computed styles in a rendered browser page.
Contrast depends on computed styles — custom properties, cascade order, media queries, and anything applied by JavaScript. Determining it needs a rendered page in a real browser. No contrast value is reported, because guessing one would be worse than reporting nothing.
Rich-result eligibility review is boundedGraph rules are checked locally; Google’s validator and display decisions remain external.
This tool checks JSON syntax, recognized Google feature contexts and documented graph requirements, including supported property alternatives for Product, software apps and stable carousels. Google’s separate carousel beta is reported as review-only because region, query and program participation cannot be established from markup. The tool does not execute Google’s Rich Results Test or guarantee display.
A single URL is not a siteQuick audits cover one response; site-wide conclusions require a bounded deep crawl.
A quick audit inspects one page. Duplicate titles across pages, broken internal links, orphan pages and indexability coverage all need a crawl, and are reported as “not checked” until a deep audit runs. Findings from a deep audit describe the pages actually crawled — never the whole site.
hreflang return links need readable alternatesReciprocity is verified only for alternate pages the bounded crawl can fetch and retain.
The standalone validator checks the supplied annotations without fetching every alternate. A deep audit checks return links on the alternates it can read within its crawl scope and budget; unread alternates remain unverified. Google can process reciprocal pairs even when some languages are omitted from the wider set. A missing return link does not by itself invalidate every other pair.
Google’s ranking algorithm is unknownThe score measures documented implementation evidence, never predicted rankings.
No tool can tell you why a page ranks where it does. The category weights here are an engineering model derived from the priorities Google’s documentation states, not a reconstruction of its algorithm. Treat the score as a measure of implementation quality against documented guidance, not a ranking prediction.
Bot protection can block the fetchA WAF or CAPTCHA failure is reported as unavailable evidence rather than a partial success.
Sites behind a WAF or CAPTCHA may refuse automated requests. When that happens the analyzer reports it plainly rather than presenting a partial audit as complete. Whether Google is allowed through the same protection is a separate question.
Comparing two audits
How the change analyzer decides that something changed, how it classifies it, and the several things it deliberately refuses to conclude.
The comparison layer answers one question — did an observed signal change between two audits — and it answers nothing else. It does not decide what a signal means: indexability comes from the same chain the report shows, conflicts from the conflict engine, scores from the scoring engine. It reads their output and asks whether it moved.
Compatibility before comparisonThe URL, audit scope and criteria version determine whether two reports are comparable.
Two reports of different URLs, or a quick audit against a deep one, are refused rather than compared — forcing it would produce a screen of differences that are all real and all meaningless. A changed criteria version does not block the comparison but is stated, because it limits how much of a score movement can be attributed to the site.
A rule change is not a site change. If the criteria version moved, a criterion may change because its threshold moved. Rule changes are reported separately and excluded from the regression-risk assessment.
NormalisationEquivalent values are compared by meaning rather than serialization.
Values are compared by meaning, not by their text. URLs ignore a trailing slash, robots directives are compared as an unordered set so index, follow equals follow, index, hreflang is compared as a mapping rather than an array, structured data as a set of types rather than as serialised JSON, and text is whitespace-normalised. Content and link counts must move by a fifth before they are reported at all.
Change classificationsCritical, important, informational and improved outcomes stay separate.
- Critical
- Can directly alter whether the URL is crawled or indexed: a status change, a noindex appearing, crawling being blocked, a canonical moving to another URL, HTTPS becoming HTTP.
- Important
- May affect how the page is understood or discovered: title, headings, a large content or internal-link reduction, hreflang or sitemap membership, structured-data types disappearing.
- Informational
- A real change with no direct search consequence — a social preview image, a snippet limit, a measurement that appeared or disappeared.
- Improvement
- Movement in the other direction, shown separately.
Evidence changes and crawl coverageA new measurement is not an improvement, and an absent page is not proof of deletion.
A metric that appears where there was none is a new measurement, not an improvement. One that disappears is a lost measurement, not a regression. A criterion moving into or out of “not checked” or “unable to verify” is neither better nor worse. Where robots.txt could not be read, permission is unknown rather than blocked.
When two deep audits reached different numbers of pages, a URL missing from the smaller crawl is not treated as removed. A page limit, a deadline, robots.txt or one failed fetch all produce the same absence, so URL-level absence is reported as absent from that crawl and never as deleted from the site.
What this cannot show
SEOLens regression risk is an engineering classification of the difference between two audit reports. It is not a prediction of ranking or traffic loss. A detected change does not prove that search results, impressions or clicks changed, and this tool cannot observe which URL Google selected as canonical, how it reacted, or whether it had indexed the page at all. Every conclusion is limited to what the two reports observed.
Internal link status evidence
A separate bounded crawl of HTTP outcomes; it does not change the audit score.
The standalone status crawler performs GET requests through the same SSRF-protected transport as an audit. Crawl mode follows same-origin anchors found in served HTML, in breadth-first discovery order, and respects robots.txt. It does not run JavaScript. Single-URL mode requests only the submitted URL and its redirect chain, so robots.txt does not apply to that user-directed diagnostic request.
How statuses and destinations are recorded
HTTP status, transport completion, body completeness and link discovery are independent observations. A status is shown only after response headers arrive. DNS, TLS, timeout, robots and safety-policy outcomes do not receive synthetic status codes. A truncated body keeps its real status but makes discovery partial; a non-HTML body keeps its status and makes discovery not applicable.
Every normalized destination is fetched at most once while repeated anchors remain source occurrences. Fragments are removed; protocol, path case, trailing slash, query order and repeated query values remain distinct. External destinations are retained as out-of-scope evidence and are not fetched by the crawl.
Internal link opportunities
A review list derived from the inspected HTML, separate from the 0–100 SEO score.
Deep audits reuse their existing fetched pages and HTTP outcomes. A browser worker builds a directed graph from observed HTTP(S) anchors and locates conservative, exact unlinked topic phrases in inspected main content. This adds no target-site requests, does not use an AI service and does not execute the target’s JavaScript.
Graph evidence and eligibility guardsWhat counts as an incoming link, a possible orphan and a usable destination.
Referring-page counts deduplicate repeated anchors and exclude self-links. Navigation links remain incoming links; contextual links are shown separately. Nofollow, sponsored, ugc and page-level directives are retained as evidence rather than silently removed. Link counts are observations, not estimates of ranking value.
A sitemap-discovered page with no observed incoming links is a possible orphan candidate in this sample, never a confirmed site-wide orphan. Incomplete extraction or capped edges make absence conclusions uncertain. Paths count observed HTML-link clicks from the actual audit starting URL; sitemap seeds and redirect hops do not add clicks. “No path observed” does not prove isolation, and a bounded crawl may miss a shorter path.
Sources and destinations require usable fetched HTML and sufficient unblocked crawl/indexability evidence. Missing self-canonical alone is allowed; ambiguous or cross-URL canonical preferences suppress confident selection. A declared canonical is not proof of Google’s selected canonical. Robots blocking is a crawling restriction, not proof that a URL is absent from Google’s index.
Phrase matching and priorityLanguage, exact-text and evidence rules keep suggestions conservative and explainable.
Phrase matching requires an explicit English (en) or Arabic (ar) HTML language declaration on both pages, including regional tags, and uses conservative Unicode tokens with exact offsets back to original text. Mixed-direction text is supported. Script alone does not establish the language: Latin is not necessarily English, and Arabic script is also used by Persian and Urdu. A proposed anchor must occur verbatim in its attributable excerpt. Existing anchors, code examples, navigation/footer text, detectable hidden or inert content and unsuitable controls are excluded. Image alt text and accessible names are not treated as editable prose. Missing, unsupported or conflicting language evidence can prevent matching while leaving graph observations available.
Pages you mark as important are your explicit preference for this session. Evidence strength and priority are explainable review labels, not probabilities, measured demand or predictions of ranking or traffic improvement. New opportunities do not add penalties to existing site findings or change the SEO score. No supported match does not mean your internal linking is perfect.
Rule version 1.0.2 calls at most 2 observed referring pages “few” and a path of at least 4 observed HTML-link clicks “long.” Both are review heuristics, not Google limits. A qualifying title or heading phrase has 2 to 8 Unicode word/number tokens and at most 160 original characters, with at least 2 distinct non-common words. Title evidence or at least 3 tokens receives the stronger evidence label. Priority sorts your important pages first, then phrase evidence, fewer contextual referrers and fewer total referrers, with URL tie-breaking.
Storage and computation limitsExact caps keep the browser worker bounded without presenting partial evidence as complete.
Product storage limits are 100 inspected pages, 500 retained internal anchors per page and 10,000 overall, 120 source blocks per page, 1,200 characters per block, 12,000 source-text characters per page, 20,000 scanned HTML elements per page and 6,000 discovery records. Serialized observations have a separate 2 MiB UTF-8 limit. The derived result also has a 2 MiB limit and at most 6,000 graph nodes; observed paths use node indexes rather than repeatedly storing full URLs. When that budget is exceeded, page outcomes, titles and eligibility are retained first, then complete internal-link lists and original text for eligible sources, ordered by byte cost with URL tie-breaking. Other complete lists, partial graph evidence and ancillary discoveries use remaining space. External and non-HTTP anchors do not consume the internal-link allowance. Incomplete source link lists never become complete through trimming; only the browser worker decides whether the retained text supports a match. Successful page loads and complete retained link lists are different measurements. Incomplete records can include failed or robots-blocked requests; zero complete lists does not mean nothing was crawled. At most 3 source suggestions per destination and 100 suggestions overall are shown. The candidate index uses at most 20 phrases per destination and 200,000 comparisons; existing link issues are capped at 500 and excerpts at 360 characters. These limits are SEOLens choices, not Google requirements.
Existing issues, cancellation and exportsHTTP evidence remains distinct, and partial computation never blocks the main audit.
Existing link issues reuse confirmed 404/410 responses and recorded redirects. Timeouts, blocked requests, unchecked URLs and server failures remain distinct. Redirects require review of the observed destination, permanence and canonical intent rather than automatic replacement. Copied HTML is an escaped suggestion, and selectors describe the fetched HTML, which can differ after hydration or edits.
The main report and independent performance measurements remain available during link computation, cancellation or failure. Full JSON and print exports retain the result; compact browser history and comparison snapshots omit the graph and source excerpts.
Supporting guidance: crawlable and contextual links, canonical preferences, robots.txt limitations, and sitemap discovery. Matching rules, evidence labels and priority heuristics are SEOLens design choices rather than Google formulas.
Every source these criteria are built on
Each criterion cites at least one of these, and none is based on third-party SEO commentary.
View all 89 source documentsGoogle, Bing, Schema.org, web standards and accessibility sources, each with its verification date.
- SEO Starter Guide: The BasicsGoogle Search Central · verified 2026-09-05
- Crawling and Indexing DocumentationGoogle Search Central · verified 2026-09-05
- What Is Googlebot — File Size LimitsGoogle Search Central · verified 2026-09-16
- Robots Meta Tags, data-nosnippet, and X-Robots-Tag SpecificationsGoogle Search Central · verified 2026-09-05
- Introduction to robots.txtGoogle Search Central · verified 2026-09-09
- How Google interprets the robots.txt specificationGoogle Crawling Infrastructure · verified 2026-09-10
- Reduce the Google crawl rateGoogle Crawling Infrastructure · verified 2026-09-09
- How to Create a robots.txt FileBing Webmaster Tools · verified 2026-09-12
- Crawl ControlBing Webmaster Tools · verified 2026-09-12
- How to Report an Issue with BingbotBing Webmaster Tools · verified 2026-09-12
- URL InspectionBing Webmaster Tools · verified 2026-09-16
- Marking Up Your Site with Structured DataBing Webmaster Tools · verified 2026-09-16
- Build and Submit a SitemapGoogle Search Central · verified 2026-09-05
- Learn About SitemapsGoogle Search Central · verified 2026-09-09
- How to Specify a Canonical URL with rel="canonical" and Other MethodsGoogle Search Central · verified 2026-09-09
- Canonicalization TroubleshootingGoogle Search Central · verified 2026-09-05
- Influencing Your Title Links in Search ResultsGoogle Search Central · verified 2026-09-05
- Control Your Snippets in Search ResultsGoogle Search Central · verified 2026-09-05
- Creating Helpful, Reliable, People-First ContentGoogle Search Central · verified 2026-09-05
- Spam Policies for Google Web SearchGoogle Search Central · verified 2026-09-05
- Intro to How Structured Data Markup WorksGoogle Search Central · verified 2026-09-05
- Structured Data Markup that Google Search SupportsGoogle Search Central · verified 2026-09-05
- Provide a Site Name to Google SearchGoogle Search Central · verified 2026-09-16
- Structured Data General GuidelinesGoogle Search Central · verified 2026-09-05
- Does Anthropic crawl data from the web, and how can site owners block the crawler?Anthropic · verified 2026-09-14
- Rich Results and Search Console FAQsGoogle Search Central Blog · verified 2026-09-14
- Article Structured DataGoogle Search Central · verified 2026-09-11
- Breadcrumb Structured DataGoogle Search Central · verified 2026-09-16
- Event Structured DataGoogle Search Central · verified 2026-09-16
- Job Posting Structured DataGoogle Search Central · verified 2026-09-16
- Local Business Structured DataGoogle Search Central · verified 2026-09-16
- Recipe Structured DataGoogle Search Central · verified 2026-09-16
- Video Structured DataGoogle Search Central · verified 2026-09-16
- Profile Page Structured DataGoogle Search Central · verified 2026-09-16
- Image License MetadataGoogle Search Central · verified 2026-09-16
- Organization Structured DataGoogle Search Central · verified 2026-09-11
- Product Snippet Structured DataGoogle Search Central · verified 2026-09-11
- Review Snippet Structured DataGoogle Search Central · verified 2026-09-11
- Software App Structured DataGoogle Search Central · verified 2026-09-11
- Course List Structured DataGoogle Search Central · verified 2026-09-11
- Carousel Structured DataGoogle Search Central · verified 2026-09-11
- Structured Data Carousels (Beta)Google Search Central · verified 2026-09-11
- Movie Carousel Structured DataGoogle Search Central · verified 2026-09-11
- Latest Google Search Documentation UpdatesGoogle Search Central · verified 2026-09-05
- Rich Results TestGoogle Search Central · verified 2026-09-16
- Tell Google About Localized Versions of Your PageGoogle Search Central · verified 2026-09-05
- Pagination and Incremental Page LoadingGoogle Search Central · verified 2026-09-05
- Mobile-First Indexing Best PracticesGoogle Search Central · verified 2026-09-05
- Secure Your Site with HTTPSGoogle Search Central · verified 2026-09-11
- Understanding Page Experience in Google Search ResultsGoogle Search Central · verified 2026-09-05
- Google Image SEO Best PracticesGoogle Search Central · verified 2026-09-05
- Link Best Practices for GoogleGoogle Search Central · verified 2026-09-09
- Understand JavaScript SEO BasicsGoogle Search Central · verified 2026-09-05
- Optimize your crawl budgetGoogle Crawling Infrastructure · verified 2026-09-21
- Redirects and Google SearchGoogle Search Central · verified 2026-09-05
- Keep a Simple URL StructureGoogle Search Central · verified 2026-09-05
- Define a Favicon to Show in Search ResultsGoogle Search Central · verified 2026-09-05
- Optimizing for Generative AI Features on Google SearchGoogle Search Central · verified 2026-09-11
- AI Features and Your WebsiteGoogle Search Central · verified 2026-09-14
- PageSpeed Insights APIGoogle for Developers · verified 2026-09-05
- Chrome UX Report APIGoogle for Developers · verified 2026-09-05
- Web Vitalsweb.dev (Google) · verified 2026-09-05
- How the Core Web Vitals Metrics Thresholds Were Definedweb.dev (Google) · verified 2026-09-05
- Largest Contentful Paint (LCP)web.dev (Google) · verified 2026-09-05
- Interaction to Next Paint (INP)web.dev (Google) · verified 2026-09-05
- Cumulative Layout Shift (CLS)web.dev (Google) · verified 2026-09-05
- Time to First Byte (TTFB)web.dev (Google) · verified 2026-09-05
- LCP Request DiscoveryChrome for Developers · verified 2026-09-11
- Browser-Level Image Lazy Loadingweb.dev · verified 2026-09-11
- Lighthouse Performance ScoringChrome for Developers · verified 2026-09-05
- Schema.org VocabularySchema.org / W3C Community Group · verified 2026-09-05
- Schema.org Markup ValidatorSchema.org / W3C Community Group · verified 2026-09-16
- Questions About Semantic MarkupsYandex Webmaster · verified 2026-09-16
- Structured Data ValidatorYandex Webmaster · verified 2026-09-16
- The Open Graph Protocologp.me · verified 2026-09-05
- Cards markup (archived copy of X’s withdrawn page)X (Twitter) Developer Platform, via the Internet Archive · verified 2026-09-21
- Web Content Accessibility Guidelines (WCAG) 2.2W3C · verified 2026-09-05
- ARIA Authoring Practices GuideW3C WAI · verified 2026-09-05
- HTML Standard — Specifying the Document’s Character EncodingWHATWG · verified 2026-09-05
- HTML Standard — The lang and xml:lang AttributesWHATWG · verified 2026-09-05
- HTML Standard — Headings and OutlinesWHATWG · verified 2026-09-05
- CSS Device Adaptation — The viewport meta elementW3C / MDN · verified 2026-09-05
- RFC 9110 — HTTP Semantics (Status Codes)IETF · verified 2026-09-05
- RFC 9111 — HTTP CachingIETF · verified 2026-09-05
- BCP 47 — Tags for Identifying LanguagesIETF · verified 2026-09-21
- HTTP Headers ReferenceMDN Web Docs · verified 2026-09-05
- Mixed contentMDN Web Docs · verified 2026-09-21
- Server Side Request Forgery Prevention Cheat SheetOWASP · verified 2026-09-05
- RFC 9309 — Robots Exclusion ProtocolIETF · verified 2026-09-05
Ready to see this applied to a real page? Run an analysis.