# Crawld: full site text Generated 2026-08-30. Canonical index: https://crawld.co/sitemap.xml --- ## Crawld: the honest scorecard for search and AI answers URL: https://crawld.co/ 129 CHECKS · 8 CATEGORIES · 2 HUMAN GATES You shipped it in a weekend. Let's see if you're findable on Google Rank higher on search. Get cited by the AI everyone's asking instead. 129 checks across SEO and GEO , and a real fix, not more homework. Your website domain Scan it for free See a real fix ↓ No signup. No card. Takes about a minute. yourdomain.com · free scan example last 12 months scanning yourdomain.com … 10 pages ☑ Total clicks 69.3K ▲ 1371% since ☑ Total impressions 6.95M ▲ 1458% since ☐ AI citations 448 across 7 engines ☐ Avg position 8.9 ▲ was 29.9 fixes merged → clicks impressions --- the run where fixes merged 71 / 100 coverage 86% · 108 of 126 pages Illustrative example, not a real customer's score. Yours is computed the same way, from what was actually crawled. Clicks and citations come from inside the product; the free scan below returns the score, the findings and the coverage. · Technical · Crawl & structure · Search appearance · Content quality · E-E-A-T & trust · Site health · AI readiness (GEO) · Performance → fetching robots.txt … ok → sitemap.xml found · 212 urls ✓ 10 pages queued (free crawl cap) → checking canonical + hreflang … 10 of 212 pages crawled advisory scan · public HTML only A demonstration of one real scan, replayed. Press the button and this panel fills with yours. 0 score 90 / 100 · grade A 10 of 212 pages measured 201 findings · 3 checks unmeasured, and named 202 more pages exist beyond the free crawl cap. ⚠ llms.txt not found: AI crawlers are guessing ⚠ 12 pages missing a canonical tag ✓ structured data valid on every product page An example, not your site. Scan yours and this fills with what was actually measured. 129 checks in the rubric 8 weighted categories 16 AI / GEO-specific checks 2 points where it waits for you The report you get Every number on this page came from a crawl. This is the site health report for kazifi.co, 217 of its 218 pages read by the engine. The partial-run banner is not a mock-up: the crawl could not fetch one page, so the report says so rather than scoring the site as though it had. Screenshot of the running product, not an illustration. WORKS WITH ANY PLATFORM Get found in every engine your customers ask. Your customers stopped searching. They started asking. Crawld reads your site the way each of these engines does, shows who's getting cited for your keywords (you, or a competitor), and hands you the fix your platform can accept, no replatforming, no SEO degree. ‹ Claude ChatGPT Gemini Perplexity Google AI Overviews Copilot Bing DeepSeek + more › Claude Anthropic's own model, increasingly used for research and shopping. Checks whether your content is structured the way Claude's search actually reads it. Citation tracking live ChatGPT The largest AI answer engine by users. Tracks whether ChatGPT's browsing and search cite your pages, not just the competitor who out-ranked you on Google. Citation tracking live Gemini Built into Search, Workspace, and Android. Tracks whether Gemini's answers reference you when it actually matters, not just that you rank. Citation tracking live Perplexity Built to cite its sources by design, which means it either credits you or it credits someone else. Tracks your citation share directly. Citation tracking live Google AI Overviews Sits above the normal results now, and pulls from a narrower slice of the page than classic ranking does. The technical checks are tuned to what it actually lifts. Citation tracking live Copilot Built into Windows, Edge, and Bing. Same rubric, same fix queue, no separate setup. Citation tracking live Bi Bing Still the index behind Copilot and half the AI crawlers on the web. The classic search checks cover it. Indexed here means visible in far more places than bing.com. Citation tracking live De DeepSeek The fastest-growing open-weight model, already embedded in apps you've never heard of. Same rubric, same tracking, whichever wrapper someone's asking it through. Citation tracking live And the rest New engines launch quarterly. When one does, the rubric gets a new check, not a rewrite. Delivered as a PR , a CMS update, or a patch file, whatever your platform allows. See how access modes work → HOW WE SCORE If we couldn't check it, it doesn't pass. PASS Actually checked, actually clean. The rule ran, the page passed it. Nothing implied. FAIL Actually checked, needs work. Goes straight to the fix queue, ranked by how many pages it affects. UNMEASURED The crawler didn't reach it, or the check couldn't run. Most tools quietly round this up to "pass." We don't. It's gray on the grid and it's gray in the score. scorecard · example_scan.crawld 129 checks · 8 categories Metadata 17 checks Structure 14 checks Crawl access 11 checks Speed 15 checks Structured data 18 checks Internal links 12 checks Content 26 checks GEO / AI 16 checks score 71 / 100 coverage 86% · gray never counts as a pass PASS FAIL UNMEASURED (remainder of each bar) THE LOOP Five stages. Two places it stops and asks you. Same shape as a pull request you already review every week, except this one writes itself, verifies itself, and waits for your call twice. crawld / remediation-run #2847 3m ago ✓ Measure · crawl and score the site passed ✓ Audit · every finding, with evidence passed ‖ Plan · ranked fix list waiting on your approval ✓ Remediate · changes generated, build verified passed ‖ Deliver · pull request opened waiting on your approval ANSWER ENGINES Ranking #1 doesn't matter if the answer doesn't mention you. 16 of the 129 checks exist for one reason: people are asking AI instead of clicking blue links, and almost nothing tells you whether you're the site being cited. citation share · "best crm for solo agencies" last 30 days ChatGPT 34% AI Overviews 27% Perplexity 21% Claude 14% Gemini 11% Bing / Copilot 8% Example query. Your dashboard tracks the share on queries you actually compete for. Crawled, indexed, cited: three different things. Indexing tools are genuinely good at the first two steps. They also stop there. Crawld is built around the third. STEP 01 → CRAWLED A bot fetched your page. Necessary, not sufficient. A crawl proves the page loads, not that anyone will ever surface it. STEP 02 → INDEXED The engine stored your page and can show it. Tools like IndexRusher get you here fast, across Google, Bing, and LLM crawlers, and this is where they stop. STEP 03 WHERE CRAWLD WORKS CITED An AI answer actually names you as its source. Indexing doesn't measure this, and doesn't guarantee it. Crawld tracks it per engine and fixes what's blocking it. THE FIX QUEUE Every fix arrives as a diff , not a suggestion. You don't need to know what a canonical tag is. You need to know whether to click merge. src/pages/pricing.astro meta-description-missing · Auto-fix 14 15 - 15 + 16 Build verified Mode A · opens PR on merge WHO THIS IS FOR Built with AI. Should still show up in search. You don't need to have read a single SEO blog post. Crawld reads the site and tells you what's wrong in language you can act on. Vibe coders Cursor, Claude Code, v0, bolt: whatever you built it with, speed usually means metadata and structure got skipped. Crawld finds the gaps and opens a PR . $ git fetch crawld $ git merge fix/meta-description Fast-forward · 1 file changed, build passed ✓ Content teams publishing at scale You publish a lot, maybe with AI's help. The real risk isn't quality. It's your own posts quietly competing with each other for the same search. Crawld flags the overlap before you hit publish. OK "Best CRMs for solo agencies" distinct topic FLAG "Top CRM tools for small teams" overlaps post #14 OK "CRM pricing compared, 2026" distinct topic WHERE THE OTHERS STOP This is where Crawld starts. Tool Crawld Them Surfer GEO + verified remediation PRs On-page SEO scoring only SEObot Overlap-checked content + verified fixes High-volume AI content generation, no verification step SEOTesting Fix → verify → re-measure loop Reporting on GSC data IndexRusher Tracks whether AI answers actually cite you Fast indexing across Google, Bing, LLMs; stops before citation Refresh Agent Full-site audit + human gates Single-article refresh MarketMuse / Clearscope Ships build-verified changes Content briefs, no delivery Semrush AI-citation share tracking + fixes Rank tracking, no remediation EARLY CUSTOMERS The score is honest. So are these. We shipped on Lovable in a weekend and then had zero organic traffic for two months. The first scan found 43 pages Google couldn’t read at all. Merged the PR on a Tuesday, indexed by Friday. score 54 → 88 MO Maya Okafor Founder, Driftlane I don’t write code and I never had to. The fix queue reads like a to-do list, and every item shows me exactly what it changes before I approve it. score 61 → 90 SR Sam Reyes Head of Content, Parcelmetrics We run fourteen client sites through it. The unmeasured category alone is worth the retainer. No more pretending a gray area is a pass in a client report. 14-site avg 58 → 84 PA Priya Anand Technical SEO lead, Copperline Running Crawld DRIFTLANE PARCELMETRICS COPPERLINE HALYARD OKTOBER.CO VELDT PRICING Free to see the problem. Paid to fix it. Free scorecard $0 Scan any URL. Full score, no signup. Start free scan Remediation from $99 one-time / site One-time audit-to-merge cleanup. Pay once, keep the fixes forever. Get started Content-ops $29 / mo Ongoing scans, fix queue, calendar, watchdog. Cancel anytime. Get started Agency from $149 / mo Multiple sites, white-label reports, volume pricing. Talk to sales FROM THE BLOG The playbook, written down. GEO basics · 6 min read Why your weekend project is invisible to ChatGPT Three gaps every AI-built site ships with, the curl command that shows each one, and roughly twenty minutes of work to close them. Read → Technical · 4 min read llms.txt, explained for humans A plain-text map of your best pages for AI crawlers. What belongs in it, what does not, and the honest answer on whether it changes anything yet. Read → Strategy · 7 min read Crawled, indexed, cited, where AI traffic begins Three gates, and passing the first two says nothing about the third. What decides whether an answer engine names you as the source. Read → FAQ The questions people actually ask. How does the free scan work? Enter a domain. We crawl your pages the way Googlebot and AI crawlers do, run all 129 checks , and show the full scorecard, every pass, fail, and unmeasured item . No signup, no card, and the results aren't held hostage behind a paywall. What does "2 human gates" actually mean? The loop stops and waits for you twice: once before any fix is generated (you approve the ranked plan ) and once before anything ships (you approve the delivery ). Nothing touches your site without those two clicks from you. What do you do with my data? We store the crawl results and the fixes we generate for you, that's it. We don't resell crawl data, don't train models on your content, and don't keep repository or CMS access beyond the scopes you grant . Revoke access any time and the data goes with it. I'm on Squarespace / Webflow / Wix, how do fixes arrive without code? Three delivery modes . With code access, you get a pull request , a proposed change you approve. With CMS access, we apply changes through your platform's own editor, listed for your approval first. With neither, you get a copy-paste patch file with exact instructions for each change. Can a fix break my site or hurt my rankings? Every change is build-verified before you ever see it, and shipped changes are re-measured on the next scan . If a fix doesn't move its check from fail to pass, it gets flagged, and anything delivered as a PR is one click to revert . What's the difference between the SEO checks and the GEO checks? The SEO checks cover what search engines rank: metadata, structure, speed, internal linking. The 16 GEO checks cover what AI answer engines cite: whether your content is structured so ChatGPT, Claude, Perplexity, and AI Overviews can lift it and attribute it to you. Find out what's actually wrong. Takes about a minute. No account, no card. Your website domain Scan it for free --- ## Pricing URL: https://crawld.co/pricing/ PRICING Free to see the problem. Paid to fix it. The diagnosis is free because a score you cannot verify is worthless to both of us. Remediation is what costs money, because remediation is what takes work. Free scorecard $0 Scan any URL. Full score, no signup. Start free scan Remediation from $99 one-time / site One-time audit-to-merge cleanup. Pay once, keep the fixes forever. Get started Content-ops $29 / mo Ongoing scans, fix queue, calendar, watchdog. Cancel anytime. Get started Agency from $149 / mo Multiple sites, white-label reports, volume pricing. Talk to sales What each tier includes Feature Free Remediation Content-ops Agency Public scorecard: score, category breakdown, coverage ✓ ✓ ✓ ✓ Full findings list, with evidence per page Email ✓ ✓ ✓ Pages per crawl 10 Full site Full site Full site Scans per domain 1 / day Unlimited Unlimited Unlimited Ranked remediation plan : ✓ ✓ ✓ Fix queue with reviewable diffs : ✓ ✓ ✓ Re-measurement after delivery : ✓ ✓ ✓ Scheduled re-scans : : ✓ ✓ Content calendar and editor : : ✓ ✓ Site memory and overlap detection : : ✓ ✓ SEO Watchdog: up to 3 competitors : : ✓ ✓ AI citation tracking per engine : : ✓ ✓ Multiple sites : : : ✓ White-label reports : : : ✓ What the free scan actually does It is worth being precise, because "free scan" means very different things across this category. Crawld's free scorecard: Crawls up to 10 pages of public HTML and reports how many it found beyond that cap Runs the full 129-check rubric against what it retrieved Returns a score, a grade, all eight category scores, and the coverage behind them Tells you how many findings there are, free, without an address Holds the findings themselves behind a verified email Produces a share link that expires after 30 days and never unlocks the findings, whoever opens it It does not write anything, and it is labelled advisory throughout, because it has no access to your repository or your CMS and therefore has never verified a build. If the scanner is unavailable, the scan fails and says so. There is no path that invents a grade. Why remediation is priced separately Producing an audit is cheap: it is a crawl and a rubric. Producing a change that builds, that a person can review as a diff, and that gets re-measured afterwards is where the cost sits. Pricing them together would mean charging everyone for work most people do not want. The retainer exists for a different reason. A site that publishes does not stay fixed: new pages introduce new findings, and the overlap between your own posts grows quietly. Monitoring is the thing that has to be ongoing. Pricing questions Is the free scan really free, and what is the catch? No card, no signup. The catch is the cap: it crawls 10 pages, one domain per day, three submissions an hour from the same caller. The score, grade, category breakdown and the number of findings are all free. Seeing what the findings actually are costs a verified email address. Why is a scan capped at 10 pages? A crawl costs real money and the free tier is a lead magnet rather than a product. The response tells you how many pages it covered against how many it found, so you always know what proportion of your site the score is a claim about. What happens if I cancel? The fixes you already merged are yours. They are in your repository or your CMS, not held in ours. You lose the ongoing scans, the queue and the monitoring. Do you charge per seat? No. Plans are priced per site, and the agency tier is priced on volume. Can I pay once instead of subscribing? That is what the remediation tier is. One audit-to-merge cleanup, paid once. The retainer exists because search is not a thing you fix once, but a one-time cleanup is a legitimate way to buy. Start with the free scan Read the full FAQ Keep reading What the free scan actually covers Why report export is not available The full FAQ --- ## The audit URL: https://crawld.co/audit/ THE AUDIT A finding you can argue with. Most audits produce four hundred rows and a bill. The useful part is not the count. It is whether each row shows its working, and whether the ordering reflects what is actually worth your afternoon. What you get Scorecard One overall score, eight category scores, and the coverage behind every one of them. The fix list is ranked by severity multiplied by pages affected, so the ordering reflects what is worth doing rather than what sounds worst. Every number states its denominator. A category score of 44 means 44% of the checks that ran in that category passed, and the coverage figure next to it tells you how many ran at all. Page audit Any crawled page, on its own terms: its own score, what passed, what failed, and, the part that usually matters, what could not be measured on it specifically. Useful when a site-wide score is masking a category of page. Twenty product pages failing identically look like one finding at the site level and twenty problems at the page level. Findings Each failing check becomes a finding attached to the page it was found on, carrying the evidence that produced it, the method that decided it, and who is allowed to fix it. A finding you cannot interrogate is an assertion. The evidence is what lets you disagree with one. Site memory What the engine knows about your corpus as a whole, which posts compete for the same intent, which keywords are already covered, and what shape your titles have converged on. This is the only surface that can see cannibalization, because cannibalization does not exist on any single page. What the link map looks like Every page as a dot, placed by how many clicks it sits from the home page, sized by how many other pages reference it. This one is kazifi.co: 218 pages, 1,378 editorial links, and 7,096 template links left out, because a footer repeated on every page says nothing about structure. Screenshot of the running product, not an illustration. What a scorecard looks like Eight categories, each showing the proportion that passed and the proportion that failed. The remainder of every bar is unmeasured, and it stays visibly grey rather than being absorbed into either side. scorecard · example_scan.crawld 129 checks · 8 categories Metadata 17 checks Structure 14 checks Crawl access 11 checks Speed 15 checks Structured data 18 checks Internal links 12 checks Content 26 checks GEO / AI 16 checks score 71 / 100 coverage 86% · gray never counts as a pass An illustrative scorecard, not a customer's. The check counts are the real rubric; the pass and fail split is an example. How findings are ranked Severity alone is the wrong sort order and it is the one almost everything uses. A missing meta description is low severity. A missing meta description on two hundred pages is the most valuable afternoon available to you, and it will sit below a single high-severity issue on a page nobody visits unless the ranking accounts for reach. So the ordering is severity multiplied by pages affected, and the list is short enough to act on. Coverage, again Every audit surface reports coverage next to its score. This is laboured on purpose: a high score at low coverage is a weaker result than a lower score at full coverage, and presenting the two identically is the single most common dishonesty in this category. Audit your site free See the full rubric Keep reading How the 129-check rubric scores a site Every check, by category Who the audit is built for --- ## The checks URL: https://crawld.co/checks/ THE CHECKS 129 checks, and what each of them is for. Every check declares two things about itself: how it was decided, and who is allowed to fix it. Both are visible wherever a finding appears, because a finding you cannot interrogate is just an assertion. The eight categories Weights are the proportion each category contributes to the overall score. They add to 100%, and they are not equal: content quality carries twice what page experience does, deliberately. Category Checks Weight Technical foundation & indexability 20 15% Crawling, rendering & site structure 16 10% Search appearance 18 10% Content quality & helpfulness 18 20% E-E-A-T & trust signals 15 15% Site content health & cannibalization 14 10% AI / answer-engine readiness (GEO) 16 12% Page experience & performance 12 8% Total 129 100% How a check is decided "Our AI found an issue" is not a finding. Every check states its method, so you can tell a fact from a judgement before you act on it. Rule A deterministic fact about the page. Present or absent, no judgement, reproducible by anyone. Graph Computed across the whole crawled site: reachability, overlap, depth. Cannot be answered from one page. API Measured by an external source. If that source is unavailable, the check is unmeasured, never a pass. Model Assessed by a language model, and labelled as such everywhere it appears so you can weigh it accordingly. Who is allowed to fix it Auto : the engine produces the change. One correct answer exists. Assist : the engine drafts, a person decides. The shape is clear, the content is a judgement. Human : only a person can decide. The engine reports and stops rather than guessing convincingly. The categories in detail Each category has its own page covering what it looks for, what commonly fails, the conditions under which its checks return unmeasured, and where to start if it scores badly. Technical foundation Whether an engine can reach the page and is permitted to keep it. 20 checks · 15% of the score Crawling & structure Whether the content is in the HTML, and whether the site hangs together. 16 checks · 10% of the score Search appearance What the result looks like once you have earned one. 18 checks · 10% of the score Content quality Whether the page answers the thing it appears to be about. 18 checks · 20% of the score E-E-A-T & trust Whether there is anything here to trust or attribute. 15 checks · 15% of the score Content health Whether your own pages are competing with each other. 14 checks · 10% of the score Answer-engine readiness Whether an AI answer can lift your content and credit you for it. 16 checks · 12% of the score Page experience Whether the page is usable once it arrives. 12 checks · 8% of the score A note on what is listed here The checks above are representative of each category, not a transcription of all 129. The category names, counts and weights are the live catalogue; the individual checks are shown to convey what a category contains and how its checks are decided. A browsable page per check (one indexable page for each of the 129, each answering the query someone actually types) is on the roadmap and is not built. This page says so rather than presenting a partial list as a complete one. Core Web Vitals, specifically Shipped: LCP, INP, CLS . Not shipped: FCP, TBT, Speed Index, accessibility score. The full Lighthouse set is not claimed anywhere on this site, because not all of it runs. Run the rubric on your site How the score is calculated Keep reading The scoring method behind the checks The 16 answer-engine checks explained How findings become reviewable diffs --- ## The fix queue URL: https://crawld.co/fix-queue/ THE FIX QUEUE Every fix arrives as a diff, not a suggestion. You don't need to know what a canonical tag is. You need to know whether to click merge, which means the change has to be in front of you, in full, before you decide. What a queue item contains Not a description of a change. The change. The diff itself : what is removed, what is added, with line numbers and surrounding context Which check it satisfies , and what that check looks for Which page or pages it applies to Its verification status , honestly labelled for the access you granted Its fix capability : whether a machine wrote this, drafted it, or refused to src/pages/pricing.astro meta-description-missing · Auto-fix 14 15 - 15 + 16 Build verified Mode A · opens PR on merge Who is allowed to fix what Every one of the 129 checks declares its fix capability, and the engine does not overrule it. This is what stops a language model from confidently rewriting something it has no basis to rewrite. Auto The engine produces the change itself. These are findings with one correct answer: a missing meta description, an absent canonical, a malformed sitemap entry, an image with no alt text where the surrounding content makes the alt obvious. You still approve the diff. "Auto" describes who wrote it, not who ships it. Assist The engine drafts and a person decides. These are findings where the shape of the fix is clear but the content is a judgement call: rewriting a title that is too long without losing what it meant, consolidating two posts that compete for one intent. The draft arrives as a proposal with the reasoning attached. Editing it before approving is the expected path. Human Only a person can decide. Anything requiring knowledge the engine does not have: whether a page should exist, whether a claim is true, whether an author is genuinely an expert on the subject. The engine will not generate a change here. It reports the finding and stops, rather than guessing convincingly. How the queue is ordered The ranking is the useful part, because the list is always longer than the time available. Input What it means Severity How much this specific problem costs on a page where it occurs. Pages affected How many crawled pages carry it. The product of the two A low-severity finding across 200 pages outranks a high-severity one on a single page, because fixing it is worth more. This is the part most audit tools get backwards by sorting on severity alone and handing you 400 rows. The two gates The queue sits between them. You approve the ranked plan before anything is generated, and you approve the delivery before anything ships. Both are explicit clicks and neither can be turned off. A change delivered as a pull request is also one click to revert, which is a property of the delivery mechanism rather than a feature, and a good reason to prefer it where you can. How changes reach your site The verification label follows the access, not the marketing. A hosted CMS has no build to run, so a change written into one is never described as build-verified. Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What runs today The free scorecard, Mode E, is live. It produces findings and a patch you apply yourself, labelled advisory throughout because it has no access to anything and has never seen a build. The connected modes above describe the delivery model the product is designed around. They are not connectors you can authorise this afternoon, and the table says which is which rather than leaving you to find out after paying. See what your site would queue Delivery and integrations Keep reading How a finding is ranked before it reaches you Which platforms can accept a change Why report export is not built yet --- ## Answer-engine readiness (GEO) URL: https://crawld.co/geo/ ANSWER ENGINES Ranking #1 doesn't matter if the answer doesn't mention you. People stopped clicking blue links and started asking. Almost nothing tells you whether you are the site being cited, which is a different measurement from rank, and needs a different instrument. Crawled, indexed, cited: three different things These get used interchangeably and they describe three states. The distance between the second and the third is where AI-referred traffic is won or lost. 01 Crawled A bot fetched the page. That is all it means: necessary, and close to worthless alone. A page can be crawled daily for a year and appear in nothing. 02 Indexed The engine stored the page and can retrieve it. A real milestone and a real bottleneck, and a whole category of tools exists to accelerate it. They also stop here, which is reasonable: the next step is a different problem. 03 Cited where Crawld works An answer engine named you as its source. Indexing makes you eligible . Whether you are actually cited depends on things indexing tools do not measure and cannot influence. What the 16 GEO checks look at Sixteen of the 129 checks sit in the AI / answer-engine readiness category, which carries 12% of the overall score. They are not a separate product bolted on. They are graded in the same rubric, with the same three outcomes, and an unmeasured GEO check is never counted as a pass. Extractability Whether a claim on the page can be lifted as an answer. Engines quote passages, not pages: a fact stated plainly under a heading that matches the question is quotable, and the same fact in the eleventh paragraph of an essay is not. Server-rendered content Whether the page ships its content in its HTML. Googlebot will render JavaScript on a second pass it schedules at its convenience. Most answer-engine crawlers will not, and an empty root div is what they index. Attribution signals Whether the page gives an engine something to credit: a named author, a date, an organisation, the structured data that ties them together. An engine that cannot characterise a source often prefers one it can. Machine-readable structure Headings in order, structured data that matches what the page actually contains, and an llms.txt that says which of your pages are worth reading. Schema that misdescribes a page is worse than none. Crawler access Whether the AI crawlers are permitted at all. A robots.txt written for Googlebot in 2019 frequently blocks half the agents that matter now, silently. The engines tracked Citation share is tracked per engine and per prompt: for the queries you actually compete on, who got cited: you, or someone else. Claude Anthropic's own model, increasingly used for research and shopping. Checks whether your content is structured the way Claude's search actually reads it. ChatGPT The largest AI answer engine by users. Tracks whether ChatGPT's browsing and search cite your pages, not just the competitor who out-ranked you on Google. Gemini Built into Search, Workspace, and Android. Tracks whether Gemini's answers reference you when it actually matters, not just that you rank. Perplexity Built to cite its sources by design, which means it either credits you or it credits someone else. Tracks your citation share directly. Google AI Overviews Sits above the normal results now, and pulls from a narrower slice of the page than classic ranking does. The technical checks are tuned to what it actually lifts. Copilot Built into Windows, Edge, and Bing. Same rubric, same fix queue, no separate setup. Bi Bing Still the index behind Copilot and half the AI crawlers on the web. The classic search checks cover it. Indexed here means visible in far more places than bing.com. De DeepSeek The fastest-growing open-weight model, already embedded in apps you've never heard of. Same rubric, same tracking, whichever wrapper someone's asking it through. What this measurement does not claim Citation share is recorded per prompt and per engine. There is currently no link from a citation back to the specific article that earned it, which means claims of the form "posts with X are cited 41% more" are not something this data can support, and you will not find them anywhere on this site. Citation is also zero-sum per answer. You are not clearing a threshold; you are competing against whichever two or three sources the model found easiest to use. A rising share for a competitor is information about you. Why this is a separate discipline Rank tracking tells you your position on a results page a growing share of your audience never sees. Index checkers tell you the page is stored. Neither tells you whether an engine named you when someone asked the question your page answers, and answering that means asking the engines, repeatedly, and recording who got credited. Scan your site free Read the long version Keep reading Why indexed and cited are different states What llms.txt is for The 16 GEO checks in the rubric --- ## How it works URL: https://crawld.co/how-it-works/ HOW IT WORKS Five stages. Two places it stops and asks you. Most tools end at a report. This one runs a loop, and closes it by measuring again, which is the part that tells you whether any of it worked. The loop Same shape as a pull request you already review every week, except this one writes itself, verifies itself, and waits for your call twice. 01 Measure The crawler fetches your pages the way Googlebot and the AI crawlers do (no browser, no patience for a framework to boot) and runs all 129 checks against what actually came back. Every check resolves to pass, fail, or unmeasured, and the third one is never rounded up. Output: A score, eight category scores, and the coverage behind each of them. 02 Audit Each failing check becomes a finding attached to the specific page it was found on, with the evidence that produced it and a statement of how it was decided: a deterministic rule, a graph analysis, a third-party API, or a model judgement. Output: A findings list you can argue with, because each one shows its working. 03 Plan waits for you Findings are ranked by severity multiplied by the number of pages affected, so a low-severity problem across two hundred pages outranks a high-severity one on a single page. The ranked plan is presented, and the loop stops. Output: A ranked remediation plan, waiting on you. Nothing proceeds until you approve it. 04 Remediate Only the approved items are worked. Each check declares who is allowed to fix it: Auto means the engine produces the change, Assist means it drafts and a person decides, Human means only a person can. The engine does not overrule that classification. Output: A set of proposed changes, each one verified as far as your access allows. 05 Deliver waits for you The changes are packaged for however much access you have granted, and the loop stops a second time. Nothing reaches your site until you approve this too. Output: A pull request, a set of drafts, or a patch file, and your decision. Then it measures again After a delivery lands, the next scan re-runs the same checks against the same pages. If a fix did not move its check from fail to pass, that gets flagged rather than quietly forgotten. This is the part almost no tool in the category does, and it is the only way to tell a change that worked from a change that merely shipped. It is also the reason the score is worth watching over time rather than screenshotting once. A single number from a single crawl is a snapshot. The same number across four crawls, with the fixes in between marked, is evidence. What the two gates are for They are a design decision, not a limitation. The engine is capable of generating a plan and generating changes without stopping; it stops because an unsupervised agent editing a production site is a different product with a different risk profile, and not the one being built here. The plan gate is where you decide what is worth doing. Approving everything is a valid answer; so is approving three items and ignoring the rest. The delivery gate is where you decide whether a specific change is correct. You are looking at a diff, not a description of a diff. Crawld is explicitly not for teams who want a fully autonomous agent with no review step. If that is what you are shopping for, the gates will feel like friction, and they are not going away. How changes reach your site How much can be verified depends on how much access you grant, and the labelling follows the access rather than the marketing. Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What runs today The free scorecard is Mode E and it is live: give it a URL, it crawls a capped slice of public HTML, scores it against the full rubric, and reports what it could not measure. It writes nothing and it verifies nothing, because it has no access to anything. The connected modes, the ones that open a pull request or write CMS drafts, are the delivery model the product is built around, not connectors you can switch on this afternoon. The table above marks which is which. Run the free scan How the scoring works Keep reading Why unmeasured is never counted as a pass The fix queue, and the two approval gates Which delivery modes exist today --- ## How we measure URL: https://crawld.co/how-we-measure/ HOW WE MEASURE Every number on your scorecard states its denominator. A score is a claim about what was actually measured. This page explains exactly what that claim covers, and what it deliberately does not. The rubric 129 checks across eight weighted categories. Each category contributes to the overall score in the proportion below, so a site that is technically flawless but thin on content cannot score its way to a clean bill of health on infrastructure alone. Category Checks Weight Technical foundation & indexability 20 15% Crawling, rendering & site structure 16 10% Search appearance 18 10% Content quality & helpfulness 18 20% E-E-A-T & trust signals 15 15% Site content health & cannibalization 14 10% AI / answer-engine readiness (GEO) 16 12% Page experience & performance 12 8% Total 129 100% Three outcomes, not two Every check on every page resolves to one of three states, and the third one is the point of this page. Pass : the rule ran against the page and the page satisfied it. Fail : the rule ran and the page did not satisfy it. It enters the fix queue. Unmeasured : the rule could not run. The crawler could not reach the page, the resource timed out, a third-party API was unavailable, or the check requires data this site does not expose. Unmeasured is not a pass. It is not rounded up, it is not quietly dropped from the denominator, and it does not silently improve your score. It appears as its own state in every view, and the coverage figure tells you how much of the rubric actually ran. This matters more than it sounds. The default in this category is to treat what could not be checked as fine, which is how tools report 95/100 for sites with real problems: they scored the third of the site they reached and presented it as the whole. What coverage means Coverage is the share of applicable checks that produced a verdict, pass or fail, rather than an unmeasured. It is reported per page, per category and for the site as a whole. A score of 71 at 86% coverage means something precise: of the checks that ran, 71% passed, and 14% of the rubric never ran at all. Both halves of that sentence are on the scorecard. A high score at low coverage is a weaker result than a lower score at full coverage, and the interface says so rather than leaving you to work it out. How each check is decided Every check declares its own method, because "our AI found an issue" is not a finding: Deterministic rule : a fact about the page. Present or absent, no judgement involved. Graph analysis : computed across your site's structure, such as internal link reachability. Third-party API : measured by an external source. Unavailable upstream means unmeasured, not pass. Model judgement : assessed by a language model, and labelled as such wherever it appears. Each check also declares who can fix it: Auto : the engine produces a verified change. Assist : the engine drafts a change and a person decides. Human : only a person can decide. The engine will not pretend otherwise. The free scan is a slice, and says so The public scorecard crawls a capped number of pages. It reports how many it covered against how many it found to exist, so the number you get is explicitly a claim about that slice. It is advisory: it inspects public HTML and does not verify changes against a build. If the scanner is unavailable, the scan fails and says so. There is no path in which a grade is produced without a crawl behind it. Scan your site free Keep reading All 129 checks, by category What a scorecard and page audit contain How rubric versions are tracked --- ## SEO Watchdog URL: https://crawld.co/seo-watchdog/ SEO WATCHDOG Find out what they changed, the week they changed it. Competitor monitoring is usually either a spreadsheet somebody stopped updating or a scraper that gets you blocked. This is a weekly diff of the pages that matter, built to be unobjectionable to the site on the other end. What it reports Tracked Why it is worth knowing Title tags The clearest signal that someone has re-targeted a page at a different query. Meta descriptions Usually a click-through experiment. Worth knowing when a competitor runs one. Headings A restructured H2 set is a rewritten argument, even when the word count barely moves. Body copy What actually changed on the page, not that something did. Pricing The change most likely to need a response the same week. Links New internal routes show what they are consolidating around; new external ones show who they are citing. Reports are diffs, not snapshots. "This H2 changed from X to Y on Tuesday" is actionable; "here is their current page" is a browser. Polite by construction The constraints below are not settings. They are how the crawler is built, and they exist because pointing automated traffic at a third party is a thing you should have to justify. Limit Detail Up to 3 competitor sites Per brand. Enough to watch the people you actually lose to. 12 pages per site It does not walk a competitor’s whole site, and it is not trying to mirror one. One request at a time No parallel fetching, ever. It is not capable of loading anyone’s server. Once a week The scheduler runs on a 168-hour interval. robots.txt obeyed Including crawl-delay. A site that disallows CrawldBot is not crawled. Identifies itself The user agent names the bot and links to a page explaining how to block it. If you have arrived here from your own access logs, the crawler page explains how to block it, and blocking works. Actionable Insights Coming soon Turning a detected change into a recommended response ("they have re-targeted this page at your primary query, here is what that implies") is not built. The in-app copy marks it coming soon and so does this page. What ships today is the change detection. Why three sites Because monitoring twenty competitors is a way of monitoring none of them. Three forces the question of who you actually lose traffic to, which for most sites is a shorter list than the one in the pitch deck. Start with your own site How the crawler behaves Keep reading What CrawldBot fetches and how to block it The crawl limits it works within Why agencies use it on a retainer --- ## Who it's for URL: https://crawld.co/who-its-for/ WHO THIS IS FOR Built with AI. Should still show up in search. You don't need to have read a single SEO blog post. The rubric reads your site and reports what is wrong in language you can act on. Vibe coders and solo builders You shipped something real in a weekend. The parts a crawler needs got skipped. Cursor, Claude Code, v0, Bolt, Lovable: whatever built it, speed came from not stopping to write metadata. The site works perfectly in your browser and arrives at a crawler as an empty div with a placeholder title. What this usually looks like Organic traffic that never started, rather than traffic that fell Pages that render client-side and ship no content in their HTML An empty on every route No sitemap, no canonical tags, no llms.txt Why this fits: You are technical enough to merge a pull request and uninterested in becoming an SEO. The fix queue is a to-do list where every item shows exactly what it changes. Content teams publishing at scale The risk is not quality. It is your own posts competing with each other. When you publish a lot, with or without AI drafting, the failure mode is cannibalization: four posts targeting one intent, splitting the signal between them, none of them winning. It is invisible from inside the CMS because each post looks fine on its own. What this usually looks like Several posts ranking weakly for the same query instead of one ranking well Rankings that move sideways when you publish more No shared view of which intents you have already covered Titles converging on the same shape without anyone deciding they should Why this fits: Site memory tracks what your corpus already covers, and overlap is flagged before you publish rather than diagnosed six months later. Agencies and consultants Audit-to-merge as a repeatable deliverable, not a bespoke project each time. The work is the same shape for every client and gets rebuilt from scratch every time. What makes it hard to package is the gap between the audit, which is easy to produce, and the remediation, which is where the hours actually go. What this usually looks like A retainer where the audit is the deliverable and the client does the work Reports that cannot honestly distinguish "we checked this" from "we could not" Competitor monitoring done by hand, or not at all No defensible way to show a client that last quarter moved anything Why this fits: The unmeasured state alone changes what you can put in a client report, because you stop having to present a gap as a pass. Who it is not for Teams who want a fully autonomous agent with no review step. The loop stops twice and asks, and that is a design decision rather than a missing feature. If the value you are buying is "nobody has to look at it", this is the wrong tool and the gates will feel like an obstacle for as long as you use it. It is also a poor fit for a site with a handful of pages and no publishing cadence. Most of what the rubric catches (crawl waste, cannibalization, coverage gaps) needs enough surface area to be worth automating. A five-page site is an afternoon of manual work, and the free scan will tell you that for nothing. Scan your site free See pricing Keep reading All audience pages, with the poor fits named What the rubric measures Pricing by tier --- ## Write URL: https://crawld.co/write/ WRITE The overlap check belongs before you publish. Cannibalization is not a content-quality problem. It is four good posts targeting one intent, splitting the signal, none of them winning, and it is invisible from inside the editor unless something is watching the whole corpus. What the editor does while you write Live GEO checklist The same answer-engine checks that run in the audit, applied to the draft in front of you: is the claim extractable, does the heading match a question someone asks, is there attribution metadata for an engine to credit. Fixing these while writing costs nothing; fixing them across two hundred published posts is a project. Live E-E-A-T checklist Whether the piece has a named author, a date, sourced claims, and the marks of someone who actually knows the subject. These are mostly Human-capability checks: the engine will tell you a claim is unsourced and will not invent a source for it. Originality against your own corpus Not plagiarism detection against the web. A similarity score against everything you have already published, because the thing most likely to outrank your new post is your old one. OK "Best CRMs for solo agencies" distinct topic FLAG "Top CRM tools for small teams" overlaps post #14 OK "CRM pricing compared, 2026" distinct topic An assistant that drafts, rewrites and retitles in place Working on the document rather than in a chat window beside it. Its suggestions carry the same classification as everything else: a retitle is Assist, and a claim about your product is Human. The calendar The content pipeline as a month view: scheduled, generating, draft, published. It exists mostly so that overlap is visible before two people commission the same piece, which is where a meaningful share of cannibalization originates. What this is not It is not a volume machine. There are tools in this category whose pitch is publishing at a rate no human reviews, and the predictable result is a corpus competing with itself: the exact failure this editor is built to catch. Writing more is not the goal; writing things that do not cancel each other out is. The overlap check is also advisory. It flags that two pieces target one intent; deciding which one should exist is a judgement about your business, and the engine does not make it. Start with a free scan How overlap is detected Keep reading How cannibalization is detected across a corpus Why content brands hit this first The content-ops tier --- ## Blog URL: https://crawld.co/blog/ FROM THE BLOG The playbook, written down. How scoring works, what the crawler reads, and what actually gets a site cited by answer engines. All posts (13) Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) Strategy · 10 min read Competitor SEO analysis without guessing at their data Three steps to build the rival list from search results, four things you can read in their HTML, and the estimates to treat as directional only. Read → Strategy · 9 min read What an SEO audit costs and how often to run one Free to five figures for the same-sounding deliverable. Five variables explain the spread, plus how often to re-run one and when to do it yourself. Read → Content strategy · 11 min read Content cannibalization: how to find it and what to do Four symptoms, one Search Console report that confirms it, and the choice between merging, differentiating and deleting the pages that compete. Read → On-page SEO · 10 min read On-page SEO: the elements worth checking on every page Six elements worth checking on every page, from titles to image attributes, plus the three answer-engine elements standard checkers still ignore. Read → Analytics · 10 min read How to check website traffic, yours and anyone else Four sources for your own numbers, four metrics worth watching, and how far an estimator can be trusted for a site you do not own. Read → Link building · 11 min read Link building strategies that still work, and three that do not The criteria that decide whether a tactic is worth running: a linkable asset, a pitch that cites real work, a live target. Four that pass, three to skip. Read → Link building · 10 min read How to check your backlinks and audit what you find Ten minutes to pull the data, longer to judge it, which sources to trust, what the authority scores mean, and why disavow is a last resort. Read → GEO basics · 8 min read Generative engine optimization, explained without the hype Ranking well does not make a page citable. What an answer engine needs instead, and the page changes that make a claim liftable. Read → Technical · 8 min read How to run a technical SEO audit that finds real problems Six passes over crawl access, rendering, architecture, structured data and Core Web Vitals, plus what to do about the checks a crawler cannot make. Read → Technical · 9 min read The SEO audit checklist that states its denominators Seven passes over indexability, rendering, structure, metadata, content health, answer-engine readiness and page experience, ordered by what breaks first. Read → Strategy · 7 min read Crawled, indexed, cited, where AI traffic begins Three gates, and passing the first two says nothing about the third. What decides whether an answer engine names you as the source. Read → Technical · 4 min read llms.txt, explained for humans A plain-text map of your best pages for AI crawlers. What belongs in it, what does not, and the honest answer on whether it changes anything yet. Read → GEO basics · 6 min read Why your weekend project is invisible to ChatGPT Three gaps every AI-built site ships with, the curl command that shows each one, and roughly twenty minutes of work to close them. Read → --- ## Compare URL: https://crawld.co/compare/ COMPARE Where the others stop. Every tool below is good at the thing it was built for. These pages are about which job you are actually trying to do, and each one ends by saying when the other tool is the right buy. Crawld vs Surfer Surfer optimises a page against the current SERP. Crawld audits the whole site and ships the fix. Read the comparison → Crawld vs SEObot SEObot produces content at volume. Crawld checks that what you publish does not compete with itself, and verifies changes before they ship. Read the comparison → Crawld vs IndexRusher IndexRusher gets you indexed. Crawld measures whether being indexed turned into being cited. Read the comparison → Crawld vs Semrush Semrush is a much broader suite. Crawld is narrower, tracks AI-citation share, and ships the fix rather than reporting it. Read the comparison → The short version Tool Crawld Them Surfer GEO + verified remediation PRs On-page SEO scoring only SEObot Overlap-checked content + verified fixes High-volume AI content generation, no verification step SEOTesting Fix → verify → re-measure loop Reporting on GSC data IndexRusher Tracks whether AI answers actually cite you Fast indexing across Google, Bing, LLMs; stops before citation Refresh Agent Full-site audit + human gates Single-article refresh MarketMuse / Clearscope Ships build-verified changes Content briefs, no delivery Semrush AI-citation share tracking + fixes Rank tracking, no remediation What actually distinguishes this Four things, and the fourth is the one worth leading with. Verified, not suggested. Changes ship as build-verified artifacts wherever the access allows it, and are labelled honestly wherever it does not. GEO-first measurement. Citation share across answer engines, not only blue-link rankings. Two human gates. Plans and deliveries pause for approval. The engine never ships unsupervised. Honest scoring. A check that could not run is never counted as a pass. This sounds like table stakes and is not: the category default is to round unmeasured up to passing, which is how tools report 95/100 for sites with real problems. A note on how these pages are written No pricing is quoted for anyone else, because prices change and a stale figure quoted about a rival is unfair. No performance claims are made about their products, because we have not benchmarked them. Each page states what the other tool is genuinely built for, and each ends with the case for buying it instead. Scan your site free How we score Keep reading The scoring method rivals are compared against What is shipped and what is not Answer-engine measurement --- ## Integrations URL: https://crawld.co/integrations/ API & INTEGRATIONS Fixes arrive the way your stack accepts them. What can be delivered, and how honestly it can be verified, depends entirely on what your platform lets a third party write. That constraint is real, so it is the organising idea here rather than a footnote. Read this first. The connectors below are not built. The delivery model is real and is what the product is designed around, but today the only thing you can run without granting any access is the free scorecard: Mode E on the ladder. Every page below carries its status, and none of them describes a connector you can authorise this afternoon. The five delivery modes Every platform lands on one of these rungs, and the verification label follows the rung rather than the marketing. Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. Build REST API Drive scans and read findings programmatically. Mode A Not built yet Next.js Fixes arrive as a pull request, built before you see it. Mode A Not built yet Webhooks Get told when a scan finishes or a fix is ready. Mode A Not built yet Zapier Route scan results into whatever your team already uses. Mode D Not built yet Platforms WordPress Changes land as unpublished drafts for you to publish. Mode D Not built yet Webflow CMS collection items updated as drafts. Mode D Not built yet Shopify Product and page metadata written through the Admin API. Mode D Not built yet Squarespace Advisory only: a checklist you apply by hand. Mode E Not built yet Ghost Post metadata and content updated as drafts. Mode D Not built yet More platforms HubSpot CMS pages and blog posts updated as drafts. Mode D Not built yet Framer Advisory only: no content-write API exists. Mode E Not built yet Notion Page properties updated where a site generator reads them. Mode D Not built yet Unicorn Platform Advisory only. Mode E Not built yet Why a CMS is never build-verified A repository can be compiled. A hosted CMS cannot. There is no build step to run, so there is nothing to verify against. A change written into WordPress or Webflow is checked in preview and labelled preview-verified, and it never carries the build-verified badge. This is the distinction the whole ladder exists to preserve. Blurring it would make the badge meaningless everywhere, including where it is earned. Platforms with no content-write API Squarespace, Framer and Unicorn Platform have no supported way for a third party to write page content or metadata. They sit on Mode E, and the findings arrive as instructions you apply in the editor. Claiming otherwise would be inventing a capability, and those pages say so plainly rather than implying a connector is coming. Run the free scan Read the API docs Keep reading How fixes arrive as reviewable diffs What is shipped and what is not API endpoints --- ## Solutions URL: https://crawld.co/solutions/ SOLUTIONS Same rubric. Different findings do the damage. There is one product and one 129-check rubric. What changes between these is which failures dominate, and every page below ends by saying where the fit is bad. Designed for These three are named in the product brief as the audiences it was built around. Content brands Enough pages that cannibalization and crawl waste are real costs, and no in-house SEO engineer. Read → Agencies and consultants Audit-to-merge as a repeatable retainer deliverable rather than a bespoke project each time. Read → Small technical teams Technical enough to merge a pull request, uninterested in becoming an SEO. Read → Also applies to Segments where the rubric does real work, written from what it actually finds on that kind of site rather than from a persona. SEO professionals A rubric that shows its working, and a coverage figure you can defend in a meeting. Startups Find out why organic traffic never started, before you spend on ads to compensate. Enterprise Read this before you start a procurement process: several answers are no. Not built yet E-commerce Coming soon: the catalogue-specific checks are not built. Not built yet Publishers Archive-scale corpora, where overlap and decay compound quietly. Nonprofits The free scorecard is a complete diagnostic and costs nothing. Higher education Sprawling multi-department sites where nobody owns the whole domain. Who it is not for Teams who want a fully autonomous agent with no review step. The loop stops twice and asks, by design. If the value you are buying is that nobody has to look at it, the gates will feel like friction for as long as you use the product, and they are not going away. Also a poor fit for a site with a handful of pages and no publishing cadence. Most of what the rubric catches (crawl waste, cannibalization, coverage gaps) needs enough surface area to be worth automating. The free scan will tell you that for nothing. Scan your site free See pricing Keep reading The rubric every segment shares Pricing by tier Who it is explicitly not for --- ## Agencies and consultants URL: https://crawld.co/solutions/agencies/ SOLUTIONS Crawld for agencies and consultants Audit-to-merge as a repeatable retainer deliverable rather than a bespoke project each time. One of the three audiences this was designed for. The others are on the solutions index . The situation The work is the same shape for every client and gets rebuilt from scratch every time. The audit is easy to produce; the hours go into turning it into changes somebody actually ships. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. The unmeasured state The single most useful thing for client reporting. You stop having to present a gap as a pass, and coverage gives you a defensible answer to "did you check everything?" Severity × pages affected A ranked list you can turn into a scope of work, rather than 400 rows the client has to triage. Competitor monitoring Weekly change reports on up to three competitor sites per brand: the recurring input a retainer needs. What the loop looks like Scan on intake to set a baseline, work the ranked plan, re-measure at the end of each cycle. The re-measurement is what makes a renewal conversation evidential rather than anecdotal. Where this is a poor fit White-label reporting is on the agency tier; multi-site management is real but the connected delivery modes that would open a PR into a client repository are not built yet. Today the deliverable is the audit and the plan. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## Content brands URL: https://crawld.co/solutions/content-brands/ SOLUTIONS Crawld for content brands Enough pages that cannibalization and crawl waste are real costs, and no in-house SEO engineer. One of the three audiences this was designed for. The others are on the solutions index . The situation You publish steadily and the library has grown past the point where anyone holds it in their head. Individual posts look fine. The corpus is where the problem lives, and nothing in your CMS can see the corpus. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. Cannibalization Several posts targeting one intent, splitting the signal so none of them ranks. Fourteen checks cover this, and it is invisible per-article by construction. Crawl waste Budget spent on pagination, faceted URLs and pages that should never have been indexed, instead of on the pages you care about. Coverage gaps What you claim to be about versus what you have actually written: computed across the whole crawled site rather than page by page. What the loop looks like Scan the library, get the overlap map, consolidate or differentiate the pages that compete, then re-measure to confirm the consolidation moved the right one up rather than killing both. Where this is a poor fit Below roughly fifty pages there is not enough surface area for corpus-level analysis to tell you much you could not work out by hand in an afternoon. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## Small technical teams URL: https://crawld.co/solutions/technical-teams/ SOLUTIONS Crawld for small technical teams Technical enough to merge a pull request, uninterested in becoming an SEO. One of the three audiences this was designed for. The others are on the solutions index . The situation You shipped something real and fast. The parts a crawler needs got skipped, because they are invisible in a browser and nothing told you they were missing. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. Client-side rendering The most common finding on sites built with AI tooling. Your page works perfectly in a browser and arrives at an answer-engine crawler as an empty div. Missing metadata Empty descriptions, placeholder titles, no canonical tags: the scaffolding defaults nobody went back to fill in. Machine-readable structure Sitemaps, llms.txt, structured data that matches what the page actually contains. What the loop looks like Scan, read the ranked list, fix the top few yourself in an afternoon. Most of what dominates here is Auto-capability work. One correct answer, no judgement required. Where this is a poor fit If you want the fixes to arrive as a pull request rather than as a list you action yourself, that is the delivery model the product is designed around and it is not built yet. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## How to run a competitor SEO analysis URL: https://crawld.co/blog/competitor-seo-analysis/ STRATEGY Competitor SEO analysis without guessing at their data Three steps to build the rival list from search results, four things you can read in their HTML, and the estimates to treat as directional only. 27 May 2026 · 10 min read Most competitor SEO analysis leans on numbers nobody can verify. The useful version starts with what you can observe directly, which turns out to be quite a lot. Pick competitors by search, not by industry Your business rivals and your search rivals overlap less than you would expect. The pages beating you for a commercial query are often a review site, a forum thread and a publisher, none of whom sell what you sell. Build the list from results, not from a board deck: Take your ten most commercially important queries. Record the top five results for each, in a private window. Count which domains appear repeatedly. That frequency list is who you actually compete with. Three to five names is enough. More than that and you are monitoring rather than analysing. What you can observe directly A competitor’s HTML, rendering behaviour and internal linking are all public, and reading them directly costs nothing. You will learn more from viewing source on the page outranking you than from any authority score, because you are seeing what they actually shipped rather than a vendor’s estimate of it. Their page structure. Open the page ranking above you and read it as a searcher. What question does it answer in the first hundred words? What subheadings does it use? What does it cover that you do not? Their HTML. View source and check the title, description, headings, structured data and canonical. You will learn more about their template strategy in five minutes than from a dashboard. What arrives without JavaScript. Fetch their page the way a crawler does and compare it to the browser view. A competitor whose content is painted client-side has a weakness that is expensive for them to fix and free for you to exploit. Their internal linking. Which pages do they link to from the navigation and from body copy? That is their own statement about what matters, and it is more reliable than any third-party authority score. What you cannot observe, and should treat carefully Traffic estimates, keyword volumes and authority scores are all inferred. Every one is a vendor’s model built on a vendor’s crawl and a purchased clickstream panel. Two rules keep them useful. Compare within one tool, never across them, because the scales are not commensurable. And trust direction over magnitude: “their organic visibility doubled since January” is usually right, while “they get 47,300 visits a month” is a point estimate with an error bar nobody prints. Run one competitor through two estimators and note the gap. That gap is the honest confidence interval on everything else the tool tells you. A worked example You sell project management software and a rival consistently outranks you for evaluation queries. Rather than starting with their backlink profile: Read their ranking page. It answers “which tool for a five-person agency” in the opening paragraph. Yours opens with a company history. Check the source. They have FAQPage structured data and you do not. They have a named author with a profile; your page is unattributed. Fetch without JavaScript. Their content is server-rendered. So is yours, so no advantage either way. Look at internal links. Every one of their comparison posts links to that page with descriptive anchor text. Yours is linked from the footer. Ask the answer engines. Four assistants, same question. They are cited by three, you by none. Four of those five findings are things you can fix this quarter, and none required a subscription. Watching for changes over time A snapshot tells you where they are. A diff tells you what they are doing. Track a small set of pages and record what changes: title tags, meta descriptions, headings, body copy, pricing and internal links. Title changes are the clearest early signal, because a rewritten title usually means a page has been re-targeted at a different query. Do it politely. If you build your own monitoring, respect robots.txt , fetch one page at a time, identify your crawler honestly, and keep the frequency low. Weekly is plenty. Anyone hammering a competitor’s server hourly will be blocked, and deserves to be. Crawld’s monitoring works within exactly those limits: up to three competitor sites, twelve pages each, one request at a time, once a week, robots.txt obeyed. The watchdog page covers what it reports , and the crawler page documents the limits and how any site can block it . The measurement nobody is running Ask the four main answer engines a question your best page answers, and record who gets cited. Repeat monthly. Almost nobody does this, which makes it the cheapest competitive advantage available. Rank tracking measures a surface a shrinking share of your audience sees, and it is silent on whether an assistant named your competitor instead of you. Crawled, indexed and cited explains why those are three different states . Turning analysis into work Competitor analysis fails when it produces admiration instead of a task list. For every finding, write the change it implies and who does it. Findings with no owner are trivia. Then check your own house first. A rival’s advantage often turns out to be that their pages can be read and yours cannot. The free Crawld scan grades your side against 129 checks and reports what it could not measure, which is the baseline any comparison needs. Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## How to find and fix content cannibalization URL: https://crawld.co/blog/content-cannibalization/ CONTENT STRATEGY Content cannibalization: how to find it and what to do Four symptoms, one Search Console report that confirms it, and the choice between merging, differentiating and deleting the pages that compete. 13 May 2026 · 11 min read Content cannibalization is when several of your own pages compete for one search intent. Each looks fine on its own, which is why it survives so long, and why the fix has to start with the corpus rather than the page. What it is, and what it is not Cannibalization happens when two or more pages target the same intent closely enough that a search engine cannot tell which one to rank. It picks one, sometimes a different one each week, and both perform worse than a single strong page would. It is not simply keyword overlap. Two pages can share a phrase and serve different intents: a page about what a canonical tag is and a page about how to implement one target readers at different moments. That is topical coverage and it is healthy. The test is intent. If a searcher would be satisfied by either page, you have one page too many. The symptoms, before you go looking Cannibalization announces itself in reporting well before anyone names the cause. Rankings oscillate between your own URLs for a single query, publishing more stops lifting traffic, several pages sit between positions 8 and 25 without breaking through, and a tangential post outranks the commercial page it was meant to support. The last one is the most commonly misread, because the tangential page looks like a success. Finding it with Search Console You do not need a tool. Search Console has enough. Open Performance , set the range to the last three months. Add a Query filter for a term you care about. Switch to the Pages tab. If more than one URL appears with meaningful impressions, those pages are sharing the query. Repeat for your top 20 commercial queries. Twenty minutes gets you a real list. For a wider sweep, export the query and page data, then group by query and count distinct pages with more than about 50 impressions. Any query with two or more is a candidate. A worked example A software company selling scheduling tools finds this for the query “team scheduling software”: URL Impressions Avg position Clicks /blog/team-scheduling-guide 4,200 14.2 61 /features/scheduling 3,800 17.6 38 /blog/how-to-schedule-shifts 1,100 22.9 7 Three pages, one intent, none on page one. The reader searching that term wants to evaluate a product, so the commercial page should own it. The resolution: /features/scheduling becomes the target. It gets the direct answer near the top, and internal links from both blog posts using descriptive anchors. /blog/team-scheduling-guide is re-angled toward a genuinely different intent, “how to build a shift rota”, and stops competing. /blog/how-to-schedule-shifts is thin and overlaps the guide. It is merged into it and 301 redirected. Two months later the commercial page holds position 6 and the combined clicks are higher than the three pages managed separately. Choosing between merge, differentiate and delete Merge when two pages say the same thing and neither is clearly better. Combine the strongest sections into one URL, redirect the other, and keep the URL with more inbound links. Differentiate when both pages have a real reason to exist but drifted together. Change the angle, the title and the opening so each serves a distinct question. This is usually the right call for a commercial page and a genuinely educational one. Delete and redirect when a page is thin, dated and earns nothing. Redirect it to the closest relevant page rather than letting it 404, and resist redirecting everything to the homepage. Whichever you pick, update the internal links pointing at the loser. A redirect that still receives dozens of internal links is a signal you did not finish the job. Why it is invisible from inside a CMS Every page in your editor looks reasonable, because it was written to be. The failure only exists in the relationship between pages, and no page-level view can show a relationship. This is why cannibalization survives content audits that grade pages individually and score well. The audit was asking the wrong unit of analysis. Stopping it before it starts The cheapest fix is at commissioning, not after publishing. Keep a list of the intents you already cover, with the URL that owns each one. Before commissioning, search your own site for the target query and read what comes back. Give every new brief a named intent and an explicit note on how it differs from the nearest existing page. Watch for converging titles. Titles drift together before content does. What good looks like One page per intent, each with a clear owner, and internal links pointing at the page you want to rank rather than scattered across near-duplicates. Fewer pages, doing more work. Detecting this needs the whole crawled site, which is why 14 of the 129 checks in the Crawld rubric sit in a content-health category rather than a page-level one. The audit page explains how overlap is computed , the editor flags it before you publish , and the free scan reports it alongside everything else, with the coverage figure that says how much of your site it actually read. Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## Crawled, indexed, cited, where AI traffic begins URL: https://crawld.co/blog/crawled-indexed-cited/ STRATEGY Crawled, indexed, cited, where AI traffic begins Three gates, and passing the first two says nothing about the third. What decides whether an answer engine names you as the source. 25 February 2026 · 7 min read Three words get used as though they mean the same thing. They describe three different states, and the distance between the second and the third is where most AI-referred traffic is won or lost. Crawled A bot fetched your page. That is all it means. Crawled is necessary and close to worthless on its own. It proves your server responded and the URL exists. It says nothing about whether the content was readable, whether it was stored, or whether anyone will ever see it. A page can be crawled every day for a year and appear in nothing. This is the step people over-index on, because it is the easiest to verify: your logs show the bot arriving, and that feels like progress. Indexed The engine stored your page and can retrieve it. This is a genuine milestone and a real bottleneck, particularly for new sites and for anything published faster than a crawler naturally revisits. A whole category of tools exists to accelerate it, pinging indexing APIs across Google, Bing and the LLM crawlers, and they work: indexing latency is a real problem with a real solution. They also stop here. Which is reasonable, because it is a different problem from the next one. Cited An answer engine named you as its source. This is the one that produces traffic now, and it does not follow from indexing. Being in the index makes you eligible to be cited. Whether you actually are depends on things indexing tools do not measure and cannot influence: Whether the answer is extractable. Engines lift passages, not pages. A claim buried in the eleventh paragraph of an essay is harder to quote than the same claim stated directly under a heading that matches the question. Whether you are attributable. Some engines cite by design: Perplexity is built around it. Others cite selectively, and prefer sources they can characterise: a named author, a date, a site that looks like it knows the subject. Whether a competitor is easier to quote. Citation is zero-sum per answer. You are not competing against a threshold; you are competing against whichever three sources the model found easiest to use. Why the gap persists Because almost nothing measures it. Rank tracking tells you your position on a results page that a growing share of your audience never sees. Index checkers tell you the page is stored. Neither tells you whether ChatGPT, Claude, Perplexity or AI Overviews named you when someone asked the question your page answers. Measuring it means asking the engines, repeatedly, for the queries you compete on, and recording who got cited. That is a different instrument from a rank tracker, and it is the one that tells you whether the work is landing. What to do about it Start by finding out where you actually stand: not on rank, on citation share for the queries that matter to you. Then work the extractability problem: headings that match real questions, answers stated plainly near the top, structure a machine can lift without inferring. None of that is a trick. It is the same advice as “write clearly”, enforced by something that reads your site the way the engines do. Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## Generative engine optimization explained URL: https://crawld.co/blog/generative-engine-optimization/ GEO BASICS Generative engine optimization, explained without the hype Ranking well does not make a page citable. What an answer engine needs instead, and the page changes that make a claim liftable. 1 April 2026 · 8 min read Generative engine optimization is the work of getting AI answers to cite you as their source. It overlaps with SEO, it is measured differently, and most sites are doing none of it. What the term actually covers Generative engine optimization (GEO) is optimising for systems that answer a question directly instead of returning ten links: ChatGPT, Claude, Perplexity, Gemini, Google’s AI Overviews. The output is a paragraph with two or three sources attached. You are either one of them or you are invisible. That last part is what makes it a separate discipline. Ranking is graded on a curve across ten positions. Citation is close to binary per answer, and the number of slots is small. Why ranking well does not carry over Ranking well does not make a page citable. Position does not transfer, because an engine composing an answer prefers whichever source it can lift a clean claim from rather than whichever sits at the top. Click-through optimisation matters less when there is no click, and rendering tolerance is lower than Googlebot’s. Position does not transfer. An engine composing an answer is not obliged to prefer the top result. It prefers the source it can most easily extract a clean claim from. Click-through optimisation is beside the point. There is often no click. Your title tag and meta description, which exist to win a click from a results page, do less work when the result is a paragraph that quotes you. Rendering tolerance is lower. Googlebot will come back and render your JavaScript on a second pass. Most answer-engine crawlers read the HTML they are given and move on. What an engine needs from a page An answer engine needs four things before it can cite a page: the content present in the HTML rather than painted in afterwards, a claim stated plainly enough to lift as written, something to attribute it to such as an author and a date, and permission to crawl the page at all. The content is in the HTML. Not painted in after hydration. The claim is extractable. Stated plainly, near a heading that resembles the question. There is something to attribute. A named author, a date, an organisation. The crawler is allowed in. Many robots.txt files predate these agents and block them by omission. Extractability, shown properly Extractability is whether one passage on the page states the answer outright. It moves more than anything else on this list and it is a writing change rather than a technical one, because an engine quoting you needs a sentence it can take as written, without stitching it together from three paragraphs. Before We are often asked about the cost of a technical audit, and the honest answer is that it varies quite a lot depending on scope, the size of the site, and whether remediation is included, though most agencies land somewhere in a fairly broad band. After A technical SEO audit typically costs between $500 and $5,000. The range depends on three things: the number of pages crawled, whether remediation is included, and whether the audit is a one-off or part of a retainer. The second version can be lifted as an answer. The first cannot, because there is no sentence in it that states the answer. Same information, same honesty, different shape. Attribution, which is cheap and often missing An engine deciding between two equally useful pages will often prefer the one it can characterise. Give it something: < article > < h1 >How answer engines choose sources</ h1 > < p class = "byline" >By Dana Whitfield, Head of Research at Corvid Analytics</ p > < time datetime = "2026-03-11" >11 March 2026</ time > </ article > That is three lines. It costs nothing and it is the difference between an anonymous page and an attributable one. Letting the right crawlers in Check your robots.txt for the agents that matter now, not the ones that mattered in 2019: User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: ClaudeBot Allow: / Blocking them is a legitimate choice. Blocking them without knowing is not a choice at all. How to measure whether any of it worked Rank tracking cannot answer this, because the surface it measures is not the surface you are being cited on. What you want is citation share: for the queries you compete on, how often each engine names you, tracked over time. A serviceable manual version costs an afternoon. Pick twenty questions your best pages answer, ask each of four engines, record who got cited, repeat monthly. That is a real baseline, and it will tell you more than a rank tracker will about this particular problem. One caution on interpretation. Citation data is recorded per prompt and per engine, so it supports statements like “our share on this cluster of questions rose.” It does not support “posts with a summary box get cited 40% more,” because nothing links a citation back to the specific change that earned it. Be careful of anyone selling you that number. What to do this week Pick your five best pages. For each: check the raw HTML contains the content, add a one-sentence direct answer under the main heading, add author and date, and confirm the answer-engine crawlers are permitted. That is a couple of hours and it covers the four properties above. The free Crawld scan runs 16 answer-engine checks as part of its rubric and reports which ones it could not measure. The GEO page covers what each check looks for , and if you want the underlying distinction spelled out, crawled, indexed and cited are three different states . Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## How to check and audit your backlinks URL: https://crawld.co/blog/how-to-check-backlinks/ LINK BUILDING How to check your backlinks and audit what you find Ten minutes to pull the data, longer to judge it, which sources to trust, what the authority scores mean, and why disavow is a last resort. 15 April 2026 · 10 min read Learning how to check your backlinks takes ten minutes. Working out which of them matter takes longer, and it is the part that changes what you do next. What a backlink check actually tells you A backlink is a link from another site to yours. Search engines treat some of them as votes, which is why an inbound link checker is a standard part of any SEO toolkit. What a backlink checker returns is a list of referring pages, the anchor text they used, and some scoring of how much each source is worth. That last number is the one to be careful with. No tool has Google’s index, so every authority score you see is a vendor’s estimate built from a vendor’s crawl. Two backlink checkers will hand you different totals for the same domain, and neither is wrong so much as differently incomplete. Where to check backlinks, free and paid Start with the one source that is not an estimate. Google Search Console is free, it is yours, and its Links report comes from Google’s own data. It shows top linking sites, top linked pages, and top anchor text. It is less detailed than commercial tools and more trustworthy about what Google actually knows. For a wider view, the commercial crawlers each offer a limited free tier: Ahrefs’ free backlink checker, Semrush, Moz Link Explorer, and Majestic. Expect a capped number of rows and a prompt to upgrade. For a small site that cap is often enough. Run the same domain through two of them before you conclude anything. If one reports 400 referring domains and another reports 180, the truth is that neither has crawled everything. The metrics, and what they are worth Every tool invents its own score. They correlate loosely with each other and none of them is a Google metric. Metric Tool What it estimates Domain Authority (DA) Moz Ranking strength of a whole domain Domain Rating (DR) Ahrefs Backlink profile strength of a domain Authority Score Semrush Blended authority, traffic and spam signals Trust Flow Majestic Proximity to a trusted seed set Spam Score Moz Likelihood a domain resembles penalised ones Two habits keep these useful. Compare within one tool, never across them. And treat relevance as a separate axis: a DR 30 link from a site about your exact subject often does more than a DR 70 link from a general directory. Judging quality without a score Open ten referring pages and look at them. You will learn more in fifteen minutes than from any dashboard. Ask four things of each: Is the page about anything? Real editorial content, or a list of outbound links with filler around it. Would a human land there? Pages with no internal links pointing at them exist for crawlers, not readers. Is the link in context? Inside a sentence that has a reason to reference you, rather than in a footer block. Is the anchor text natural? A run of exact-match commercial anchors is a pattern, and patterns are what get looked at. A worked example Say you run a small accounting SaaS and your checker returns three new links this month: A regional business magazine , linking from an article about invoicing rules, anchor text “cloud accounting tools”. Editorial, relevant, natural anchor. Keep, and consider whether the journalist covers your subject often enough to be worth a relationship. A general “top 200 business resources” page with 340 outbound links and no visible author, anchor text “best accounting software 2026”. Low value. Harmless, but it will never do anything for you. A gambling site in another language , linking from a comment thread, anchor text “accounting software cheap”. This is the pattern to watch: irrelevant subject, foreign language, commercial anchor, user-generated placement. Only the third is worth further thought, and even then the answer is usually to leave it alone. Which brings us to the tool everyone reaches for too quickly. Disavow is a last resort, not routine maintenance Google’s disavow tool tells the search engine to ignore specified links. It exists mainly for sites recovering from a manual action, and Google has said repeatedly that most sites never need it. The risk runs one way. Disavowing a bad link that was already being ignored gains you nothing. Disavowing a good link you misjudged removes a real signal. Given that asymmetry, the bar for using it should be high: You have a manual action in Search Console naming unnatural links, or You bought links in the past and want to distance yourself from them. An unfamiliar domain in your report is not a reason. Spam links pointing at you without your involvement are, in the ordinary case, simply discounted. Monitoring, once the audit is done An audit is a snapshot. What you want ongoing is a small, boring routine: Check Search Console monthly for sudden changes in linking domains. Watch for lost links to pages that convert, since those are worth an email asking why. Note new links to pages you did not promote, because that is where unplanned demand shows up. What this has to do with the pages you control Backlinks are one input. The other is whether the pages earning them can be crawled, read and quoted in the first place. A page with strong inbound links that ships no content in its HTML is a page that spends its authority on nothing, and answer engines that do not execute JavaScript will simply skip it. That half is fully in your control and is what the free Crawld scorecard measures: 129 checks across indexability, structure, content health and answer-engine readiness, with the coverage figure alongside the score. It does not check backlinks, and the transparency page lists exactly what it does and does not do . For the on-page side of the same problem, the SEO audit checklist walks through it . Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## How to check website traffic URL: https://crawld.co/blog/how-to-check-website-traffic/ ANALYTICS How to check website traffic, yours and anyone else Four sources for your own numbers, four metrics worth watching, and how far an estimator can be trusted for a site you do not own. 29 April 2026 · 10 min read There are two versions of how to check website traffic , and confusing them is why so many reported figures are wrong. Measuring your own site is arithmetic. Estimating someone else’s is inference. Your own site: measured, not estimated If you own the site, install analytics and read real numbers. Google Analytics 4 is free and the default. It reports sessions, users, engaged sessions, and events. Note that GA4 dropped the classic bounce rate in favour of engagement rate, so old benchmarks do not map cleanly onto it. Google Search Console answers a different question: how you performed in search. Impressions, clicks, average position and the queries behind them. Analytics tells you what visitors did; Search Console tells you how they found you. Server logs are the neglected third source. They record every request including the ones that never run JavaScript, which means they see crawler traffic your analytics misses entirely. If you want to know whether an answer-engine crawler visited, this is the only place it shows up. Privacy-first alternatives (Plausible, Fathom, Matomo) trade some depth for simpler reporting and lighter data collection. For most content sites the depth you lose is depth nobody read. The metrics worth watching Four metrics carry most of the meaning in a traffic report. Users rather than sessions when you care about reach, engaged sessions to filter out three-second bounces, pages per session to show whether internal links are working, and conversions, which is the only one the business will act on. Everything else is diagnostic. Time on page in particular is measured badly by every analytics tool, because the last page of a visit has no exit event to measure against. Someone else’s site: estimates, and how they are built You cannot measure a site you do not own. Traffic estimator tools infer it, and knowing how changes how much you trust the output. Most blend three inputs: clickstream data bought from browser extensions and ISPs, their own keyword rankings database multiplied by estimated click-through rates, and statistical modelling to fill the gaps. That method has predictable failure modes: Small sites are unreliable. Below roughly 5,000 visits a month the panel data is too thin. Non-search traffic is invisible. A site living on a newsletter or TikTok will read as near-dead. Direction beats magnitude. “Their traffic tripled since March” is usually right. “They get 47,300 visits” is a guess with a confidence interval nobody shows you. Run any competitor through two estimators and you will often see a twofold difference. Use them to compare sites within one tool, and to watch trends. Do not put the absolute number in a board deck. A worked example You are considering a guest post on an industry blog and want to know if it is worth the effort. Estimator check. Similarweb suggests roughly 60,000 monthly visits; Semrush suggests 25,000. Wide gap, so treat both as soft. Direction. Both show the trend rising over twelve months. That agreement is worth more than either figure. Source mix. The estimator attributes 70% to organic search, which means the audience arrives with intent rather than from a single viral post. Reality check. Their recent posts have comments and their newsletter has visible subscriber numbers. Signs of actual readers. Decision. Worth pitching. Not because of 60,000, which you do not believe, but because three independent signals agree the site is alive and growing. Traffic sources, and what each is telling you The split matters more than the total. Organic search is compounding and slow. Direct is people who already know you, though it is also where untagged campaigns and some app traffic land, so a sudden spike is often a tracking problem. Referral shows which relationships pay. Social is spiky by nature and rarely compounds without email capture behind it. A common misreading: direct traffic jumping 40% overnight almost never means brand awareness improved overnight. Check your campaign tagging first. What traffic will not tell you Traffic is a lagging indicator of things you did months ago, and it is silent about a growing share of how people find answers. Someone who asks an AI assistant about your subject, gets an answer citing a competitor, and never clicks anything, does not appear in your analytics as a lost visitor. They do not appear at all. That gap is why citation share is now measured separately from traffic. The GEO page explains what is tracked and how , and crawled, indexed and cited covers why being in the index is not enough . For the pages themselves, the free Crawld scan checks whether they can be crawled, read and quoted. It reports its own coverage rather than a bare score, and it does not estimate traffic for you or anyone else. Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## Why your weekend project is invisible to ChatGPT URL: https://crawld.co/blog/invisible-to-chatgpt/ GEO BASICS Why your weekend project is invisible to ChatGPT Three gaps every AI-built site ships with, the curl command that shows each one, and roughly twenty minutes of work to close them. 14 January 2026 · 6 min read A site built in a weekend usually works. It loads, it looks right, the buttons do what they should. What it often cannot do is explain itself to a machine that arrives without a browser, without patience, and without any interest in your animations. That machine is what decides whether an answer engine can cite you. Here are the three gaps that show up most often, in the order they cost you. 1. The page is empty until JavaScript runs Most AI crawlers fetch your HTML and read it. They do not wait for a framework to boot, hydrate and fill in the content. If your page ships an empty <div id="root"> and paints everything client-side, what the crawler sees is an empty div. Googlebot will render JavaScript, eventually, on a second pass it schedules when it feels like it. Most answer-engine crawlers will not. The fix: server-render or pre-render the pages that matter. In practice this means static generation for anything that is not user-specific. Check what you actually ship with curl -s https://yourdomain.com | head -100 , if your headline is not in there, neither is it in the index. 2. There is no metadata to lift An answer engine quoting you needs a title, a description, and enough structure to know which part of the page answers the question. Site builders and scaffolding tools generate a placeholder title and an empty description, and almost nobody goes back to fill them in. An empty <meta name="description"> is worse than a missing one. A missing tag lets the engine compose something from your content. An empty one is an explicit statement that there is nothing to say. The fix: a real title and description per page, headings in order, and Article or FAQPage structured data where it genuinely applies. Not on every page: schema that lies about what a page contains is its own problem. 3. There is no llms.txt llms.txt is a plain-text file at your site root that tells AI crawlers what your site is and which pages are worth reading. It is young, it is not universally honoured, and it takes about ten minutes to write. It is not a ranking trick. It is closer to a README for a machine: this is the product, these are the pages that explain it, here is what is a marketing page and here is what is documentation. The fix: write one. Keep it short and honest: a summary paragraph and a linked list of your genuinely useful pages beats an exhaustive dump of every URL. What this adds up to None of these three is hard. All three are invisible, which is why they survive launch: nothing in your browser tells you that the page a crawler receives is different from the page you see. That gap between what you see and what a crawler gets is the entire problem, and it is what a scan is for. Crawld’s free scorecard fetches your pages the way those crawlers do and reports what came back: including, explicitly, the checks it could not run. A check that could not run is never counted as a pass. Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## Link building strategies that still work URL: https://crawld.co/blog/link-building-strategies/ LINK BUILDING Link building strategies that still work, and three that do not The criteria that decide whether a tactic is worth running: a linkable asset, a pitch that cites real work, a live target. Four that pass, three to skip. 22 April 2026 · 11 min read Most link building strategies fail at the same point: the pitch. The tactic is rarely the problem, and the email almost always is. Start with something worth linking to Every strategy below assumes a page a stranger would reference without being paid. If you do not have one, outreach is just asking for a favour. Linkable assets tend to share a property: they contain something the linker cannot get elsewhere. Original survey data. A tool that does one job. A genuinely comprehensive reference on a narrow subject. A strong opinion, argued properly, by someone qualified to hold it. A rewritten summary of the top five results is not one of these. It is the most commonly produced piece of content in the industry and the least linked. Guest posting, done without embarrassment Guest posting still works when the post is good enough that the host would have commissioned it anyway. Find targets by looking at where people in your field actually publish, not by searching for “write for us”. Then pitch a specific piece, not a willingness to write. Here is a pitch that gets replies, which you can adapt directly: Subject: Pitch: what 40 accounting firms told us about invoice chasing Hi Rina, Your piece on late payment culture last month made a point I have data on. We surveyed 40 small accounting firms this quarter about how they chase overdue invoices, and the split between automated and manual follow-up was far wider than I expected. I could write that up as 1,200 words with the chart, exclusive to you. Happy to send an outline first. Marcus Feld, Ledgerpoint Three things make it work: a specific reference to their actual work, something they cannot get elsewhere, and a low-commitment next step. No paragraph about how much you admire the blog. Broken link building Find a dead page that people still link to, then offer your live equivalent. Pick a resource page or a well-linked article in your field. Run it through a link checker to find outbound 404s. Check whether you have, or could write, a genuine replacement. Email the page owner pointing out the dead link and offering yours. The honest framing is what makes this land. You are doing them a small favour first. Keep the email to three sentences and mention the broken link before you mention yourself. Its weakness is scale. Good targets are uncommon, and the tactic rewards patience over volume. Resource pages Some pages exist to list useful things. Getting onto one is a legitimate ask, provided you belong there. Search for pages collecting tools or reading in your subject, check that they are maintained (a page last updated in 2019 will not respond), and send a short note explaining which section you fit and why. Expect a low response rate and a decent conversion rate among those who reply. The skyscraper technique, with the caveat The original method: find a well-linked piece, make something better, then ask everyone linking to the original to link to yours instead. It worked well enough to be copied into the ground. Two adjustments make it viable now. Better has to mean different, not longer. Adding 2,000 words to a competitor’s structure produces a longer version of a thing that already exists. Add what is missing: original data, a working example, a section nobody covers. Ask people with a reason to switch. Someone who linked to a 2021 piece with outdated figures has a reason. Someone who linked to an evergreen essay does not. Digital PR, which is where the real links are Journalists need numbers. If you have data nobody else has, you have the ingredient national coverage is built from. You do not need much. A survey of 200 people in your industry, run properly and written up honestly, is enough for a trade publication story. The rules are simple: publish the methodology, do not overclaim what the sample supports, and give the journalist the chart already made. Three strategies to skip Buying links. It violates Google’s guidelines, the sites selling them sell to everyone, and the footprint is exactly what link spam detection is built to find. Mass directory submission. Business directories worked in 2009. Now the general ones are ignored and the niche ones are worth doing individually, which is a different activity. Comment and forum link drops. Almost universally nofollow , and a reliable way to be remembered as the person who spammed the community. Measuring whether any of it worked Count referring domains, not links. Fifty links from one site is one relationship. Track: New referring domains per month, which is the honest measure of outreach. Whether the pages earning links are the pages you want ranking. Reply rate per campaign, because a 2% reply rate is a pitch problem, not a volume problem. Where links stop being the constraint Links raise a page’s ceiling. They do not fix a page that cannot be crawled, renders empty without JavaScript, or competes with three of your own posts for the same query. Those failures cap what any amount of outreach can achieve, and they are cheaper to fix. The free Crawld scan covers that side: 129 checks across indexability, structure and content health, including whether your own pages are competing with each other. It does not build or measure links, and the transparency page is explicit about scope . If cannibalisation is the constraint, the audit page explains how it is detected . Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## llms.txt, explained for humans URL: https://crawld.co/blog/llms-txt-explained/ TECHNICAL llms.txt, explained for humans A plain-text map of your best pages for AI crawlers. What belongs in it, what does not, and the honest answer on whether it changes anything yet. 3 February 2026 · 4 min read llms.txt is a Markdown file you put at the root of your site, at /llms.txt , describing what your site is and which pages are worth reading. It is the newest member of a small family that already includes robots.txt and sitemap.xml , and it exists because those two answer the wrong questions for a language model. robots.txt says what a crawler may fetch. sitemap.xml says what exists. Neither says what any of it means , or which twelve of your four hundred URLs actually explain the product. What goes in it Keep it short. The format is deliberately plain: # Crawld > An SEO and GEO scorecard that ships verified fixes behind your approval. ## Docs - [ How we measure ]( https://crawld.co/how-we-measure ): the scoring method and why unmeasured is never counted as a pass - [ The crawler ]( https://crawld.co/bot ): what our bot fetches and how to block it ## Product - [ Pricing ]( https://crawld.co/pricing ): plans and what each includes A heading with your site name, a blockquote summarising it in one sentence, then grouped links with a short gloss on each. That is the whole specification worth caring about. What does not go in it Every URL you have. That is what a sitemap is for. This file is a curation, and its value comes entirely from what you left out. Marketing copy. A model reading this is trying to work out what you do. “Supercharge your workflow” tells it nothing. Anything untrue. It is a public file. Claims in it are claims. Does it actually work? Honestly: partially, and unevenly. Adoption is real but not universal, and no major engine has committed to honouring it the way they honour robots.txt . You should treat it as cheap insurance rather than a lever. What makes it worth the ten minutes is that the work is not wasted if adoption stalls. Writing an llms.txt forces you to answer “which pages on this site actually explain what we do?”, and most teams discover the answer is fewer pages than they thought, and that two of them are out of date. Why your site builder didn’t make one Because it is new, because it is not required, and because nothing breaks without it. The same reasons your meta description is empty. Generating one is straightforward if you already know your page inventory and which pages matter, which is the same inventory a site audit produces. That is the connection: this file is a by-product of knowing your own site structure, and most sites do not. Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## On-page SEO checklist URL: https://crawld.co/blog/on-page-seo-checklist/ ON-PAGE SEO On-page SEO: the elements worth checking on every page Six elements worth checking on every page, from titles to image attributes, plus the three answer-engine elements standard checkers still ignore. 6 May 2026 · 10 min read On-page SEO is everything you control inside a single page: the title, the headings, the copy, the links out of it. This is the list worth working through, and the order that saves the most time. Where on-page ends and technical begins The line is useful because the fixes live in different places. Technical SEO asks whether an engine can reach and render the page at all, which is usually a server or template problem. On-page asks whether the page, once retrieved, is understandable and worth ranking. That is usually a writing problem. Run the technical pass first. Optimising the heading structure of a page that returns 404 to crawlers is a wasted afternoon. The technical audit walkthrough covers that half . Titles, which do the most work per character The title element is the single highest-leverage tag on the page. Three rules cover most of it: Front-load the distinctive part. Search results truncate around 60 characters, and mobile truncates earlier. Make it unique across the site. Duplicate titles are a reliable sign of duplicate intent, which is a bigger problem than the tag. Describe the page, not the site. “Pricing” tells a searcher nothing. “Pricing: plans, limits and what the free tier covers” tells them whether to click. A quick way to find the weak ones: export your titles, sort alphabetically, and read the duplicates. They tend to arrive in clusters from one template. Meta descriptions, which do less than people think The description is not a ranking factor. It is ad copy for a result you already earned, and search engines rewrite it perhaps 60% of the time anyway. Write one regardless, because an empty one is worse than a missing one. A missing tag lets the engine compose something from the page. An empty tag is a statement that there is nothing worth saying. Aim for 140 to 160 characters, include the term someone would have searched, and make it a sentence rather than a keyword list. Headings that describe the argument One h1 per page, matching what the page is about. Then h2 for each distinct section, h3 for subdivisions, in order and without skipping levels. The reason to care is not that a crawler counts them. It is that headings are how a machine finds the part of the page that answers a question, which matters more every year as answers get lifted rather than linked. Before < h1 >Welcome</ h1 > < h3 >Our Services</ h3 > < h2 >About</ h2 > After < h1 >Bookkeeping services for UK contractors</ h1 > < h2 >What we handle each month</ h2 > < h3 >VAT returns</ h3 > < h2 >Pricing</ h2 > The second version can be skimmed by a person and parsed by a machine. The first can be done by neither. Content that answers the query it targets Length is the wrong target. Coverage is the right one. A page ranks when it answers the question a searcher had, including the follow-up questions they were about to have. Two checks that catch most problems: Does the page state its answer plainly, early? If someone has to read four paragraphs to find it, so does an answer engine, and it will usually pick a page that does not make it work. Does the page cover the obvious next question? Search the target term, read the “people also ask” box, and check whether your page addresses any of it. Internal links, which most pages under-use Internal links do two jobs: they pass authority and they tell an engine what a page is about, through the anchor text. Link with descriptive anchors. “Our pricing page” beats “click here”, and beats a bare URL. Link from high-authority pages to the ones that need help, not only the reverse. Give every page at least three inbound internal links, so nothing depends on the sitemap alone. Check for orphans, which are pages nothing links to. Images, where the cheap wins are Alt text that describes the image, written for someone who cannot see it. If it is decorative, use alt="" rather than stuffing a keyword into it. Width and height attributes , which prevent layout shift and are the most common cause of a poor Cumulative Layout Shift score. Modern formats and lazy loading below the fold. The elements most on-page checkers still miss Standard tools grade the list above and stop. Three more matter now: Content present without JavaScript. A page that renders client-side arrives at most answer-engine crawlers as an empty container. Fetch your own page with curl and read what comes back. Attribution on the page. A named author, a date, an organisation. An engine deciding between two useful pages often prefers the one it can characterise. An extractable answer. One sentence, near a matching heading, that states the thing directly. Generative engine optimization covers what that looks like . Working through it without drowning Do not fix page by page. Group by cause. Two hundred pages missing a description is one template change, not two hundred tasks. Order by severity multiplied by pages affected, and the cheap site-wide wins come first. The free Crawld scan grades these elements as part of a 129-check rubric and reports what it could not measure rather than counting it as a pass. The full rubric is published by category , and the audit page explains what a finding contains . Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## The SEO audit checklist URL: https://crawld.co/blog/seo-audit-checklist/ TECHNICAL The SEO audit checklist that states its denominators Seven passes over indexability, rendering, structure, metadata, content health, answer-engine readiness and page experience, ordered by what breaks first. 4 March 2026 · 9 min read Most audits hand you four hundred rows and a score. This SEO audit checklist works the other way round: it tells you what was checked, what could not be, and what to fix first. What an SEO audit is actually for An audit is a measurement, and a measurement without a denominator is a rumour. Before you look at any score, ask what proportion of the site it describes. A 94 from a crawler that reached forty pages of a four-hundred-page site is not a better result than a 71 from a crawl that reached everything. It is a smaller claim wearing a bigger number. That is the single habit worth taking from this checklist. Every section below ends with a note on what makes the check impossible to run, because a check that could not run is not a pass. Indexability comes first Nothing else on the list matters if an engine cannot reach and keep the page. Work through these before touching content: The page returns a 200, not a soft 404 or a redirect chain ending nowhere. robots.txt permits the crawlers you want, including the answer-engine ones. No conflicting noindex between the meta robots tag and the X-Robots-Tag header. The canonical tag points at a real, indexable URL rather than a 404 or a redirect. HTTPS with a valid certificate and no mixed content. Cannot be checked when: the crawler is blocked outright, or the page sits behind authentication. Record those pages as unmeasured and count them in the denominator. Rendering: what actually arrives A page can be perfectly indexable and still arrive empty. Fetch your own page the way a crawler does and read what comes back: curl -s https://yourdomain.com/pricing | grep -c "<h1" If that returns zero, your headline is being painted by JavaScript after the fact. Googlebot will render it eventually, on a second pass it schedules at its convenience. Most answer-engine crawlers will not. Structure and internal links Site structure exists only across a full crawl, which is why a page-by-page audit cannot see it. Check that every page you care about is reachable by an internal link rather than through the sitemap alone, that click depth from the homepage stays shallow, and that filtered URLs do not multiply without limit. Every page you care about is reachable by at least one internal link, not only through the sitemap. Click depth from the homepage stays shallow. Pages six clicks deep get crawled rarely. The sitemap lists pages that exist and return 200, and nothing else. Faceted or filtered URLs do not generate near-infinite combinations. Metadata that survives contact with a template Metadata fails in templates rather than on individual pages. A title describing the site instead of the page, or an empty description left behind by scaffolding, repeats across every URL that template generates. Fixing the template fixes hundreds of pages at once, which is why this is the cheapest category to recover. Before < title >Pricing</ title > < meta name = "description" content = "" > After < title >Pricing: plans, limits and what the free tier covers</ title > < meta name = "description" content = "Compare plans, see exactly what the free scan includes, and what each paid tier adds." > An empty description is worse than a missing one. A missing tag lets the engine compose something from the page. An empty one is a statement that there is nothing worth saying. Content health, measured across the corpus Individual pages look fine on their own. The damage happens between them: Two or more pages targeting a single search intent, splitting the signal so none of them ranks. Near-duplicate titles, which are usually a symptom of near-duplicate content. Thin pages that dilute the site rather than extending it. Stale pages whose performance has been declining for months. A quick way to find the first one: export your top queries, group by landing page, and look for a query where two of your own URLs both appear with weak positions. That pair is competing with itself. Answer-engine readiness Answer-engine readiness asks whether an AI answer can lift your content and credit you for it. Being indexed makes a page eligible; being cited depends on whether a passage states the answer directly, whether there is an author and a date to attribute it to, and whether those crawlers are permitted at all. It is the newest section of any SEO audit checklist and the one most sites skip. The answer is extractable: stated plainly under a heading that matches the question, not buried mid-essay. There is something to attribute, such as a named author, a date, and an organisation. An llms.txt exists and curates what is worth reading. The format is short and worth writing . Answer-engine crawlers are permitted, which a robots.txt written years ago often blocks by accident. Page experience, weighted honestly Speed matters and it matters less than the industry implies. A fast page with nothing to say still has nothing to say. Check Largest Contentful Paint, Interaction to Next Paint and Cumulative Layout Shift, then stop. Most layout shift on content sites comes from images without width and height attributes, which is a ten-minute fix. Cannot be checked when: field data is unavailable for a low-traffic page. Lab data is a substitute, not a replacement, and the report should say which one it used. How to order the work Severity alone is the wrong sort order, and it is the one almost everything uses. Rank by severity multiplied by the number of pages affected instead. A missing meta description is low severity. A missing meta description on two hundred pages is the most valuable afternoon available to you, and a severity-sorted list will bury it under a single high-severity issue on a page nobody visits. Turning the checklist into a habit Run the full list quarterly, and the indexability and rendering sections after any template change or migration. Record the coverage figure every time, so you can tell a real improvement from a crawl that simply reached more pages. If you would rather have it run for you, the free Crawld scorecard applies this checklist as 129 checks across eight weighted categories, and reports every check it could not run instead of rounding it up to a pass. The full rubric is published here , and the scoring method explains how coverage is calculated . Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## What an SEO audit costs, and how often URL: https://crawld.co/blog/seo-audit-cost-and-frequency/ STRATEGY What an SEO audit costs and how often to run one Free to five figures for the same-sounding deliverable. Five variables explain the spread, plus how often to re-run one and when to do it yourself. 20 May 2026 · 9 min read SEO audit pricing ranges from nothing to five figures for what sounds like the same deliverable. The spread is real, and it comes down to two questions: how much was checked, and who fixes it. What you are actually paying for An audit has three separable parts, and quotes rarely say which they include. The crawl is cheap. Software walks the site and records what it finds. This is the part free tools do. The interpretation is where cost starts. Someone decides which of 400 findings matter for your business, in what order, and which are false alarms. Software cannot do this well, because it does not know that the 200 pages failing a check are a deprecated section you are about to delete. The remediation plan is the expensive part. Not a list of problems, but specific instructions someone can act on, scoped to your platform and your team. A £500 audit that produces a crawl export is not overpriced. It is a different product from a £5,000 audit that produces a prioritised plan, and comparing them on price alone is comparing the wrong attribute. Rough bands, with the caveats Type Typical range What you get Free automated £0 A crawl, a score, generic recommendations Paid tool subscription £80 to £400 per month Ongoing crawls, dashboards, no interpretation Freelance audit £500 to £2,500 A crawl plus a human reading it, one pass Agency audit £2,500 to £10,000 Full technical and content review, prioritised plan Enterprise audit £10,000 upward Multi-site, multi-locale, stakeholder workshops Treat these as orientation, not quotes. Rates vary by market, and a specialist charging at the top of a band is often cheaper per useful finding than a generalist at the bottom. What moves the number Five variables account for most of the spread in audit pricing. Site size, how unusual the platform is, whether scope covers content and competitive work or technical only, what the deliverable actually is, and whether remediation is included. That last one is the difference between a report and a result. The question to ask any quote “What happens to the findings?” A report that ends at a list transfers the work to you. If you have an engineer who will action it, that may be exactly what you want and you should not pay for more. If you do not, you are buying a document that will sit in a shared drive, and the audit was never the constraint. Ask a second one too: “What did you check, and what could you not?” An auditor who cannot tell you the second half either checked everything, which is unlikely, or is not tracking it, which tells you how much the score is worth. How often to run one Frequency depends on how fast the site changes, not on the calendar. Full audit: annually for a stable site, twice a year if you publish weekly or run a catalogue. Technical spot-check: after every migration, redesign or template change. These are where sudden damage comes from, and the damage is usually invisible for weeks. Indexability check: monthly. Ten minutes in Search Console. Coverage report, sudden drops, new errors. Content health: quarterly. Overlap accumulates as you publish, so this is a function of output rather than time. A useful trigger to add: audit whenever organic traffic moves more than 20% in a month without an obvious cause. That is a symptom, and symptoms are cheaper to investigate early. When to do it yourself Doing it yourself makes sense when three things are true: the site is small enough to hold in your head, someone on the team can implement fixes, and you mainly need to know what is wrong rather than be persuaded. A capable person with a free crawler and a checklist will find most of what a mid-range paid audit finds. What they will miss is the prioritisation, which is the part experience actually buys. The honest split: pay for an audit when you need judgement or you need to convince someone. Do it yourself when you need a list and you were going to do the work anyway. The SEO audit checklist is the list . What not to buy An audit with no coverage figure. If the report gives a score without saying what proportion of the site it examined, the number is unbounded. A 94 from a crawl that reached a third of your pages is not better than a 71 from a complete one. An audit that counts what it could not check as a pass. This is the category default and it is how tools report flattering numbers for sites with real problems. A retainer that repeats the same audit quarterly without re-measuring whether the last round of fixes moved anything. The free Crawld scorecard runs the full 129-check rubric and reports coverage alongside the score, so you can see the denominator before you decide whether to pay anyone. The scoring method is published in full , and pricing for the paid tiers is here . Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## How to run a technical SEO audit URL: https://crawld.co/blog/technical-seo-audit/ TECHNICAL How to run a technical SEO audit that finds real problems Six passes over crawl access, rendering, architecture, structured data and Core Web Vitals, plus what to do about the checks a crawler cannot make. 18 March 2026 · 8 min read A technical SEO audit looks for the failures a content review cannot see: pages an engine cannot reach, cannot render, or cannot make sense of. Here is how to run one without drowning in findings. Where a technical audit differs from a content audit A content audit asks whether a page deserves to rank. A technical audit asks whether it is capable of ranking at all. The two fail differently. Thin content underperforms quietly over months. A misconfigured X-Robots-Tag header removes a section of your site from the index in a week, and nothing in your analytics explains why. Run the technical pass first. Fixing prose on a page that no crawler can reach is effort spent on a page nobody will see. Start with what the crawler receives Not what your browser receives. The gap between the two is where most of the surprises live. Pick five representative URLs (a homepage, a category, a product or article, a paginated page, a filtered page) and fetch each without JavaScript. Compare the word count to what you see in a browser. If the raw HTML is dramatically thinner, everything downstream in this audit is measuring the wrong document. Crawl access, in the order it fails robots.txt allows the agents you want. Check the answer-engine crawlers specifically, since older files predate them. Meta robots and the header agree. A page with index in the tag and noindex in the X-Robots-Tag is removed, and the tag is the one people look at. Canonicals resolve. A canonical pointing at a redirect, a 404, or a different domain tells the engine to index something other than the page you are auditing. Status codes are honest. A soft 404 returning 200 keeps a dead page in the index and spends crawl budget on it every visit. Architecture, which only exists site-wide Site architecture can only be assessed across a full crawl, because it describes relationships between pages rather than properties of one. Three findings matter most: orphan pages that nothing links to, click depth pushing commercial pages out of frequent crawling, and crawl traps generating URLs without limit. Orphans. Pages with no internal links pointing at them. These depend entirely on your sitemap being read and trusted, which is a thinner thread than most people assume. Depth. Count clicks from the homepage. Crawl frequency drops sharply with depth, so anything commercially important sitting five or six clicks deep is being visited rarely. Crawl traps. Faceted navigation that generates combinations without limit. A store with five filters and a sort parameter can produce more URLs than it has products, and every one of them consumes budget that indexable pages needed. Structured data that matches reality Validate it, then read it. Valid schema describing the wrong thing is worse than no schema, because it is a confident false statement. The common failure is a template applied too broadly. A Product type on a category listing, or Article on a tag archive, will validate cleanly and misdescribe the page. Spot-check one page per template rather than trusting a site-wide pass rate. Core Web Vitals, without the theatre Core Web Vitals reduce to three metrics that are genuine ranking inputs: Largest Contentful Paint, Interaction to Next Paint and Cumulative Layout Shift. Anything beyond those is diagnostic, useful to an engineer, and not something to report to a stakeholder as though it affected rankings. Layout shift is usually the cheapest to fix. Images without width and height attributes are the most common cause, and adding them is mechanical: <!-- shifts --> < img src = "/hero.webp" alt = "Scorecard showing category scores" > <!-- reserved --> < img src = "/hero.webp" alt = "Scorecard showing category scores" width = "1200" height = "630" > The checks people skip Some technical findings turn up constantly and rarely appear on a standard checklist. Pagination that loses its trail, a staging subdomain left indexable, redirect chains surviving a migration, and duplicate content on parameterised URLs are all common, all quiet, and all cheap to fix once someone looks. Pagination that loses its trail. Page two onward with no self-referencing canonical and no link back to page one. Staging left indexable. A staging. or dev. subdomain with no noindex , competing with production for its own content. Redirect chains after a migration. Three hops still resolve, and they leak a little at each one. Duplicate content on parameterised URLs. Session identifiers and tracking parameters generating distinct URLs for one page. What to do about what you cannot check Some checks will not run. A third-party API is down, a page times out, an area sits behind a login. Record those as unmeasured and keep them in the denominator. This matters more than it sounds. A technical SEO audit that quietly drops what it could not reach produces a flattering score and hides the section of your site most likely to be broken, since unreachable pages and broken pages correlate heavily. Turning findings into a plan Group by cause, not by page. Two hundred pages missing a canonical is one template fix, not two hundred tasks. Then order by severity multiplied by pages affected, so the cheap site-wide wins come before the expensive one-off ones. The free Crawld scan runs this pass as part of a 129-check rubric and reports its coverage alongside the score. The technical foundation and crawling categories are documented in full , and if you want to know what the crawler does to your site while it works, the crawler page explains it . Scan your site free All posts Keep reading The 16 answer-engine checks All 129 checks, by category More writing on GEO and SEO --- ## Crawld vs IndexRusher URL: https://crawld.co/compare/indexrusher/ COMPARE Crawld vs IndexRusher IndexRusher gets you indexed. Crawld measures whether being indexed turned into being cited. What IndexRusher is built for Getting pages indexed fast across Google, Bing and LLM crawlers. The difference Indexing latency is a real problem with a real solution, and tools in this category solve it. For a new site, or one publishing faster than a crawler naturally revisits, getting into the index quickly is a genuine bottleneck. The thing to understand is that indexing and citation are different states. Being in the index makes you eligible to be cited. Whether you actually are depends on whether the answer is extractable from your page, whether there is enough attribution for an engine to credit you, and whether a competitor is simply easier to quote. Crawld is built around that third step. It tracks citation share per engine and per prompt for the queries you compete on, and the 16 GEO checks measure what is blocking a citation rather than what is blocking an index entry. These are complementary rather than opposed. Fast indexing plus nothing worth citing gets you into an index nobody quotes from. Side by side Feature Crawld IndexRusher Problem solved Whether engines cite you Whether engines have stored you Measurement Citation share per engine, per prompt Index status Remediation Ranked fixes delivered as diffs Not the focus Full-site audit 129 checks Not the focus When IndexRusher is the better choice If your pages are not getting indexed at all, fix that first, and a dedicated indexing tool will do it faster than a rubric will. Citation share is a question you can only ask once you are in the index. How to check for yourself Run the free scorecard against your own site and compare it to whatever you are using now. The number worth comparing is not the score. It is the coverage next to it. A 94 at 40% coverage and a 71 at 95% coverage are not measuring the same thing, and only one of them tells you what it did not look at. Run the free scan All comparisons Keep reading How we score, and why it reads lower All comparisons Answer-engine measurement --- ## Crawld vs Semrush URL: https://crawld.co/compare/semrush/ COMPARE Crawld vs Semrush Semrush is a much broader suite. Crawld is narrower, tracks AI-citation share, and ships the fix rather than reporting it. What Semrush is built for A broad marketing suite: rank tracking, keyword research, backlinks, competitive intelligence, site audit. The difference Semrush is an order of magnitude broader than Crawld and it would be silly to pretend otherwise. Keyword research, backlink analysis, paid intelligence, rank tracking across locales | Crawld does none of it and is not trying to. Two differences are substantive rather than a matter of scope. The first is what happens after a finding. A site audit that produces a prioritised issue list ends there; the fixing is your weekend. Crawld generates the change, verifies it where it can, and puts it in front of you as a diff behind two approval gates. The second is measurement of AI answers. Rank tracking tells you your position on a results page that a growing share of your audience never sees. Crawld tracks whether ChatGPT, Claude, Perplexity, Gemini and AI Overviews named you as a source for the prompts you compete on, which is a separate instrument. There is also a scoring philosophy difference worth naming: a check that could not run is never counted as a pass here, and the coverage figure is shown next to every score. That makes Crawld scores look lower and mean more. Side by side Feature Crawld Semrush Scope Narrow: audit, remediation, GEO Broad marketing suite After a finding A verified diff you approve A prioritised report AI citation share Tracked per engine and prompt Not the focus Unmeasured handling Its own state, with coverage shown Differs by report Keyword research, backlinks, paid Not offered Extensive When Semrush is the better choice If you need one tool covering the whole discipline: keyword research through backlinks through rank tracking | Crawld does not replace it and is not intended to. It replaces the part of your week that goes on turning an audit into merged changes. How to check for yourself Run the free scorecard against your own site and compare it to whatever you are using now. The number worth comparing is not the score. It is the coverage next to it. A 94 at 40% coverage and a 71 at 95% coverage are not measuring the same thing, and only one of them tells you what it did not look at. Run the free scan All comparisons Keep reading How we score, and why it reads lower All comparisons Answer-engine measurement --- ## Crawld vs SEObot URL: https://crawld.co/compare/seobot/ COMPARE Crawld vs SEObot SEObot produces content at volume. Crawld checks that what you publish does not compete with itself, and verifies changes before they ship. What SEObot is built for High-volume automated content generation and publishing. The difference SEObot is built around throughput: generate articles, publish them, repeat, with as little human involvement as the buyer wants. If the constraint you are working against is that nobody has time to write, it addresses that constraint directly. Crawld is built around the opposite constraint. Its two human gates are deliberate, and the engine will not generate a change for a check classified as Human-only. It stops and asks, twice, by design. The substantive disagreement is about what happens after volume. Publishing at rate without a corpus-level view produces cannibalization: several posts targeting one intent, splitting the signal so none of them ranks. That failure is invisible per-article and only exists at the level of the whole site, which is where Crawld measures. There is also a verification difference. Crawld ships changes as build-verified artifacts where it holds a repository, and labels them preview-verified or advisory where it does not. Content generated and published without a verification step is a different risk profile, and worth choosing knowingly. Side by side Feature Crawld SEObot Primary output Verified fixes to an existing site New articles at volume Human review Two mandatory gates Optional by design Overlap with your own content Flagged before publish Not the focus Verification Build-verified where access allows, labelled honestly where not No build step Technical audit 129 checks across 8 categories Not the focus When SEObot is the better choice If you genuinely need a lot of pages quickly and you have a plan for reviewing them, volume tooling does something Crawld does not do at all. Crawld does not write your content calendar for you. How to check for yourself Run the free scorecard against your own site and compare it to whatever you are using now. The number worth comparing is not the score. It is the coverage next to it. A 94 at 40% coverage and a 71 at 95% coverage are not measuring the same thing, and only one of them tells you what it did not look at. Run the free scan All comparisons Keep reading How we score, and why it reads lower All comparisons Answer-engine measurement --- ## Crawld vs Surfer URL: https://crawld.co/compare/surfer/ COMPARE Crawld vs Surfer Surfer optimises a page against the current SERP. Crawld audits the whole site and ships the fix. What Surfer is built for On-page content optimisation: scoring a draft against what currently ranks for a keyword. The difference Surfer is a content editor. You give it a target keyword, it analyses what is ranking, and it scores your draft against those pages as you write. That is a genuinely useful loop and it is tight: the feedback arrives while you are still writing. Crawld starts from the other end. It crawls the site you already have, scores it against a fixed 129-check rubric rather than against whoever currently ranks, and produces a ranked list of what to change. The output is a diff you approve, not a score to write toward. The deeper difference is what each treats as the unit of work. For Surfer it is a page. For Crawld it is a site, which is why cannibalization between your own posts is something one of them can see and the other structurally cannot. They also answer to different search surfaces. Surfer optimises against the ranked results page. Crawld weights 16 of its 129 checks toward whether an answer engine can lift and attribute your content, which is a different question with a different answer. Side by side Feature Crawld Surfer Unit of analysis The whole crawled site A page against its target keyword Scoring basis A fixed 129-check rubric with published weights What currently ranks for the query Unmeasured handling Reported as its own state, never a pass Not applicable: different model Cannibalization Detected across your corpus Out of scope Answer-engine readiness 16 checks, citation share per engine Not the focus Delivery A reviewable diff behind two human gates Guidance in the editor When Surfer is the better choice If your problem is that you are writing a specific article and want to know whether it covers what the ranking pages cover, Surfer is a better tool for that job and Crawld is not competing for it. Plenty of teams should run both. How to check for yourself Run the free scorecard against your own site and compare it to whatever you are using now. The number worth comparing is not the score. It is the coverage next to it. A 94 at 40% coverage and a 71 at 95% coverage are not measuring the same thing, and only one of them tells you what it did not look at. Run the free scan All comparisons Keep reading How we score, and why it reads lower All comparisons Answer-engine measurement --- ## Docs URL: https://crawld.co/docs/ DOCS The public API. What is documented here is the unauthenticated scan surface: the one behind the free scorecard. It is the part that is stable enough to build against. Response envelope Every response uses the same wrapper, success or failure. { "success": true, "message": "Scan", "data": { … }, "error": null, "timestamp": "2026-08-12T06:00:00Z", "requestId": "req_…" } On failure, success is false , data is null , and error carries { code, details } . The one exception is rate limiting, which is returned by the limiter ahead of the application and does not use the envelope: worth handling explicitly. Endpoints POST /api/v1/public/scans Queue a scan. Returns 202 with a scan id: the result is not ready yet. Request { "url": "https://example.com", "captchaToken": "…" } Response { "scanId": "scan_…", "status": "queued", "reused": false } Worth knowing Rate limited to 3 submissions an hour per caller. One real crawl per domain per day; inside that window reused comes back true and you get that day’s scan. CAPTCHA is verified before anything else, before URL validation, before any DNS lookup, and fails closed in production. The URL is checked against an SSRF guard. Addresses that resolve into private space are rejected. GET /api/v1/public/scans/{scanId} Read a scan. Poll until status leaves queued or running. Response { "status": "done", "score": 71, "grade": "C+", "categories": [{ "name": "…", "score": 62 }], "findingsCount": 12, "pagesScanned": 10, "pagesTotal": 47, "remainingPages": 37, "unlocked": false, "advisory": "Advisory scan of public HTML only." } Worth knowing findings is omitted entirely until an email is verified: omitted, not blanked, because a locked field still present in the payload is one devtools tab away from unlocked. A scan that never ran has no grade at all rather than a default one. No traffic metrics are returned. There are no clicks, impressions, positions or citation counts in this payload. GET /api/v1/public/scorecards/{slug} Read a shared scorecard by its share slug. Response Same shape as a scan, always with `unlocked: false`. Worth knowing Always locked, even if the owner has verified their email: otherwise one verified address would publish the findings to anyone holding the link. Expires 30 days after the scan. POST /api/v1/public/scans/{scanId}/claim Trade an email address for the findings. Sends a verification link. Request { "email": "you@example.com" } Response 202 accepted Worth knowing Rate limited separately from scan submission, because it sends mail rather than starting a crawl. Error codes Code Meaning CAPTCHA_REQUIRED The verification challenge was not solved. Fails closed in production. VALIDATION_FAILED The URL was rejected: malformed, or resolving somewhere it should not. SERVICE_UNAVAILABLE The scanning engine is not available. No grade is invented in its absence. RESOURCE_NOT_FOUND Unknown scan id, or a scorecard slug that has expired. 429 Rate limited. Handled outside the envelope, so parse defensively. A worked example # queue curl -sX POST https://api.crawld.co/api/v1/public/scans \ -H 'Content-Type: application/json' \ -d '{"url":"https://example.com"}' # poll until status is done or failed curl -s https://api.crawld.co/api/v1/public/scans/scan_abc123 The authenticated API Not documented yet The brand-scoped surface: fix queue, audits, watchdog, citations: exists but is not publicly documented, which means it is not stable enough to build against and should not be treated as though it were. When it is documented it will be here. Webhooks and the platform integrations are described on the integrations pages , including which of them are built. Try it on your own site Crawler behaviour Keep reading Queue and scan limits in detail What the platform actually runs How the crawler identifies itself --- ## Queue & scan limits URL: https://crawld.co/limits/ PLATFORM Every limit, and why it exists. Limits published rather than discovered. Each one has a reason, and the reasons are more useful than the numbers. Limit Value Why Pages per free scan 10 A crawl costs real money and the free tier is a diagnostic, not a product. The result always reports how many pages exist beyond the cap. Scans per domain 1 per 24 hours A repeat request returns that day’s scan. This is the control that survives IP rotation, because a free scan attaches to a domain rather than a visitor. Submissions per caller 3 per hour Deliberately not the main defence: VPNs defeat it and office NAT makes it punish real users. A cheap first hurdle. Shared scorecard lifetime 30 days A share link is a snapshot, not a permanent free product. It never exposes findings, whoever opens it. Email verification token 24 hours A verification link is a password-equivalent sitting in an inbox. Competitor monitoring 3 sites, 12 pages each Per brand. It does not walk a competitor’s whole site. Monitoring frequency Every 168 hours Once a week, one request at a time, robots.txt obeyed. What happens when you hit one Rate limited : a 429, returned by the limiter ahead of the application, so it does not use the standard response envelope. Parse defensively. Domain cooldown : not an error. You get the existing scan with a flag saying it was reused. Expired share link : treated as not found. Shared scorecards are not resurrected. Paid plans Full-site crawling and unlimited scans per domain are what the paid tiers buy. The per-caller rate limit and the politeness constraints on competitor monitoring are not negotiable at any tier: they exist to protect the sites being crawled, not to meter you. See pricing API documentation Keep reading Why the free scan is capped The API these limits apply to Rules on what you may scan --- ## Platform overview URL: https://crawld.co/platform/ PLATFORM What is actually running. A map of the surfaces, with the ones you can build against today separated from the ones that exist but are not stable enough to publish. Surfaces Surface Status What it is Public scan API Available now Unauthenticated. Queue a scan, poll it, read the scorecard. Documented in full. Scanning engine Available now Fetches public HTML the way search and answer-engine crawlers do, and runs the rubric against what came back. Rubric catalogue Available now 129 checks across 8 weighted categories, each declaring its method and fix capability. Brand-scoped API Not built yet Fix queue, audits, watchdog and citations exist behind auth but are not publicly documented, which means they are not stable enough to build against. Webhooks Not built yet Outbound notification on scan completion and queue events. Delivery connectors Not built yet The access ladder: pull requests, CMS drafts. Designed, not built. How a scan actually runs 01 Submit A URL arrives. Verification runs before anything else (before URL validation, before any DNS lookup) so the endpoint cannot be used as an unauthenticated probe. The host is then checked against an SSRF guard, because this is a surface where a stranger causes us to fetch something. 02 Queue Accepted with a scan id. A repeat request for the same domain inside 24 hours returns that day's scan rather than crawling again, which also stops the endpoint being pointed repeatedly at a third party. 03 Crawl and score A capped crawl of public HTML, then the rubric. Every check resolves to pass, fail or unmeasured, and coverage is recorded alongside the score. 04 Report Score, grade, per-category scores and finding count are returned. The findings themselves are withheld until an email is verified: omitted from the payload entirely rather than blanked, because a locked field that is still present is one devtools tab away from unlocked. If the engine is unavailable The scan fails and says so. There is no path that returns a grade without a crawl behind it. This used to be untrue and was fixed, and it is the single most important property of the platform to understand. API documentation Queue and scan limits Keep reading API endpoints and the response envelope Every cap that applies What a scan needs from your site --- ## System requirements URL: https://crawld.co/system-requirements/ PLATFORM What a scan needs from your site. There is nothing to install. The requirements are about what your site exposes and what your infrastructure permits. For the site being scanned A publicly reachable URL over HTTPS. Pages behind authentication are never fetched. A resolvable public hostname. Addresses in private space are rejected by the SSRF guard. robots.txt that permits CrawldBot , if you want it crawled. A disallow is obeyed and the scan will report reduced coverage. Server-rendered content, ideally. Not a requirement, but a page that paints entirely client-side is what most answer-engine crawlers see as empty, which is itself one of the most common findings. Allowlisting the crawler If a WAF or bot filter sits in front of your site, identify us by user agent: CrawldBot/1.0 (+https://crawld.co/bot) We do not crawl under a browser user agent, and we do not rotate identities to get around filtering. If we are blocked, the scan reports lower coverage rather than working around it. For using the dashboard A current version of any major browser. Nothing exotic is required. JavaScript enabled for the application. This marketing site works without it. Cookies enabled for authenticated sessions. For the API Anything that can make an HTTPS request and parse JSON. The ability to poll: scan submission returns immediately with an id, not a result. In production, a CAPTCHA token on scan submission. The check fails closed, so a scan without one is rejected. Run a scan Crawler behaviour Keep reading What CrawldBot fetches The platform surfaces available How to tell an outage from a limit --- ## Analytics | Crawld blog URL: https://crawld.co/blog/category/analytics/ ANALYTICS Analytics 1 article on analytics. All posts Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) Analytics · 10 min read How to check website traffic, yours and anyone else Four sources for your own numbers, four metrics worth watching, and how far an estimator can be trusted for a site you do not own. Read → --- ## Content strategy | Crawld blog URL: https://crawld.co/blog/category/content-strategy/ CONTENT STRATEGY Content strategy 1 article on content strategy. All posts Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) Content strategy · 11 min read Content cannibalization: how to find it and what to do Four symptoms, one Search Console report that confirms it, and the choice between merging, differentiating and deleting the pages that compete. Read → --- ## GEO basics | Crawld blog URL: https://crawld.co/blog/category/geo-basics/ GEO BASICS GEO basics 2 articles on geo basics. All posts Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) GEO basics · 8 min read Generative engine optimization, explained without the hype Ranking well does not make a page citable. What an answer engine needs instead, and the page changes that make a claim liftable. Read → GEO basics · 6 min read Why your weekend project is invisible to ChatGPT Three gaps every AI-built site ships with, the curl command that shows each one, and roughly twenty minutes of work to close them. Read → --- ## Link building | Crawld blog URL: https://crawld.co/blog/category/link-building/ LINK BUILDING Link building 2 articles on link building. All posts Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) Link building · 11 min read Link building strategies that still work, and three that do not The criteria that decide whether a tactic is worth running: a linkable asset, a pitch that cites real work, a live target. Four that pass, three to skip. Read → Link building · 10 min read How to check your backlinks and audit what you find Ten minutes to pull the data, longer to judge it, which sources to trust, what the authority scores mean, and why disavow is a last resort. Read → --- ## On-page SEO | Crawld blog URL: https://crawld.co/blog/category/on-page-seo/ ON-PAGE SEO On-page SEO 1 article on on-page seo. All posts Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) On-page SEO · 10 min read On-page SEO: the elements worth checking on every page Six elements worth checking on every page, from titles to image attributes, plus the three answer-engine elements standard checkers still ignore. Read → --- ## Strategy | Crawld blog URL: https://crawld.co/blog/category/strategy/ STRATEGY Strategy 3 articles on strategy. All posts Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) Strategy · 10 min read Competitor SEO analysis without guessing at their data Three steps to build the rival list from search results, four things you can read in their HTML, and the estimates to treat as directional only. Read → Strategy · 9 min read What an SEO audit costs and how often to run one Free to five figures for the same-sounding deliverable. Five variables explain the spread, plus how often to re-run one and when to do it yourself. Read → Strategy · 7 min read Crawled, indexed, cited, where AI traffic begins Three gates, and passing the first two says nothing about the third. What decides whether an answer engine names you as the source. Read → --- ## Technical | Crawld blog URL: https://crawld.co/blog/category/technical/ TECHNICAL Technical 3 articles on technical. All posts Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) Technical · 8 min read How to run a technical SEO audit that finds real problems Six passes over crawl access, rendering, architecture, structured data and Core Web Vitals, plus what to do about the checks a crawler cannot make. Read → Technical · 9 min read The SEO audit checklist that states its denominators Seven passes over indexability, rendering, structure, metadata, content health, answer-engine readiness and page experience, ordered by what breaks first. Read → Technical · 4 min read llms.txt, explained for humans A plain-text map of your best pages for AI crawlers. What belongs in it, what does not, and the honest answer on whether it changes anything yet. Read → --- ## CrawldBot: what our crawler does URL: https://crawld.co/bot/ THE CRAWLER CrawldBot, and how to stop it. If you found this page from your access logs, this is everything about what our crawler does and how to turn it off. How to identify it Every request we make carries this user agent: CrawldBot/1.0 (+https://crawld.co/bot) The URL in the string points here, so the bot always identifies itself and always links to instructions for blocking it. We do not crawl under a browser user agent. What it fetches Public HTML, and the resources a page references that are needed to judge it: stylesheets, the robots.txt , the sitemap, structured data. It does not submit forms, does not follow links behind authentication, and does not attempt to access anything a signed-out visitor could not. Why it might be visiting you There are two reasons, and they behave differently: Someone scanned their own site. A customer, or a visitor running the free scorecard on a domain they control. This is a one-off crawl of a capped number of pages. Someone is monitoring you as a competitor. Crawld's SEO Watchdog lets a customer track up to three competitor sites for changes to titles, meta descriptions, headings, body copy, pricing and links. If that is why we are here, the limits below are what constrain it. The limits it works within 12 pages per site on a competitor crawl. It does not walk your whole site. One request at a time. No parallel fetching, ever. It is not capable of loading your server. Once a week. The scheduler runs on a 168-hour interval. robots.txt is obeyed , including crawl-delay. How to block it Add this to your robots.txt and we will stop on the next scheduled run: User-agent: CrawldBot Disallow: / To allow most of the site but keep it out of one area: User-agent: CrawldBot Disallow: /internal/ Disallow: /staging/ If you would rather not wait for the next run, or you believe the crawler is misbehaving, contact us and we will add the domain to a permanent exclusion list. What we do with what we collect Crawl results are stored against the account that requested them and are used to produce that account's scorecard or change report. We do not resell crawl data and we do not train models on the content we fetch. Keep reading The monitoring that uses this crawler What we do with crawl data What you may point the scanner at --- ## Content health checks URL: https://crawld.co/checks/content-health/ THE RUBRIC Site content health & cannibalization Whether your own pages are competing with each other. 14 checks · 10% of the score Checks 14 11% of 129 Weight 10% of the overall score Share of score Why it is weighted this way The category that is invisible from inside a CMS, because every post looks fine on its own. It only exists at the level of the corpus. What this category covers Intent overlap between pages Near-duplicate titles across the site Coverage gaps against your own keyword set Stale pages with decaying performance Internal link distribution across competing pages Orphaned commercial pages Representative checks A sample of what sits in this category, with how each one is decided and who is allowed to fix it. The full list of 129 lives in the product. Check Method Fix Intent overlap between pages Four posts targeting one query split the signal and none of them ranks. Graph Assist Near-duplicate titles Convergent titles are usually a symptom of convergent content. Graph Assist Coverage gaps against your own keyword set What you claim to be about, versus what you have actually written. Graph Human Stale pages with decaying performance Depends on a third-party data source; unavailable upstream means unmeasured, not pass. API Human What usually goes wrong Four posts targeting one intent, splitting the signal so none of them ranks A tangential blog post outranking the commercial page it was meant to support Titles converging on the same shape without anyone deciding they should Internal links scattered across near-duplicates instead of pointing at one winner When these checks return unmeasured A check in this category that cannot run is reported as unmeasured and counted in the denominator. It is never rounded up to a pass, which is the difference between a score that bounds its own claim and one that implies it examined everything. The crawl covered too small a slice of the site for corpus-level comparison to mean anything A third-party performance source needed for decay detection was unavailable Where to start Check your top twenty queries in Search Console for more than one URL earning impressions. Pick the page that should own each intent, then merge or re-angle the rest. Repoint the internal links at the winner. Previous category E-E-A-T & trust Next category Answer-engine readiness Scan your site free All eight categories Keep reading How coverage and scoring work All eight categories How a finding becomes a reviewable diff --- ## Content quality checks URL: https://crawld.co/checks/content-quality/ THE RUBRIC Content quality & helpfulness Whether the page answers the thing it appears to be about. 18 checks · 20% of the score Checks 18 14% of 129 Weight 20% of the overall score Share of score Why it is weighted this way The heaviest category at 20%, and the one with the most model-judged checks. Weighted highest because it is what everything else exists to deliver. What this category covers Substance against the query the page targets Heading structure, and whether it matches the argument Whether the answer is stated plainly and early Duplicate body content across URLs Readability and formatting for scanning Freshness where the subject demands it Representative checks A sample of what sits in this category, with how each one is decided and who is allowed to fix it. The full list of 129 lives in the product. Check Method Fix Substantive content, not a stub Thin pages dilute a site rather than extending it. Model Human Heading structure matches the argument Headings are how a machine finds the part of the page that answers a question. Rule Assist Answers the query its title implies Model-judged and labelled as such, because it is a judgement. Model Human No duplicated body content across URLs Two URLs with the same content compete with each other and neither wins. Graph Assist What usually goes wrong A page that takes four paragraphs to reach the answer its title promised Thin pages that dilute the site rather than extending it Headings that decorate rather than describe, so nothing can be skimmed or lifted The same body content served on two URLs When these checks return unmeasured A check in this category that cannot run is reported as unmeasured and counted in the denominator. It is never rounded up to a pass, which is the difference between a score that bounds its own claim and one that implies it examined everything. The model-judged checks could not run because the provider was unavailable The page returned no readable body content to assess Where to start Put the direct answer in the first hundred words of your top twenty pages. Fix heading order before rewriting any prose. Consolidate or delete the thin pages rather than expanding them. Previous category Search appearance Next category E-E-A-T & trust Scan your site free All eight categories Keep reading How coverage and scoring work All eight categories How a finding becomes a reviewable diff --- ## Crawling & structure checks URL: https://crawld.co/checks/crawling-structure/ THE RUBRIC Crawling, rendering & site structure Whether the content is in the HTML, and whether the site hangs together. 16 checks · 10% of the score Checks 16 12% of 129 Weight 10% of the overall score Share of score Why it is weighted this way Separated from indexability because a page can be perfectly indexable and still arrive empty. This is where client-side rendering does its damage. What this category covers Whether content is present without JavaScript Internal link reachability and orphan detection Click depth from the homepage Sitemap accuracy against what exists Crawl traps from faceted or parameterised URLs Pagination and its canonical handling Representative checks A sample of what sits in this category, with how each one is decided and who is allowed to fix it. The full list of 129 lives in the product. Check Method Fix Content present without JavaScript The single most common finding on sites built with AI tooling. Most answer-engine crawlers do not execute your bundle. Rule Assist Reachable from an internal link An orphan page depends entirely on the sitemap being read and trusted. Graph Assist Click depth from the homepage Depth correlates with crawl frequency. Pages six clicks deep get visited rarely. Graph Assist Sitemap exists, is valid, and matches reality A sitemap listing pages that 404 spends your crawl budget on nothing. Rule Auto No crawl traps Faceted URLs generating infinite combinations consume budget that indexable pages needed. Graph Assist What usually goes wrong A page that paints entirely client-side, arriving at most answer-engine crawlers as an empty container Orphan pages that depend entirely on the sitemap being read and trusted Faceted navigation generating more URLs than the site has products Commercially important pages sitting five or six clicks from the homepage When these checks return unmeasured A check in this category that cannot run is reported as unmeasured and counted in the denominator. It is never rounded up to a pass, which is the difference between a score that bounds its own claim and one that implies it examined everything. The crawl is capped before reaching a section, which the coverage figure reports A rendering timeout prevents comparing the raw and rendered document The sitemap is unreachable, so its accuracy cannot be assessed Where to start Compare raw HTML against the rendered page on five representative URLs. Find the orphans, then link to them from somewhere a reader would plausibly click. Cap or block the parameter combinations that generate URLs without limit. Previous category Technical foundation Next category Search appearance Scan your site free All eight categories Keep reading How coverage and scoring work All eight categories How a finding becomes a reviewable diff --- ## E-E-A-T & trust checks URL: https://crawld.co/checks/eeat/ THE RUBRIC E-E-A-T & trust signals Whether there is anything here to trust or attribute. 15 checks · 15% of the score Checks 15 12% of 129 Weight 15% of the overall score Share of score Why it is weighted this way Weighted at 15% and rising in importance, because attribution is precisely what an answer engine needs in order to cite you rather than someone else. What this category covers Named authorship and author profiles Published and updated dates Organisation identity and contact details Whether claims are sourced Editorial policy and correction signals Author expertise relative to the subject Representative checks A sample of what sits in this category, with how each one is decided and who is allowed to fix it. The full list of 129 lives in the product. Check Method Fix Named author with a real profile An engine that cannot characterise the source often prefers one it can. Rule Human Published and updated dates present Freshness cannot be assessed without them. Rule Auto Organisation identity and contact details Anonymous sites are a weaker citation for a model to stand behind. Rule Human Claims are sourced Model-judged, advisory, and never auto-fixed. The engine has no basis to invent a citation. Model Human What usually goes wrong No named author anywhere, leaving an engine with nothing to characterise Missing dates, which makes freshness impossible to assess Statistics quoted without a source A contact page with a form and no organisation identity behind it When these checks return unmeasured A check in this category that cannot run is reported as unmeasured and counted in the denominator. It is never rounded up to a pass, which is the difference between a score that bounds its own claim and one that implies it examined everything. Author or organisation markup is absent entirely, so there is nothing to evaluate A model-judged expertise check could not run Where to start Add author, date and organisation to your top pages. Three lines of markup. Source the statistics you already quote. Give authors real profile pages rather than a name in plain text. Previous category Content quality Next category Content health Scan your site free All eight categories Keep reading How coverage and scoring work All eight categories How a finding becomes a reviewable diff --- ## Answer-engine readiness checks URL: https://crawld.co/checks/geo/ THE RUBRIC AI / answer-engine readiness (GEO) Whether an AI answer can lift your content and credit you for it. 16 checks · 12% of the score Checks 16 12% of 129 Weight 12% of the overall score Share of score Why it is weighted this way Twelve per cent, and the reason this rubric differs from a conventional SEO audit. Graded on the same scale as everything else, so an unmeasured GEO check is never a pass. What this category covers Whether an answer is extractable from a passage llms.txt presence and usefulness Answer-engine crawler permissions in robots.txt Attribution metadata an engine can cite Question-shaped headings Server-rendered content availability Representative checks A sample of what sits in this category, with how each one is decided and who is allowed to fix it. The full list of 129 lives in the product. Check Method Fix llms.txt present and useful A curated list of what is worth reading, not a dump of every URL. Rule Auto Answer extractable from a passage Engines quote passages. A claim buried mid-essay is harder to lift than the same claim under a matching heading. Model Assist AI crawlers permitted in robots.txt Blocking them is a valid choice, but it should be a choice. Rule Assist Attribution metadata present Author, date, organisation: the things a citation is built from. Rule Auto What usually goes wrong A claim buried mid-essay, where no passage states the answer directly A robots.txt permitting Googlebot while silently blocking the answer-engine agents No llms.txt, so nothing curates which pages are worth reading Anonymous pages with nothing an engine can attribute a quote to When these checks return unmeasured A check in this category that cannot run is reported as unmeasured and counted in the denominator. It is never rounded up to a pass, which is the difference between a score that bounds its own claim and one that implies it examined everything. The page could not be fetched without JavaScript, so extractability cannot be judged A model-judged extractability check could not run Where to start Add a one-sentence direct answer under the main heading of your best pages. Check robots.txt for OAI-SearchBot, PerplexityBot and ClaudeBot explicitly. Write an llms.txt. It takes about ten minutes. Previous category Content health Next category Page experience Scan your site free All eight categories Keep reading How coverage and scoring work All eight categories How a finding becomes a reviewable diff --- ## Page experience checks URL: https://crawld.co/checks/page-experience/ THE RUBRIC Page experience & performance Whether the page is usable once it arrives. 12 checks · 8% of the score Checks 12 9% of 129 Weight 8% of the overall score Share of score Why it is weighted this way The lightest category at 8%, deliberately. Performance is a real ranking input and a much smaller one than the industry implies. A fast page with nothing to say still has nothing to say. What this category covers Largest Contentful Paint Interaction to Next Paint Cumulative Layout Shift Image sizing and lazy loading Render-blocking resources Mobile viewport and tap target sizing Representative checks A sample of what sits in this category, with how each one is decided and who is allowed to fix it. The full list of 129 lives in the product. Check Method Fix Largest Contentful Paint Shipped. Measured from field or lab data depending on availability. API Assist Interaction to Next Paint Shipped. API Assist Cumulative Layout Shift Shipped. Most commonly caused by images without dimensions, which is an Auto fix. API Auto Images sized and lazily loaded below the fold The cheapest performance win on most sites. Rule Auto What usually goes wrong Images without width and height, which is the most common cause of layout shift Render-blocking third-party scripts loaded before first paint Fonts fetched from a third-party CDN, adding two cross-origin round trips before text appears Tap targets too small or too close together on mobile When these checks return unmeasured A check in this category that cannot run is reported as unmeasured and counted in the denominator. It is never rounded up to a pass, which is the difference between a score that bounds its own claim and one that implies it examined everything. Field data is unavailable for a low-traffic page, so only lab data exists The third-party performance API was unavailable during the crawl Where to start Add width and height to every image. Mechanical, and it usually fixes layout shift outright. Self-host fonts rather than fetching them at request time. Defer third-party scripts that do not need to run before first paint. Previous category Answer-engine readiness Scan your site free All eight categories Keep reading How coverage and scoring work All eight categories How a finding becomes a reviewable diff --- ## Search appearance checks URL: https://crawld.co/checks/search-appearance/ THE RUBRIC Search appearance What the result looks like once you have earned one. 18 checks · 10% of the score Checks 18 14% of 129 Weight 10% of the overall score Share of score Why it is weighted this way Ten per cent because it affects click-through rather than eligibility. It is also the category with the highest proportion of Auto fixes, which makes it the cheapest ground to recover. What this category covers Title presence, uniqueness and length Meta description presence and quality Open Graph and Twitter card completeness Structured data validity, and whether it matches the page Breadcrumb markup Favicon and site name signals Representative checks A sample of what sits in this category, with how each one is decided and who is allowed to fix it. The full list of 129 lives in the product. Check Method Fix Title present, unique, and within length Truncation is a smaller problem than duplication across a hundred pages. Rule Assist Meta description present and non-empty An empty description is worse than a missing one. It explicitly says there is nothing to say. Rule Auto Structured data valid and matching the page Schema that misdescribes the page is a liability, not a bonus. Rule Assist Open Graph and Twitter tags complete Governs how the page appears when a human shares it, which is where a lot of early traffic comes from. Rule Auto What usually goes wrong An empty meta description, which is worse than a missing one because it states that there is nothing to say Duplicate titles, which arrive in clusters from a single template Schema that validates cleanly while describing the wrong page type Titles that describe the site rather than the page When these checks return unmeasured A check in this category that cannot run is reported as unmeasured and counted in the denominator. It is never rounded up to a pass, which is the difference between a score that bounds its own claim and one that implies it examined everything. A page could not be fetched, so its tags cannot be read Structured data validation depends on an external service that is unavailable Where to start Export every title, sort alphabetically, and read the duplicates. Fill the empty descriptions before writing new ones from scratch. Spot-check one page per template for schema type accuracy. Previous category Crawling & structure Next category Content quality Scan your site free All eight categories Keep reading How coverage and scoring work All eight categories How a finding becomes a reviewable diff --- ## Technical foundation checks URL: https://crawld.co/checks/technical-foundation/ THE RUBRIC Technical foundation & indexability Whether an engine can reach the page and is permitted to keep it. 20 checks · 15% of the score Checks 20 16% of 129 Weight 15% of the overall score Share of score Why it is weighted this way Weighted heavily because everything else is conditional on it. A page that cannot be indexed cannot rank, cannot be cited, and cannot be improved by fixing anything else on this list. What this category covers Status codes, redirects and redirect chains robots.txt directives, including the answer-engine agents Meta robots and X-Robots-Tag agreement Canonical tags, and whether they resolve HTTPS, certificate validity and mixed content Sitemap presence, validity and accuracy Representative checks A sample of what sits in this category, with how each one is decided and who is allowed to fix it. The full list of 129 lives in the product. Check Method Fix Page returns a 200 A soft 404 or a redirect chain ending nowhere is invisible to everything downstream. Rule Human Not blocked by robots.txt Including the AI crawlers, which a robots.txt written years ago frequently blocks by accident. Rule Assist No conflicting noindex A meta robots tag and an X-Robots-Tag header disagreeing is a common and silent cause of disappearance. Rule Auto Canonical resolves to itself or a real page A canonical pointing at a 404 or a redirect tells the engine to index nothing. Rule Auto HTTPS, valid certificate, no mixed content Mixed content downgrades trust signals and breaks rendering in ways that are hard to see. Rule Human What usually goes wrong A canonical pointing at a redirect or a 404, which tells the engine to index something other than the page A soft 404 returning 200, keeping a dead page in the index and spending crawl budget on it every visit A robots.txt written before the answer-engine crawlers existed, blocking them by omission A staging subdomain left indexable, competing with production for its own content When these checks return unmeasured A check in this category that cannot run is reported as unmeasured and counted in the denominator. It is never rounded up to a pass, which is the difference between a score that bounds its own claim and one that implies it examined everything. The crawler is blocked outright by robots.txt or a firewall The page sits behind authentication The server times out or returns a 5xx during the crawl Where to start Fix conflicting index directives first. They are cheap to correct and they remove pages silently. Resolve canonicals that point somewhere invalid. Audit robots.txt against the list of crawlers you actually want. Next category Crawling & structure Scan your site free All eight categories Keep reading How coverage and scoring work All eight categories How a finding becomes a reviewable diff --- ## FAQ URL: https://crawld.co/faq/ FAQ The questions people actually ask. Grouped, and answered without hedging. If an answer is 'not yet', it says that. The free scan How does the free scan work? Enter a domain. We crawl your pages the way Googlebot and AI crawlers do, run all 129 checks , and show the full scorecard, every pass, fail, and unmeasured item . No signup, no card, and the results aren't held hostage behind a paywall. Why only 10 pages? Because a crawl costs money and the free tier is a lead magnet rather than a product. The result always reports how many pages it covered against how many it found, so the score is explicitly a claim about that slice rather than about your whole site. Can I rescan immediately? One real crawl per domain per day. A repeat request inside that window returns the scan that already ran rather than crawling again, which also stops the endpoint being used to point repeated traffic at somebody else’s site. Why do you want my email for the findings? The score, grade, category breakdown and the number of findings are free and always visible, that is the diagnosis. The findings themselves are what an address buys. A shared scorecard link never unlocks them, whoever opens it. How long does a share link last? Thirty days. A shared scorecard is a snapshot, not a permanent free product. Scoring What does "unmeasured" mean, and why does it matter so much? It means the check could not run: the crawler could not reach the page, a resource timed out, or a third-party source was unavailable. It is not counted as a pass. The industry default is to quietly round it up, which is how tools report 95/100 for sites with real problems. Why is my score lower here than on other tools? Usually because they counted what they could not check as fine. Compare the coverage figures rather than the scores: a 94 at 40% coverage and a 71 at 95% coverage are not describing the same thing. What's the difference between the SEO checks and the GEO checks? The SEO checks cover what search engines rank: metadata, structure, speed, internal linking. The 16 GEO checks cover what AI answer engines cite: whether your content is structured so ChatGPT, Claude, Perplexity, and AI Overviews can lift it and attribute it to you. Do you use AI to decide whether my content is good? For some checks, yes, and those are labelled Model wherever they appear so you can weigh them accordingly. Most of the rubric is deterministic rules and graph analysis over your crawled site. A model judgement is never silently presented as a fact. Fixes and delivery What does "2 human gates" actually mean? The loop stops and waits for you twice: once before any fix is generated (you approve the ranked plan ) and once before anything ships (you approve the delivery ). Nothing touches your site without those two clicks from you. I'm on Squarespace / Webflow / Wix, how do fixes arrive without code? Three delivery modes . With code access, you get a pull request , a proposed change you approve. With CMS access, we apply changes through your platform's own editor, listed for your approval first. With neither, you get a copy-paste patch file with exact instructions for each change. Can a fix break my site or hurt my rankings? Every change is build-verified before you ever see it, and shipped changes are re-measured on the next scan . If a fix doesn't move its check from fail to pass, it gets flagged, and anything delivered as a PR is one click to revert . What does "build-verified" actually mean? That a build ran against the change before you saw it. It applies to the modes where we hold or mirror a repository, and to nothing else. A change written into a hosted CMS is preview-verified, and a patch you apply yourself is advisory. Those labels are never blurred. Which platform integrations can I use today? The free scorecard, which needs no access at all. The connected delivery modes (pull requests, CMS drafts) are the model the product is designed around and are not connectors you can authorise today. The integrations pages mark the status of each one. Data and the crawler What do you do with my data? We store the crawl results and the fixes we generate for you, that's it. We don't resell crawl data, don't train models on your content, and don't keep repository or CMS access beyond the scopes you grant . Revoke access any time and the data goes with it. How will I recognise your crawler in my logs? It identifies itself as CrawldBot and the user agent links to a page explaining what it does and how to block it. It never crawls under a browser user agent. Someone is monitoring my site with this. How do I stop it? Disallow CrawldBot in your robots.txt and it stops on the next scheduled run. Competitor monitoring is capped at 12 pages per site, one request at a time, once a week, and robots.txt is obeyed. Do you train models on my content? No, and we do not resell crawl data. Repository or CMS access is never kept beyond the scopes you grant, and revoking it takes the data with it. Scan your site free How we measure Keep reading The scoring method in full Crawler behaviour and how to block it Where to send a bug or a security report --- ## Help centre URL: https://crawld.co/help/ HELP Start here. A short index rather than a ticket queue. Most questions are answered by one of these seven pages. Understanding your score What pass, fail and unmeasured mean, and why coverage is next to every number. Read → What the checks look for All eight categories, how each check is decided, and who is allowed to fix it. Read → Using the free scan What it covers, why it is capped, and what a verified email unlocks. Read → Limits and quotas Crawl depth, cooldowns, rate limits and share-link expiry. Read → The API Endpoints, response envelope and error codes. Read → Our crawler on your site What it fetches, how to identify it, how to block it. Read → Is something broken? How to tell an outage from a limit, and how to report one. Read → Still stuck Contact us with what you submitted, roughly when, and the requestId from the response if you have one. Every API response carries one and it is the fastest route to an answer. Read the full FAQ Contact us Keep reading Understanding your score API endpoints and error codes How to tell an outage from a limit --- ## REST API URL: https://crawld.co/integrations/api/ BUILD Crawld and REST API Drive scans and read findings programmatically. Status: Not built yet . this connector is not built. The page describes how delivery works on REST API given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a REST API site without granting any access at all. How fixes would arrive You call the API and decide what to do with the result. Nothing is written on your behalf. The API is the substrate the other integrations sit on rather than an integration itself. What it can deliver depends entirely on what access you have granted alongside it. Delivery mode Mode A · Repo write Verification Build-verified The strongest rung. Because we hold the repository, the change can be built before you ever see it, and the pull request carries the result of that build. Nothing merges without your click. What the rubric can see here Queue a scan for a URL and poll it to completion Read the scorecard: overall score, per-category scores, coverage Read the findings list once a scan is unlocked Limits worth knowing before you plan around this The public scan surface is unauthenticated, rate limited to 3 submissions an hour per caller, and capped at 10 pages per crawl One real crawl per domain per day: repeat calls return that day’s scan rather than re-crawling The authenticated API is not documented publicly yet Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to REST API at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your REST API site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Framer URL: https://crawld.co/integrations/framer/ MORE PLATFORMS Crawld and Framer Advisory only: no content-write API exists. Status: Not built yet . this connector is not built. The page describes how delivery works on Framer given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Framer site without granting any access at all. How fixes would arrive A findings list you apply in the Framer editor. Framer has no public API for writing page content or metadata. Claiming otherwise would be inventing a capability, so this sits on the bottom rung. Delivery mode Mode E · URL only Verification Advisory only What the free scorecard runs on today. It reads public HTML, scores it, and hands back findings. It changes nothing and claims nothing about a build, because it has never seen one. What the rubric can see here Everything visible in the published HTML Per-page instructions naming what to change in the editor Limits worth knowing before you plan around this Advisory only Framer sites are heavily client-rendered, which is itself one of the most common findings here Not build-verified. Framer sits on Mode E, so a change delivered here is advisory only: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Framer at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Framer site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Ghost URL: https://crawld.co/integrations/ghost/ PLATFORMS Crawld and Ghost Post metadata and content updated as drafts. Status: Not built yet . this connector is not built. The page describes how delivery works on Ghost given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Ghost site without granting any access at all. How fixes would arrive Draft revisions through the Admin API, published by you. Ghost’s Admin API writes posts and their metadata cleanly, which makes drafts straightforward. Delivery mode Mode D · CMS drafts Verification Preview-verified There is no build to verify against on a hosted CMS, so the change is checked in preview instead. This rung is deliberately never labelled build-verified, because it is not. What the rubric can see here Post titles, excerpts, meta titles and descriptions Canonical URLs set per post Structured data derived from post fields Limits worth knowing before you plan around this Never build-verified Handlebars theme templates are not edited Not build-verified. Ghost sits on Mode D, so a change delivered here is preview-verified: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Ghost at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Ghost site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## HubSpot URL: https://crawld.co/integrations/hubspot/ MORE PLATFORMS Crawld and HubSpot CMS pages and blog posts updated as drafts. Status: Not built yet . this connector is not built. The page describes how delivery works on HubSpot given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a HubSpot site without granting any access at all. How fixes would arrive Draft updates through the CMS API, published on your action. HubSpot’s CMS API exposes page and blog content, including the metadata fields that matter most. Delivery mode Mode D · CMS drafts Verification Preview-verified There is no build to verify against on a hosted CMS, so the change is checked in preview instead. This rung is deliberately never labelled build-verified, because it is not. What the rubric can see here Page titles, meta descriptions and canonical URLs Blog post metadata and body content Alt text on hosted files Limits worth knowing before you plan around this Never build-verified HubL templates and modules are outside scope Which fields are writable depends on your HubSpot tier Not build-verified. HubSpot sits on Mode D, so a change delivered here is preview-verified: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to HubSpot at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your HubSpot site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Next.js URL: https://crawld.co/integrations/nextjs/ BUILD Crawld and Next.js Fixes arrive as a pull request, built before you see it. Status: Not built yet . this connector is not built. The page describes how delivery works on Next.js given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Next.js site without granting any access at all. How fixes would arrive A branch with the change, a pull request, and the result of a real build attached to it. A Next.js site is a repository, so the change can be compiled before it is proposed. That is what makes the build-verified label honest here and not on a hosted CMS. Delivery mode Mode A · Repo write Verification Build-verified The strongest rung. Because we hold the repository, the change can be built before you ever see it, and the pull request carries the result of that build. Nothing merges without your click. What the rubric can see here Metadata exported from generateMetadata and static metadata objects Canonical tags, Open Graph and structured data emitted from layouts App Router and Pages Router both, including per-route metadata Whether a route ships content in its HTML or paints it client-side Limits worth knowing before you plan around this Content coming from a headless CMS is fixed at the CMS, not in the repo, that is a Mode D problem living inside a Mode A site A fix that depends on runtime data cannot be verified by a build alone Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Next.js at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Next.js site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Notion URL: https://crawld.co/integrations/notion/ MORE PLATFORMS Crawld and Notion Page properties updated where a site generator reads them. Status: Not built yet . this connector is not built. The page describes how delivery works on Notion given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Notion site without granting any access at all. How fixes would arrive Property and block updates through the Notion API. Notion is a source of content rather than a website. What reaches the web depends on whichever generator sits in front of it, and that is what actually gets scored. Delivery mode Mode D · CMS drafts Verification Preview-verified There is no build to verify against on a hosted CMS, so the change is checked in preview instead. This rung is deliberately never labelled build-verified, because it is not. What the rubric can see here Page properties a generator maps to titles and descriptions Heading structure inside the page body Limits worth knowing before you plan around this Never build-verified The rubric scores the published site, not the Notion workspace: findings that live in the generator’s templates cannot be fixed here Not build-verified. Notion sits on Mode D, so a change delivered here is preview-verified: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Notion at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Notion site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Shopify URL: https://crawld.co/integrations/shopify/ PLATFORMS Crawld and Shopify Product and page metadata written through the Admin API. Status: Not built yet . this connector is not built. The page describes how delivery works on Shopify given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Shopify site without granting any access at all. How fixes would arrive Metafield and page updates staged for your review. The Admin API covers products, collections and pages. Theme Liquid is a separate, riskier surface and is left alone. Delivery mode Mode D · CMS drafts Verification Preview-verified There is no build to verify against on a hosted CMS, so the change is checked in preview instead. This rung is deliberately never labelled build-verified, because it is not. What the rubric can see here Product and collection titles, descriptions and SEO fields Page metadata and handles Structured data emitted from product metafields Limits worth knowing before you plan around this Never build-verified Liquid templates are not edited: a theme-level problem is reported, not fixed App-injected markup often cannot be attributed to a source you control Not build-verified. Shopify sits on Mode D, so a change delivered here is preview-verified: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Shopify at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Shopify site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Squarespace URL: https://crawld.co/integrations/squarespace/ PLATFORMS Crawld and Squarespace Advisory only: a checklist you apply by hand. Status: Not built yet . this connector is not built. The page describes how delivery works on Squarespace given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Squarespace site without granting any access at all. How fixes would arrive A findings list with the exact change for each page, applied by you in the editor. Squarespace’s public API covers commerce, not page content. There is no supported way to write a page’s metadata programmatically, so this platform sits on the bottom rung and the site says so. Delivery mode Mode E · URL only Verification Advisory only What the free scorecard runs on today. It reads public HTML, scores it, and hands back findings. It changes nothing and claims nothing about a build, because it has never seen one. What the rubric can see here Everything the crawler can see from the public HTML Per-page instructions naming the field to change in the editor Limits worth knowing before you plan around this Advisory only. Nothing is written, and nothing is verified If Squarespace opens a content-write API this moves up a rung; until then it does not Not build-verified. Squarespace sits on Mode E, so a change delivered here is advisory only: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Squarespace at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Squarespace site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Unicorn Platform URL: https://crawld.co/integrations/unicorn-platform/ MORE PLATFORMS Crawld and Unicorn Platform Advisory only. Status: Not built yet . this connector is not built. The page describes how delivery works on Unicorn Platform given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Unicorn Platform site without granting any access at all. How fixes would arrive A findings list you apply in the builder. No supported content-write API, so the honest rung is the bottom one. Delivery mode Mode E · URL only Verification Advisory only What the free scorecard runs on today. It reads public HTML, scores it, and hands back findings. It changes nothing and claims nothing about a build, because it has never seen one. What the rubric can see here Everything visible in the published HTML Per-page instructions for the builder Limits worth knowing before you plan around this Advisory only Not build-verified. Unicorn Platform sits on Mode E, so a change delivered here is advisory only: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Unicorn Platform at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Unicorn Platform site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Webflow URL: https://crawld.co/integrations/webflow/ PLATFORMS Crawld and Webflow CMS collection items updated as drafts. Status: Not built yet . this connector is not built. The page describes how delivery works on Webflow given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Webflow site without granting any access at all. How fixes would arrive Draft updates to CMS collection items, published on your action. Webflow’s CMS API can write collection fields. The Designer canvas is a separate surface the API does not reach. Delivery mode Mode D · CMS drafts Verification Preview-verified There is no build to verify against on a hosted CMS, so the change is checked in preview instead. This rung is deliberately never labelled build-verified, because it is not. What the rubric can see here CMS collection fields, including SEO title and description per item Alt text stored on CMS image fields Slug and canonical settings exposed on collection items Limits worth knowing before you plan around this Never build-verified Static pages built in the Designer are not writable through the API. those come back as advisory findings Publishing to a custom domain is an explicit step you take Not build-verified. Webflow sits on Mode D, so a change delivered here is preview-verified: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Webflow at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Webflow site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Webhooks URL: https://crawld.co/integrations/webhooks/ BUILD Crawld and Webhooks Get told when a scan finishes or a fix is ready. Status: Not built yet . this connector is not built. The page describes how delivery works on Webhooks given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Webhooks site without granting any access at all. How fixes would arrive An HTTP POST to an endpoint you control, when something happens worth reacting to. Webhooks carry notifications, not changes. They are how you hang your own automation off the loop. Delivery mode Mode A · Repo write Verification Build-verified The strongest rung. Because we hold the repository, the change can be built before you ever see it, and the pull request carries the result of that build. Nothing merges without your click. What the rubric can see here Scan completed, with the score and coverage A fix entered the queue and is waiting on your approval A delivery was made, or failed Limits worth knowing before you plan around this Delivery is at-least-once, so your handler has to tolerate a repeat A webhook never carries the findings themselves, only the fact that they exist Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Webhooks at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Webhooks site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## WordPress URL: https://crawld.co/integrations/wordpress/ PLATFORMS Crawld and WordPress Changes land as unpublished drafts for you to publish. Status: Not built yet . this connector is not built. The page describes how delivery works on WordPress given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a WordPress site without granting any access at all. How fixes would arrive A draft revision written through the REST API. It goes live when you publish it, not before. WordPress exposes a full content API, so metadata and body copy can be written. There is no build step to verify against, so the change is checked in preview instead. Delivery mode Mode D · CMS drafts Verification Preview-verified There is no build to verify against on a hosted CMS, so the change is checked in preview instead. This rung is deliberately never labelled build-verified, because it is not. What the rubric can see here Titles, meta descriptions and canonical tags, where the SEO plugin stores them in post meta Heading structure and body content in the editor Alt text on media library items Limits worth knowing before you plan around this Never build-verified. There is no build Theme-level output, functions.php and hardcoded template markup are outside what the API can reach Which SEO plugin you run changes where metadata lives, and therefore what can be written Not build-verified. WordPress sits on Mode D, so a change delivered here is preview-verified: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to WordPress at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your WordPress site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Zapier URL: https://crawld.co/integrations/zapier/ BUILD Crawld and Zapier Route scan results into whatever your team already uses. Status: Not built yet . this connector is not built. The page describes how delivery works on Zapier given the constraints of the platform, and what verification would honestly be available. Today you can run the free scorecard against a Zapier site without granting any access at all. How fixes would arrive A trigger fires in Zapier and you decide where it goes: Slack, a ticket, a spreadsheet. Zapier moves information between tools. It can carry a finding to the person who should act on it; it cannot compile anything. Delivery mode Mode D · CMS drafts Verification Preview-verified There is no build to verify against on a hosted CMS, so the change is checked in preview instead. This rung is deliberately never labelled build-verified, because it is not. What the rubric can see here Scan completed triggers New fix queued triggers Mapping findings into a tracker or a channel Limits worth knowing before you plan around this Nothing routed through Zapier is build-verified, because nothing is built Rate limits are Zapier’s, and they are stricter than the API’s Not build-verified. Zapier sits on Mode D, so a change delivered here is preview-verified: never build-verified. That badge means a build actually ran, and on this platform there is no build to run. Where this sits on the ladder Mode What we get How changes arrive Verification Status A · Repo write A connected repository we can push a branch to Opens a pull request. You review the diff and merge it. Build-verified Not built yet B · Read-only repo Read access, plus somewhere to push Pushes to a fork or a branch you own. Build-verified Not built yet C · Cloud mirror A one-time snapshot or export Works on a clone in our cloud, returns a patch or a PR. Build-verified (mirror) Not built yet D · CMS drafts OAuth into a hosted CMS Writes unpublished drafts. You publish them. Preview-verified Not built yet E · URL only Nothing: a public URL A scorecard and a patch you apply yourself. Advisory only Available now Build-verified means a build actually ran. It applies to modes A, B and C and to nothing else. A hosted CMS has no build to run, so mode D is preview-verified and mode E is advisory. Those labels are never blurred, because the honesty of the verification is the product. What you can do today Scan the site. The free scorecard needs no access to Zapier at all. It reads the published HTML the way a search or answer engine does, runs the full 129-check rubric against it, and reports what it could not measure rather than counting it as a pass. Everything it finds is actionable by hand regardless of whether a connector ever exists, because the finding names the page and the change rather than a button to press. Scan your Zapier site All integrations Keep reading The five delivery modes explained How fixes arrive as reviewable diffs What is shipped and what is not --- ## Download report URL: https://crawld.co/report/ REPORTS Report export is not built yet. Listed because people look for it, and answered honestly because promising an export that does not exist is the sort of thing this product is supposed to be against. No PDF, CSV or JSON export. Not built yet Data export is on the roadmap and is explicitly flagged internally as something not to promise yet. There is no download button hiding behind a paywall: the feature does not exist. What you can take away today A share link. Every free scan produces one. It shows the score, grade, category breakdown and coverage, expires after 30 days, and never exposes the findings themselves regardless of who opens it. The API. GET /api/v1/public/scans/{scanId} returns the whole scorecard as JSON. If you need a report today, this is the honest route. It is the same data any export would contain. The docs cover it. The screen. Unsatisfying, and true. What an export will have to get right The reason this is not trivial: a report that shows a score without its coverage figure is exactly the dishonesty the rubric exists to avoid. Any export has to carry the denominators: how many checks ran, how many pages were reached against how many exist, and which findings were model-judged rather than deterministic. A prettier version of a bare number would be worse than nothing. Agency reporting White-label reporting sits on the agency tier and is subject to the same constraint. Until export exists, the API is the integration point for anyone building client reporting on top of this. Read the API docs What else is not built Keep reading The API that returns the same data What else is not built yet Queue and scan limits --- ## E-commerce URL: https://crawld.co/solutions/ecommerce/ SOLUTIONS Crawld for e-commerce Coming soon: the catalogue-specific checks are not built. Status: Not built yet : read the cautions below before planning around this. The rubric applies; the segment-specific work does not exist yet. The situation Product catalogues fail in ways a content site does not: faceted navigation generating near-infinite URLs, thin variant pages competing with each other, and product schema that has to match inventory that changes daily. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. What works today The general rubric applies: indexability, crawl waste, metadata, structured-data validity, page experience. A store is still a site. What is missing Catalogue-aware checks: variant consolidation, faceted-URL policy, inventory-linked schema freshness, and category-page cannibalization at catalogue scale. Delivery The Shopify connector that would write metafields is not built. Findings are advisory and applied by hand. What the loop looks like The free scan is genuinely useful on a store today for crawl waste and indexability. The catalogue-specific work is not there yet. Where this is a poor fit Marked coming soon because it is. Do not buy this expecting catalogue-aware analysis in the current version. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## Enterprise URL: https://crawld.co/solutions/enterprise/ SOLUTIONS Crawld for enterprise Read this before you start a procurement process: several answers are no. Status: Not built yet : read the cautions below before planning around this. The rubric applies; the segment-specific work does not exist yet. The situation You are evaluating tooling against a security review, a compliance checklist and an SLA. It is cheaper for both of us to establish now which of those this product currently fails. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. The gates are mandatory The loop stops twice for human approval and that cannot be turned off. WEBSITE.md is explicit: this is not for teams wanting a fully autonomous agent. If the value you are buying is "nobody has to look at it", this is the wrong tool. No certifications No SOC 2, no ISO 27001, no published penetration test, no contractual uptime guarantee. These are stated plainly on the security page rather than discovered during review. Content leaves our infrastructure Model-judged checks send page content to a large-language-model provider. That is disclosed on the privacy page and is a real consideration for a data-classification review. What the loop looks like The free scorecard needs no access and no contract, so a technical evaluation costs nothing and commits nothing. Where this is a poor fit If your procurement requires certification, an SLA or a signed DPA today, this is not yet a product you can buy. That is a straight answer rather than a "contact sales". Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## Higher education URL: https://crawld.co/solutions/higher-education/ SOLUTIONS Crawld for higher education Sprawling multi-department sites where nobody owns the whole domain. The situation Dozens of departments publishing independently under one domain, several CMSs, and orphaned pages going back years. No single person has ever seen the whole site. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. Orphaned and unreachable pages Graph analysis over the crawled site finds what nothing links to: the characteristic failure of a large devolved site. Duplicate and competing pages Three departments publishing their own version of the same information, competing with each other under one domain. Coverage honesty On a site this size, a score without a coverage figure is close to meaningless. This is the segment where the unmeasured state matters most. What the loop looks like Scan to get a map of the actual domain rather than the assumed one. The first useful output is usually the inventory, before any fix. Where this is a poor fit Devolved publishing means findings often need routing to whoever owns that department, and there is no workflow for that. The output is a list, not an assignment system. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## Nonprofits URL: https://crawld.co/solutions/nonprofits/ SOLUTIONS Crawld for nonprofits The free scorecard is a complete diagnostic and costs nothing. The situation A site maintained by whoever has time, often on a hosted platform, usually with no budget line for search tooling and no in-house specialist. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. Indexability and structure The findings that matter most here are also the cheapest to fix, and most are Auto-capability: one correct answer, no judgement. Trust signals Organisation identity, contact details and sourced claims. Weighted heavily, and usually straightforward for a nonprofit to satisfy properly. Platform constraints Squarespace, Wix and similar have no content-write API, so findings arrive as instructions rather than changes. That is the honest ceiling. What the loop looks like Run the free scan, action the top findings by hand in the editor. For many organisations that is the entire useful engagement, and it is free. Where this is a poor fit There is no nonprofit pricing tier today. The free scorecard is genuinely free and uncapped in features. It is capped in pages. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## Publishers URL: https://crawld.co/solutions/publishers/ SOLUTIONS Crawld for publishers Archive-scale corpora, where overlap and decay compound quietly. The situation A large archive accumulated over years, much of it still ranking, some of it competing with itself, and no practical way to audit all of it by hand. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. Cannibalization at scale A decade of coverage on a recurring subject produces dozens of pages targeting the same intent. This is the segment where corpus-level analysis pays for itself fastest. E-E-A-T signals Named authors, dates, sourced claims and organisation identity: heavily weighted, and the things answer engines use to decide whether you are attributable. Answer-engine citation Publishers are among the most cited and least measured. Citation share per engine is the metric this segment currently flies blind on. What the loop looks like Scan, work the overlap map across the archive, re-measure. Archive consolidation is the highest-leverage work and the easiest to get wrong without a before-and-after. Where this is a poor fit The crawl covers what it can reach; a very large archive will report partial coverage, and the coverage figure is what tells you how much of it the score actually describes. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## SEO professionals URL: https://crawld.co/solutions/seo-professionals/ SOLUTIONS Crawld for seo professionals A rubric that shows its working, and a coverage figure you can defend in a meeting. The situation You already know what is wrong. What you lack is an instrument that reports honestly enough to hand to someone else, and any measurement at all of whether answer engines cite you. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. Method transparency Every check declares whether it is a deterministic rule, a graph analysis, a third-party API or a model judgement. You can weigh a finding before acting on it. Citation share Per engine and per prompt. A separate instrument from rank tracking, measuring a surface rank tracking cannot see. Honest denominators Coverage next to every score. A 94 at 40% coverage and a 71 at 95% are not the same claim, and only one tool tells you which you are looking at. What the loop looks like Use it as the measurement layer under whatever process you already run. The ranked plan is a starting point you are expected to argue with. Where this is a poor fit This is not a replacement for a broad suite. There is no keyword research, no backlink analysis and no paid intelligence, and none is planned. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## Startups URL: https://crawld.co/solutions/startups/ SOLUTIONS Crawld for startups Find out why organic traffic never started, before you spend on ads to compensate. The situation Launch happened, the site looks right, and organic never arrived. The usual diagnosis is "SEO takes time" when the actual cause is that the pages were never readable by a crawler in the first place. Which findings dominate here The product is the same product for every segment: one rubric, 129 checks, the same three outcomes. What changes is which findings do the damage. Indexability Whether the pages can be reached and stored at all. Necessary before anything else on the list matters. Answer-engine readiness A disproportionate share of early discovery now happens through AI answers rather than results pages, and new sites are the least likely to be structured for it. Thin and duplicate pages Programmatic or templated pages that dilute the site rather than extending it. What the loop looks like The free scan is enough to answer "is this a content problem or a plumbing problem", which is the question actually blocking you. Where this is a poor fit If the site is five pages, the free scan will tell you that in a minute and there is likely nothing here worth paying for yet. Scan your site free All solutions Keep reading The rubric every segment shares All audience pages Pricing by tier --- ## About URL: https://crawld.co/about/ ABOUT Built around one idea. A check that could not run is never counted as a pass. Almost everything else about this product follows from taking that seriously. The problem Search tooling splits into two halves that both stop short. Reports end at a number: a score, a PDF, four hundred rows, and a bill, with the fixing left as your weekend. Content tools end at a brief. They tell you what to write, hand it back, and nothing gets published or measured. Underneath both sits a quieter problem: the numbers are frequently soft. Checks that could not run get counted as passing, so a 94/100 can mean "we looked at a third of your site and liked what we saw." The idea Three outcomes, not two. Pass, fail, and unmeasured , and unmeasured is never rounded up. Coverage is reported next to every score, so a number is always a bounded claim about what was actually examined rather than an implied claim about everything. This sounds like table stakes and is not. It is the reason a Crawld score reads lower than a competitor's on the same site, and the reason it means more. What follows from it Every check declares its method (deterministic rule, graph analysis, third-party API, or model judgement) so you can weigh a finding before acting on it. Every check declares who may fix it. Where only a person can decide, the engine reports and stops rather than guessing convincingly. Verification labels follow the access. Build-verified means a build ran. A hosted CMS has no build, so it is never given that badge. The loop closes. Changes are re-measured, because a fix that shipped and a fix that worked are different things. What it refuses to do The loop stops twice for human approval and that cannot be turned off. An unsupervised agent editing a production site is a different product with a different risk profile, and not this one. If the value you want is that nobody has to look at it, this will feel like friction permanently. It also does not promise rankings. Nobody can. The rubric measures things within your control; where you place is not among them. Where it currently is The measurement half is real and running: the free scorecard crawls, scores against the full rubric, and fails honestly when the engine is unavailable. The delivery half (pull requests, CMS drafts, the access ladder) is designed and not built. The transparency page lists which is which, and every integration page carries its own status. Publishing that distinction is not modesty. A product whose entire claim is honest measurement cannot be vague about its own state. Scan your site free How we measure Keep reading What is shipped and what is not The scoring idea the product is built on Who it is for, and who it is not --- ## Changelog URL: https://crawld.co/changelog/ CHANGELOG What changed, and why your score might have moved. A score is only comparable to itself if the thing producing it holds still. When it doesn't, that has to be visible: otherwise last quarter's number and this quarter's are two different measurements wearing the same label. Why a rubric needs a version If a category gains three checks, every score in that category shifts, and none of your pages changed. A site owner comparing this month to last month would be reading a difference that came from us rather than from them. So the rubric carries a version, it is shown in the footer of every page, and a scorecard records the version it was produced under. Two scores are only directly comparable when they share one. What gets logged Change Why it is logged A check is added, removed or reweighted Changes what a score means, so it changes the rubric version. A check changes its method A rule becoming a model judgement is a change in how much the result should be trusted. A check changes its fix capability Auto becoming Assist means a machine stopped being allowed to decide something. An engine is added to citation tracking Changes the denominator of citation share. A crawler behaviour changes Anyone who has allowed or blocked CrawldBot needs to know. Interface changes and bug fixes that do not move a number are not logged here. This is a record of things that change what a score means. Notable changes These are substantiated by the source. They are listed without dates because the source does not record them, and a date invented for a changelog is worse than no date at all. The free scan stopped inventing grades Public scorecard The public scan endpoint previously wrote a finished result the moment it was called (status done, grade C+, score 72 from the demo seed, canned category rows, eight findings) and then returned "queued". Nothing was ever fetched, for any URL, and the fabricated result was publicly shareable. A scan now either runs on the real engine or reports honestly that scanning is unavailable. There is no remaining path that produces a grade without a crawl behind it. This is the single change most relevant to anyone deciding whether to trust a number on this site. The 129-check catalogue was restored Rubric The check catalogue was believed lost and had been regenerated with placeholder names. A surviving reference table turned out to hold the originals, and the catalogue was restored from it verbatim: real names, methods, severities and fix capabilities, with per-category counts of 20/16/18/18/15/14/16/12. Those counts are what every category score on this site is a proportion of. A dated public feed Not published yet A running, dated changelog with an RSS feed is not published. The honest position is that the product does not currently maintain a public release history, and this page is the beginning of one rather than a backfill of one. If you are integrating against the rubric and need to be told when it moves, that is exactly what the version in the footer and this page are for, and it is worth asking us directly until the feed exists. See the current rubric How scores are calculated Keep reading The current rubric How scores are calculated What is shipped and what is not --- ## Transparency URL: https://crawld.co/transparency/ TRANSPARENCY What is real, and what is not. A product built on refusing to round unmeasured up to a pass cannot be vague about its own state. So here is the ledger. Shipped Available now The free public scorecard, end to end: real crawl, real rubric, honest failure when the engine is down The 129-check catalogue with its eight weighted categories Per-check method declaration: rule, graph, third-party API, or model judgement Per-check fix capability: Auto, Assist, Human Coverage reporting on every score Competitor monitoring, weekly, with published politeness limits Core Web Vitals: LCP, INP and CLS Not shipped Not built yet Described on this site, in some cases in detail, because the design is real and worth explaining. Not available to you today, and every page that describes one says so on the page. The access ladder: modes A to E. Designed, and present in the codebase only as static demo reference data Every platform connector: WordPress, Webflow, Shopify, Ghost, HubSpot, Notion and the rest Pull-request delivery, CMS draft writing, and anything labelled build-verified in practice Outbound webhooks A documented authenticated API Actionable Insights in competitor monitoring Data export and downloadable reports FCP, TBT, Speed Index and an accessibility score A per-check browsable reference. One indexable page for each of the 129 Any public status or uptime reporting Things we will not claim Per-article citation lift. Citation share is recorded per prompt and per engine, with no link back to the specific article that earned it. Claims of the form "posts with X are cited 41% more" are not something this data supports, so you will not find one here. Build-verified on a hosted CMS. There is no build to run. That badge appears only where a repository is held or mirrored. A ranking outcome. Nobody can promise one. Certification. No SOC 2, no ISO 27001, no penetration-test report, no uptime guarantee. Where content goes Some checks are model-judged, and running them sends content from the pages you scan to a large-language-model provider. We do not train models on it and do not resell it, but it does leave our infrastructure, and that is worth stating on its own rather than in a sub-processor footnote. The privacy page covers it in full. A correction we had to make The public scan endpoint once wrote a finished result the moment it was called (a grade, a score, canned category rows) and returned "queued". Nothing was ever fetched, for any URL, and the fabricated result was publicly shareable. It now either runs on the real engine or reports honestly that scanning is unavailable. It is on this page because a company that claims honest measurement should publish the time it failed at exactly that, not just the policy that followed. Scan your site free Changelog Keep reading How rubric changes are versioned The scoring method in full Which connectors are built --- ## Careers URL: https://crawld.co/careers/ CAREERS No open roles listed. Rather than a page of aspirational culture copy with nothing to apply to, here is the actual position. Right now There are no published openings. If that changes, roles will be listed here with a real description and a real process rather than an open-ended talent-pool form. If you want to get in touch anyway The most useful thing you can send is not a CV. Run the free scan on a site you know well and tell us where the rubric is wrong: a check that fires when it should not, a finding whose evidence does not support it, a category weight you would argue with. That is genuinely useful to us and it demonstrates more than a covering letter does. Contact details are here. What the work is like The hard problems are not the crawling. They are deciding what a check may honestly assert, keeping model judgement clearly separated from deterministic fact, and resisting the pressure to let an unmeasured result quietly become a pass, which is the entire product thesis and also the thing that breaks first under deadline. Get in touch What we are building Keep reading What we are building How to get in touch The scoring thesis behind the work --- ## Contact URL: https://crawld.co/contact/ CONTACT Where to send it. No contact form. There is no endpoint behind one yet, and a form that quietly drops what you write is worse than an address. If it is… Send it to Worth including Something is broken support@crawld.co Include what you submitted, roughly when, and the requestId from the response if you have it. A security issue security@crawld.co Please report before disclosing publicly. See the responsible disclosure page. Our crawler is hitting your site crawler@crawld.co Or block it yourself in robots.txt, that works immediately and needs nothing from us. Sales, agency or volume pricing sales@crawld.co Worth reading the enterprise page first; several answers there are no. Press or anything else hello@crawld.co If you found us in your access logs You probably want the crawler page , which explains exactly what CrawldBot fetches, the limits it works within, and how to block it. Blocking works on the next scheduled run and needs no permission from us. What we cannot help with We cannot tell you why a specific page ranks where it does: nobody can, and anyone who says otherwise is selling something. The rubric measures things within your control; placement is not among them. Run a free scan Read the FAQ Keep reading How to tell an outage from a limit Crawler behaviour and how to block it Reporting a vulnerability --- ## Cookie policy URL: https://crawld.co/cookies/ COOKIE POLICY What this site stores in your browser. Short, because there is very little to describe. Working draft. Accurate about current behaviour; not reviewed by counsel. Re-verify before launch, and re-verify again the moment any analytics or advertising script is added, because the answer changes completely at that point. This marketing site rr-theme : localStorage, not a cookie. Remembers whether you chose the light or dark theme. Never leaves your browser. Cloudflare Turnstile : set only if you submit a scan, and only to distinguish a person from a script. It is a verification challenge, not a tracking mechanism. There is no analytics script on this site, no advertising or remarketing tags, and no third-party trackers. Fonts and brand icons are served from our own origin rather than a CDN, so loading a page here does not announce you to anyone else. Because nothing here is used for tracking or profiling, there is no consent banner. If that changes, a banner appears and this page changes with it. The application Session cookies : required to keep you signed in. Removing them signs you out. Preference storage : interface state such as your chosen theme. Clearing them Every browser lets you clear site data or block storage per site. Doing so costs you your theme preference and your session; nothing else is lost, because nothing else is kept there. Privacy Security Keep reading What we collect and keep The controls that are in place The terms in plain language --- ## Data processing agreement URL: https://crawld.co/dpa/ DATA PROCESSING Data processing agreement. The substance of the processing relationship, written plainly. The executable document does not exist yet, and this page says so rather than implying otherwise. No signable DPA exists today. Not built yet This page describes the processing accurately so you can assess it, but there is no counter-signed agreement to put in front of a procurement team. If your process requires one, that is a genuine blocker: see the enterprise page , which says the same thing. Roles For the site data you ask us to crawl, you are the controller and we are the processor. For your account and billing records, we are the controller. What is processed Public page content from the sites you submit: the text, markup and metadata a crawler retrieves. Derived findings : the checks, their outcomes and the evidence supporting them. Account data : the address you signed up with and your billing records. Abuse-prevention metadata : the requesting address, stored only as a truncated keyed hash. The scanner reads public HTML. It does not submit forms, does not authenticate, and does not reach anything a signed-out visitor could not. It is not designed to process personal data, though public pages sometimes contain it. Sub-processors Processor Purpose What it sees Amazon Web Services Hosting and object storage Crawl artifacts and uploaded files Polar Payments and subscriptions Billing identity and transaction records Resend Transactional email Recipient addresses and message contents Cloudflare Verification challenge and edge delivery Request metadata for abuse prevention A large-language-model provider Model-judged checks Content from the pages you scan The last row is the one that matters most for a data-classification review: running model-judged checks sends page content outside our infrastructure. We do not train models on it and contract for the same, but it does leave. Retention and deletion Free scan results and share links expire after 30 days. Account crawl data and findings are kept while the account is open and deleted with it. Repository or CMS access is limited to the scopes you grant; revoking access removes the associated data. Security measures TLS in transit, hashed verification tokens, keyed and truncated requester hashes, scoped access, and audit logging written out of band so the trail does not depend on the request succeeding. The security page lists the controls and, more usefully, what is absent: no SOC 2, no ISO 27001, no penetration-test report, no uptime guarantee. Your rights Export and deletion are available from the account. Statutory rights of access, correction, portability and erasure apply where the law grants them, and those are the mechanisms. Privacy Security Keep reading What we collect and keep Security controls, and what is absent Why enterprise procurement may stall --- ## Privacy URL: https://crawld.co/privacy/ PRIVACY What we collect, and what we do with it. Written from what the system actually does, so you can check it against the behaviour rather than against a template. This is a working draft, not a published policy. The factual descriptions below reflect how the system behaves. The legal framing has not been reviewed by counsel and this page must not be relied on as a privacy policy until it has been. The short version We store the crawl results and the fixes we generate for you. That is the product. We do not resell crawl data. We do not train models on your content. We do not keep repository or CMS access beyond the scopes you grant. Revoke access at any time and the associated data goes with it. The free scan specifically You can run a scan without an account, and the design tries hard not to accumulate anything about you in the process. The requesting address is never stored raw. It is stored as a keyed HMAC, truncated: enough to recognise a repeat caller, not enough to be a log of who visited. Results expire after 30 days, including any share link, and a share link never exposes the findings regardless of who opens it. An email address is only collected if you choose to unlock the findings, and the verification token is stored hashed. Content sent to model providers Some checks in the rubric are model-judged, and those are labelled Model everywhere they appear. To run them, content from the pages you ask us to scan is sent to a large-language-model provider. This is a real disclosure and worth stating plainly rather than burying in a sub-processor table, if you scan a page, the text of that page leaves our infrastructure. "We do not train models on your content" means we do not use it to train models and we contract for the same; it does not mean the content never goes anywhere. Scanning a site you do not control, which the free scorecard permits, sends that site's public HTML the same way. Sub-processors Processor Purpose What it sees Amazon Web Services Object storage and hosting Files you upload and generated artifacts Polar Payments and subscriptions Billing identity and transaction records Resend Transactional email Your address, and the contents of mail we send you Cloudflare CAPTCHA (Turnstile) and edge delivery Request metadata for abuse prevention A large-language-model provider Model-judged checks Page content from the sites you ask us to scan Fonts on this site are self-hosted, so loading a page here does not send your address to a font CDN. Brand icons are inlined at build time for the same reason. There is no analytics script on this site today. Retention Data Kept Why Free scan results 30 days A shared scorecard is a snapshot, not a permanent product. Requester IP for a free scan Stored only as a truncated HMAC Enough to block a repeat abuser, not a log of who visited. Email verification tokens Stored hashed, valid 24 hours The column is reachable from an unauthenticated surface; a raw token there would be the link itself. Account crawl data and findings While the account is open Deleted with the account. Repository and CMS access Only the scopes you grant Revoking access takes the associated data with it. Your rights You can export your data, and you can delete your account, which deletes the crawl data and findings associated with it. If you are in a jurisdiction with statutory rights of access, correction, portability or erasure, those apply and the mechanisms above are how they are exercised. Being crawled by us If you are here because CrawldBot appeared in your logs rather than because you are a customer, the crawler page explains exactly what it fetches, the limits it works within, and how to block it. Blocking works and takes effect on the next scheduled run. What our crawler does Security Keep reading The controls that are in place What this site stores in your browser The processing terms in substance --- ## Responsible disclosure URL: https://crawld.co/responsible-disclosure/ RESPONSIBLE DISCLOSURE Reporting a vulnerability. Report it to us before disclosing it publicly, and we will not pursue you for good-faith research. What we commit to We will not pursue legal action for good-faith research conducted within the boundaries below. We will acknowledge your report and tell you what we found, including if we disagree that it is a vulnerability. We will credit you if you want credit, and not if you do not. What we ask Test against your own account and your own domains. Not another customer's data, and not a third-party site through our scanner. Do not degrade the service. No load testing, no denial of service, no mass automated scanning. Do not exfiltrate data. Demonstrate access; do not collect. If you encounter someone else's data, stop and tell us. Give us time to fix it before publishing. What we are most interested in The unauthenticated scan surface is the part a stranger can reach, so it is where a flaw costs the most: The SSRF guard. Anything that gets our fetcher to reach an address it should refuse: private space, cloud metadata endpoints, redirect chains that escape validation. The verification path. Anything that submits a scan without solving the challenge, or that turns the endpoint into a probe for whether a hostname resolves internally. The email gate. Anything that returns findings for a scan whose address was never verified, or that exposes them through a share link. Token handling. Verification tokens are stored hashed; anything that recovers or replays one. Tenancy boundaries. Anything that reads one brand's data from another's session. Out of scope Missing security headers with no demonstrated impact. Rate limits being generous. They are deliberately not the primary control. Reports generated by an automated scanner with no verification or proof of exploitability. Social engineering of our people or our vendors. Bounties None offered There is no bug-bounty programme and no payment. Saying so plainly is fairer than letting a researcher assume otherwise and invoice afterwards. How to report Email security@crawld.co with steps to reproduce, what you were able to access, and what you did not do. Include a requestId from any response envelope if one is relevant. Security practices Usage policy Keep reading Security controls in place What you may point the scanner at Where to send the report --- ## Security URL: https://crawld.co/security/ SECURITY What is actually in place. A security page that lists aspirations is worse than none. These are controls the system implements; the section at the bottom is what it does not have. Working draft. The controls described are implemented. The page has not been through a formal review and should not be treated as a security attestation. Controls SSRF protection on the scan endpoint The public scanner takes a URL from a stranger and causes a fetch, which makes it a request-forgery surface by construction. Submitted hosts are resolved and checked before anything is fetched, and addresses in private space are rejected. The CAPTCHA check runs before validation specifically so the endpoint cannot be used as an unauthenticated probe for whether a hostname resolves internally: the rejection reasons would otherwise differ and leak that. Abuse controls on the free tier In the order they matter: Turnstile, which fails closed once enabled so an outage at Cloudflare is not a bypass; a per-domain cooldown, which survives IP rotation because a free scan attaches to a domain rather than a visitor; and a per-caller rate limit, which is deliberately not the main defence because VPNs defeat it and office NAT makes it punish real users. Credential and token handling Email verification tokens are stored hashed, not raw: the column is reachable from an unauthenticated surface, and a raw token there would be the link itself. Requester addresses are stored as truncated keyed HMACs rather than in the clear. Access scoping Repository and CMS access is limited to the scopes you grant and is not retained beyond them. Revoking access removes the associated data. Audit logging Requests are logged through middleware into a queue and written in batches, so the audit trail does not depend on the request path succeeding. Transport and isolation TLS in transit. Database access runs through a connection pooler with read-replica support; caches and queues are separated by function rather than sharing one instance. How the crawler behaves Relevant if you are assessing us as a third party pointing traffic at your infrastructure. Identifies itself as CrawldBot with a URL in the user agent Obeys robots.txt, including crawl-delay One request at a time: never parallel 12 pages per site on competitor monitoring, once a week Never submits forms and never attempts authenticated areas Full detail, including how to block it, is on the crawler page . Reporting a vulnerability Report it to us directly before disclosing it publicly, and we will not pursue you for good-faith research conducted against your own account and without degrading service for others. We are particularly interested in anything touching the unauthenticated scan surface (the SSRF guard, the CAPTCHA path, the share-link and email-gate logic) because that is the part of the system a stranger can reach. What we do not have Not in place Stated because omitting it is how security pages mislead: No SOC 2, ISO 27001 or equivalent certification is held or claimed. No published penetration-test report. No formal bug-bounty programme, and no bounty payments. No contractual uptime guarantee. If your procurement process requires any of those, that is a legitimate reason not to buy this yet, and it is better established now than during a security review. Privacy Crawler behaviour Keep reading How to report a vulnerability What we collect and keep Crawler behaviour and limits --- ## Sitemap URL: https://crawld.co/sitemap/ SITEMAP Every page on this site. 93 destinations, grouped. Every page here is one hop from any indexed page, which is the point of a sitemap that humans can read. Start here Home How it works How we measure Pricing FAQ Sitemap Comparisons Each ends by saying when the other tool is the better buy. All comparisons vs Surfer vs SEObot vs IndexRusher vs Semrush Platform Overview API documentation Queue & scan limits System requirements Status Crawler / robots.txt Company About Transparency Changelog Careers Contact Help centre Legal and policies Privacy policy Terms of service Security practices Cookie policy Data processing agreement Responsible disclosure Usage policy Blog by category Strategy (3) Technical (3) GEO basics (2) Link building (2) Analytics (1) Content strategy (1) On-page SEO (1) Product The audit The checks The fix queue Answer engines (GEO) Write SEO Watchdog Who it's for Download report The rubric Eight weighted categories, one page each. All eight categories Technical foundation Crawling & structure Search appearance Content quality E-E-A-T & trust Content health Answer-engine readiness Page experience Solutions One rubric; what changes is which findings dominate. All solutions Content brands Agencies and consultants Small technical teams SEO professionals Startups Enterprise E-commerce Publishers Nonprofits Higher education Integrations How delivery would work per platform. None of these connectors is built. All integrations REST API Next.js Webhooks Zapier WordPress Webflow Shopify Squarespace Ghost HubSpot Framer Notion Unicorn Platform Blog All posts Competitor SEO analysis without guessing at their data What an SEO audit costs and how often to run one Content cannibalization: how to find it and what to do On-page SEO: the elements worth checking on every page How to check website traffic, yours and anyone else Link building strategies that still work, and three that do not How to check your backlinks and audit what you find Generative engine optimization, explained without the hype How to run a technical SEO audit that finds real problems The SEO audit checklist that states its denominators Crawled, indexed, cited, where AI traffic begins llms.txt, explained for humans Why your weekend project is invisible to ChatGPT For machines Four artefacts, generated at build time from one page list by scripts/generate-seo.mjs , so they cannot disagree with each other or with this page. /sitemap.xml : every URL with its last-modified date, change frequency and priority. A single file; the 50,000-URL threshold that would force a sitemap index is a long way off. /robots.txt : crawl rules, pointing at the sitemap. /llms.txt : a curated index for language models, in the format described here . Curated rather than exhaustive; the value is in what it leaves out. /llms-full.txt : the readable text of every page in one file, so a model can ingest the site in a single fetch rather than crawling it. One route is excluded from all of them: /scorecard/[slug] , the shared-scorecard view. Its slugs are created at runtime, it renders on demand and it carries noindex , so the meta tag, robots.txt and the sitemap all say the same thing about it. Every carries a trailing slash, matching each page's own exactly. That is not cosmetic: this site is served directory-style, so a sitemap entry without the slash would put every canonical URL behind a redirect. Scan your site free Help centre Keep reading Where to find an answer All 129 checks The blog --- ## Status URL: https://crawld.co/status/ STATUS There is no status dashboard yet. A status page that is not driven by real monitoring is theatre. This one says what is actually true instead. No automated uptime reporting. Not built yet There is no public health endpoint, no incident history and no subscribe-to-updates. Publishing a green tick that nothing verifies would be worse than publishing nothing. How to tell whether scanning is working Submit a scan. The system is built so that failure is explicit rather than silent, if the engine is unavailable you get an error saying so, not a fabricated grade. That property is the actual status check. "The scanner is offline right now" : the engine is unavailable. Nothing was crawled and nothing was scored. "Taking longer than expected" : the scan was accepted but has not finished inside the polling window. It may still complete. A 429 . you are rate limited, which is working as designed rather than an outage. A result appears : the whole path is healthy, because a result cannot be produced without a crawl behind it. What is not an outage A repeated scan returning yesterday's result. That is the 24-hour domain cooldown. Lower coverage than last time. Usually the target site changed, blocked the crawler, or was slower to respond. An expired share link. Shared scorecards last 30 days by design. Reporting a problem Tell us what you submitted, roughly when, and what came back. If you have the requestId from a response envelope, include it. Every response carries one and it is the fastest way to find the request in our logs. Report an issue Queue and scan limits Keep reading Every cap that applies Where to report a problem What is shipped and what is not --- ## Terms URL: https://crawld.co/terms/ TERMS The terms, in plain language. The operational limits below are the ones the system enforces, so they are worth reading even if the legal sections are not yet final. Working draft, not a contract. The service limits described are accurate and enforced. The legal terms have not been reviewed by counsel and this page must not be relied on as terms of service until they have been. What the service does Crawld crawls sites you are entitled to have crawled, scores them against a published rubric, and, on paid plans and where you have granted access, proposes changes for your approval. It does not publish changes without your explicit approval at two separate points. Service limits Limit Value Free scan crawl depth 10 pages per scan Free scan frequency One real crawl per domain per day Free scan submissions 3 per hour per caller Shared scorecard lifetime 30 days Competitor monitoring Up to 3 sites, 12 pages each, weekly What we ask of you Only submit sites you are entitled to have crawled. The scanner causes real requests to real infrastructure. Pointing it at a third party you have no relationship with is a misuse of it, and the per-domain cooldown exists partly to limit the damage that misuse could do. Do not attempt to circumvent the abuse controls. Rotating addresses to bypass the domain cooldown, or automating around the CAPTCHA, is grounds for a block. Review what you approve. The gates are not decorative. A change you approve is a change you have chosen to make to your own site. What we do not promise Being specific here is more useful than a disclaimer paragraph: No ranking outcome is promised. Nobody can promise one. The rubric measures things that are within your control; where you rank is not among them. A score is a claim about what was measured. Coverage is reported next to every score precisely so that the claim is bounded rather than implied to be total. Advisory means advisory. A finding delivered without repository access has not been verified against a build, and the label says so. Applying it is your decision. Availability is not contractually guaranteed. If the scanning engine is unavailable, a scan fails and reports that rather than returning a fabricated result. Payment and cancellation Plans are billed per site rather than per seat. Cancelling stops future billing and ends ongoing scans, monitoring and queue access. Changes you already merged are yours. They live in your repository or your CMS, not in ours, and cancelling does not reach into them. Your content You keep ownership of your content and of the changes generated for your site. We use crawl data to operate the service for you. We do not resell it and we do not train models on it. Content is sent to a large-language-model provider in order to run model-judged checks: see Privacy , which states that plainly rather than in a sub-processor footnote. Ending the relationship You can delete your account, which deletes the crawl data and findings held against it, and you can revoke repository or CMS access independently at any time. Revocation takes the associated data with it. Privacy Security Keep reading What we collect and why The controls that are actually in place What you may point the scanner at --- ## Usage policy URL: https://crawld.co/usage-policy/ USAGE POLICY What you may point the scanner at. This tool causes real HTTP requests to real infrastructure that may not be yours. That is the whole reason this page exists. Working draft. Behaviourally accurate; not reviewed by counsel. The one rule that matters Only submit sites you are entitled to have crawled. Your own, your client's, or a competitor you are monitoring within the published limits. Submitting a URL here makes us fetch it, from our addresses, on your behalf. Not permitted Using the scanner as a traffic source. Pointing repeated scans at a third party to generate load. The per-domain cooldown exists partly to bound this, and circumventing it is a ban. Circumventing the abuse controls. Rotating addresses past the cooldown, or automating around the verification challenge. Scanning to find vulnerabilities in someone else's site. This is a content and structure rubric, not a security scanner, and using it as reconnaissance is a misuse. Reselling raw scan output as your own product without adding anything. Agency use, running client sites and reporting on them, is expected and fine. Submitting URLs that resolve into private address space. The SSRF guard rejects these; attempting it repeatedly is treated as an attack. Competitor monitoring, specifically Monitoring a competitor is legitimate and is a product feature. It is bounded to three sites per brand, twelve pages each, one request at a time, once a week, with robots.txt obeyed. Those limits are not configurable, and a site that disallows CrawldBot is not crawled at all. If you are on the receiving end Disallow CrawldBot in your robots.txt and monitoring stops on the next scheduled run. You need no permission and no account. The crawler page has the exact syntax. Enforcement Abuse is met with a block. Requester addresses are stored as truncated one-way hashes: enough to recognise a repeat offender, not enough to be a log of who visited. Terms Crawler behaviour Keep reading The terms in plain language Crawler limits and politeness Queue and scan limits