NAVINES SEO LAB

Server log analysis for SEO: verify bots and find crawl problems

Turn server and CDN logs into a crawl diagnosis: verify Googlebot, group status codes, compare important URLs and document what logs cannot prove.

Direct answer

Logs can show which verified agents requested which URLs and how the server responded; they cannot show every crawl decision, rendered result or index state.

Collect the evidence at the layer that served the request

Start with a dated, access-controlled sample from the CDN and origin. A cached request may never reach the origin, so an origin-only export can undercount successful delivery. Document the time zone, retention, sampling and cache behavior before comparing days.

  • Useful fields: timestamp, normalized path, status, method, response time, cache outcome and verified bot category.
  • Exclude cookies, authorization headers and sensitive query values from the working dataset.
  • Keep raw client IPs only where required for bot verification under your retention policy; use aggregated categories in the analysis report.

Verify Googlebot before interpreting crawl activity

A user-agent string is self-declared. Verify requests using Google's published IP ranges or its documented reverse-DNS and forward-DNS procedure. For DNS verification, the forward lookup must return the original address; a suggestive hostname alone is insufficient.

Keep verified Googlebot, other verified Google fetchers, unverified claims and ordinary visitors in separate groups. Search crawling and a user-triggered fetch answer different questions. An unexpected request should not become an SEO conclusion until its identity and context are established.

Prioritize requests by page importance and outcome

PatternInvestigationRetest
Important pages repeatedly return 5xxCheck deployment and origin errors in the same time windowRequest the same paths after recovery and inspect new logs.
Many requests follow redirect chainsFind internal links still using old destinationsVerify direct links and one-hop genuine moves.
Many parameter variants are fetchedInspect filters and internal URL generationSample the revised link graph and subsequent requests.
A useful page has no observed requestsCheck sampling, sitemap membership and incoming linksCombine inventories and inspect the URL; do not assume exclusion.

Worked example and report boundaries

Illustrative incident: product pages appear healthy when opened manually, but the CDN sample records short bursts of 503 responses around each deployment. Group requests by minute and template rather than averaging the entire week. Compare the burst with deployment events and retest the same URLs after a release fix.

The report should state the affected interval, verified bot category, request count, response distribution and proposed acceptance test. A request log cannot show Google's final canonical selection, ranking or every rendering step. Use URL Inspection and search performance as separate evidence.

Common mistakes

  • Trusting user-agent strings alone.
  • Logging sensitive query values.
  • Equating requests with indexing.

Limitations

  • CDNs and sampling can omit events.
  • Bot verification and retention add operational cost.

Primary sources

Related guides and tools

WhatsApp