Every other SEO tool gives you a report about crawling. Server logs are the crawling itself — every request your server actually received, with a timestamp, a URL and a response code. It's the only unmediated evidence available, and it's the only place a growing category of crawler is visible at all.
Start here: do you actually need this?
Unusual placement for a caveat, but it saves people considerable time and most guides on this topic bury it or omit it.
Crawl budget is not a practical constraint for small sites. If you run a few hundred pages that are being indexed normally, log analysis will probably confirm that everything is fine, which is a real answer but an expensive one. The effort is better spent elsewhere.
Log analysis earns its cost in four situations:
- Large sites — tens of thousands of URLs and up, where crawl allocation genuinely matters.
- E-commerce with faceted navigation, where filter combinations can generate effectively unlimited URLs.
- After a migration, to verify that redirects are being followed and new URLs discovered.
- Diagnosing a specific problem — pages not indexing, unexplained traffic loss, suspected server issues.
Outside those, treat this article as background rather than a task list.
Verify the crawler before you analyse anything
The step that's skipped most often and invalidates everything downstream when it is.
User agent strings are trivially spoofable. Anything can claim to be Googlebot, and plenty does — scrapers, competitors' tools, and assorted automated traffic that finds the disguise convenient. A meaningful share of requests in raw logs identifying as major crawlers are not those crawlers.
Which means an unverified analysis may be a detailed study of scraper behaviour, presented as insight about search engines. That's worse than no analysis, because it produces confident wrong conclusions about where crawl effort is going.
The reliable method is a reverse DNS lookup on the requesting IP, confirming it resolves to a legitimate crawler hostname, then a forward lookup on that hostname to confirm it maps back to the same IP. Major search engines also publish IP range files you can match against, which is faster at scale.
The rule The user agent is a claim, not evidence. Verify by IP or you're analysing whoever chose to impersonate a search engine that week.
Getting the logs is often the hard part
The practical barrier that stops most log analysis projects, and it's worth knowing before you promise anyone a deliverable.
If you're behind a CDN, origin logs are incomplete. Cached requests are served at the edge and may never reach your origin server, so origin logs will systematically under-represent crawler activity. You need the CDN's logs, which are usually available but often on a different plan tier or requiring configuration to enable.
Multiple servers mean multiple log sets that need consolidating, and load balancers can obscure the original requesting IP unless forwarding headers are configured and logged.
Retention is frequently short. Many default configurations keep days rather than months. If you want to analyse a period, you often need to have decided in advance — which is a good argument for turning on longer retention now, before you need it.
Access usually requires someone else. Logs live with infrastructure, not marketing, so plan for a request rather than a download.
Rankmath
Most popular WordPress SEO plugin with powerful on-page optimization features built in
Best for: WordPress SEO Plugin
What logs answer that nothing else can
Search Console is a report on outcomes. Crawl simulators tell you what a crawler could find. Logs tell you what actually happened.
| Question | Why other tools can't |
|---|---|
| How often is each specific URL crawled? | Search Console reports aggregate patterns, not per-URL frequency over time |
| Where is crawl effort concentrated? | Requires the full request set, not a sample |
| Are redirect chains being hit, and how long? | Simulators show the chain; logs show whether crawlers actually traverse it |
| Are there intermittent server errors? | Sampled reporting misses transient 5xx spikes entirely |
| Are orphan pages being crawled? | A crawler following your links can't find pages nothing links to |
| Which AI systems are fetching your content? | Nothing else reports this at all |
That last row is the newly important one.
The AI crawler layer
The strongest current reason to look at logs, and the reason this old technique became relevant again.
Search Console covers Google Search. It says nothing about the systems fetching your content for AI answers, assistant responses and agent workflows. Those requests hit your server like any other, with their own user agents — and your logs are the only place that activity is visible.
What you can establish from them: which AI systems fetch you and how often, whether their volume is rising or falling, which sections of the site they concentrate on, and whether they're getting successful responses or hitting errors and blocks.
Two practical uses. First, if you've made decisions about which crawlers to allow — via robots.txt or bot management at your CDN — the logs tell you whether those decisions are actually taking effect, which is frequently not what people assume. Second, this is behavioural evidence rather than vendor claims, which is exactly the standard applied in assessing whether llms.txt does anything — what crawlers request beats what anyone says they read.
Worth checking your CDN's bot management settings alongside this. Default configurations can block or challenge crawlers you'd rather allow, and the block is invisible unless you look — a site can be inadvertently excluding itself from the retrieval layer described in generative engine optimisation without anyone deciding to.
What you'll usually find
Findings cluster in predictable places, which makes the first pass faster than it sounds.
Crawl effort concentrated somewhere unproductive. The headline finding on most large sites. Parameter URLs, faceted navigation combinations, internal search results, endless pagination, calendar pages. Requests spent there aren't reaching your commercial pages — an issue particularly common on e-commerce sites, where the interaction between filtering and product page structure generates URL combinations nobody intended to publish.
Important pages crawled far less than assumed. Compare crawl frequency against commercial importance. The mismatch is often startling, and it usually traces to internal linking — pages buried deep get visited rarely, which is one of the more concrete arguments for the structure work in building a site structure that ranks.
Redirect chains longer than anyone believed. Chains accumulate silently across site changes, and logs show whether crawlers are actually traversing them — the maintenance issue covered in redirects and migrations.
Intermittent errors. A 5xx spike at 3am on Tuesdays, invisible in any sampled report, that has been quietly costing you crawl access for months.
Crawling of things that should be gone. Sections you blocked, pages you deleted, a staging path that was never restricted. Crawlers keep trying long after you've forgotten.
Orphan pages receiving crawl requests. Pages nothing links to but crawlers still find — usually from old links, sitemaps, or external references. A crawl simulator can't find these because it follows links; only logs reveal them.
A first pass, in order
- Verify the crawlers. Filter to confirmed bots by IP. Everything after depends on this.
- Response code distribution by crawler. What proportion of requests get 200s versus 3xx, 4xx and 5xx? A high redirect share means effort is being spent traversing rather than reading.
- Top requested URLs. Sort by request count. Then ask whether those are the pages that matter commercially. Frequently they aren't.
- Requests by directory or section. Where is the effort going at a structural level?
- Parameter analysis. What share of requests include query strings, and are those URLs worth crawling?
- Compare against your URL inventory. Which important pages appear rarely or never? Which crawled URLs aren't in your sitemap?
- Non-search crawlers. Who else is fetching you, how much, and is that what you intended?
Steps two and three answer most questions on their own. If crawlers are getting 200s on your commercial pages at reasonable frequency, your crawl situation is fine and the problem is elsewhere — which is itself a useful, time-saving conclusion.
Tooling, briefly
You don't need a specialist platform for a first look.
Spreadsheets handle a week of logs from a modest site perfectly well once parsed, though they struggle beyond a few hundred thousand rows.
Command line tools — grep, awk, sort, uniq — process large files fast and are the practical choice for anyone comfortable with a terminal. Filtering to a verified crawler and counting requests per URL is a couple of commands.
Dedicated log analysis platforms handle verification, visualisation and cross-referencing against crawl data automatically, and are worth it for ongoing monitoring on large sites. Note that most content on this topic is published by these vendors, which is fine but does shape the framing toward "everyone needs continuous log monitoring."
The honest sequence: do one manual pass first. If it finds nothing actionable, you've answered the question cheaply. If it finds a lot, you now have a business case for tooling.
What log analysis won't tell you
Worth setting expectations, since the technique attracts a certain mystique.
- Why a page ranks or doesn't. Crawling is necessary, not sufficient. A well-crawled page can rank poorly for entirely unrelated reasons.
- Whether content was indexed. Logs show the request, not the outcome. Search Console covers that side.
- What the crawler made of the page. A 200 response tells you the file was served, not that the content was parsed usefully — which matters given that many crawlers don't execute JavaScript.
- Anything about rankings directly. Crawl frequency correlates loosely with importance and is not a ranking signal you can manipulate.
That third point deserves emphasis. If your content is assembled client-side, crawlers can receive a perfectly successful 200 response containing almost nothing useful. Logs will look healthy while the content is effectively invisible — which is why they're a complement to the checks in a technical SEO audit rather than a replacement for them.
Server response time in logs is also worth a glance, since slow responses reduce how much a crawler will fetch — the crawl-side consequence of the work in making a site faster.
Acting on what you find
Findings map to a small number of fixes.
Crawl waste on parameters and facets: handle via robots.txt disallow, canonical tags, or removing internal links to those combinations. Be careful — blocking URLs that carry links or rankings causes different problems.
Important pages under-crawled: improve internal linking and reduce click depth. This is the most reliable lever available.
Redirect chains: collapse to single hops and update internal links to point at final destinations directly.
Server errors: an engineering conversation, and one where log evidence is unusually persuasive because it's specific and timestamped.
Unwanted crawler activity: robots.txt for cooperative crawlers, bot management for the rest. Decide deliberately which AI systems you allow rather than inheriting a default.
If the practical obstacle is that nobody can get you the logs, or that the first pass needs someone who's read a few before, that's a specialist task rather than a strategic one — and it's where an SEO partner with technical depth resolves in an afternoon what otherwise sits in a backlog for a quarter.
The short version
Logs are the only unmediated record of what crawlers actually did, and they're now the only place AI crawler activity is visible at all — which is the strongest current reason to look. Verify crawlers by reverse DNS or published IP ranges before analysing anything, because user agents are spoofable and an unverified analysis may be describing scrapers. Expect the acquisition to be the hard part, particularly behind a CDN where origin logs are incomplete. Most findings cluster in one place: crawl effort concentrated on parameters, facets and pagination rather than commercial pages. And be honest about whether you need this — for a few hundred well-indexed pages, the answer is usually no.
Pages not getting indexed and nobody can say why?
We read the logs, find where crawl effort is actually going, and fix what's blocking you.
Explore Web Development →Frequently asked questions
What can log files tell you that Search Console cannot?
Logs record every request your server actually received, which makes them the only unmediated evidence of crawler behaviour rather than a summarised report about it. They show which specific URLs were requested and how often, the response code returned each time, the exact sequence of redirect hops, and crucially the activity of crawlers that Search Console does not cover at all — including AI systems fetching your content. Search Console tells you the outcome of crawling; logs tell you the crawling itself.
How do you verify that traffic claiming to be Googlebot really is?
Never trust the user agent string, because it is trivially spoofable and a meaningful share of requests identifying as Googlebot are not. The reliable method is a reverse DNS lookup on the requesting IP address, checking that it resolves to a legitimate crawler hostname, followed by a forward lookup on that hostname to confirm it maps back to the same IP. Search engines also publish IP range files that can be matched against. Skipping this step means your analysis may be describing scrapers rather than search engines.
Does every website need log file analysis?
No, and this is worth saying plainly because most guides imply otherwise. Crawl budget is not a practical constraint for small sites, so a few hundred pages that are being indexed normally rarely justify the effort. Log analysis earns its cost on large sites, on e-commerce with faceted navigation generating many URL combinations, after a migration, and whenever you are diagnosing a specific problem such as pages not being indexed. Outside those situations the time is usually better spent elsewhere.
Why are AI crawlers a reason to look at logs now?
Because logs are the only place their activity is visible. Search Console reports on Google Search crawling and says nothing about systems fetching your content for AI answers or agent workflows. Server logs record every one of those requests with a timestamp, a user agent and a URL, which lets you see which systems are fetching you, how frequently, and which sections they take. That is behavioural evidence rather than vendor claims, and it is currently the most reliable way to understand your exposure to that layer.
What are the most common findings in an SEO log file analysis?
Crawl effort concentrated somewhere unproductive is the usual headline — parameter URLs, faceted navigation combinations, internal search results, pagination, or old redirect chains absorbing requests that should be reaching commercial pages. Other frequent findings include important pages crawled far less often than assumed, intermittent server errors that never appear in sampled reporting, redirect chains longer than expected, and crawlers spending significant time on sections that were supposed to have been blocked or removed.