SEO Log File Analyzer

Upload your server access logs and see how search engine and AI crawlers really move through your site: which bots come, how often, where the crawl budget goes.

  • Daily limit 0/3
  • Monthly limit 0/5
Free preview: only the first 1,000 lines were analysed. Sign in with a Pro plan to analyse the whole file. See pricing

Upload or paste server access logs (Apache/Nginx combined format) to see how search-engine bots crawl your site. Parsing happens on the server and nothing is stored.

Your crawler tells you what is on your site. Your server logs tell you what Google actually does with it. SEO Log File Analyzer reads your access logs and shows which bots visit, how often and which URLs they hit. It also shows how much crawl budget disappears into errors. Analyse your logs now: upload a file or paste a few thousand lines, and nothing is stored anywhere.

This is the one dataset in SEO that cannot be guessed at. Rankings, impressions and crawl stats are all reported to you second hand. Access logs are the raw record of every request your server actually answered, written by your own machine.

What the logs tell you

Upload the access log and the tool separates real crawler traffic from everything else. It then breaks that down into figures you can act on.

  • Which crawlers came. Every known bot with its hit count, unique URLs and the mix of status codes it received.
  • How often they came. Crawl frequency by day, so a drop or a spike is visible rather than inferred.
  • What they spent time on. Your most crawled URLs, which is often not the list you would have guessed.
  • Where crawling went astray. Redirects, 404s, rate limits and server errors, each as a share of requests with the URLs behind them.

AI crawlers get their own view

GPTBot, ClaudeBot and PerplexityBot now account for a real share of automated traffic, and they behave nothing like Googlebot. They are listed separately with hit counts and the date they were last seen.

AI crawler activity shows whether known AI-related crawlers are discovering and fetching your content. Where the operator documents it, each crawler is labelled as a training crawler, an AI-search crawler or a user-triggered fetcher. A crawler visit does not guarantee inclusion or citation in AI-generated answers.

Crawl efficiency and where crawling goes

Crawl efficiency is only worth discussing with evidence, and the log is that evidence. Every response is classed by what it means. Successful, not modified (304 is not waste), redirects, 404 and 410, rate limiting, server errors. Each class is a share of the bot's requests, with the URLs behind it.

On a large site this is usually where the surprise lives. URL parameters, listing and archive patterns and old redirects can absorb a large share of crawling. The tool groups them by pattern and shows the share, the request count and the number of unique URLs. It marks them as potential inefficiencies, not as confirmed problems: that judgement stays with you.

Compare the logs with a crawl

If you also use DiagnoSEO Website Audit, pick a crawled project. The analyser then lines the two datasets up against each other. Two lists come out of that comparison, and both are useful.

URLs the bot requested that the crawl did not discover are potential orphan URLs, and only that. The tool sorts them into likely categories, such as parameter URLs or old redirects, and leaves the conclusion to you. URLs found by the crawl but not requested are reported as not seen by that bot in this log period. A short log window says nothing about whether a page is ignored.

The feature is optional. Without Website Audit installed the control simply does not appear and log analysis works exactly the same.

Where to find your access logs

Most hosting panels expose them under a name like «raw access logs» or «access_log». On a server you manage, Apache writes to /var/log/apache2/access.log and Nginx to /var/log/nginx/access.log by default.

The format needs to be the standard combined log format, which is what both ship with. Cloudflare, a CDN or a load balancer in front of your origin changes what the origin sees. Pull the logs from whichever layer actually answers the crawler.

What happens to the file

Parsing runs on the server in memory. The log itself is not written to disk or to a database. No project is saved and there is no history. The one thing kept is a short-lived cache of crawler IP verification results. It holds no URLs, user agents or request data.

Files above 12 MB are truncated to the first 12 MB. On a mid-sized site that is usually several days of traffic. If your upload is rejected outright, the limit is your server's post_max_size rather than anything in the tool.

Log File Analyzer compared with other tools

Log analysis is usually either a spreadsheet exercise or a module buried inside an expensive platform. This one is a single page that reads a file and gives you the tables. «Depends» means the capability exists in some tools or in higher tiers.

CapabilityDiagnoSEO Log File AnalyzerOther tools
Upload or paste, results on the same page✅⚠️ depends
Nothing stored, no project to create✅⚠️ depends
Separate view for AI crawlers (GPTBot, ClaudeBot, PerplexityBot)✅❌
Wasted crawl budget ranked by URL and status✅✅
Crawl frequency by day✅✅
Cross-reference against your own crawl✅⚠️ depends
Orphan detection from the comparison✅⚠️ depends
Interface in 33 languages✅❌

Frequently asked questions

  • The standard combined format that Apache and Nginx write by default. Each line needs an IP, a timestamp, the request, a status code and a user agent. Lines it cannot parse are skipped and counted, so you can see whether the file was understood.

  • No. Parsing happens on the server in memory and the result is returned to your browser. Nothing is written to disk or to a database, and there is no history because the tool keeps no state.

  • Up to 12 MB per run, and up to 300 000 lines. Larger files are truncated rather than rejected. If the upload fails before it reaches the tool, your server's post_max_size is the limit to raise.

  • Bots are identified by user agent, which anyone can spoof. Treat the numbers as a strong signal rather than proof. If a single bot shows implausible volume, verify a sample of those IPs with a reverse DNS lookup before acting on it.

  • Usually the logs come from the wrong layer. If Cloudflare or a CDN sits in front of your origin, it answers most crawler requests itself and your origin never sees them. Pull the logs from the edge instead.

  • No. Log analysis is complete on its own. Website Audit only adds the cross-reference, which compares the logs against a crawl and lists potential orphan URLs and URLs not seen in the log period. Without it the control is hidden.

  • Log analysis is available from the Pro plan upwards. Open the tool to see the current state of your account.

  • Monthly is enough for a stable site. Look sooner after a migration, a redesign or a sudden traffic change, because the log shows what the crawler did about it before the rankings catch up.

Rankings tell you the outcome. Logs tell you the behaviour that produced it. Upload a log file and see what the crawlers have been doing.

Unlock Higher Rankings and Quality Traffic

Grow your business with the #1 AI-powered full stack software for SEO and content marketing.

Upgrade to Pro