-
Home
-
User Agent Directory
- YandexPagechecker
YandexPagechecker
Yandex •
Web Crawler
What is YandexPagechecker?
Yandex lists YandexPagechecker as the robot that accesses a page when it is checked with the Structured Data Validator. It is separate from the main Yandex indexing bot. Yandex says this agent does take general robots.txt rules into account, so blocked pages may not be available for validation.
User Agent String
Mozilla/5.0 (compatible; YandexPagechecker/1.0; +http://yandex.com/bots)
How to Control YandexPagechecker
Block Completely
To prevent YandexPagechecker from accessing your entire website, add this to your robots.txt file:
# Block YandexPagechecker
User-agent: YandexPagechecker
Disallow: /
Block Specific Directories
To restrict access to certain parts of your site while allowing others:
User-agent: YandexPagechecker
Disallow: /admin/
Disallow: /private/
Disallow: /wp-admin/
Allow: /public/
Set Crawl Delay
To slow down the crawl rate (note: not all bots respect this directive):
User-agent: YandexPagechecker
Crawl-delay: 10
How to Verify YandexPagechecker
Verification Method:
Use Yandex's documented forward-confirmed reverse-DNS check for the source IP; the user-agent alone does not authenticate the request.
Use Yandex's documented forward-confirmed reverse-DNS check for the source IP; the user-agent alone does not authenticate the request.
Learn more in the official documentation.
Detection Patterns
Multiple ways to detect YandexPagechecker in your application:
Basic Pattern
/YandexPagechecker/i
Strict Pattern
/^Mozilla\/5\.0 \(compatible; YandexPagechecker\/1\.0; \+http\:\/\/yandex\.com\/bots\)$/
Flexible Pattern
/YandexPagechecker[\s\/]?[\d\.]*?/i
Vendor Match
/.*Yandex.*YandexPagechecker/i
Implementation Examples
// PHP Detection for YandexPagechecker
function detect_yandexpagechecker() {
$user_agent = $_SERVER['HTTP_USER_AGENT'] ?? '';
$pattern = '/YandexPagechecker/i';
if (preg_match($pattern, $user_agent)) {
// Log the detection
error_log('YandexPagechecker detected from IP: ' . $_SERVER['REMOTE_ADDR']);
// Set cache headers
header('Cache-Control: public, max-age=3600');
header('X-Robots-Tag: noarchive');
// Optional: Serve cached version
if (file_exists('cache/' . md5($_SERVER['REQUEST_URI']) . '.html')) {
readfile('cache/' . md5($_SERVER['REQUEST_URI']) . '.html');
exit;
}
return true;
}
return false;
}
# Python/Flask Detection for YandexPagechecker
import re
from flask import request, make_response
def detect_yandexpagechecker():
user_agent = request.headers.get('User-Agent', '')
pattern = r'YandexPagechecker'
if re.search(pattern, user_agent, re.IGNORECASE):
# Create response with caching
response = make_response()
response.headers['Cache-Control'] = 'public, max-age=3600'
response.headers['X-Robots-Tag'] = 'noarchive'
return True
return False
# Django Middleware
class yandexpagecheckerMiddleware:
def __init__(self, get_response):
self.get_response = get_response
def __call__(self, request):
if self.detect_bot(request):
# Handle bot traffic
pass
return self.get_response(request)
// JavaScript/Node.js Detection for YandexPagechecker
const express = require('express');
const app = express();
// Middleware to detect YandexPagechecker
function detect_yandexpagechecker(req, res, next) {
const userAgent = req.headers['user-agent'] || '';
const pattern = /YandexPagechecker/i;
if (pattern.test(userAgent)) {
// Log bot detection
console.log('YandexPagechecker detected from IP:', req.ip);
// Set cache headers
res.set({
'Cache-Control': 'public, max-age=3600',
'X-Robots-Tag': 'noarchive'
});
// Mark request as bot
req.isBot = true;
req.botName = 'YandexPagechecker';
}
next();
}
app.use(detect_yandexpagechecker);
# Apache .htaccess rules for YandexPagechecker
# Block completely
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} YandexPagechecker [NC]
RewriteRule .* - [F,L]
# Or redirect to a static version
RewriteCond %{HTTP_USER_AGENT} YandexPagechecker [NC]
RewriteCond %{REQUEST_URI} !^/static/
RewriteRule ^(.*)$ /static/$1 [L]
# Or set environment variable for PHP
SetEnvIfNoCase User-Agent "YandexPagechecker" is_bot=1
# Add cache headers for this bot
<If "%{HTTP_USER_AGENT} =~ /YandexPagechecker/i">
Header set Cache-Control "public, max-age=3600"
Header set X-Robots-Tag "noarchive"
</If>
# Nginx configuration for YandexPagechecker
# Map user agent to variable
map $http_user_agent $is_yandexpagechecker {
default 0;
~*YandexPagechecker 1;
}
server {
# Block the bot completely
if ($is_yandexpagechecker) {
return 403;
}
# Or serve cached content
location / {
if ($is_yandexpagechecker) {
root /var/www/cached;
try_files $uri $uri.html $uri/index.html @backend;
}
try_files $uri @backend;
}
# Add headers for bot requests
location @backend {
if ($is_yandexpagechecker) {
add_header Cache-Control "public, max-age=3600";
add_header X-Robots-Tag "noarchive";
}
proxy_pass http://backend;
}
}
Should You Block This Bot?
Recommendations based on your website type:
| Site Type | Recommendation | Reasoning |
|---|---|---|
| E-commerce | Optional | Evaluate based on bandwidth usage vs. benefits |
| Blog/News | Allow | Increases content reach and discoverability |
| SaaS Application | Block | No benefit for application interfaces; preserve resources |
| Documentation | Selective | Allow for public docs, block for internal docs |
| Corporate Site | Limit | Allow for public pages, block sensitive areas like intranets |
Advanced robots.txt Configurations
E-commerce Site Configuration
User-agent: YandexPagechecker
Crawl-delay: 5
Disallow: /cart/
Disallow: /checkout/
Disallow: /my-account/
Disallow: /api/
Disallow: /*?sort=
Disallow: /*?filter=
Disallow: /*&page=
Allow: /products/
Allow: /categories/
Sitemap: https://example.com/sitemap.xml
Publishing/Blog Configuration
User-agent: YandexPagechecker
Crawl-delay: 10
Disallow: /wp-admin/
Disallow: /drafts/
Disallow: /preview/
Disallow: /*?replytocom=
Allow: /
SaaS/Application Configuration
User-agent: YandexPagechecker
Disallow: /app/
Disallow: /api/
Disallow: /dashboard/
Disallow: /settings/
Allow: /
Allow: /pricing/
Allow: /features/
Allow: /docs/
Quick Reference
User Agent Match
YandexPagechecker
Robots.txt Name
YandexPagechecker
Category
seo
Respects robots.txt
Yes
Copied to clipboard!