Cookie Consent by Free Privacy Policy Generator IbouBot User Agent - Ibou Bot Details | CL SEO

IbouBot

Ibou • Web Crawler
Search Respects robots.txt
#search engine #indexing #france

What is IbouBot?

Ibou says its crawler discovers public pages and indexes them for its search engine, with links back to sources. The operator says it does not use the crawled data to train AI models. IbouBot follows robots.txt and Crawl-delay, and the company publishes reverse-DNS and IP-range checks for verifying traffic.

User Agent String

Mozilla/5.0 (compatible; IbouBot/1.0; [email protected]; +https://ibou.io/iboubot.html)

How to Control IbouBot

Block Completely

To prevent IbouBot from accessing your entire website, add this to your robots.txt file:

# Block IbouBot User-agent: IbouBot Disallow: /

Block Specific Directories

To restrict access to certain parts of your site while allowing others:

User-agent: IbouBot Disallow: /admin/ Disallow: /private/ Disallow: /wp-admin/ Allow: /public/

Set Crawl Delay

To slow down the crawl rate (note: not all bots respect this directive):

User-agent: IbouBot Crawl-delay: 10

How to Verify IbouBot

Verification Method:
Forward-confirm the source IP's reverse DNS under ibou.io, or compare it with Ibou's current published IP ranges; a user-agent match is insufficient.

Learn more in the official documentation.

Detection Patterns

Multiple ways to detect IbouBot in your application:

Basic Pattern

/IbouBot/i

Strict Pattern

/^Mozilla\/5\.0 \(compatible; IbouBot\/1\.0; \+bot@ibou\.io; \+https\:\/\/ibou\.io\/iboubot\.html\)$/

Flexible Pattern

/IbouBot[\s\/]?[\d\.]*?/i

Vendor Match

/.*Ibou.*IbouBot/i

Implementation Examples

// PHP Detection for IbouBot function detect_iboubot() { $user_agent = $_SERVER['HTTP_USER_AGENT'] ?? ''; $pattern = '/IbouBot/i'; if (preg_match($pattern, $user_agent)) { // Log the detection error_log('IbouBot detected from IP: ' . $_SERVER['REMOTE_ADDR']); // Set cache headers header('Cache-Control: public, max-age=3600'); header('X-Robots-Tag: noarchive'); // Optional: Serve cached version if (file_exists('cache/' . md5($_SERVER['REQUEST_URI']) . '.html')) { readfile('cache/' . md5($_SERVER['REQUEST_URI']) . '.html'); exit; } return true; } return false; }
# Python/Flask Detection for IbouBot import re from flask import request, make_response def detect_iboubot(): user_agent = request.headers.get('User-Agent', '') pattern = r'IbouBot' if re.search(pattern, user_agent, re.IGNORECASE): # Create response with caching response = make_response() response.headers['Cache-Control'] = 'public, max-age=3600' response.headers['X-Robots-Tag'] = 'noarchive' return True return False # Django Middleware class iboubotMiddleware: def __init__(self, get_response): self.get_response = get_response def __call__(self, request): if self.detect_bot(request): # Handle bot traffic pass return self.get_response(request)
// JavaScript/Node.js Detection for IbouBot const express = require('express'); const app = express(); // Middleware to detect IbouBot function detect_iboubot(req, res, next) { const userAgent = req.headers['user-agent'] || ''; const pattern = /IbouBot/i; if (pattern.test(userAgent)) { // Log bot detection console.log('IbouBot detected from IP:', req.ip); // Set cache headers res.set({ 'Cache-Control': 'public, max-age=3600', 'X-Robots-Tag': 'noarchive' }); // Mark request as bot req.isBot = true; req.botName = 'IbouBot'; } next(); } app.use(detect_iboubot);
# Apache .htaccess rules for IbouBot # Block completely RewriteEngine On RewriteCond %{HTTP_USER_AGENT} IbouBot [NC] RewriteRule .* - [F,L] # Or redirect to a static version RewriteCond %{HTTP_USER_AGENT} IbouBot [NC] RewriteCond %{REQUEST_URI} !^/static/ RewriteRule ^(.*)$ /static/$1 [L] # Or set environment variable for PHP SetEnvIfNoCase User-Agent "IbouBot" is_bot=1 # Add cache headers for this bot <If "%{HTTP_USER_AGENT} =~ /IbouBot/i"> Header set Cache-Control "public, max-age=3600" Header set X-Robots-Tag "noarchive" </If>
# Nginx configuration for IbouBot # Map user agent to variable map $http_user_agent $is_iboubot { default 0; ~*IbouBot 1; } server { # Block the bot completely if ($is_iboubot) { return 403; } # Or serve cached content location / { if ($is_iboubot) { root /var/www/cached; try_files $uri $uri.html $uri/index.html @backend; } try_files $uri @backend; } # Add headers for bot requests location @backend { if ($is_iboubot) { add_header Cache-Control "public, max-age=3600"; add_header X-Robots-Tag "noarchive"; } proxy_pass http://backend; } }

Should You Block This Bot?

Recommendations based on your website type:

Site Type Recommendation Reasoning
E-commerce Allow Essential for product visibility in search results
Blog/News Allow Increases content reach and discoverability
SaaS Application Block No benefit for application interfaces; preserve resources
Documentation Allow Improves documentation discoverability for developers
Corporate Site Allow Allow for public pages, block sensitive areas like intranets

Advanced robots.txt Configurations

E-commerce Site Configuration

User-agent: IbouBot Crawl-delay: 5 Disallow: /cart/ Disallow: /checkout/ Disallow: /my-account/ Disallow: /api/ Disallow: /*?sort= Disallow: /*?filter= Disallow: /*&page= Allow: /products/ Allow: /categories/ Sitemap: https://example.com/sitemap.xml

Publishing/Blog Configuration

User-agent: IbouBot Crawl-delay: 10 Disallow: /wp-admin/ Disallow: /drafts/ Disallow: /preview/ Disallow: /*?replytocom= Allow: /

SaaS/Application Configuration

User-agent: IbouBot Disallow: /app/ Disallow: /api/ Disallow: /dashboard/ Disallow: /settings/ Allow: / Allow: /pricing/ Allow: /features/ Allow: /docs/

Quick Reference

User Agent Match

IbouBot

Robots.txt Name

IbouBot

Category

search

Respects robots.txt

Yes

Quick Actions

Official Docs
Copied to clipboard!