As I mentioned in my Dead Software Walking: The ongoing evolution of relayd(8) and httpd(8)
post, development of
relayd(8) and httpd(8) was revived. In diesem Beitrag würde
ich gerne über einige Featuers sprechen die in relayd(8) und httpd(8) einzuhalten werden.
httpd: add custom HTTP header support #
OpenBSD’s httpd can now manipulate HTTP response headers directly in httpd.conf. Until now, adding
security headers or stripping headers from a FastCGI backend meant changing the application code or
putting relayd in front of httpd. For a simple setup, that was often more work than it should be.
The new header directive has three options. set adds a header or replaces an existing one with the
same name. add appends a header even if one with that name already exists. remove suppresses a
header, whether httpd set it or the backend sent it or it was inherited from the server context.
By default, headers only apply to 2xx and 3xx responses. Add always to include error responses as well, which you usually want for security headers. Headers set in the server context are inherited by location blocks, and a location can override or remove them by name.
Here is a practical example: a personal blog with static pages, cacheable assets, a drafts folder, and some PHP.
server "blog.sizoefovid.org" {
listen on * tls port 443
tls {
certificate "/etc/ssl/blog.sizeofvoid.org.crt"
key "/etc/ssl/private/blog.sizeofvoid.org.key"
}
hsts max-age 31536000 # built-in, no "header set" needed
root "/htdocs/blog"
# Security headers for every response, including 404/500 pages
header set "Content-Security-Policy" "default-src 'self'; img-src 'self' data:; frame-ancestors 'none'" always
header set "X-Content-Type-Options" "nosniff" always
header set "Referrer-Policy" "strict-origin-when-cross-origin" always
header set "Permissions-Policy" "camera=(), microphone=(), geolocation=()" always
# Fingerprinted assets can be cached forever
location "/assets/*" {
header set "Cache-Control" "public, max-age=31536000, immutable"
}
# Drafts: shareable by link, but not indexed and not cached
location "/drafts/*" {
header set "X-Robots-Tag" "noindex, nofollow" always
header set "Cache-Control" "no-store" always
}
# Don't reveal the PHP version
location "*.php" {
fastcgi socket "/run/php-fpm.sock"
header remove "X-Powered-By"
}
}httpd: add header block/drop rules for request filtering #
httpd can now reject incoming requests based on request headers. This is useful for keeping AI
scrapers, vulnerability scanners, and other unwanted clients away, without relayd or a firewall rule
that only knows IP addresses.
Of course, it’s not perfect, and if the user-agent isn’t set, there’s not much we can do. But for a
large number of AI scrapers, this seems to be helping at the moment!
I’d like to refer you to the email from purplerain () secbsd ! org, who carried out an analysis on this and requested this feature – and didn’t just write an initial version of it. Thanks again!
Following those tests, the latest patch is now running in production. It has been
running without issues for at least 24 hours so far.
I will continue performing stress tests and will send additional results over the
next few days.
RESULTS
Duration: 2h 51m 11s
Requests sent: 10000000
Dropped: 4820945 48.2%
Blocked: 182191 1.8%
Served: 4996864 50.0%
Redirected: 0 0.0%
Network errors: 0 0.0%
BY CATEGORY
FAKE: 5003136 sent · 4820945 dropped (96.4%) · 182191 responded (3.6%)
LEGITIMATE: 4996864 sent · 0 dropped (0.0%) · 4996864 responded (100.0%)
HTTP STATUS
200: 4996864
403: 45392
404: 136799
Here are some other use cases:
$ cat security-tools.conf
header drop "user-agent" "commix*"
header drop "user-agent" "dav*"
header drop "user-agent" "dirbuster*"
header drop "user-agent" "feroxbuster*"
header drop "user-agent" "ffuzz*"
header drop "user-agent" "gobuster*"
header drop "user-agent" "sqlmap*"
header drop "user-agent" "whatweb*"
header drop "user-agent" "wfuzz*"
header drop "user-agent" "wpscan*"
# Drop requests with and empty User-Agent by default.
header drop "user-agent" ""
# ncrack
# lbd
# legion
Purple Rain– https://marc.info/?l=openbsd-tech&m=179086980940833&w=2
There are two new options. header block answers a matching request with an HTTP status code and then
closes the connection. For 3xx codes, you must give a target URL, which is sent as the Location
header. For all other codes, you can give an optional label that shows up in the log, so you can see
which rule matched. header drop closes the connection silently, without any response. This is the
better choice for scanners: they get no information back, and your server does no extra work.
Both the header name and the value are glob patterns (*, ?, […]) and are matched case-insensitively.
Here is the blog from the previous example, with some request filtering added:
server "blog.sizoefovid.org" {
listen on * tls port 443
tls {
certificate "/etc/ssl/blog.sizeofvoid.org.crt"
key "/etc/ssl/private/blog.sizeofvoid.org.key"
}
# AI scrapers: answer with 403 and a log label per rule
header block "User-Agent" "*GPTBot*" 403 "ai-gptbot"
header block "User-Agent" "*ClaudeBot*" 403 "ai-claudebot"
header block "User-Agent" "*CCBot*" 403 "ai-ccbot"
header block "User-Agent" "*Bytespider*" 403 "ai-bytespider"
# Known scanners: no response at all
header drop "User-Agent" "*zgrab*"
header drop "User-Agent" "*masscan*"
header drop "User-Agent" "*Nuclei*"
# Log4Shell probes can hide in any header
header drop "*" "*${jndi:*"
# Legacy browsers: send them to a static fallback site
header block "User-Agent" "*MSIE*" 302 "https://legacy.example.org/"
# ... see example above
}For inspiration, here’s a list from Purple Rain that he sent me. It contains his research. You can
save lists like this in a file and include them in your httpd config using include "ua.conf".
# GENERIC USER-AGENTS
header drop "user-agent" "*bot*"
header drop "user-agent" "go*"
header drop "user-agent" "modat*"
header drop "user-agent" "axios*"
header drop "user-agent" "curl*"
header drop "user-agent" "crawler*"
header drop "user-agent" "headlesschrome*"
header drop "user-agent" "httpie*"
header drop "user-agent" "java*"
header drop "user-agent" "libredtail*"
header drop "user-agent" "libwww*"
header drop "user-agent" "lwp*"
header drop "user-agent" "node*"
header drop "user-agent" "okhttp*"
header drop "user-agent" "php*"
header drop "user-agent" "puppeteer*"
header drop "user-agent" "*python*"
header drop "user-agent" "*request*"
header drop "user-agent" "ruby*"
header drop "user-agent" "scrapy*"
header drop "user-agent" "selenium*"
header drop "user-agent" "wget*"
# AI AND TRAINING
header drop "user-agent" "*anthropic*"
header drop "user-agent" "aiwebindex*"
header drop "user-agent" "amazon*"
header drop "user-agent" "amzn*"
header drop "user-agent" "anomura*"
header drop "user-agent" "apify*"
header drop "user-agent" "aranet*"
header drop "user-agent" "awario*"
header drop "user-agent" "azureai*"
header drop "user-agent" "bigsur*"
header drop "user-agent" "bytespider*"
header drop "user-agent" "chatglm*"
header drop "user-agent" "chatgpt*"
header drop "user-agent" "claude*"
header drop "user-agent" "cloudflare*"
header drop "user-agent" "cohere*"
header drop "user-agent" "cotoyogi*"
header drop "user-agent" "cragcrawler*"
header drop "user-agent" "crawl4ai*"
header drop "user-agent" "crawlspace*"
header drop "user-agent" "cursor*"
header drop "user-agent" "datenbank*"
header drop "user-agent" "deepseek*"
header drop "user-agent" "devin*"
header drop "user-agent" "exa*"
header drop "user-agent" "facebook*"
header drop "user-agent" "factset*"
header drop "user-agent" "firecrawl*"
header drop "user-agent" "friendlycrawler*"
header drop "user-agent" "geisthaus*"
header drop "user-agent" "iask*"
header drop "user-agent" "img2dataset*"
header drop "user-agent" "imagespider*"
header drop "user-agent" "isscyberriskcrawler*"
header drop "user-agent" "kagi-fetcher*"
header drop "user-agent" "kangaroo*"
header drop "user-agent" "kimi*"
header drop "user-agent" "klaviyo*"
header drop "user-agent" "kunatocrawler*"
header drop "user-agent" "laion*"
header drop "user-agent" "lcc*"
header drop "user-agent" "lightpanda*"
header drop "user-agent" "linguee*"
header drop "user-agent" "manus*"
header drop "user-agent" "meta*"
header drop "user-agent" "mistralai*"
header drop "user-agent" "netestate*"
header drop "user-agent" "newsai*"
header drop "user-agent" "notebooklm*"
header drop "user-agent" "novaact*"
header drop "user-agent" "omgili*"
header drop "user-agent" "openai*"
header drop "user-agent" "opencode*"
header drop "user-agent" "operator*"
header drop "user-agent" "panscient*"
header drop "user-agent" "perplexity*"
header drop "user-agent" "poggio*"
header drop "user-agent" "poseidon*"
header drop "user-agent" "querit*"
header drop "user-agent" "shap*"
header drop "user-agent" "sidetrade*"
header drop "user-agent" "terra*"
header drop "user-agent" "tiktok*"
header drop "user-agent" "trae*"
header drop "user-agent" "twinagent*"
header drop "user-agent" "useai*"
header drop "user-agent" "velenpublicwebcrawler*"
header drop "user-agent" "webzio*"
header drop "user-agent" "yaK*"
header drop "user-agent" "yandex*"
# SECURITY TOOLS
# Drop reconnaissance tools, scanners, fuzzers, brute-force tools, etc.
# This is not and exhaustive list of tools.
# Dirbuster, feroxbuster, and gobuster are examples, but these could be covered
# by a single "*buster*" pattern.
# The same applies to fuzzers as ffuf and wfuzz, both use the string "fuzz" in
# their User-Agent, so they could be covered by a single "fuzz*" pattern.
header drop "user-agent" "*buster*"
header drop "user-agent" "commix*"
header drop "user-agent" "dav*"
header drop "user-agent" "dirbuster*"
header drop "user-agent" "feroxbuster*"
header drop "user-agent" "fuff*"
header drop "user-agent" "fuzz*"
header drop "user-agent" "gobuster*"
header drop "user-agent" "sqlmap*"
header drop "user-agent" "whatweb*"
header drop "user-agent" "wfuzz*"
header drop "user-agent" "wpscan*"
# Drop requests with and empty User-Agent by default.
header drop "user-agent" ""
# ncrack
# nuclei
# lbd
# legion
# SEARCH ENGINE AND CRAWLERS
header drop "user-agent" "baidu*"
header drop "user-agent" "bing*"
header drop "user-agent" "seekport*"
header drop "user-agent" "slurp*"
header drop "user-agent" "sogou*"
header drop "user-agent" "yahoo*"
# SEO TOOLS AND SCRAPERS
header drop "user-agent" "ahrefs*"
header drop "user-agent" "deepcrawl*"
header drop "user-agent" "majestic*"
header drop "user-agent" "screamingfrog*"
header drop "user-agent" "sitebulb*"!!! This list is very aggressive. It also blocks search engines, link previews, curl, and uptime monitors. Read it before you use it and take only what fits your site.
relayd: add patters(7) support and improve glob(7) documentation #
relayd filter rules for cookie, header, path, query, and url now accept an optional pattern keyword
before the key or value. With it, the string is read as a patterns(7) expression, the same syntax
httpd uses for location match. Without it, relayd uses glob(7) as before, so existing configs keep
working.
patterns(7) gives you character classes like %d, anchors like ^ and $, and repetition. This lets you
write rules that glob cannot express, like matching only numeric IDs. The glob(7) rules in relayd
were never really documented. The man page now explains what they can and cannot do.
What’s next in pattern patching syntax world #
I am also working on a follow-up change. Next to pattern matching, you will be able to write glob and glob ignorecase. glob stays the default. glob ignorecase makes matching case-insensitive on any field,
including cookie values, paths, and query strings, which are otherwise compared exactly as sent. The
change is not committed yet, so the syntax may still change.
I’m not entirely sure yet, but I could imagine extending this matching syntax to include additional elements and maybe httpd.
devel #
Development happens primarily on the Gothub instance. The other locations are kept in sync:
Primary: https://rsadowski.gothub.org/
Mirror: https://codeberg.org/rsadowski/relayd
Mirror: https://github.com/sizeofvoid/relayd
As always, the source of truth is OpenBSD -current.A special thanks to all who support my work.