{"id":5582,"date":"2026-09-16T14:01:59","date_gmt":"2026-09-16T14:01:59","guid":{"rendered":"https:\/\/redmonk.com\/jgovernor\/?p=5582"},"modified":"2026-09-16T14:09:30","modified_gmt":"2026-09-16T14:09:30","slug":"the-agents-are-coming-for-the-web-and-the-web-isnt-ready","status":"publish","type":"post","link":"https:\/\/redmonk.com\/jgovernor\/the-agents-are-coming-for-the-web-and-the-web-isnt-ready\/","title":{"rendered":"The agents are coming for the web and the web isn\u2019t ready."},"content":{"rendered":"<p><a href=\"http:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-scaled.jpeg\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-5583\" src=\"http:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-1024x683.jpeg\" alt=\"People playing whack game at carnival in Coquitlam BC Canada.\" width=\"1024\" height=\"683\" srcset=\"https:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-1024x683.jpeg 1024w, https:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-300x200.jpeg 300w, https:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-768x512.jpeg 768w, https:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-1536x1024.jpeg 1536w, https:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-2048x1365.jpeg 2048w, https:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-480x320.jpeg 480w, https:\/\/redmonk.com\/jgovernor\/files\/2026\/09\/AdobeStock_693835607-941x627.jpeg 941w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/a><\/p>\n<p>Just riffing off an interesting post by Per Buer of Varnish Software &#8211; <a href=\"https:\/\/info.varnish-software.com\/blog\/the-agents-are-coming-for-the-web-and-the-web-isnt-ready\">The agents are coming for the web and the web isn\u2019t ready<\/a>.<\/p>\n<p>The problem statement is that public infrastructure is being swamped by crawlers and agents. As Buer says:<\/p>\n<blockquote><p>Konstantin Ryabitsev&#8217;s post <a href=\"https:\/\/people.kernel.org\/monsieuricon\/creepy-crawlies\">Creepy crawlies<\/a> puts numbers on something anyone running a public git host has felt for a while: agents are an increasingly heavy burden on anyone publicly hosting information. So an ever increasing number of people are sending out more and more agents on longer and longer goose-chases and the traffic is increasing at an alarming rate.<\/p><\/blockquote>\n<p>&#8220;Anyone running a public git host&#8221; &#8211; yeah, GitHub knows something about that.<\/p>\n<p>Buer&#8217;s post is specifically about git.kernel.org, which is indeed being swamped with inbound traffic, and not just inbound traffic but stupid inbound traffic. Agents don&#8217;t much care for efficiency. In the post Per shows how Varnish Cache Language (VCL) might help.<\/p>\n<p>But the specific problem is described in Ryabitsev&#8217;s <a href=\"https:\/\/people.kernel.org\/monsieuricon\/creepy-crawlies\">post<\/a>:<\/p>\n<blockquote><p>At the time of writing, linux.git is about 1.48 million commits. Oh, and we have about 922 forks of it on git.kernel.org \u2014 but don&#8217;t worry, it&#8217;s actually extremely efficient on the backend, since it&#8217;s mostly the same objects in every fork.<\/p>\n<p>Unless, of course, you&#8217;re a scraper, in which case you have, oh, several BILLION valid URLs you can scrape, only to get 922 duplicates of the same 1.48 million commits \u2014 which is exactly what the scrapers are doing.<\/p>\n<p>But wait, it&#8217;s not just commits itself. You can also ask for patches, plain renders, diffs between arbitrary commits \u2014 cgit is happy to let you, which was perfect for the times when the Internet was for humans or crawlers who obeyed robots.txt, and is AWFUL right about now, because we can generate 1.2 METRIC BAJILLION valid URLs just for a single fork of linux.git.<\/p><\/blockquote>\n<p>kernel.org has effectively been playing whack a mole with a mole (well maybe not quite <em>6.02 x 10<sup>23<\/sup><\/em> but you get the idea) of moles, and the details are legit fascinating. Agentic and agent-driven problem solving is always a trip to see in detail. And of course humans are more than capable of building shitty CI systems and crawlers.<\/p>\n<blockquote><p><strong>Block them<\/strong><\/p>\n<p>Initially, this was the solution \u2014 look through the logs, find out which IPs are obvious scraper bots, and fail2ban them. At first, this was easy, because the bots helpfully told you who they were via their user-agent. Then, they wised up and started pretending that they were random vanilla browsers.<\/p>\n<p>So, we started banning them by IP \u2014 after all, it&#8217;s easy to figure out that an IP that is trying to grab every possible commit in a 8-year-old abandoned fork of linux is not really some lone Chrome on Windows user who is just furiously clicking every link that comes across their screen.<\/p>\n<p>The bots then started fanning out to entire subnets, but this was still meh, because obviously an IP coming from Google Compute is just pretending to be a Firefox user. Banning the whole ASN was justified, even if this occasionally caught a random legitimate instance trying to automate link checking in commits.<\/p>\n<p><strong>Enter&#8230; your TV?<\/strong><\/p>\n<p>And&#8230; that&#8217;s when things turned really, really ugly. Suddenly, the crawlers were coming from millions of random residential or mobile IPs, all pretending to be random modern browsers. An IP like that would make 4-5 requests and then never show up in the logs again. There was no point in banning them, because by the time you figured out that they were bots, they were already done with you. You just needlessly ballooned your firewall ruleset by adding IPs that would never be back.<\/p>\n<p>They descended like swarms of locust, hit hard and fast until the system fell over and then moved on to the next target until you recovered. Then, they returned. Rinse. Repeat.<\/p><\/blockquote>\n<p>Anyway you should just read the post. The point I wanted to make is that agents and frontier model companies are incredibly noisy neighbours &#8211; they don&#8217;t care about your robots.txt file. Varnish isn&#8217;t alone here &#8211; Cloudflare and Fastly are both working the problem. Cloudflare quite vocally &#8211; it built a Block all Bots switch, but yesterday <a href=\"https:\/\/cloudflare.net\/news\/news-details\/2026\/Cloudflare-Helps-End-the-Search-or-AI-Training-Tradeoff\/default.aspx\">announced a more granular approach<\/a>, allowing &#8220;good&#8221; and &#8220;bad&#8221; LLM traffic, with different settings for search, AI training, and AI agents. Cloudflare is calling for &#8220;Accountable AI Crawling&#8221; &#8211; with allowances for responsible behaviour from the likes of Apple, Google and Microsoft. The CDN and next gen Cloud platforms are due a moment in the sun if they can help shape and manage these new traffic patterns.<\/p>\n<p>But, yeah, in the current ethical wasteland of the industry when it comes to data collection and content aggregation for model training, this will continue to be a problem. And it&#8217;s worth stressing that while this post began talking about git hosting and caching, agents are going to be generating load in every domain &#8211; and not just across content but across transaction systems. Agents are incredibly hungry and they don&#8217;t need to sleep. Anything public facing is going to face incredible scaling demands. It&#8217;s very early days in the agent era. GitHub agent traffic has far surpassed the company&#8217;s expectations, which has led to reliability issues. But it&#8217;s going to be alone. Agent traffic increases aren&#8217;t going to be a couple of 100 per cent. We&#8217;re talking orders of magnitude. And that&#8217;s something we&#8217;re all going to need to deal with. First came the web, then came e-business, then self-service, and then mobile traffic. Well it&#8217;s now the agent era.<\/p>\n<p>The agents are coming for the web and the web isn\u2019t ready.<\/p>\n<p>&nbsp;<\/p>\n<p>Cloudflare, GitHub, and Fastly are clients.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&#8220;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Just riffing off an interesting post by Per Buer of Varnish Software &#8211; The agents are coming for the web and the web isn\u2019t ready. The problem statement is that public infrastructure is being swamped by crawlers and agents. As Buer says: Konstantin Ryabitsev&#8217;s post Creepy crawlies puts numbers on something anyone running a public<\/p>\n","protected":false},"author":5,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[1],"tags":[],"class_list":["post-5582","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"acf":[],"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p9wfjh-1s2","jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/posts\/5582","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/users\/5"}],"replies":[{"embeddable":true,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/comments?post=5582"}],"version-history":[{"count":2,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/posts\/5582\/revisions"}],"predecessor-version":[{"id":5585,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/posts\/5582\/revisions\/5585"}],"wp:attachment":[{"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/media?parent=5582"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/categories?post=5582"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/tags?post=5582"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}