How many machines are on the web: reading bot and agent figures

This is a live Practicum page for Volumes 1-3. Verified: 2026-08-18. These figures can become outdated within months, which is why they do not belong in a printed book. The book explains the principle. This page holds current values.

Discipline: almost every measurement comes from a company that sells bot protection or traffic access. These are industry estimates from interested parties, not a census of the internet. Each figure below says who measured it and what the measurement included.

"Bots are already more than half the internet" makes a strong headline, but it is usually imprecise. This page explains what sits behind the claim and how to read the next headline yourself.


Main figures for 2025-2026

What was measured Value Who measured it, and what was included Verified
Share of automated requests to web pages 57.5%, compared with 42.5% from people Cloudflare Radar, announced by CEO Matthew Prince on Jun 3, 2026. This covers the Cloudflare network and HTML requests to web pages, not email, video, or games. 2026-08
How much earlier than forecast About 18 months. Prince had expected the crossover by the end of 2027. His public forecast from Mar 2026 2026-07
Annual growth in agent traffic About 80 times, or 7,851% HUMAN Security, "State of AI Traffic" 2026. These are agents that click and complete forms, not search crawlers. 2026-07
Growth in scrapers 597% in one year Same source 2026-07
Automated traffic growth compared with human traffic About eight times faster Same source 2026-07
What the crawlers came for 52% of crawler requests were for AI training as of Jun 2026, up from 22% in spring 2025 Cloudflare. The most useful single line in the table: the machines are not mostly indexing you for search any more 2026-08
Bot share in other measurements 51% using 2024 data, and 53% in an estimate that dates the crossover to 2023 Imperva and Thales both sell bot protection. The gap with Cloudflare is a definition gap: Cloudflare counts HTML requests, Imperva counts all web traffic including app and API calls. 2026-08

How to read the table. Each company sees a different part of the internet and uses its own definitions. The exact figures do not match. The direction does: machine traffic is growing, and more of that growth now comes from agents acting for specific people, rather than routine system automation.


Three things these figures do not mean

They do not mean that neural networks made more than half of online text and images. Traffic counts requests to websites, not content. A person may publish model-written text through a normal browser. A search crawler may open a page a thousand times without creating any content.

They do not mean that every machine is harmful. The group includes search crawlers, monitoring services, price collectors, security scanners, and agents that are honestly finding a supplier for their owner. Harmful bots are a separate, smaller part of the total.

They do not mean the crossover happened on one exact day. The measurement companies say they cannot establish an exact date. It depends on the network being measured and the definition of a bot.


Why this is an economic change

The commercial web assumes that a visitor is a person. Advertising impressions, sales funnels, and page-view analytics all use that assumption. It stops working when a large share of visitors are programs.

The new economy is still small, and the gap is revealing:

  • Pitchbook estimates that agents currently handle about 1% of the potential volume of work;
  • Forsy estimates annual "agent gross product" at about $36 billion, which is small next to the global economy.

The reason is the same one discussed in Volume 3's chapter on agent money: payments and accountability are not solved. A machine can visit, read, and compare. It cannot yet answer for the purchase. Agent presence is already large, but the agent economy is still small.

Some platforms are much further ahead. Stripe says that agents make about 70% of calls to its interface. Brokerage service Alpaca says the agent share of calls grew from low single digits to about 30% in one quarter. These are the companies' own claims.


The part that changed in 2026: machines started getting billed

Until recently the only choices were to let a crawler in or block it. A third option arrived, and the infrastructure companies are competing to own it: charge the machine for reading you.

  • Pay per crawl. Cloudflare lets a site attach a price to access instead of a simple allow or block. A crawler that reaches paid content can be answered with HTTP 402 Payment Required, a status code that sat almost unused in the standard for thirty years and now has a job.
  • Machine-readable licence terms. RSL, or Really Simple Licensing, adds licence and price terms to a site in a form a machine can parse. The RSL Collective is backed by Reddit, Yahoo, Medium, and O'Reilly. Cloudflare's Content Signals Policy and the IETF AIPref working group are parallel efforts toward the same thing.
  • A default that flips soon. From Sep 15, 2026, new sites on Cloudflare are set to allow ordinary search crawling while blocking AI training and agent use on advertising-supported pages. Check the current policy before you plan around it.
  • A market forming. Cloudflare acquired Human Native in Jan 2026, which points at a paid-access data marketplace rather than free scraping.

Why this matters more than the traffic percentage. A share of visitors being machines is a fact about plumbing. A price attached to machine access is a fact about who captures the value of what you publish. If half of crawler traffic is now collecting training data, the question stops being "how do I block them" and becomes "on what terms do I let them in." That is the same output-to-outcome move this series describes, arriving at the level of a website.

This layer is young and the terms will move. Treat every date here as a starting position.


The measurement problem

This part rarely makes the headline, but it changes how the numbers should be used.

Detectors catch agents that identify themselves as bots. An agent that behaves like a person can avoid the filter. The true machine share is therefore probably higher than the measured figure, but nobody knows by how much.

False positives block people. A University of Bamberg study estimates a 7% to 15% false-positive rate in real traffic. These are human visitors that the system classified as bots.

Analytics can become wrong without an obvious warning. Seer Interactive reports that agentic browsers inflate engagement, reduce measured bounce rates, and distort visit duration. If you make decisions from web analytics, some of the figures may no longer describe people.


What a reader can do with this

If other people or systems verify you, see Volume 3, Chapters 11-12. The first reader of your page may increasingly be a program. It has no intuition and does not assume good intent. It will not fill in a missing fact out of politeness. State what matters to the decision: the responsible person's name, the date, the number, and the source. Mood and an image are not enough. Use the Trust Footprint Map.

If you sell, see Volume 2, Chapter 6. The book explains how to build for machine readers and why an agent does not care about a button's color. Use Agentic browsers for the tool map and the worksheet for visibility in AI search.

If you are reading a feed, see Volume 3, Chapter 1. The share of machine traffic does not tell you whether a person wrote the specific text in front of you. Those are different questions. Mixing them is the most common mistake in conversations about a "dead internet."


How to check the next headline

Ask four questions:

  1. Whose network was measured? Nobody measures the entire internet. A company measures its own infrastructure.
  2. What did it count? Page requests, all traffic, and unique visitors are different measures.
  3. Who counted, and what does it sell? Bot protection? Traffic access? Add the conflict label.
  4. What did it call a bot? A search crawler, a scraper, and an agent acting for a person have different effects.

Sources (verified 2026-08-18)

Every measurement organization listed here, except the university research team, has a commercial interest in this topic. That is not a reason to ignore the data. It is a reason to read the figures as industry estimates.

How many machines are on the web: reading bot and agent figures