Login
Login is restricted to DCN Publisher Members. If you are a DCN Member and don't have an account, register here.

Digital Content Next

Menu

InContext / An inside look at the business of digital content

Blocking AI bots is not an AI strategy

Publishers can’t make informed decisions about which AI systems to block, allow, or license unless they can easily understand who is consuming their content and why.

September 28, 2026 | By Joseph Varvara, VP Sales & Marketing – SupertabConnect on
-bot traffic among humans concept; what to block?-

Publishers are being pushed to make increasingly consequential decisions about AI. Allow this crawler. Block that one. License another. Update robots.txt. Protect premium content. Negotiate with an AI platform. Yet many commercial teams at media and news companies are being asked to make these decisions without a clear picture of the non-human traffic already consuming their content.

I’ve spent a lot of time recently talking with publishers and AI companies about content access and licensing. Those conversations quickly get into fair use, grounding vs. training, fair market pricing, and future demand. But rarely does anyone start with the more basic question:

What are these machines actually doing on my site?

We need to stop treating every bot the same

Many publishers already have some of this data. They may receive reports from their CDN or keep a list of their largest sources of bot traffic. The harder question is what to do with it.

“Bot traffic” has become a convenient catch-all for automated activity, but it hides important differences. An AI training crawler, an AI search agent, a traditional search indexer, and a scraper can all request the same article for completely different reasons.

A search crawler might index a story and drive readers back to the publisher. An AI system might retrieve it to answer a user’s question. A training crawler might collect it for model development. An unidentified scraper might repeatedly access an archive. It’s the same article, with very different behavior, and very different value.

Bot classification needs to become as fundamental to publisher analytics as understanding referrals, engagement and subscriptions. Knowing that a request came from a machine isn’t enough. Publishers need to know which agents, crawlers and AI systems are consuming their content, what they access, how frequently they return, and how their behavior changes over time.

Publishers have spent decades getting better at understanding human audiences. We know where readers came from, what they read, whether they came back, and whether they subscribed. 

An AI system doesn’t behave like a reader. It might retrieve a single article, request specific information, or systematically access one section of a site. That consumption may never appear as a conventional session, referral, or subscription conversion, yet it may still have significant value.

Publishers need to start asking different questions. Which AI systems are accessing our content? What are they requesting? Which systems are increasing their activity? What content do they consume most frequently? Without that visibility, publishers are building an AI strategy with only part of the picture.

Blocking is a tool, not a strategy

The first phase of the publisher response to AI understandably focused on control: “Who should we block?” That’s an important question. But blocking bots is not a monetization strategy. In some cases, blocking may absolutely be the right decision, but it should be an informed decision.

A search crawler that generates discovery shouldn’t necessarily be treated like an AI training crawler. An AI company with an existing licensing agreement shouldn’t necessarily be treated like an unidentified scraper. And an AI agent repeatedly retrieving valuable content might represent a commercial opportunity rather than simply unwanted traffic.

Classification turns a blanket bot policy into an informed access strategy. Publishers can decide which systems to allow, which to restrict, which to block, and which might warrant a licensing conversation.

You can’t price what you can’t see

This becomes even more important when the conversation moves from access to AI content licensing. One thing I’ve observed in licensing discussions is how difficult it can be to establish the value of content without understanding how it is actually being consumed.

Knowing an AI company visited your site is one thing. Knowing that it repeatedly retrieves your financial reporting, product reviews, or local news is something else entirely. Now you have commercial intelligence you can bring into a licensing conversation.

Instead of negotiating exclusively around theoretical value, publishers can begin to understand actual machine consumption: what is being requested, by whom, and how often. And different consumption may deserve different economics. Search indexing, AI training and real-time retrieval create different use cases. Publishers should have the information necessary to decide whether those uses deserve different access rules, licensing terms, or pricing.

The machine audience is already here

I don’t think publishers should view every AI system consuming their content as an adversary. I also don’t think they should assume every request deserves free access. The more useful position sits between those extremes. Understand the consumption first. Decide the economics second.

AI systems increasingly consume content one request at a time. That creates an opportunity to move toward consumption-based commerce, where machine consumption can eventually be metered, priced, governed, and settled based on what is actually used. 

But none of that works without visibility. The first step in an AI strategy isn’t deciding what to block. It’s knowing who is consuming your content, what they’re taking, and how that consumption should be treated.

Liked this article?

Subscribe to the InContext newsletter to get insights like this delivered to your inbox every week.