The Scrapers vs. Websites: Why Hide the Data in the First Place?

· Paul Kolomiets

This week I took a little trip to the wonderful land of ToS and public content that you can simultaneously access and not access.

I decided not to write a boring report about what I did and what didn’t work.

Instead, let’s think about how we ended up here.

Yes, I realize reading legal documents for fun is a strange habit. Please forgive this particular vice.

Websites protect their content. But in practice, they are often less concerned about the data itself than about losing control over how people access it.

The data can be publicly available to a human using a browser, while automatically obtaining exactly the same data is prohibited.

Alternative access can break advertising, analytics as a source of user data, and, to some extent, paid features.

If you want a concrete example, dust off hiQ v. LinkedIn.

Very roughly, the argument that publicly accessible data should therefore be available to any client was weakened by the idea that the site’s ToS, as a contract with its users, can impose restrictions on how that access is used.

I am deliberately simplifying the legal conclusions here. If you actually care about the legal details, please spend some extra time reading the case. The real legal picture is considerably less clear-cut.

And that’s more or less where we’ve been ever since.

Everything is public. And somehow, you still can’t use it.

Funny.

This arrangement worked reasonably well until we started throwing robots into the picture.

AI agents don’t fit neatly into this model.

Imagine an agent working alongside you while technically respecting the site’s ToS. It could, for example, passively collect information while you independently browse the pages.

Which, of course, is something you are very unlikely to do.

Manually browsing several hundred pages just to collect enough data for an analysis is, at best, tedious.

But even if you somehow manage to do everything “right”, the user’s behavior has fundamentally changed.

You request more data. You spend less time on each page. You ignore the ads.

In other words, many of the things that motivated websites to introduce these restrictions in the first place stop working for users who actively rely on AI agents.

And there is an even bigger problem.

The same thing happens to people who simply trust AI search.

The first group might still be small enough for companies to ignore. The second one is much larger. And, more importantly, these users may never interact with the website at all.

They can get most of what they need without ever looking at the feed.

If this trend takes hold, I can see two possible outcomes: websites either start relaxing their access restrictions, or we get more deals between search engines and large content providers, along the lines of what we’re already seeing with Reddit and Google.

I suspect the second scenario eventually runs into either the search engine’s monopoly position and regulatory intervention, or simply becomes too expensive for content providers.

I’d rather believe we end up with new ways of monetizing content than with an even more closed web.

There is one factor that could change the equation: paid access.

What if a paid account came with explicit, expanded rights for AI agents to collect and process information?

That could be a sufficiently attractive deal for both content providers and power users. Instead of trying to distinguish between “good” and “bad” automation, you could simply make different levels of access part of the product.

Maybe my predictions sound strange. My crystal ball is clearly having a bad day.

But I want to emphasize the part I consider most important:

The user’s behavior is changing.

And when the way people consume information changes, websites have to adapt.

My guess is that, eventually, that adaptation will mean more openness, not less.

If this kind of problem is on your desk — more about me, or get in touch. In the worst case, we will just have a topic to discuss, and each goes back to our own code.