Cloudflare Lets Websites Block AI Training Without Losing Search

Cloudflare

Connectivity cloud company Cloudflare has launched a setting that enables website owners to remain visible in search results while blocking crawlers from using their content for artificial intelligence training, the company said in a Tuesday (Sept. 15) press release.

    Get the Full Story

    Complete the form to unlock this article and enjoy unlimited free access to all PYMNTS content — no additional logins required.

    Subscribe to our daily newsletter, PYMNTS Today.

    By completing this form, you agree to receive marketing communications from PYMNTS and to the sharing of your information with our sponsor, if applicable, in accordance with our Privacy Policy and Terms and Conditions.

    Cloudflare’s new Disallow AI Training setting is designed to solve a problem presented by today’s mixed-use crawlers, which collect content for both a search index and AI training. Until now, website owners who blocked these crawlers from AI training also blocked them from search, according to the release.

    Now, if a website owner uses Disallow AI Training to refuse AI training, mixed-use crawlers will be allowed to access that site only if they honor that choice, per the release.

    Apple, Google and Microsoft have each either met the criteria required by the setting or have committed to do so. Several other companies operate separate crawlers for search and training, which allows Cloudflare to block their training crawlers without affecting the search ones, according to the release.

    “This is how we make the internet better: preserving the openness that makes search valuable while giving the people and businesses behind the web meaningful control over how their work is used,” Cloudflare Co-Founder and CEO Matthew Prince said in the release. “We look forward to continued collaborative engagement with companies like Apple, Google and Microsoft as we work together to build a healthy ecosystem.”

    Bryan Becker, director of product, appsec integrity and trust solutions at Cloudflare, said in a Tuesday blog post that Cloudflare has been talking with operators of crawlers about this issue since July.

    “The response has been encouraging: Almost all agreed that site owners should have control and transparency into how their content is used, and reassurance that their choices will be respected,” Becker said.

    The PYMNTS Intelligence report “How Enterprises Can Build a ‘Know Your Agent’ Defense: Digital Identity Verification in the Age of Bots“ found that nearly 90% of organizations said managing bot activity is a challenge.

    As AI-powered agents interact with platforms alongside human users, weak or outdated controls allow malicious bots to slip through while legitimate users are mistakenly blocked, the report found.