Publishers will be able to opt out of AI Search, thanks to new regulation

This is interesting and seems to have flown under the radar:

And here are the details from Google:

Obviously, Google has been very quiet about this, but I am surprised that it hasn’t been reported more widely. This doesn’t just impact big publishing houses, it changes the game for everyone who publishes on the web.

Of course, for this to really work, we need OpenAI, Meta, Anthropic, X and everyone else to join and agree.

This is a no-win situation for publishers, though.

If the vast majority of Google users are relying on Google’s AI Overview, then publishers who opt out will be effectively invisible.

Essentially the same thing happened when EU copyright law gave publishers the right to opt out of having excerpts used in Google News.

4 Likes

Maybe, but at least it will be their choice - not Google’s.

Actually, I would argue the opposite. Google has all the power here. By giving publishers a “choice” that it is business suicide, the UK government is failing to address the much more difficult questions about how to protect the rights of copyright holders in the age of genAI. They aren’t putting limits on Google, they’re forcing publishers to decide how much of a beating they’re willing to take.

5 Likes

So what would be your solution? What should the government do (even just a top level idea)?

[edit]

For some reason this dropped off my reply:

“how much of a beating they’re…” The competition to exposure and publicity has existed since the first directories appeared on the web - actually, when the first BBSes appeared, then SEO arrived and everyone was chasing the top spot in search results. If you didn’t invest in SE optimisation, you were effectively “invisible”, but it was your choice.

Only difference now is that the results can be either AI generated or “traditional”, the difference is that before AI results, the level of visibility was up to the publishers (with that I mean everyone publishing on the web), then it was taken out of your hands, and now this gives it back to your control to a degree it ever was.

The sad fact is that only a handful of mega corporations control what people find on the web and how they find it and I can’t see any single government being able to change that. EU as a block may to a degree, but even they will face an uphill struggle to change the status quo. Move away from Microsoft Office EU wide is a good example how hard it is.

[/edit]

A start might be to clarify the rights of copyright holders regarding use of their work in training data sets.

5 Likes

I doubt Musk will ever agree to anything that might impede his plans for world domination.

They have begun that process already: Report on Copyright and Artificial Intelligence - GOV.UK

But here is the tricky part: how? How do you enforce something local that is global?

So how can we enforce our laws on companies operating in countries who have different laws without stopping publishing anything in them?

We already see how weak even “global laws” or “international laws” as some like to call them, are; in China where intellectual property rights don’t exists and they don’t give a damn about stealing/ copying and selling their customers’ designs.

I don’t disagree with you at all, but it’s way too easy to say what should be done, lot harder to tell how. In my opinion this is one step closer to taking back control from these mega corporations.

1 Like

Thanks to the link to the UK report. Here in the US, the Copyright Office has also been exploring copyright and AI issues. Interestingingly, or perhaps forebodingly, part three of their report, the part on generative-AI training, has been in pre-publication for over a year. The document says,

The Office is releasing this pre-publication version of Part 3 in response to congressional inquiries and expressions of interest from stakeholders. A final version will be published in the near future, without any substantive changes expected in the analysis or conclusions.

So what’s taking so long?

Anyway, good discussion in the report even if I don’t agree with (or even deeply understand) it all. I do like the discussion of fair use and AI (fair use in the US is similar to fair dealing UK).

As to how to enforce any regulations around AI and copyright, I guess the Berne Convention/Union might play a part? Maybe…

Considering that the convention existed long before the AI circus started, and WIPO Copyright Treaty (WCT) of 1996 extended it to include digital environments addressing issues like digital rights management and online communications, what do you think?

So the question still stands: how?

I don’t have an answer therefore I can’t in good conscience criticise things that governments do to improve the situation. They may not be perfect and most certainly don’t resolve the problem in one go, but each of these is a small step forward.

The President fired the Librarian of Congress last year, and has been in a separation of powers fight with the Library ever since. That might be a little disruptive to the normal functioning of the office. (The Copyright Office is part of the Library of Congress. As the name suggests, the Library is arguing that it is actually part of the legislative branch and therefore not answerable to the President.)

A good start would be to enforce the existing statutory penalties for infringement. The LLM companies seem to think “if we paid all that money, it would bankrupt us” is a thing that other people should care about.

3 Likes

Yes, that would be, but how do you do that to a company that has no business presence in the country?

The simple fact is that this is a global, not a local problem and you can not solve it with local legislation. See OFCOM vs 4Chan ( 4Chan responds to £520,000 Ofcom fine with AI picture of hamster - BBC News & 4chan launches legal case against Ofcom in US federal court - BBC News)

I believe Google decided to adhere to the UK regulation for two reasons:

  1. They have business presence in the country (legal peril) and

  2. it’s good PR for them at the time there is more resentment against AI companies behaviour and use of AI

I can only speculate, but I think they are also calculating that paying for publishers to use their content is such a small cost that it is worth it as long as they can somehow be able to keep the competition out of it, see Google’s deal with Reddit (https://www.reuters.com/technology/reddit-ai-content-licensing-deal-with-google-sources-say-2024-02-22/)

Not only for other content holders to agree, but AI trainers can base a set of survers in another country, one that doesn’t sign into the agreement. I doubt you could demonstrated that an AI had been trained on a particular chunk of data.

Pretty likely they have a business presence in at least one of the signatories to the Berne Convention.

Actually previous lawsuits have succeeded in showing exactly that. In some cases, the trail has begun with the LLM company publicly bragging about the sources of their data, but there have been other approaches, too.

1 Like

Even if all signatory countries would get their act together, the fines would be just another cost of doing business for these companies. None of them are high enough to really hurt them. We see this in fines to Microsoft, Google, Meta and multitude of other large corporations who are getting fined in regular intervals and they just move on - after they have challenged those fines on every possible court and dragged the process for years on end.

If the worst comes to worst, they start “buying” the training data from companies conveniently located in non-signature countries and have no knowledge what-so-ever what data these companies use for the model training, where they get it or how they get it.

I do hope at some point in time there is a global response to stop these abuses, but I am not holding my breath considering that even the global minimum tax is far from becoming reality and only just over 140 countries have signed up for it.

https://www.oecd.org/en/topics/global-minimum-tax.html

The G7 have reached agreement on a path forward for the global minimum tax and Pillar 2 of the G20 / OECD Inclusive Framework project on Base Erosion and Profit Shifting.

The agreement seeks to maintain the core objectives of Pillar 2 - combatting multinational tax avoidance—while promoting a stable global tax environment that supports fair competition. Recent discussions have considered U.S. Treasury concerns with the application of the rules alongside the U.S minimum tax system.

G7 partners have reached an understanding on a possible solution that would allow the US minimum tax system to operate alongside the Pillar 2 rules but take steps to ensure any substantial risks with respect to the level playing field or base erosion and profit shifting are addressed.

The G7 will now discuss and develop this understanding, and the principles upon which it is based, within the Inclusive Framework of over 140 countries and jurisdictions, while making clear that the removal of proposed retaliatory tax measures in U.S. legislation is essential for this further progress to be made.

https://www.reuters.com/business/finance/global-tax-deal-risk-first-pillar-teeters-2023-06-16/

“Pillar I has a much rockier path forward. Indeed it is quite likely that it will ultimately fail,” said tax lawyer Peter Barnes, who heads industry forum the International Fiscal Association.

“Pillar 1 will not be implemented by the United States. It will not pass Congress,” said one official close to the discussions at the Paris-based Organisation for Economic Cooperation and Development.

And of digital services tax:

The spread of digital services taxes could be a red flag for Republicans in Congress, who introduced a bill last month for a reciprocal levy on countries that target U.S. companies with “unfair taxes” under the global minimum tax.

U.S. Treasury Secretary Janet Yellen told CNBC last week that the bill had little chance of passing and that the United States would get on board with the global minimum.

So yeah, I don’t hold my breath waiting for any global solution to how to control AI companies. At least this regulation makes a small difference by giving back control to the publishers.

Isn’t this an issue to raise with your elected representatives?

2 Likes

I actually think this is a pretty big deal. The interesting part isn’t the UK specifically; it’s the precedent of publishers finally getting a real “no” button instead of just hoping AI companies respect their wishes. The real test is whether OpenAI, Meta, Anthropic, etc. end up having to honor the same choice.

Good luck trying to get electorate to vote candidates campaigning on this issue in. In all countries. Besides, it’s a legal matter so even if you got those politicians in, the next step is to change the law or to get courts to issue higher fines.

On a more positive note, here is another step forward to make a difference:

Transparency obligations for providers and deployers of certain AI systems

  • Marking and labelling of AI-generated content: certain AI-generated or manipulated content must be clearly and visibly labelled and include machine-readable marks. The EU has created a set of icons that can be used for this purpose. This applies to

    • images, audio, and video content that resemble existing persons, objects, places, entities, or events (deepfakes)

    • emotion recognition and biometric categorisation tools

    • text published to inform the public on matters of public interest where there has been no human review or editorial control.

  • Transparency when interacting with an AI system: Users must be clearly informed when they are not interacting with a real person, but an AI system, for example a chatbot, AI agent, and avatar.

…enforcing the transparency rules and may issue fines

  • up to €15 million, or 3% of global annual turnover, for companies

  • up to €750k for EU institutions, bodies, and agencies

  • with proportionality taken into account for small and medium-sized enterprises (SMEs) and small mid-cap companies (SMCs)

https://digital-strategy.ec.europa.eu/en/factpages/quick-facts-transparency-rules-ai-systems

I don’t see anywhere how this can be applied on individuals posting or publishing AI content anonymously, but again, if they can reign in big players (not just AI companies, but also publishers and news organisations etc.) it may lead to them starting to require this from content creators.

Since 2025 China has had a similar legislation (dual explicit/implicit labeling with a three-tier detection classification) and South Korea enabled AI Basic Act in January this year (stricter human-visible-only rules for deepfakes). India enabled a law in February that requires 10% visual label requirement and an 3-hour takedown window.

The USA has not, although California, Texas and Colorado have state level laws since the beginning of this year.

And of course:

in the US, tech industry group NetChoice and allied lobbyists are pushing for federal preemption of state labeling laws, arguing a 50-state patchwork harms innovation, while tech, labor, and civil rights groups have successfully blocked two congressional preemption attempts (one failing 99–1). Canada’s AIDA died from unspecified “overwhelming critique” in committee.

There are some question about the EU legislation also:

some legal scholars flag technical/ definitional weaknesses in the EU’s own deepfake definition, and a Stanford study questions whether labeling actually reduces misinformation’s persuasive power

Even if these are not 100% effective (they aren’t), they are proof that countries are starting to take the AI generated content seriously and are trying to do something about it.

Then again, there is also this:

Summary of above:

Companies are buying up used books to try to avoid model collapse. LLM outputs are getting worse, not better, as the amount of LLM-generated content in the training sets increases.

1 Like

I gave up trying to point this out to the internets about a year ago. There is nothing in the mathematics of neural nets – of which LLMs are a type – that says that approaching completion of the model – towards 100%, whatever the model – will produce factual results.

And, for completion, as a model tends to 100% – and we’re nowhere close to that – the number of nodes – read GPUs – tends to infinity. That’s your data centres right there.

For those interested, read up on the universal approximation theorem, as a starter, and go from there.

2 Likes