Two loud predictions have both been wrong for about three years now.
The first: generative models replace writers, marketers and editors outright, and content becomes a button you press. The second, the backlash version: it's all slop, serious people aren't really using it, the whole thing deflates. Neither survives contact with the measurements. A 2025 Science Advances study of more than 15 million PubMed abstracts found that at least 13.5% of 2024 biomedical abstracts showed the statistical fingerprints of LLM processing, rising to 40% in some subcorpora. That is not a button nobody presses, and it is also not a profession that stopped existing — those abstracts still have authors, institutions and accountability attached.
What actually moved is where the expensive part of the work sits. Drafting got cheap. Deciding whether a draft is true, original and safe to publish did not get cheap at all, and as of 2 August 2026 that step is no longer only an editorial preference. It is written into European law as a named exemption.
The carve-out that just became law
Article 50 of the EU AI Act (Regulation 2024/1689) became applicable on 2 August 2026. Three of its obligations touch anyone shipping content:
| Obligation | What it requires | From |
|---|---|---|
| Art. 50(1) | Tell people when they're interacting with an AI system, unless that's obvious | 2 Aug 2026 |
| Art. 50(2) | Mark synthetic audio, image, video and text in a machine-readable format | 2 Aug 2026; provisionally pushed to 2 Dec 2026 for systems already on the market |
| Art. 50(4) | Disclose deepfakes, and disclose AI-generated text published to inform the public on matters of public interest | 2 Aug 2026 |
Read 50(4) closely, because it contains the whole argument of this article. The disclosure duty for public-interest text does not apply where the content has undergone human review or editorial control and a person holds editorial responsibility for its publication.
The law does not care whether a machine helped write it. It cares whether an identifiable human took responsibility for publishing it. That is the same line every functioning newsroom has drawn for a century, now with a regulation number attached.
The 50(2) marking obligation is the one to actually plan around, because it can't be satisfied by a byline. It requires provenance to be embedded in the artifact. The two shipped mechanisms are Google DeepMind's SynthID, which watermarks images, video and audio and adjusts token probabilities to watermark text out of the Gemini app, and C2PA Content Credentials, an open manifest standard now at spec 2.3 whose steering committee includes Adobe, Google, Meta, Microsoft, OpenAI, Sony, TikTok and the BBC. The two work differently and fail differently: SynthID is embedded in the pixels and is designed, in DeepMind's words, "to stand up to modifications like cropping, adding filters, changing frame rates, or lossy compression," while a C2PA manifest is metadata riding alongside the file, which any pipeline that strips EXIF also strips. Provenance is a chain of custody either way, not a lie detector, and neither survives a system that never applied it in the first place.
Google has been saying a version of this since 2023
The SEO panic about "AI content penalties" has always misread the policy. Google's own documentation, last updated December 2025, says that "using generative AI tools or other similar tools to generate many pages without adding value for users may violate Google's spam policy on scaled content abuse." The trigger is many pages and without adding value. Authorship is not the criterion, and the spam policy applies to scaled unoriginal content no matter how it's produced — a content farm with fifty underpaid humans breaks the same rule.
So the two regimes converge on the same test from opposite directions. Google asks whether anyone added value. The AI Act asks whether anyone took responsibility. Both are asking about the review step, and neither is asking who typed the first draft.
The failure mode that hasn't gone away
The reason unreviewed output stays dangerous is structural, and OpenAI's own researchers published the clearest account of it. In Why Language Models Hallucinate (2025), they argue that models confabulate because standard training and evaluation reward guessing over admitting uncertainty: under the 0-1 scoring that most benchmarks use, a model that always answers beats an otherwise identical model that sometimes says "I don't know," because an abstention scores exactly as badly as a wrong answer. The fix they propose is to change the scoring, not to add another layer of prompt engineering.
Sit with what that means for content work. The systems are optimised to produce a fluent, confident answer under uncertainty. Fluent confident wrongness is precisely the failure mode publishing cannot absorb, because it's invisible to every check except knowing the subject.
The legal profession has become the public test case. The AI Hallucination Cases database, maintained by legal researcher Damien Charlotin, collects court decisions in which filings contained AI-fabricated citations; it has grown from a couple of hundred entries to well over a thousand in roughly a year. Every one of those was a professional with malpractice exposure, filing into a system that verifies citations for a living. Fabricated references are the one hallucination class that is trivially checkable — you open the case — and they still shipped, because nobody opened the case.
The "AI voice" is measurable, which is stranger than it sounds
The Science Advances work is worth understanding properly, because it does something detectors can't. Kobak and colleagues didn't classify individual abstracts. They tracked excess frequency of style words across a 15-million-document corpus and found an abrupt jump after ChatGPT's release in late 2022, concentrated in a small vocabulary: four adjectives (crucial, comprehensive, intricate, pivotal) and four verbs (delve, underscore, utilize, align).
That is a population-level measurement, and it's robust for the same reason it's useless on your single blog post: excess frequency only means something across thousands of documents. One writer who happens to like the word "pivotal" is noise. Ten thousand of them at once is a signal.
Two things follow. First, "AI voice" is a real, quantified phenomenon rather than a vibe. Second, it is trivially avoidable by anyone who reads their own draft out loud, which is why the marker set is a moving target rather than a stable fingerprint.
Common mistakes
Treating output as a draft you proofread rather than one you verify
Proofreading catches tone and typos. It does not catch a confidently invented statistic, a misattributed quote or a plausible-looking citation to a paper that doesn't exist. Those are the errors these systems produce, and they're specifically the ones that read well. The rule that works: every factual claim, number and reference gets opened and checked against the source, and anything you can't open gets cut. That is slower than writing the sentence yourself, and it is the actual cost of AI-assisted drafting.
Running "AI at scale" as an SEO strategy
This is the mistake the policy was written to catch. Publishing hundreds of generated pages to blanket a keyword space is the textbook definition of scaled content abuse in Google's spam policies, and it's judged on the pattern rather than the tooling. The economics also stopped working: generation costs collapsed for everyone simultaneously, so mass-produced pages compete against an infinite supply of identical pages. Whatever advantage existed there was arbitraged away by 2024.
Assuming nobody can tell, and writing in the marker vocabulary anyway
You can't be reliably caught by a classifier. You can absolutely be clocked by a reader who has seen four hundred of these — the uniform paragraph lengths, the tricolon in every third sentence, the closing paragraph that restates the article. The cost isn't a penalty, it's that a skimming reader decides in six seconds nobody was home when this was written, and leaves.
Treating provenance metadata as somebody else's problem
If you generate synthetic media in or for the EU market, Article 50(2) makes machine-readable marking a shipping requirement — with a backstop of 2 December 2026 for systems already on the market, under the omnibus deal provisionally agreed between Council and Parliament on 7 May 2026. Practically, that means knowing whether your generation pipeline emits C2PA manifests or SynthID watermarks, and whether your CDN, image optimiser or CMS strips them. Most image pipelines strip metadata by default as an optimisation. That default is now a compliance bug.
Two predictions I'm willing to be wrong about
The marker words rotate while the share keeps climbing. Rerun Kobak's excess-vocabulary method on the 2026 PubMed corpus and I expect the assisted-writing lower bound to be materially above 13.5%, while "delve" and "underscore" have flattened and a different style vocabulary has spiked in their place — because the 2024 markers got publicly named, and writers and model post-training both react to being named. The method is published with code, so this is checkable rather than rhetorical. If the same eight words are still the top signal, I'm wrong about how fast the feedback loop runs.
Provenance replaces detection as the operative question, and detection never gets good. By the end of 2027 the practical answer to "was this AI-generated" will come from checking a manifest or a watermark, not from a classifier, and no statistical text detector will have published an independently replicated sub-1% false-positive rate on mixed human-and-AI text. Mixed text is the honest test case, because mixed is what real workflows produce. If a detector vendor publishes that result and someone outside the company reproduces it, I'm wrong.
FAQ
Will AI replace content writers entirely?
No, and the shape of the 2026 evidence says why. The work that got automated is first-draft production. The work that didn't is verification, original reporting, subject-matter judgement and taking responsibility for a claim in public — which the EU AI Act now treats as a named human function under Article 50(4). What is genuinely shrinking is the market for undifferentiated first drafts, which was a real job for a lot of people and is not coming back.
How do you tell if content was AI-written?
Honestly: for a single document, you often can't, and you should be suspicious of anyone selling certainty. Statistical detectors have documented false-positive problems on formal, plain and non-native-English writing, and they degrade badly on lightly edited hybrid text. Our own AI content detector carries that warning in its documentation, because it applies to every detector on the market: treat a verdict as one weak signal, never as the basis for an academic, employment or legal decision. The signals that hold up better are provenance metadata (a C2PA manifest, a SynthID watermark), and old-fashioned verification — checking whether the specifics in the piece are actually true.
Is AI-generated content bad for SEO?
Not per se, and Google's documentation says so directly. What is penalised is producing many pages without adding value for users, which is a scale-and-value test rather than an authorship test. A single well-researched, verified, genuinely useful page that a model helped draft is fine. Four hundred thin pages are a spam-policy problem whether a model or a freelancer produced them.
Is the copyright question settled?
No, and don't plan as if it is. Litigation over training data remains active across multiple jurisdictions, with outcomes pointing in different directions — some rulings have treated training itself as transformative while treating the acquisition of pirated corpora as a separate and expensive problem. If your content strategy depends on a particular legal outcome, that's an unhedged bet on an open question.
Where this leaves you
The useful mental model isn't "AI writes, humans edit." It's that generation became a commodity input, and everything that makes content worth publishing — accuracy, originality, a point of view, someone's name on it — stayed exactly where it was, and got more valuable because the surrounding supply got cheaper. The regulators arrived at the same conclusion from the compliance side, which is a strange kind of validation but a useful one: the law's escape hatch from disclosure is a human who reviewed the thing and is willing to be responsible for it.
That's the whole job now. It was always most of the job.
Cover photo by Ron Lach on Pexels.
