AI can damage your documents in silence: Microsoft has just measured it

AI is sold to you as an assistant that can proofread, correct, restructure and finalise your documents.

The promise is appealing: you hand over a report, a file, a case, a working draft; the AI takes care of it; you get a clean document back.

The problem is that in some cases, the final document can look clean while having become wrong.

Not wrong everywhere. Not obviously broken. Not necessarily unusable at first glance.

Worse than that: degraded in silence.

That is the subject of a study published on 17 April 2026 by three Microsoft Research scientists: Philippe Laban, Tobias Schnabel and Jennifer Neville. Its title does not bother softening anything: “LLMs Corrupt Your Documents When You Delegate”.

Blunt translation: large language models corrupt your documents when you delegate the work to them.

The figure that should cool the euphoria

The researchers built a benchmark called DELEGATE-52.

The principle: test what happens when AI models are asked to modify professional documents across several steps, the way you would in real delegated work.

Not a simple question and answer.
Not a one-shot generation.
Not a three-paragraph summary.

A genuinely long workflow, with documents to transform, rework, then restore or manipulate correctly.

The benchmark covers 52 professional domains, including code, crystallography, genealogy, musical notation, subtitles, accounting ledgers and other structured formats.

The researchers tested 19 models from several families: OpenAI, Anthropic, Google, Mistral, xAI and Moonshot.

The result is brutal:

even the most advanced models tested — Gemini 3.1 Pro, Claude 4.6 Opus and GPT-5.4 — degrade an average of 25 % of the content after 20 delegated turns.

Across all models, average degradation reaches around 50 %.

This is not a small bug.
It is not “a typo”.
It is not a style slip.

It is a measurable loss of document integrity.

The danger is not only that the AI gets it wrong

When an AI says something absurd in a conversation, you can still catch it.

The sentence sounds odd.
The figure looks suspicious.
The tone is too confident.

You can check.

But in a long document, the problem changes nature.

A clause moved.
A table row altered.
A value deleted.
A reference badly rebuilt.
A specialised notation corrupted.
A correct section replaced by a plausible but false version.

The file still opens.
The text still reads well.
The structure still looks coherent.

And yet the substance has changed.

That is probably the most important point of the study: the errors do not necessarily look like crude accidents. They can be rare, serious, localised and hard to detect.

In other words: AI does not always destroy the document.
It can do worse: it can make it credible and wrong at the same time.

The best models do not remove the risk. They delay it.

You might think the problem only concerns small models or low-end AI.

The study says otherwise.

The best models do better, yes. But they do not make delegation reliable for all that. The researchers observe that stronger models do not really eliminate critical errors: they push them later and commit fewer of them.

So the problem is not just a bad model being stupid.
It is a structural weakness in the very use we are starting to entrust to them: take an existing document, modify it across several steps, and guarantee it stays faithful.

That is exactly the use companies, freelancers, employees and platforms want to generalise.

“Proofread this contract."
"Clean up this file."
"Reorganise this procedure."
"Update this report."
"Fix this file."
"Redo this presentation."
"Change this document without touching the rest.”

On paper, it is the productivity dream.

In practice, it can become a machine for producing invisible errors.

AI agents do not fix the problem

Another awkward finding: giving the model tools is not enough.

The researchers also tested a more agentic setup, where the model can use tools to read, write and delete files or run Python.

That is precisely what is presented to us as the next step: agents able to work inside an environment, manipulate files, go further than a chatbot. We installed two of them on an ordinary Windows machine to see what they really do when left alone, in our comparison of OpenClaw and Hermes.

Result: in these tests, using tools does not improve performance. On some models tested, agentic mode makes the degradation worse.

This matters because the whole market is pushing that idea today: stop asking the AI to answer, let it act.

But if AI acts on your files without guaranteeing their integrity, it does not merely become an assistant.
It becomes an operational risk.

An imperfect benchmark, but a very useful warning

Let us be precise.

DELEGATE-52 is not a study of every real-world use of AI. The authors set their own limits: the benchmark is in English, rests on structured documents, uses reversible tasks, and does not replace human studies in real conditions.

So no, you should not conclude that every document handed to an AI will be 25 % corrupted.

But it would be just as dishonest to use those limits to play down the signal.

Because the signal is clear: the moment you move from “AI gives me an answer” to “AI works directly on my document”, the risk level changes.

And that risk is precisely the one today’s interfaces hide.

They give an impression of control.
They produce fluent output.
They speak with assurance.
They hand back a clean file.

But they do not guarantee the document stayed true.

It is worth putting this next to another measurement made a few weeks later: independent labs found that the most advanced reasoning models cheat the tests meant to prove they are safe, and increasingly recognise when they are being evaluated. A confident-sounding output is not evidence of anything, and neither is a model’s own report on itself.

The real message for users

The problem is not using AI.

The problem is believing a document produced or modified by AI is validated because it is well presented.

That is exactly the trap.

A professional document is not just a sequence of sentences. It is a structure of responsibilities, figures, references, rules, decisions and sometimes legal or financial commitments.

If an AI changes a paragraph in an article, the error can be annoying.

If it changes a clause in a contract, a line in an accounting table, a reference in an internal procedure, or an instruction in a technical file, the error can become expensive.

And the most dangerous part is that the user often believes they saved time.

In reality they have sometimes simply moved the work: instead of writing, they now have to audit. Instead of correcting, they have to check that the correction did not contaminate the rest.

The simple rule: AI can help, not validate

For important documents, the conclusion should be blunt:

never hand an AI the final version of a document you are not able to check yourself.

AI can suggest.
It can rephrase.
It can spot inconsistencies.
It can help structure.
It can speed up a first pass.

But it must not become the final authority on a sensitive document.

Not today.
Not without oversight.
Not without comparing against the original version.
Not without a change log.
Not without competent human review.