What is generative AI: generative AI explained for a claims and counter-fraud audience
Generative AI is a class of machine learning model that produces new content on request, such as text, images, video, audio or documents, by synthesising statistical patterns learned from large volumes of existing material. Nothing is retrieved from a library. Every output is assembled.
Strip out the marketing and the mechanism is fairly plain. The model predicts what plausibly comes next, one token or one region of pixels at a time, conditioned on whatever prompt it was given. That single behaviour explains both the usefulness and the risk, because plausible and true are different tests.
What does generative AI mean for claim evidence?
Three properties change the shape of the problem for anyone assessing a file.
- There is no original: a generated image has no scene, no camera and no moment behind it. When you ask where a file came from, there is nothing to go back to.
- The marginal cost is near zero: the tenth fabricated invoice costs what the first one cost. Volume attacks become affordable in a way that hand editing never was.
- The skill floor collapsed: the tooling is consumer grade and sits in ordinary app stores. Capability that used to need a specialist now needs a sentence.
What has not changed is the underlying claim behaviour. Exaggeration, staged loss and invented ownership all predate the technology. Generative tools make the supporting material easier to produce, which is a different thing from inventing a new category of fraud.
Generative AI, deepfakes and shallowfakes are not interchangeable
The terms get used loosely, and the distinction matters when you write up a case. Generative AI is the underlying technology. A deepfake is one product of it, synthetic media depicting something that never happened. A shallowfake is a genuine file that somebody altered with an everyday photo editor, and no model was involved at all.
Most manipulated evidence reaching insurers is still the third kind. The edited real photo remains the high-volume submission because it always was the easiest one to make.
Generative AI explained: a worked example
A theft claim arrives with a clear photograph of a high-value watch on a kitchen worktop and a matching retailer receipt, both as email attachments. Both were produced from a prompt in the same sitting, with a plausible date and a plausible total.
The handler cannot ask the file where it came from, because it never came from anywhere. The only durable questions left are about provenance: how did this material reach us, through what route, and can that route be checked later.
What the reported figures actually show
Aviva reported more than 18,400 suspect claims worth 233 million pounds across its brands in 2025, equivalent to around 638,000 pounds detected per day, and said a growing number were supported by AI-generated images and manipulated documents, mostly in motor. The value of detected motor fraud rose 39 percent that year.
Read that as a detection figure. A rise in detected fraud reflects more fraud and better detection at the same time, and no public number tells you the size of what nobody caught. The wider picture sits in AI-generated insurance fraud.
Does the law require generative output to be labelled?
In part. Article 50(2) of the EU AI Act requires providers of AI systems that generate synthetic audio, image, video or text to mark the output in a machine-readable format so it is detectable as artificially generated or manipulated. That obligation applies from 2 August 2026, with a limited grace period to 2 December 2026 for systems already on the market before that date.
Useful, but do not build a control on it. A marking obligation binds providers in scope. It does not bind a model run locally, one built outside the scope, or a file stripped of its metadata on the way to you. The absence of a watermark proves nothing about a file, and its presence only tells you what a compliant tool chose to declare.
Where this leaves a counter-fraud function
- Detection scores are inputs: classifiers return probabilities that shift every time generation models move. A score earns a closer look. It does not settle a claim.
- Provenance outlasts appearance: asking how a file reached you is a more durable question than asking whether it looks generated, because generation quality improves every year and the question of origin does not get easier to dodge.
- Reduce unverifiable intake: material recorded inside a controlled capture route, with a server-side receipt time and a cryptographic hash, gives a reviewer something to check months later. An emailed attachment gives them a file and a hope.
- Keep the human central: technical signals such as an unusual device history or a virtual camera present during capture are reasons to look, never verdicts. Genuine claimants trip signals for ordinary reasons.
One caveat worth repeating internally: a receipt timestamp proves when your organisation received a submission. It says nothing about when the loss occurred. Overstating that is a reliable way to lose an argument you should have won.