Identifiable, grounded, graded: three gates before generative output publishes itself
Published · Updated
At 8 articles a day in 7 languages, no human reads every asset. Three gates decide what ships: identifiability on the asset, automated validation of key claims, and evals that mirror the real task distribution.
On one platform we publish 8 SEO articles every day across 7 languages and 4 separately indexed country domains, with 0 manual intervention, alongside multi-language generative video campaigns. At that rate nobody reads every asset before it goes live. The useful question stops being whether the model writes well and becomes much narrower: what has to be true about a generated asset before an unsupervised pipeline is allowed to publish it?
Identifiability became a shipping requirement
The AI Act entered into force on 1 August 2024 and became applicable on 2 August 2026, with some exceptions. Two provisions bear directly on generative pipelines. When using AI systems such as chatbots, humans should be made aware that they are interacting with a machine so they can take an informed decision. And providers of generative AI have to ensure that AI-generated content is identifiable. Starting on 2 December 2027, high-risk AI systems face strict obligations before they can be put on the market.
Read as an engineering constraint rather than a legal footnote, identifiability is a property of the output path. It has to be decided where assets are generated and stored, because a pipeline running with 0 manual intervention has no human standing at the end to attach a label. If the marker is not produced by the same step that produces the article or the video, it will not exist on the asset that reaches a reader.
Provenance records how, not whether
At the heart of the C2PA specification is the Content Credential, a cryptographically bound structure that records an asset's provenance. Content Credentials, also known as a C2PA Manifest, contain one or more assertions: statements about the asset such as its origin, meaning when and where it was created, its modifications, meaning what happened using what tools, and its use of AI, meaning how it was authored. For generative video campaigns, that is an audit trail attached to the file itself.
The specification is explicit about what it does not do. Content Credentials do not provide value judgements about whether a given set of provenance data is true, and the approach is not a cure-all for misinformation, but instead seeks to mitigate against its threats in the digital domain. So provenance answers how an asset was made. It says nothing about whether the claims inside the article are correct. That is a second gate, not the same one.
The gate provenance cannot be
OWASP's Gen AI work names the failure precisely. Misinformation occurs when LLMs produce false or misleading information that appears credible. Overreliance occurs when users place excessive trust in LLM-generated content, failing to verify its accuracy. In a daily autoblog driven by search trends, the person doing the over-trusting is the reader, who has no view of the pipeline and no reason to suspect a confident paragraph of being invented.
Two mitigations map onto a publishing pipeline without much argument. Use retrieval-augmented generation to improve reliability by retrieving relevant and verified information from trusted external databases during response generation. And implement tools and processes to automatically validate key outputs, especially output from high-stakes environments. The operative word is automatically. At 8 articles a day in 7 languages, a validation step that depends on someone being available is a validation step that does not run.
Evals sized to the real distribution
Anthropic's guidance on building evaluations is blunt about where test sets go wrong: design evals that mirror your real-world task distribution, and structure questions to allow for automated grading, for example multiple-choice, string match, code-graded or LLM-graded. For an estate of 7 languages across 4 country domains, the real distribution is multilingual and per-market. An English-only eval set measures a task the pipeline does not actually perform.
The same guidance pushes against the instinct to hand-grade: more questions with slightly lower-signal automated grading is better than fewer questions with high-quality human hand-graded evals. And most use cases need multidimensional evaluation along several success criteria. For generated publishing, the criteria separate cleanly: are the claims grounded in retrieved sources, is the language correct for the market it is served to, and is the AI marker present on the asset.
The practical shape is three gates wired into the generation step rather than bolted on afterwards: identifiability attached where the asset is produced, automated validation on the claims that matter, and an eval set that mirrors the real task distribution and grades along several criteria at once. None of that requires a larger model or a longer prompt. It requires deciding what a pipeline is permitted to publish while nobody is watching, and encoding that decision in code.
Sources
European Commission (Shaping Europe’s digital future) — AI Act — https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
OWASP Gen AI Security Project — LLM09:2025 Misinformation — https://genai.owasp.org/llmrisk/llm092025-misinformation/
Claude Platform Docs (Anthropic) — Define success criteria and build evaluations — https://platform.claude.com/docs/en/test-and-evaluate/develop-tests
C2PA — C2PA and Content Credentials Explainer (2.4) — https://spec.c2pa.org/specifications/specifications/2.4/explainer/Explainer.html
Neurolinks case study — Four markets, one codebase — https://neurolinks.be/work/tatano-energy
Working on a project where these methods apply?