
As AI-generated content becomes more common across social platforms, the way it is labelled has taken on urgent importance. Driven by impending regulatory mandates like the EU AI Act, which requires clear disclosure of synthetic content starting 2 August 2026, social media platforms have begun to implement standardised tagging systems. These frameworks are designed to inject transparency into the media ecosystem, operating on the assumption that a visible warning will naturally cultivate user trust and mitigate the risk of digital deception.

Source: ChatGPT/ EU
Labels are meant to help users recognise when they are engaging with synthetic or altered media, but when applied inaccurately or without context, they risk becoming a new source of misinformation themselves. A team of researchers from the CISPA Helmholtz Centre for Information Security, Ruhr University Bochum (RUB), and the Max Planck Institute looked into how well AI labels work. They surveyed 1,300 people to see how they reacted to these labels in reducing misinformation about AI-generated content. They discovered that while AI labels can help identify content created by AI, there’s a significant concern about the potential for incorrect or missing labels, which could weaken trust in the whole system. This tension between transparency and misrepresentation makes labelling central to the broader challenge of synthetic media literacy.
How are Social Media Platforms Approaching AI Labels?
Social media platforms generally disclose AI-generated or AI-modified content in two ways: creator-initiated (voluntary) disclosure, where users disclose their own use of AI, and platform-initiated (automated) disclosure, where platforms apply labels after detecting technical indicators such as embedded metadata or Content Credentials. These disclosures are typically presented through text labels, icons or badges, and expandable information panels that provide additional context.
Creator-Initiated (Voluntary) Disclosure
This approach allows creators to voluntarily disclose when their content incorporates AI-generated or AI-modified material. While the platform provides the disclosure mechanism, the label is applied based on the user’s declaration. Platforms such as TikTok and YouTube enable creators to identify content containing synthetic or significantly altered media during the upload process.
Platform-Initiated (Automated) Disclosure
Unlike voluntary disclosure, automated disclosure occurs when a platform independently detects indicators that content has been created or modified using AI and applies a label without requiring any action from the creator. This detection may rely on embedded metadata, Content Credentials, or other technical signals. For example, Meta automatically applies its “AI Info” label when qualifying AI indicators are detected.

Icons and Badges
Many platforms use visual icons or badges to communicate AI disclosures at a glance. Rather than serving as the disclosure itself, these icons act as visual indicators that AI-related information is available. For example, LinkedIn displays a “CR” (Content Credentials) badge on images carrying verified Content Credentials, while Meta uses an “AI Info” icon to indicate that additional information about AI use is available.
Contextual Information
Beyond the initial label or icon, some platforms provide users with additional information explaining why content has been labelled and, where available, how it was created or modified. On LinkedIn, selecting the Content Credentials badge reveals details such as whether generative AI was used, the software or camera involved, the creation date, and the organisation that issued the credentials. Similarly, Meta’s “AI Info” panel provides further context about why content was labelled, including whether it was created or only modified using AI tools
Possible Risks Associated with these AI Labels
While platforms have taken steps to proactively flag AI-generated content, the way these labels are applied is not always straightforward. One challenge lies in how broadly these categories are defined. A fully synthetic deepfake created from scratch is not the same as a photograph with minor AI-based edits, yet both may be grouped under the same “AI-generated” tag. This lack of nuance can make it harder for audiences to grasp the actual degree of manipulation involved. Instead of clarifying, the label may flatten important distinctions.
In some cases, authentic or lightly edited media also ends up being flagged as fully AI-generated. When that happens, the label itself risks becoming misleading, raising concerns that original content is being unfairly co-opted or its credibility undermined. Rather than improving transparency, inaccurate or overly broad labels can distort how content is interpreted before users have engaged with it.
Research by CISPA researchers suggests that labeling does not “simply increase truthfulness, but instead changes how people evaluate information.” Labels act as cognitive shortcuts, guiding attention and shaping trust, often more strongly than the content itself. The researchers found that labels make people more cautious, but not necessarily more accurate in their judgments. This means that when labels are inaccurate or lack nuance, they can shape perceptions in ways that do not reflect the content itself.


The implications extend beyond individual pieces of content. For creators, inaccurate labels may cast doubt on genuine work, potentially eroding trust in their output. For audiences, the blurred line between “AI-generated” and “AI-altered” may distort how media is interpreted, leaving users unsure about what they are actually seeing. Over time, repeated inaccuracies could weaken confidence not only in AI labeling systems but also in the authenticity of digital content more broadly.
Building AI Literacy Skills
While AI labels provide some guidance, they do have their limitations. As this article has shown, inaccurate or overly broad labels can sometimes create confusion rather than clarity, particularly when they fail to distinguish between different levels of AI involvement. Developing AI literacy is therefore essential because users need to understand what AI labels mean, what they don’t mean, and how to interpret them alongside other signals such as provenance, source credibility and contextual information.
Recognizing AI-generated media requires specific skills, such as noticing unnatural movements and inconsistencies in visual content, as well as assessing the reliability of the account sharing it. By honing these skills, users can make more informed judgments about the content they encounter. Structured literacy education has an important role to play in helping people understand and critically evaluate AI-generated and AI-modified media.
One interesting approach is gamification. Gamified learning transforms the user from a passive recipient of truth into an active investigator. By simulating real-world scenarios, game-based frameworks leverage “error-driven learning,” allowing users to misidentify manipulated media in a risk-free environment.
The Dubawa AI game, available through the Dubawa WhatsApp Chatbot, is a good example of this. By playing, users practice spotting signs of manipulation and learn to distinguish between authentic, altered and fully AI-generated images. Because it is interactive, the game does not simply tell people what to look for; it allows them to experience the process of discovery themselves.
Gamification is not the only option. Short tutorials embedded into social platforms, community fact-checking exercises and school-based AI literacy modules could all complement this kind of learning. Together, these approaches point to a broader ecosystem of transparency and accountability, which is becoming increasingly necessary as social media platforms continue to play a central role in the creation and spread of both information and misinformation.

Beyond the game, the Dubawa Chatbot provides a broader literacy ecosystem by allowing users to verify claims against a database of fact-checks from the International Fact-Checking Network (IFCN) and report suspicious media content to fact-checkers, serving as an early warning system for emerging false narratives.
Recommendations for Improvement
For AI content labeling on social platforms to go beyond being just a tag, labels alone are not enough. Labels should be accurate, provide meaningful context and help users understand the extent of AI involvement rather than simply signalling that AI was used. To improve transparency and foster user trust, platforms should consider the following enhancements:
- Improve Label Accuracy and Granularity: Distinguish between fully AI-generated content, AI-assisted content and content with minor AI-based edits to reduce over-labelling and minimise the risk of authentic content being misrepresented.
- Implement Feedback Mechanisms: Establish systems for users to report labeling inaccuracies, ensuring a continuous process of improvement.
- Provide Richer Context: Use AI labels, icons and contextual information panels to explain why content has been labelled, how AI was used and, where available, provide provenance information such as Content Credentials.
- Collaborate Across the Information Ecosystem: Partner with creators, influencers, fact-checking organisations, media literacy advocates and accountability CSOs to promote awareness and support the effective implementation of AI disclosure standards.
- Integrate AI Literacy into Labeling Systems: Rather than functioning solely as warnings, AI labels should also serve as educational tools. Platforms can leverage contextual information linked to AI icons and labels to help users understand synthetic media, transforming disclosures into opportunities for learning while reinforcing transparency and trust.

