From Comic Panel to Training Data: How African Comic Art Becomes AI’s Raw Material

Imagine an artist in Lagos, Douala, Accra or Nairobi finishing an illustration after several days of work. Perhaps it is a dramatic splash page. Perhaps it is the first public image of a new superhero. Perhaps it is a quiet panel whose architecture, clothing and dialogue place the story unmistakably inside an African city.

She exports the file, writes a caption, and posts it. The purpose is simple: visibility. A publisher might see it. Readers might share it. A client might offer a commission. For a creator working in an industry with limited physical distribution, publishing online is not a lifestyle choice. It is how the work enters the world.

What happens next is harder to see.

A fan may repost the image. A blog may embed it in an article. Another account may remove the caption, crop out the signature, or upload a screenshot. A web crawler may encounter one of those copies. A dataset builder may pair the picture with whatever text sits beside it. A model developer may later download that dataset, convert the image into numbers, and use it to train a generator. Months or years afterwards, somebody types a prompt and receives a new image carrying traces of the original character, composition or visual language.

The artist sees only the first stage and the last. The post she uploaded and the synthetic imitation that eventually comes back.

To follow that journey concretely, this piece follows one illustrative case, assembled from patterns described by artists across the region rather than any single documented incident. Call her Tomiwa, a Lagos-based illustrator self-publishing a superhero webcomic on Instagram. Her panels resurface throughout, because it is easier to trace what happens to a single image than to an entire industry.

Not every illustration posted online enters a training dataset, and not every image generator is built the same way. Some companies train on licensed libraries. Some use public-domain material. Some collect what is available on the open web. Many disclose too little for outsiders to reconstruct the mixture. But the route from a public image to a generative system is not speculation. Significant parts of it have been described by the platforms, dataset organisations and model developers themselves.

The Two Routes from Post to Dataset

The first route may begin and end inside the platform where the artist posts.

Meta has acknowledged that publicly shared Facebook and Instagram posts, including photographs and text, formed part of the information used to train some of its generative AI models. The company states that private posts and private messages between friends and family were not used in the same way.

This means an illustration does not necessarily need to be seized by some unknown external scraper before it becomes training material. The platform already hosting it may be the organisation using it.

That single fact should change how African comic creators read the familiar instruction to post consistently. Public distribution is essential to discovery, but it now performs two functions at once. The image attracts human readers while becoming potential machine-learning material. The artist enters the platform seeking an audience; the platform may interpret that same public act through the far broader permissions written into its terms. The legal permission a company claims and the meaningful consent a creator believes she has given are not the same thing.

Source: Forbes/Getty Images

The second route begins when the image escapes the original post.

African comics already move through digital networks because the formal systems for printing, retail and distribution remain uneven. Platforms such as Comic Republic, Al’Khariqun, Kugali and Zebra Comics (rebranding to ZEBRA) distribute work through websites, apps and digital stores. ZEBRA describes its platform as a way for readers to reach African comics, manga and webtoons on phones, tablets and computers. Kugali offers free and paid digital comics alongside its animation and augmented-reality work. Digital circulation is not an accident of the African comics economy. It is one of the infrastructures holding that economy up.

These are legitimate, artist-serving platforms, named here because they show how central digital distribution has become, not because any of them is accused of mishandling creator work. The risk this piece describes sits downstream of any single platform, in the many hands an image passes through once it leaves.

Circulation also creates copies. The panel Tomiwa uploaded to Instagram may reappear on a fan page, a visual-search platform, an online magazine, a portfolio site, a piracy archive. Each repost can separate the image from its creator. The signature may remain visible, but the name, title, licensing terms and source link tend to vanish. The image that began as “Page 14 of a forthcoming Nigerian graphic novel by a named artist” slowly becomes “African superhero art”, and then simply “fantasy character”.

At that point, the creator has not only lost attribution. The work has acquired a new, machine-readable identity.

Inside the Machine: From Common Crawl to a Generated Image

To see why that matters, start with Common Crawl.

The non-profit maintains an open repository of information gathered from the web, with billions of pages accumulated over years. Its crawlers visit publicly accessible pages and preserve page content, requests and crawl metadata in standard archive formats. It was built to make large-scale web research possible, not to collect African comic art. But material that enters a general web archive can be processed by other organisations for purposes the original uploader never anticipated.

LAION-5B provides a documented example of what can happen next. Released in 2022, its creators processed Common Crawl records, located image links and the alternative text associated with them, and used OpenAI’s CLIP system to measure how closely each image matched its accompanying words (Schuhmann et al., 2022). After filtering, the project released information representing approximately 5.85 billion image-text pairs. While this enormous collection helped researchers reproduce or train influential image-and-language systems, subsequent safety audits and removal actions prompted a major revision, culminating in the 2024 launch of Re-LAION-5B to address concerns around unfiltered subsets.

There is an important technical distinction here. LAION describes its released datasets as indexes: image URLs, captions or alt text, and computed metadata, rather than a warehouse holding every original picture. It downloaded images temporarily to assess and score them, then discarded the files. A developer wanting to train on a LAION subset must reconstruct it by following the URLs and downloading whatever is still reachable.

For Tomiwa, that distinction offers limited comfort. The index makes her image discoverable and gives downstream users a route back to it. The dataset organisation can say it supplied links, not pictures. The web archive can say it collected public pages. The model developer can say it downloaded from a dataset assembled by someone else. The platform can say the work was public.

Every participant describes a different technical action. Together, those actions form one economic chain in which a creator’s labour becomes a useful input.

The text paired with the image matters almost as much as the picture itself. If the original post names the artist and the character, those names can become part of the relationship the model learns. If a repost replaces the caption with “African mythology” or “Afrofuturist warrior”, the individual creator disappears while the cultural category survives. Dialogue written in Yorùbá, Swahili, Amharic or Camfranglais may be misread, ignored, or replaced by an English description supplied by somebody else. The Stable Diffusion v1 model card warned that its training relied mainly on English descriptions, that non-English communities and cultures were consequently represented less adequately, and that Western culture was often treated as the default.

This creates a peculiar danger for African comics. A system can learn to produce the appearance of an African world without learning the names, authorship or cultural reasoning behind it. It may attach a ceremonial object to the wrong community, treat distinct architectural traditions as interchangeable, or collapse several countries into one visual category.

The art survives as pattern. The artist and the context are filtered out. Representation appears to rise while attribution weakens.

Why Circulation Builds the Dataset

It is worth pausing on the mechanism, because it explains why African comics are unusually exposed.

A model does not learn from one image. It learns from repetition across many images. For a comic character, that repetition is built into the art form itself. A character appears across dozens, sometimes hundreds, of panels: the same face at different angles, the same costume under different lighting, the same weapons, the same movement, the same colour palette, the same environment. A webcomic published consistently produces exactly the kind of consistent, captioned, publicly indexed image set that a training pipeline rewards.

In other words, the discipline African creators are told will build their audience — publish often, keep the character consistent, keep the captions descriptive — is the same discipline that makes the work legible to a dataset.

This is why the loss is not only reputational. An African comic character can be the seed of a much larger property. Comic Republic describes its work across digital comics, animation, games and other formats; ZEBRA has pursued licensing and adaptation across publishing, film, television, animation, games and merchandise. A panel is not valuable only as the image on a screen. It may be the first asset in a franchise.

When training-data debates treat such panels as disposable web content, they ignore the economic future the creator is trying to build out of them.

Memory, Style and the Value of a Recognisable Character

During training, a generator does not ordinarily file each source picture into a searchable album. In a latent diffusion system, images are converted into compressed mathematical representations, the accompanying text is encoded, and the model is trained to reverse a process that adds noise to images. Across vast numbers of examples, its internal parameters adjust to associate words with shapes, colours, textures and compositions. When a user later types a prompt, the system starts from noise and builds an image from those learned relationships.

This is why the claim that “the model doesn’t store its images” settles nothing. Copies may still have been made while assembling and processing the data, and models can memorise particular examples. Researchers presenting at the 2023 USENIX Security Symposium extracted more than 1,000 training images from several diffusion models, including photographs and company logos. Stable Diffusion’s own model card also acknowledged some memorisation where duplicated images were present in its training data. The model is not simply a collage machine, but neither is memorisation an imaginary concern.

Exact reproduction is only one possible injury. A model may never regenerate a panel pixel for pixel and still erode the value around it. Comic art is not a collection of unrelated images. Repeated panels build a character’s face, costume, weapons, movement, palette and world, a visual identity readers learn to recognise. If enough public images of the same character circulate, a system becomes better able to produce something that resembles it, or occupies the same commercial space.

What the Law Does and Doesn’t Protect

Nigeria’s Copyright Act 2022 makes one thing clear: publishing an illustration online does not place it in the public domain. Copyright protection does not require registration. The owner of copyright in an artistic work holds the exclusive right to reproduce and publish it, among other protected acts, and unauthorised use of the whole work or a substantial part, including forms recognisably derived from the original, is potential infringement.

Source: Tekedia

The harder question is how those rights apply to machine learning. The Act does not expressly say whether copying protected images into a commercial training dataset is permitted. Its fair-dealing provisions require weighing the purpose and character of the use, the nature of the work, the amount used, and the effect on the potential market or value of the work. Those factors could matter in a dispute, but there is little Nigerian case law explaining how they should apply when millions or billions of works are copied for model training.

It would be premature, then, to claim every training use is automatically unlawful, just as it would be wrong to assume that anything visible online is automatically free to train on.

The Act also restricts the unauthorised removal or alteration of rights-management information where it is done knowingly to facilitate or conceal infringement. That matters here, because attribution and licensing information routinely vanish as an image is reposted. Whether a particular caption, watermark or embedded record meets the statutory definition would depend on the facts, but the law recognises the principle underneath: information linking a work to its owner is part of the protection ecosystem, not decoration.

The greatest practical problem is evidence. A creator may discover her work in a public dataset such as LAION, because public search tools can expose the link. She is far less likely to know whether a closed commercial model used the image, whether a contractor downloaded it, whether a dataset was copied before a removal request, or whether the work survives inside a model that has already been trained.

In one large audit of more than 1,800 text datasets, the Data Provenance Initiative identified serious gaps in licensing and attribution. Although that study concerned text rather than comic art, the warning is relevant: datasets can be repackaged and reused faster than reliable provenance travels with them. If specialists struggle to trace where training material came from, an independent illustrator working with a phone and an inconsistent internet connection has little chance of reconstructing the chain alone.

Existing protections each address a piece of the problem, and none is complete. A website owner can use a robots.txt file to ask identified crawlers to stay away, but the internet standard itself states that these rules are not access authorisation. An artist posting on Instagram does not control Instagram’s robots.txt. Content Credentials can attach verifiable information about an image’s origin and editing history, but provenance does not by itself stop a company collecting the image. Services such as Have I Been Trained can help artists search certain public datasets and register objections, but an opt-out only works where the relevant dataset builder or developer recognises it.

Source: Tech Cult

/Glaze and Nightshade approach from another direction. Rather than relying on a polite request to crawlers, they alter the image before publication so that a model has more difficulty learning the artist’s style, or learns an incorrect association from an unauthorised copy. They create friction. They also add a task: every image must be processed, checked and published in protected form. For African artists already writing, drawing, lettering, marketing, distributing and financing their own work, self-defence becomes one more unpaid production role.

Building a Fairer Chain

The chain does not have to work this way. Adobe states that its current Firefly models were trained on licensed material, including Adobe Stock, alongside public-domain content, and that Stock contributors’ participation is governed by contributor agreements. Although these licensing arrangements, introduced during the commercial rollout in June 2023 and later scrutinised following April 2024 disclosures regarding third-party AI-generated content within the stock pool, have faced criticism from digital creators over consent and data purity, the model proves a broader point: indiscriminate web collection is not a technical necessity. Dataset composition remains fundamentally a business and governance decision.

African comics platforms can make their own decisions before waiting on foreign regulators or technology companies. Contracts, platform hygiene, studio records, and collective infrastructure must all be considered.

Publisher and platform agreements should state clearly whether uploaded artwork may be used for machine learning, whether that permission extends to third parties, and whether creators will be paid. Websites can publish machine-readable restrictions, block known bulk scrapers where feasible, and preserve creator identification when generating thumbnails and promotional copies. Studios can keep secure archives of layered files, sketches, publication dates and agreements that help establish provenance. None of these measures is perfect, but together they make authorship more difficult to erase.

Furthermore, a searchable registry of African comics, characters, artists and authorised licences could help creators establish ownership and give responsible AI developers somewhere to ask permission. Publishers could build a consent-based training library in which creators opt in, specify permitted uses and receive compensation. Universities could work with comic organisations to test detection and protection tools against African illustration styles, languages and internet conditions. The Nigerian Copyright Commission and its equivalent across Africa could issue specific guidance on commercial text-and-data mining rather than leaving individual creators to interpret a general statute against some of the world’s largest companies.

ZEBRA has already drawn one boundary in public: it says its published comics remain human-created, while AI is used in supporting areas such as research, analytics and administration. That position recognises the real choice, not between rejecting every form of AI and surrendering the creative process, but between using technology and letting it price you out of your own work.

The burden, though, cannot rest entirely on the artist. It is unreasonable to tell creators they must become copyright lawyers, cybersecurity specialists and machine-learning auditors before they can safely show anything. Platforms know what their systems collect. Dataset builders know which pages and links they process. Model developers know which datasets they acquire. Those organisations are better placed to document the chain, honour refusals and negotiate licences.

African comics need visibility. Their creators need readers, commissions, publishers, adaptation deals and international discovery. The answer cannot be to keep every panel offline until it is commercially successful, because being seen is often what makes success possible.

But visibility should not be treated as the surrender of ownership.

Ad: This story is sponsored by Fiztech Ironclad

Why “Poisoning” Is Only Half the Story

There is a temptation, in conversations like this one, to reach for the language of sabotage. If the machines are coming, the logic runs, then artists should poison the well, feed the crawlers corrupted images, teach the models the wrong thing, make the theft unprofitable.

The instinct is understandable, and the tools technically exist. But it misdiagnoses the problem, and it places the burden in the wrong place.

Poisoning assumes the artist knows which images are being taken, when, and by whom. In practice she doesn’t. The whole difficulty described in this piece is that the chain is opaque: she can see the post and she can see the imitation, and almost nothing in between. A defence you can only deploy against an enemy you can see is not much of a defence against one you can’t.

It also quietly concedes the argument. If the price of showing your work is that every panel must be armour-plated before publication, then the cost of visibility has been shifted onto the creator, the same shift that runs through the entire story. The question is not how artists should fight the scrapers. It is why the systems that do the scraping are permitted to make the fight necessary.

The real problem is not that an illustration travels. Art has always travelled, through readers, exhibitions, print, adaptation, influence. The problem is that a comic panel can now be copied, relabelled, indexed, downloaded, processed and monetised through a chain that remains almost completely invisible to the person who made it.

A fair creative ecosystem would let Tomiwa see that chain, and refuse it, or set a price for it. Until that exists, the journey from comic panel to training dataset will remain less like discovery and more like disappearance.

Sources

  1. Meta Platforms — disclosure that publicly shared Facebook and Instagram posts were used to train generative AI models (Meta AI blog / privacy updates).
  2. Common Crawl Foundation — documentation on crawl scope, formats and archive structure. https://commoncrawl.org
  3. LAIONLAION-5B: Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., & Jitsev, J. (2022). LAION-5B: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35, 25278–25294. https://openreview.net/forum?id=M3Y74vmsMcY
  4. Radford et al. (OpenAI)Learning Transferable Visual Models From Natural Language Supervision (CLIP), 2021.
  5. Diffusion-model memorisation study — Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tramèr, F., Balle, B., Ippolito, D., & Wallace, E. (2023). Extracting training data from diffusion models. arXiv. https://doi.org/10.48550/arxiv.2301.13188
  6. Data Provenance Initiative — Longpre, S., Mahari, R., Chen, A., Obeng-Marnu, N., Sileo, D., Brannon, W., Muennighoff, N., Khazam, N., Kabbara, J., Perisetla, K., Wu, X., Shippole, E., Bollacker, K., Wu, T., Villa, L., Pentland, S., & Hooker, S. (2024). A large-scale audit of dataset licensing and attribution in AI. Nature Machine Intelligence, 6(8), 975–987. https://doi.org/10.1038/s42256-024-00878-8 Cited by: 295
  7. Nigeria Copyright Act 2022 — Federal Republic of Nigeria. https://www.copyright.gov.ng
  8. Adobe Firefly — Adobe Inc. (2023, June 8). Adobe brings Firefly and Express to enterprises. Adobe Newsroom. https://news.adobe.com/news/news-details/2023/adobe-brings-firefly-and-express-to-enterprises
  9. Have I Been Trainedhttps://haveibeentrained.com
  10. Glaze — Shan, Cryan, Wenger, et al., University of Chicago. https://glaze.cs.uchicago.edu
  11. Nightshade — Shan, Ding, Passananti, et al., University of Chicago. https://nightshade.cs.uchicago.edu
  12. Content Credentials / C2PA — Coalition for Content Provenance and Authenticity. https://c2pa.org
  13. Robots Exclusion Protocol — RFC 9309, IETF.
  14. Comic Republichttps://comicrepublic.com
  15. Kugali Mediahttps://kugali.com
  16. Zebra Comics / ZEBRA — platform and licensing statements.

Written by Seyi Adedokun

Edited by Mujeeb Jummah

——————————————————–

——————————————————–

AI Use at TheACE
TheACE uses artificial intelligence tools to support research, drafting and analysis across Africa’s creative industries. All content is verified, edited and approved by our human editorial team to ensure accuracy, clarity and responsible storytelling. AI assists our work; it does not replace human judgment.

Share Post:

Join the Empire

Get the stories, insights, and behind-the-scenes knowledge you won’t find anywhere else, delivered straight to your inbox