The platform Cara faced three scraping attacks for AI training; the person behind the first offered to help protect artists
The platform Cara faced three scraping attacks in August 2026 (12 million works/12 TB, 8.5 million links, 123 thousand images). After apologizing, the person behind the first attack decided to collaborate on an open-source tool to protect artists.
The platform for artists Cara, which positions itself as a safer alternative to networks like Instagram and offers protective features including Glaze, a tool against AI imitation of artistic styles, faced a series of three large-scale scraping attacks in August 2026. According to the Wired report, on 13 August, the user MandarinDawnPoppy994 began by posting a 12-terabyte archive containing approximately 12 million works on Reddit (the subreddit r/DefendingAIArt) — practically the entire publicly accessible library of Cara. By his own account, downloading the data cost him less than 10 USD; Zhang, the founder of Cara, learned about the attack from users who alerted her to the post.
After the first archive was published, two more attacks followed, which Zhang describes as “copies” inspired by the first. The user CaptiveDreamer downloaded 8.5 million links to works along with metadata (usernames, titles, tags) and uploaded them to Hugging Face. Although Hugging Face agreed to delete personal metadata after a wave of removal requests, it refused to remove the links to the works themselves on the grounds that it hosts no copies of the works and that the links point to the originals published on Cara. The third attack took place on 22 August, when an unknown attacker published 123 000 images along with text posts and user profiles containing personal data on Academic Torrents.
In response, Zhang launched a GoFundMe fundraiser with a target of 120 000 USD for legal fees associated with defending against scraping; according to Wired, it had raised over 100 000 USD by the Thursday before the article was published. Cara introduced temporary measures such as a login gate, but Zhang considers them only a partial solution to a long-term problem affecting the entire internet — she points out that Cara had probably been scraped before and that larger platforms are scraped even more, so leaving Cara does not in itself guarantee safety.
An unexpected development is that the person behind the first attack, who goes by the nickname “Heft” (a student in North America with a background in software and an interest in digital archiving), regretted his actions after being confronted by Zhang, deleted the dataset and apologized. By his own account, he originally intended the attack only as a technical project, but decided to publish it as “ragebait” on Reddit, underestimating its impact on the artist community. He has now decided to work with Zhang on developing a new open-source tool to protect creators.
Why it matters
The case shows that even a platform built with copyright protection in mind cannot effectively prevent mass scraping and that large data hosts like Hugging Face are not always willing to remove links to scraped content. For content creators, the practical consequence is that moving to a “safer” platform does not reduce the risk to zero, while for the platforms themselves, it represents a concrete financial and legal burden that they must address through ad hoc measures.
Two audiences, two different impacts
What this means
For individuals
Creators publishing their work on platforms like Cara have no reliable protection against their work being downloaded en masse for AI training, even on sites that position themselves as a safer alternative to Instagram — switching to another platform does not eliminate the risk of scraping.
For a business
A smaller platform like Cara faces high legal and operating costs associated with defending against scraping (increased server fees, a target of 120 000 USD for legal assistance) and encounters limits to enforceability — Hugging Face refused to delete links to scraped data on the grounds that it does not host any copies itself.
Risks and complianceCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.