Sebastian Raschka published a tutorial on building an AI text detector from scratch
On the Ahead of AI blog, Sebastian Raschka described how to build an AI text detector from scratch using a fine-tuned DistilBERT model. The educational project also trains a small language model to bypass detection and highlights the risk of false positives.
Sebastian Raschka, author of the Ahead of AI blog, published a tutorial describing how to build an AI-generated text detector from scratch. According to the author, the motivation was a recently launched AI detection feature in the Substack user interface and recurring reader questions about interesting local projects demonstrating the capabilities of small language models (SLM). The tutorial aims to explain how AI detectors work while also showing a more general approach to building a scorer or verifier for LLM applications beyond common tasks such as mathematics and code.
The detector is designed to return a score from 0 to 100 expressing the estimated probability that a text was generated by AI, with this being an estimate from a classifier trained on a specific data distribution, rather than a general probability of authorship. The implementation relies on fine-tuning a DistilBERT classifier. According to the author, the methodology is inspired by the approach of Pangram models, which, according to available information, are said to underpin the AI detection feature in Substack.
The project also includes using the detector as a verifier to train a small language model whose task is to produce text that bypasses detection. The author uses this to illustrate the nature of AI text detection as a continuous game of cat and mouse between new detection techniques and new models that learn to evade them. He also points out that detectors are prone to false positives, meaning that human-written text can be flagged as AI-generated.
The tutorial is intended to result in a working detector API that can be used by both people and agents, along with a user interface for local deployment. The source article describes the details of implementation, training and deployment.
Why it matters
The tutorial clearly shows how AI text detection tools work, as they are increasingly used for content moderation (e.g. on Substack) and for assessing the authenticity of writing. It also demonstrates their vulnerability to specifically trained models and the risk of falsely flagging human-written text, which is relevant both to authors using text editing tools and to platform operators relying on such detection.
Two audiences, two different impacts
What this means
For individuals
Anyone who writes text and uses tools such as ChatGPT to edit it risks having the resulting excessively edited text classified as AI-generated and flagged as spam, despite having written it themselves.
For a business
The tutorial shows that AI text detectors of the type used by Substack can, in principle, be bypassed by a specifically trained model, which is relevant to companies relying on such solutions for content moderation or compliance.
Risks and complianceCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.