Skip to content
→ All work

Case study 24 / 26

AI Slop Detector

An end-to-end AI-writing detector: scraped data, synthetic generation, a LoRA-tuned BERT classifier, and an honest benchmark.

Collaborative project — the dataset and model weights are published on Hugging Face by a co-author.

Status
Research
Period
2026
Domain
ai
Language
Python
Last push
01 APR 2026
License
MIT
Source of claims
README.md, training/src/train.py, raid/src/inference.py, run_detector.py

Human paragraphs scraped from Wikipedia are paired with GPT-5 Nano rewrites to build a balanced dataset; a LoRA adapter on bert-base-cased learns to separate them, and the model is evaluated on the external RAID benchmark to measure where it generalises — and where it does not.

01/The problem

Detecting machine-generated text needs paired data that differs only in authorship. Most public detectors are black boxes with unpublished limits.

02/The system

Human paragraphs scraped from Wikipedia are paired with GPT-5 Nano rewrites to build a balanced dataset; a LoRA adapter on bert-base-cased learns to separate them, and the model is evaluated on the external RAID benchmark to measure where it generalises — and where it does not.

Inference path5 components
  • Input to Tokenizer
  • Tokenizer to BERT + LoRA
  • BERT + LoRA to Softmax
  • Softmax to Result

03/Implementation

  1. 01Crawled Wikipedia and stored two random paragraphs per page — 10,001 human-written paragraphs.
  2. 02Generated AI counterparts in two passes with separate contexts (summarise, then rewrite) using GPT-5 Nano through the Batch API.
  3. 03Fine-tuned bert-base-cased with a LoRA adapter (r = 16, rank-stabilised, all linear layers) and a classification head; 95/5 train/validation split, 5 epochs, fp16, gradient accumulation of 8.
  4. 04Inference returns P(AI) from a two-class softmax; text above 0.5 is labelled AI-generated. The benchmark path scores each paragraph separately.
  5. 05Evaluated on the RAID benchmark training set by domain and generator family, reporting TPR at 1% FPR.

04/Engineering

Two-pass generation

Summarising and rewriting in separate contexts stops the model from paraphrasing the source sentence by sentence, which would make the task artificially easy.

Parameter-efficient tuning

LoRA trains a small adapter instead of the full network, so the published artefact is light and the base model stays untouched.

Measure transfer, not just accuracy

In-distribution validation accuracy is reported alongside an external benchmark, so the model's limits are explicit.

05/Interface

No product screenshots are published for this project. The visual above is a code-driven representation of how it behaves, built from the repository source — not a screenshot.

06/Tech stack

  • Python
  • PyTorch
  • Hugging Face Transformers
  • PEFT / LoRA
  • BERT
  • OpenAI Batch API
  • uv

07/Result

Verified outcomes

  • 96% validation accuracy on the held-out split (as reported in the repository).
  • On RAID, better than random across all generator families and most domains.
  • TPR at 1% FPR: about 90% on reviews; 30–40% on Wikipedia, arXiv abstracts and news; 1–20% on books, recipes, poetry and Reddit.

Known limitations

  • Trained only on Wikipedia-style text generated by a single model.
  • 512-token context, one paragraph at a time.

08/Links

Next case study

ML Trading Bot →