Building a free plagiarism checker on 214M open documents
Announcements · 6 min read · By Bibek Howlader, Founder, Labrix
If your university doesn't give you Turnitin or iThenticate, checking a paper before submission usually means paying per check or not checking at all. We wanted a checker any researcher could run on a draft for free and actually trust. So we built one, together with the document collection behind it.
This post covers what's in the collection, how the matching works, what our own tests show, and what the checker still can't see. You should know the limits before you rely on a score.
What we check against
The checker looks for your text in 214,051,046 documents. Every one of them is either openly licensed or public domain, or we store only a fingerprint of it.
| Source | Documents |
|---|---|
| OpenAlex abstracts (scholarly works across every field) | 155.2 million |
| Wikipedia in 334 languages other than English | 44.8 million |
| English Wikipedia | 6.9 million |
| PubMed Central full-text papers (open licences) | 5.3 million |
| PubMed Central author manuscripts | 974,372 |
| Open-access computer science and AI papers, 2025–2026 | 379,675 |
| Full texts hosted on Zenodo, 2025–2026 | 224,695 |
| ACL Anthology papers | 84,030 |
| Project Gutenberg books | 77,127 |
| ICLR accepted papers | 8,847 |
Two more things get searched on top of the collection:
- The live web. A Free check runs about one web search for every 120 words of your text, up to 40 per check. Paid plans search every sentence, up to 400 searches per check (about a 30-page paper).
- Your own library. Papers you've saved in Labrix are checked against your drafts only. Nobody else's check ever sees them.
What we don't check against
This is the part most checkers leave out.
- Paywalled full text. We don't store papers that publishers keep behind a paywall. For many of them we have the abstract through OpenAlex, but not the body. Some publishers have also withdrawn their abstracts from OpenAlex, so even abstract coverage has gaps.
- Text whose licence doesn't allow reuse. Labrix is a commercial product, so non-commercial licences (CC BY-NC) are left out. That's about a million PubMed Central papers. So are most arXiv preprints, whose default licence doesn't cover this kind of use.
- Other people's submissions. We don't keep what you check, so we can't catch a copy of another student's unpublished thesis, unless it's online somewhere. Institutional checkers with a submission repository can, and that's a real gap.
- Heavy paraphrasing. The matching finds copied and lightly edited text (see below). A passage rewritten sentence by sentence, by a person or a paraphrasing tool, can get past it.
- Some languages. Languages written without spaces between words, like Chinese, Japanese and Thai, aren't covered yet. Results on smaller-language Wikipedias are also weaker, because many of their articles are near-identical stubs generated by bots.
How the matching works
The method is winnowing, the fingerprinting idea behind MOSS, the code-plagiarism detector that universities have used for decades:
- Your text is split into words, and every overlapping run of 5 words is turned into a number (a hash).
- Out of every 4 consecutive hashes, we keep the smallest. That shrinks each document's fingerprint a lot. Any run of 8 or more words you share with a source still leaves at least one fingerprint in common.
- Every document's fingerprints are stored in one sorted index, 8 bytes per entry: the hash plus a document id. The index for the main 167-million-document set is about 178 GB. All of it lives in a storage bucket that costs a few dollars a month.
- A check looks up your fingerprints, ranks candidate documents by how many they share with you, and then fetches the real text of the top candidates. It lines up the passages word by word, so every match in the report shows you the source's actual wording, not just a score.
A few details make the results more honest:
- Stock phrases are dropped. A fingerprint that appears in more than 1,000 documents ("the results of this study suggest that…") is removed from the index, so ordinary academic phrasing isn't flagged as copying.
- Cited passages are marked as cited. If a matched passage carries a citation to the source it came from, the report shows it that way instead of counting it as unattributed copying.
- Duplicates are merged. When the same paper exists as full text and as an abstract, the report points to the full text.
Each match also has a Cite this source button that looks the source up and copies a ready-made APA reference.
What we measured
We tested each part of the collection by planting passages from real documents into unrelated text, running the check, and counting how often the exact source came back.
| Part of the collection | Planted passages found, exact source |
|---|---|
| Main set: PubMed Central + English Wikipedia + OpenAlex | 98.9% |
| Other-language Wikipedias + Project Gutenberg | 98.3% |
| Computer science and AI papers, 2025 | 100 of 100 |
| Computer science and AI papers, 2026 | 100 of 100 |
| PubMed Central author manuscripts | 100 of 100 |
| ACL Anthology | 100 of 100 |
| Zenodo full texts | 96 of 100 |
| ICLR papers (smaller test) | 40 of 40 |
Before launch, we planted one more passage per set and ran all of them through the same code path a real check uses. 8 of 8 came back with the exact source.
Some finer detail:
- Longer copies are easier. Copied runs of 20–40 words were found 95.5–99.5% of the time. Very short runs of 8–12 words dropped to 80–98%. At this scale the same sentence often exists in several documents, and the checker can point at a different copy than the one you took it from.
- Light rewording still gets caught. With one word in every six to ten swapped, 96.7–100% of passages were still found.
- Original text isn't 0%. Unrelated, original papers typically score 0–0.5% on most sets and around 3% against the main 167-million-document set, where common academic phrasing turns up somewhere. So a few percent on your own work is normal. Read the matched passages, not just the number.
These are our own tests, not an independent benchmark. We'd welcome anyone who wants to test the checker on their own planted passages.
The bug that taught us to test the real path
Our first evaluations reported perfect scores on several full-text sets, and the live checker matched nothing in them. The evaluation script read source texts from local samples, but the publishing step never uploaded those texts to the bucket the live checker reads from. Every index lookup found the right document, then the text fetch came back empty and the match was silently dropped.
We fixed the upload. Since then, a set only goes live after planted passages come back through the exact function a user's check calls. The preflight that deploys the backend refuses to run while any connected set is missing its texts.
Privacy
- You need a free account with a verified email address. This keeps bots from draining the web-search budget everyone shares.
- Your text is not added to any shared database. No one else's check is ever compared against your draft.
- If you check the same text again within 30 days, you get the saved result back instantly, and it doesn't use any of your allowance. A result is never reused after the collection or the matching engine changes.
How much is free
| Plan | Checked per month |
|---|---|
| Free | 7,500 words (30 pages) |
| Go | 25,000 words (100 pages) |
| Max | 50,000 words (200 pages) |
| Team | 250,000 words (1,000 pages), shared |
A page is 250 words. Paid plans also search the web on every sentence instead of a sample, up to 400 searches per check. Prices are on the pricing page.
Try it
Paste a draft into the plagiarism checker. If you know of an openly licensed collection in your field that we're missing, tell us. Sources that researchers actually copy from are the ones most worth adding.