Forensic detection and benchmarking of LLM text watermarks
WMTrace never downloads model weights during scan or through the web API.
Models must be selected, license-audited, downloaded at an exact revision, and
staged as immutable local bundles. Model files belong under artifacts/, which
is ignored by Git.
Downloading weights does not make a detector calibrated. Without a valid empirical calibration artifact, WMTrace returns E2 diagnostics or abstains; it does not apply an upstream demo threshold.
python -m venv .venv
. .venv/bin/activate
pip install -e ".[dev,ml,model-setup]"
The tested runtime and download-client versions are pinned in pyproject.toml.
Do not run hf download manually for normal WMTrace setup; the explicit sync
command reads exact repositories, revisions, destinations, and enabled states
from the same validated local configuration used by the services.
Start from the checked-in example:
export WMTRACE_CONFIG=configs/wmtrace.models.example.json
wmtrace models sync --dry-run
The example enables only D3. D1 and D2 are present but disabled because their reference weights are much larger. Review repositories, revisions, licenses, destinations, disk requirements, and device policy in the dry-run output and configuration before enabling them.
Download and generate the checksummed manifest explicitly:
wmtrace models sync
To use a configuration elsewhere or restrict the run:
wmtrace models sync --config /path/to/wmtrace.local.json --dry-run
wmtrace models sync --config /path/to/wmtrace.local.json --detector likelihood-rank
Normal wmtrace scan and wmtrace serve never invoke this workflow.
DistilGPT2 is relatively small and useful for exercising the D3 pipeline. It is not, by itself, a validated production detector.
The example configuration downloads distilbert/distilgpt2 at revision
2290a62682d06624634c1f46a6ad5be0f47f38aa into its D3 bundle when
wmtrace models sync runs.
Approximate download size is a few hundred MB. Confirm the pinned model card, files, and Apache-2.0 license before redistribution or deployment.
The upstream Fast-DetectGPT local demo defaults to GPT-Neo 2.7B. WMTrace pins
the adapter behavior to upstream commit
971b05202bac2bb504d60c0ac0812fea7a8f7c82 (MIT).
When D1 is enabled, the example configuration downloads
EleutherAI/gpt-neo-2.7B at revision
e24fa291132763e59f4a5422741b424fb5d59056.
Expect roughly 10–11 GB of weights and more runtime memory. A same-model sampling/scoring configuration reuses one loaded runtime. Stronger upstream pairs are substantially larger and require a separate reproduction.
Record the model-card license and the upstream GPT-Neo code/data licenses separately; a code license does not automatically establish a model-weight or dataset license.
The published Binoculars configuration uses Falcon-7B and
Falcon-7B-Instruct. WMTrace pins adapter behavior to upstream commit
c8ae2f90d50ee696418bc71d8d9e5020e5f9d7b8 (BSD-3-Clause).
When D2 is enabled, the example configuration downloads:
tiiuae/falcon-7b at ec89142b67d748a1865ea4451372db8313ada0d8;tiiuae/falcon-7b-instruct at
8782b5c5d8c9290412416618f36a133653e85285.The pair requires roughly 28–30 GB of weight storage plus inference overhead. A capable GPU or a high-memory CPU machine is recommended. Confirm the pinned model cards and Apache-2.0 terms before use.
WMTrace does not reuse the reference repository’s Falcon threshold. A local calibration is mandatory.
MELD v5 is a single 395M-parameter encoder (~1.6 GB FP32), not a model pair. It
runs on CPU or one GPU and accepts at most 2,048 tokens per call. The v5
checkpoint is MIT-licensed and was released 2026-07-31, after the May
preprint; its model card states that scores from earlier repository revisions
are not comparable, so pin a specific commit rather than tracking main.
The bundle stages one extra file beyond the standard transformers layout:
meld_config.json — the scoring head’s dimensions and published score
offsets. Without it the checkpoint still satisfies every transformers
requirement, so ml/bundles.py requires it explicitly for meld and rejects
a bundle that omits it as bundle_invalid at inspection rather than letting
it fail later at load time.Standard loaders must not be used: pipeline() and
AutoModelForSequenceClassification discard the custom detector weights and can
silently return random-model outputs. ml/meld_backend.py builds the head
explicitly and loads it with strict=True.
Unlike D1–D3, MELD scores without a local calibration, reporting the checkpoint’s published threshold at grade E2 with the authors’ interpretation rules attached. That published number never grants E3. For E3:
wmtrace models calibrate --detector meld --metric raw_score --direction higher \
--language en --domain your-domain --controls path/to/human-corpus
Note the direction: for MELD a higher raw score is more machine-like, the
opposite of the likelihood metrics above. The model card recommends at least
100 words per input and reports that below 50 words the checkpoint flags most
genuine human text; WMTrace enforces both as capability gates. Original line
breaks should be preserved, so the detector reads the nfc representation
rather than nfkc.
Every operational bundle needs:
calibration.json inside the bundle;D3 loads a model directly from the bundle root. D1 uses sampling_dir and
scoring_dir; D2 uses observer_dir and performer_dir. Pair directories must
remain inside the bundle root and produce identical token IDs and compatible
vocabulary shapes.
Detailed contracts:
wmtrace models sync downloads each exact revision and atomically writes the
public manifest after hashing downloaded files. It skips disabled detectors,
rejects non-commit revisions and paths escaping the bundle root, and records
every repository/revision/destination in JSON or terminal output.
Calibration remains an artifact-building workflow: after creating
calibration.json, run wmtrace models sync again so the calibration file is
included in the immutable manifest. Scanning never creates calibration data.
Example D3 configuration:
{
"schema_version": 1,
"detectors": {
"likelihood-rank": {
"bundle": "distilgpt2-d3-v1",
"device": "cpu",
"target_fpr": 0.01,
"min_tokens": 32
}
},
"bundles": {
"distilgpt2-d3-v1": {
"manifest": "artifacts/bundles/distilgpt2-d3-v1/manifest.json",
"calibration": "calibration.json",
"models": [{
"role": "model",
"repository": "distilbert/distilgpt2",
"revision": "2290a62682d06624634c1f46a6ad5be0f47f38aa",
"path": "."
}],
"licenses": {"model": "Apache-2.0", "tokenizer": "Apache-2.0"},
"context_length": 1024,
"precision": "float32",
"device_requirements": ["cpu", "cuda"],
"calibration_valid_through": "2027-12-31"
}
}
}
Save this as wmtrace.local.json, or set WMTRACE_CONFIG to its path.
Relative manifest paths resolve relative to the configuration file.
The complete D1–D3 source declarations are available in
configs/wmtrace.models.example.json.
wmtrace scan --language en --domain general document.txt --json
wmtrace serve --port 8177
Availability progresses through explicit states:
model_missing -> bundle_missing/bundle_invalid -> uncalibrated -> available
For an available detector, the evidence bundle records the local bundle, manifest, calibration, and resolved WMTrace configuration identities.
Before claiming detector performance:
Do not commit model files, private datasets, credentials, or unlicensed text to the repository.