WMTrace

Forensic detection and benchmarking of LLM text watermarks

View the Project on GitHub gtesei/llm-watermark

Installing local models for D1, D2, and D3

WMTrace never downloads model weights during scan or through the web API. Models must be selected, license-audited, downloaded at an exact revision, and staged as immutable local bundles. Model files belong under artifacts/, which is ignored by Git.

Downloading weights does not make a detector calibrated. Without a valid empirical calibration artifact, WMTrace returns E2 diagnostics or abstains; it does not apply an upstream demo threshold.

1. Install optional tooling

python -m venv .venv
. .venv/bin/activate
pip install -e ".[dev,ml,model-setup]"

The tested runtime and download-client versions are pinned in pyproject.toml. Do not run hf download manually for normal WMTrace setup; the explicit sync command reads exact repositories, revisions, destinations, and enabled states from the same validated local configuration used by the services.

2. Select enabled services and preview downloads

Start from the checked-in example:

export WMTRACE_CONFIG=configs/wmtrace.models.example.json
wmtrace models sync --dry-run

The example enables only D3. D1 and D2 are present but disabled because their reference weights are much larger. Review repositories, revisions, licenses, destinations, disk requirements, and device policy in the dry-run output and configuration before enabling them.

Download and generate the checksummed manifest explicitly:

wmtrace models sync

To use a configuration elsewhere or restrict the run:

wmtrace models sync --config /path/to/wmtrace.local.json --dry-run
wmtrace models sync --config /path/to/wmtrace.local.json --detector likelihood-rank

Normal wmtrace scan and wmtrace serve never invoke this workflow.

3. D3 smoke-test model

DistilGPT2 is relatively small and useful for exercising the D3 pipeline. It is not, by itself, a validated production detector.

The example configuration downloads distilbert/distilgpt2 at revision 2290a62682d06624634c1f46a6ad5be0f47f38aa into its D3 bundle when wmtrace models sync runs.

Approximate download size is a few hundred MB. Confirm the pinned model card, files, and Apache-2.0 license before redistribution or deployment.

4. D1 Fast-DetectGPT reference baseline

The upstream Fast-DetectGPT local demo defaults to GPT-Neo 2.7B. WMTrace pins the adapter behavior to upstream commit 971b05202bac2bb504d60c0ac0812fea7a8f7c82 (MIT).

When D1 is enabled, the example configuration downloads EleutherAI/gpt-neo-2.7B at revision e24fa291132763e59f4a5422741b424fb5d59056.

Expect roughly 10–11 GB of weights and more runtime memory. A same-model sampling/scoring configuration reuses one loaded runtime. Stronger upstream pairs are substantially larger and require a separate reproduction.

Record the model-card license and the upstream GPT-Neo code/data licenses separately; a code license does not automatically establish a model-weight or dataset license.

5. D2 Binoculars reference pair

The published Binoculars configuration uses Falcon-7B and Falcon-7B-Instruct. WMTrace pins adapter behavior to upstream commit c8ae2f90d50ee696418bc71d8d9e5020e5f9d7b8 (BSD-3-Clause).

When D2 is enabled, the example configuration downloads:

The pair requires roughly 28–30 GB of weight storage plus inference overhead. A capable GPU or a high-memory CPU machine is recommended. Confirm the pinned model cards and Apache-2.0 terms before use.

WMTrace does not reuse the reference repository’s Falcon threshold. A local calibration is mandatory.

5b. MELD v5 encoder

MELD v5 is a single 395M-parameter encoder (~1.6 GB FP32), not a model pair. It runs on CPU or one GPU and accepts at most 2,048 tokens per call. The v5 checkpoint is MIT-licensed and was released 2026-07-31, after the May preprint; its model card states that scores from earlier repository revisions are not comparable, so pin a specific commit rather than tracking main.

The bundle stages one extra file beyond the standard transformers layout:

Standard loaders must not be used: pipeline() and AutoModelForSequenceClassification discard the custom detector weights and can silently return random-model outputs. ml/meld_backend.py builds the head explicitly and loads it with strict=True.

Unlike D1–D3, MELD scores without a local calibration, reporting the checkpoint’s published threshold at grade E2 with the authors’ interpretation rules attached. That published number never grants E3. For E3:

wmtrace models calibrate --detector meld --metric raw_score --direction higher \
  --language en --domain your-domain --controls path/to/human-corpus

Note the direction: for MELD a higher raw score is more machine-like, the opposite of the likelihood metrics above. The model card recommends at least 100 words per input and reports that below 50 words the checkpoint flags most genuine human text; WMTrace enforces both as capability gates. Original line breaks should be preserved, so the detector reads the nfc representation rather than nfkc.

6. Generated immutable bundle

Every operational bundle needs:

  1. a complete local model and tokenizer snapshot;
  2. exact repository, model, and tokenizer revisions;
  3. SHA-256 for every declared file;
  4. model, tokenizer, code, and dataset license records;
  5. context length, precision, and supported devices;
  6. a checksummed calibration.json inside the bundle;
  7. an expiration date.

D3 loads a model directly from the bundle root. D1 uses sampling_dir and scoring_dir; D2 uses observer_dir and performer_dir. Pair directories must remain inside the bundle root and produce identical token IDs and compatible vocabulary shapes.

Detailed contracts:

wmtrace models sync downloads each exact revision and atomically writes the public manifest after hashing downloaded files. It skips disabled detectors, rejects non-commit revisions and paths escaping the bundle root, and records every repository/revision/destination in JSON or terminal output.

Calibration remains an artifact-building workflow: after creating calibration.json, run wmtrace models sync again so the calibration file is included in the immutable manifest. Scanning never creates calibration data.

7. Configure WMTrace

Example D3 configuration:

{
  "schema_version": 1,
  "detectors": {
    "likelihood-rank": {
      "bundle": "distilgpt2-d3-v1",
      "device": "cpu",
      "target_fpr": 0.01,
      "min_tokens": 32
    }
  },
  "bundles": {
    "distilgpt2-d3-v1": {
      "manifest": "artifacts/bundles/distilgpt2-d3-v1/manifest.json",
      "calibration": "calibration.json",
      "models": [{
        "role": "model",
        "repository": "distilbert/distilgpt2",
        "revision": "2290a62682d06624634c1f46a6ad5be0f47f38aa",
        "path": "."
      }],
      "licenses": {"model": "Apache-2.0", "tokenizer": "Apache-2.0"},
      "context_length": 1024,
      "precision": "float32",
      "device_requirements": ["cpu", "cuda"],
      "calibration_valid_through": "2027-12-31"
    }
  }
}

Save this as wmtrace.local.json, or set WMTRACE_CONFIG to its path. Relative manifest paths resolve relative to the configuration file.

The complete D1–D3 source declarations are available in configs/wmtrace.models.example.json.

8. Verify status and run

wmtrace scan --language en --domain general document.txt --json
wmtrace serve --port 8177

Availability progresses through explicit states:

model_missing -> bundle_missing/bundle_invalid -> uncalibrated -> available

For an available detector, the evidence bundle records the local bundle, manifest, calibration, and resolved WMTrace configuration identities.

9. Calibration and release gate

Before claiming detector performance:

Do not commit model files, private datasets, credentials, or unlicensed text to the repository.