WMTrace

Forensic detection and benchmarking of LLM text watermarks

View the Project on GitHub gtesei/llm-watermark

Testing WMTrace

How to run the suite, what each layer of it proves, how to add tests for new detectors, and — just as important — what the suite deliberately does not prove.

WMTrace makes claims about evidence. A test suite that only checked “the code runs” would be worthless here: the failure mode that matters is a detector that runs perfectly and produces a miscalibrated answer. Most of what follows exists to catch that.


Quick start

python -m venv .venv && . .venv/bin/activate    # or: uv venv && . .venv/bin/activate
pip install -e ".[dev]"                          # or: uv pip install -e ".[dev]"

pytest                              # whole suite
pytest -v                           # per-test names
pytest tests/test_statistics.py     # one file
pytest -k zero_width                # one topic across files
pytest tests/test_golden_examples.py::test_example_a_watermarked_statistics   # one test
pytest --durations=5                # find slow tests

Configuration lives in pyproject.toml ([tool.pytest.ini_options]): testpaths = ["tests"], addopts = "-q". No plugins; tests/conftest.py holds only shared fixtures. Run the bare pytest command — python -m pytest adds the repo root to sys.path and can hide import errors that contributors will hit.

The suite is fast (well under a second) and has no network, no fixtures on disk, and no external services. If a test ever needs any of those, that is a design smell worth pushing back on.


The layers, and what each one is for

Layer File Proves
Statistics tests/test_statistics.py The math is right and calibrated: exact binomial matches direct summation, stays stable in the far tail, and the null behaves — P(p ≤ α) ≤ α under the null hypothesis.
Golden examples tests/test_golden_examples.py The arithmetic published in docs/RESEARCH.md is reproduced exactly. Documentation cannot drift from code.
Detectors tests/test_detectors.py Embed→detect round-trips, negative controls (clean text, wrong key), abstention on short and non-ASCII input, Unicode findings.
CLI tests/test_cli.py Every subcommand exits 0, JSON output parses, and the human output carries its scope note.
API tests/test_api.py Endpoint contracts, structured errors, input limits, static assets served.

Three more files arrive with the remediation plan (findings/IMPLEMENTATION_PLAN.md): tests/test_imports.py (import hygiene), tests/test_calibration.py (gamma fidelity and false-positive rate), tests/test_contracts.py (invariants), tests/test_attacks.py and tests/test_robustness.py (the P3 harness).


The techniques, and why each one is used

1. Golden tests pin published numbers

docs/RESEARCH.md publishes worked examples with exact statistics. Those are executable:

def test_example_a_watermarked_statistics():
    result = run_kgw(EXAMPLE_A_WATERMARKED)
    assert result["usable_tokens"] == 49
    assert result["green_tokens"] == 25
    assert math.isclose(result["statistic"], 4.2064, abs_tol=1e-4)
    assert math.isclose(result["p_value"], 8.03e-5, rel_tol=1e-2)

Rule: if a change breaks a golden test, reconcile the document and the test together. Never loosen the assertion to make it pass. These are the anchor that makes refactoring safe — a “harmless cleanup” that moves a published number fails here immediately.

2. Negative controls, not just round-trips

A round-trip test alone is nearly worthless: it proves the detector finds what the embedder put there, which a return True detector also achieves. Every positive test is paired with controls that must come back negative:

3. Calibration is measured, not assumed

The most dangerous defect in this project is a detector whose declared null disagrees with its actual behavior — it produces confident false accusations. Two complementary tests catch it:

# direct: does the keyed PRF accept at the declared rate?
rate = sum(is_green(KEY, [], w, gamma) for w in words) / len(words)
assert abs(rate - gamma) < 0.02

# end-to-end: do unwatermarked documents stay unflagged?
flagged = sum(detector_flags(random_document()) for _ in range(40))
assert flagged == 0

This pairing exists because the second test alone is too coarse: a green-rate skew below ~0.02 hides inside sampling noise at 40 documents, while the first test catches it directly. (A real defect of exactly this kind — round(1/gamma) as a modulus — flagged 34 of 40 clean documents at γ=0.4. See F6 in findings/FINDINGS.md.)

4. Subprocess isolation for import hygiene

Import-order bugs are invisible to an ordinary test run, because once any test imports wmtrace.core, sys.modules caches it and the cycle never re-forms:

@pytest.mark.parametrize("module", MODULES)
def test_module_imports_standalone(module):
    result = subprocess.run([sys.executable, "-c", f"import {module}"],
                            capture_output=True, text=True)
    assert result.returncode == 0, result.stderr

A fresh interpreter per module is the whole point. Do not “optimize” this into a plain importlib.import_module loop — that reintroduces the masking which let this class of bug ship in the first place.

5. Invariants over examples

Where a rule holds for all inputs, test it that way — a randomized sweep with fixed seeds, asserting properties rather than values:

@pytest.mark.parametrize("seed", range(8))
def test_every_detector_returns_a_coherent_bundle(seed):
    for bundle in analyze(random_text(seed=seed)):
        assert 0 <= bundle.result["green_tokens"] <= bundle.result["usable_tokens"]
        assert 0.0 <= bundle.result["p_value"] <= 1.0
        assert bundle.result["adjusted_p_value"] >= bundle.result["p_value"]
        assert bundle.scope["does_not_support"]

This is property-based testing without taking on hypothesis as a dependency. Seeds are fixed so failures reproduce exactly.

6. Shape assertions where exact values are brittle

Measured curves (robustness, estimation precision) assert direction and magnitude, not exact numbers:

zs = [point["z"] for point in substitution_curve()]
assert zs == sorted(zs, reverse=True)          # monotone degradation
assert curve[-1]["fraction_of_baseline"] < 0.30
assert unigram["lift"] > 2.0 and contextual["lift"] < 1.0

An exact-value assertion here would fail on any generator tweak while proving nothing extra. The measured margins (3.25× vs 0.38×) are wide enough that the shape assertion still catches a real regression.

7. Contracts, checked in optimized mode

Preconditions use require() (→ ValueError, the caller’s fault); postconditions and invariants use ensure() (→ AssertionError, our bug). Neither is a bare assert, because python -O strips those — and these guard the project’s central promise. Verify the guarantee survives:

python -OO -c "
from wmtrace.core.evidence import EvidenceBundle, Namespace
try:
    EvidenceBundle(Namespace.WATERMARK, {}, {},
                   {'grade': 'E3', 'decision': 'human_written'},
                   {'does_not_support': ['x']})
    print('FAIL: contract was stripped')
except ValueError as exc:
    print('ok, contract survives -OO:', exc)
"

Determinism rules

Randomized tests must reproduce exactly on every machine and every run.


Adding a detector? Test it like this

A new detector is not done until it has all five:

  1. Round-trip — embed (or generate) under a known key, detect it.
  2. Negative control — clean text, and where the scheme is keyed, the wrong key. Both must return no_scheme_evidence.
  3. Abstention — input too short / wrong language / missing key must return E0 or E1 with a reason, never a guess.
  4. Grade cap — unkeyed channels must not exceed E2; nothing in this build may claim E4/E5. Add it to the sweep in tests/test_contracts.py.
  5. Scope statement — the bundle must carry a non-empty does_not_support.

If the detector reproduces something published in docs/RESEARCH.md, add a golden test too, and cite the section in the docstring.


Verifying the remediation plan, task by task

When executing findings/IMPLEMENTATION_PLAN.md, each task’s fix is proven by one specific command:

# F1 circular import
python -c "import wmtrace.detectors.greenlist"     # must not raise
pytest tests/test_imports.py -v

# F6 gamma miscalibration
pytest tests/test_calibration.py -v

# F2 normalization claim
pytest tests/test_golden_examples.py -k zero_width -v

# F5 capacity
echo "The board approved the budget." | wmtrace embed --scheme zero-width --capacity
pytest tests/test_cli.py -k capacity -v

# F3/F4 attack harness
pytest tests/test_attacks.py tests/test_robustness.py -v
wmtrace attack steal --scheme unigram-toy --docs 20 --length 120
wmtrace attack robustness --length 200

# Task 7 contracts + refactor safety
pytest tests/test_contracts.py -v
pytest tests/test_golden_examples.py -v            # no published number moved
git diff --stat HEAD                               # src/ line count flat or down

Every task ends with a full pytest run. A task is not complete while the suite is red.


What this suite does not prove

Stated plainly, because a testing document that only lists strengths is a marketing document.


See also