Forensic detection and benchmarking of LLM text watermarks
How to run the suite, what each layer of it proves, how to add tests for new detectors, and — just as important — what the suite deliberately does not prove.
WMTrace makes claims about evidence. A test suite that only checked “the code runs” would be worthless here: the failure mode that matters is a detector that runs perfectly and produces a miscalibrated answer. Most of what follows exists to catch that.
python -m venv .venv && . .venv/bin/activate # or: uv venv && . .venv/bin/activate
pip install -e ".[dev]" # or: uv pip install -e ".[dev]"
pytest # whole suite
pytest -v # per-test names
pytest tests/test_statistics.py # one file
pytest -k zero_width # one topic across files
pytest tests/test_golden_examples.py::test_example_a_watermarked_statistics # one test
pytest --durations=5 # find slow tests
Configuration lives in pyproject.toml ([tool.pytest.ini_options]):
testpaths = ["tests"], addopts = "-q". No plugins; tests/conftest.py holds
only shared fixtures. Run the bare pytest command — python -m pytest adds the
repo root to sys.path and can hide import errors that contributors will hit.
The suite is fast (well under a second) and has no network, no fixtures on disk, and no external services. If a test ever needs any of those, that is a design smell worth pushing back on.
| Layer | File | Proves |
|---|---|---|
| Statistics | tests/test_statistics.py |
The math is right and calibrated: exact binomial matches direct summation, stays stable in the far tail, and the null behaves — P(p ≤ α) ≤ α under the null hypothesis. |
| Golden examples | tests/test_golden_examples.py |
The arithmetic published in docs/RESEARCH.md is reproduced exactly. Documentation cannot drift from code. |
| Detectors | tests/test_detectors.py |
Embed→detect round-trips, negative controls (clean text, wrong key), abstention on short and non-ASCII input, Unicode findings. |
| CLI | tests/test_cli.py |
Every subcommand exits 0, JSON output parses, and the human output carries its scope note. |
| API | tests/test_api.py |
Endpoint contracts, structured errors, input limits, static assets served. |
Three more files arrive with the remediation plan
(findings/IMPLEMENTATION_PLAN.md):
tests/test_imports.py (import hygiene), tests/test_calibration.py (gamma
fidelity and false-positive rate), tests/test_contracts.py (invariants),
tests/test_attacks.py and tests/test_robustness.py (the P3 harness).
docs/RESEARCH.md publishes worked examples with exact statistics. Those are
executable:
def test_example_a_watermarked_statistics():
result = run_kgw(EXAMPLE_A_WATERMARKED)
assert result["usable_tokens"] == 49
assert result["green_tokens"] == 25
assert math.isclose(result["statistic"], 4.2064, abs_tol=1e-4)
assert math.isclose(result["p_value"], 8.03e-5, rel_tol=1e-2)
Rule: if a change breaks a golden test, reconcile the document and the test together. Never loosen the assertion to make it pass. These are the anchor that makes refactoring safe — a “harmless cleanup” that moves a published number fails here immediately.
A round-trip test alone is nearly worthless: it proves the detector finds what
the embedder put there, which a return True detector also achieves. Every
positive test is paired with controls that must come back negative:
no_scheme_evidenceno_scheme_evidenceThe most dangerous defect in this project is a detector whose declared null disagrees with its actual behavior — it produces confident false accusations. Two complementary tests catch it:
# direct: does the keyed PRF accept at the declared rate?
rate = sum(is_green(KEY, [], w, gamma) for w in words) / len(words)
assert abs(rate - gamma) < 0.02
# end-to-end: do unwatermarked documents stay unflagged?
flagged = sum(detector_flags(random_document()) for _ in range(40))
assert flagged == 0
This pairing exists because the second test alone is too coarse: a green-rate
skew below ~0.02 hides inside sampling noise at 40 documents, while the first
test catches it directly. (A real defect of exactly this kind — round(1/gamma)
as a modulus — flagged 34 of 40 clean documents at γ=0.4. See F6 in
findings/FINDINGS.md.)
Import-order bugs are invisible to an ordinary test run, because once any test
imports wmtrace.core, sys.modules caches it and the cycle never re-forms:
@pytest.mark.parametrize("module", MODULES)
def test_module_imports_standalone(module):
result = subprocess.run([sys.executable, "-c", f"import {module}"],
capture_output=True, text=True)
assert result.returncode == 0, result.stderr
A fresh interpreter per module is the whole point. Do not “optimize” this
into a plain importlib.import_module loop — that reintroduces the masking
which let this class of bug ship in the first place.
Where a rule holds for all inputs, test it that way — a randomized sweep with fixed seeds, asserting properties rather than values:
@pytest.mark.parametrize("seed", range(8))
def test_every_detector_returns_a_coherent_bundle(seed):
for bundle in analyze(random_text(seed=seed)):
assert 0 <= bundle.result["green_tokens"] <= bundle.result["usable_tokens"]
assert 0.0 <= bundle.result["p_value"] <= 1.0
assert bundle.result["adjusted_p_value"] >= bundle.result["p_value"]
assert bundle.scope["does_not_support"]
This is property-based testing without taking on hypothesis as a dependency.
Seeds are fixed so failures reproduce exactly.
Measured curves (robustness, estimation precision) assert direction and magnitude, not exact numbers:
zs = [point["z"] for point in substitution_curve()]
assert zs == sorted(zs, reverse=True) # monotone degradation
assert curve[-1]["fraction_of_baseline"] < 0.30
assert unigram["lift"] > 2.0 and contextual["lift"] < 1.0
An exact-value assertion here would fail on any generator tweak while proving nothing extra. The measured margins (3.25× vs 0.38×) are wide enough that the shape assertion still catches a real regression.
Preconditions use require() (→ ValueError, the caller’s fault);
postconditions and invariants use ensure() (→ AssertionError, our bug).
Neither is a bare assert, because python -O strips those — and these guard
the project’s central promise. Verify the guarantee survives:
python -OO -c "
from wmtrace.core.evidence import EvidenceBundle, Namespace
try:
EvidenceBundle(Namespace.WATERMARK, {}, {},
{'grade': 'E3', 'decision': 'human_written'},
{'does_not_support': ['x']})
print('FAIL: contract was stripped')
except ValueError as exc:
print('ok, contract survives -OO:', exc)
"
Randomized tests must reproduce exactly on every machine and every run.
random.Random(seed), rng_seed=11,
estimate_green_set(..., seed=0). Never rely on the global random state.assert flagged == 0 over 40 documents
at a target FPR of 1e-4 is safe; assert precision > 0.81 when the measured
value is 0.812 is not — leave margin (> 0.6).), never as
a literal character. A literal will be silently mangled by editors, and this
project’s own inspector exists to find exactly that.from tests.test_x import y only
resolves when the repo root happens to be on sys.path — true under
python -m pytest, false under a bare pytest. Anything shared goes in
tests/conftest.py as a fixture. Verify with bare pytest, which is what
contributors and CI actually run.A new detector is not done until it has all five:
no_scheme_evidence.tests/test_contracts.py.does_not_support.If the detector reproduces something published in docs/RESEARCH.md, add a
golden test too, and cite the section in the docstring.
When executing findings/IMPLEMENTATION_PLAN.md, each task’s fix is proven by
one specific command:
# F1 circular import
python -c "import wmtrace.detectors.greenlist" # must not raise
pytest tests/test_imports.py -v
# F6 gamma miscalibration
pytest tests/test_calibration.py -v
# F2 normalization claim
pytest tests/test_golden_examples.py -k zero_width -v
# F5 capacity
echo "The board approved the budget." | wmtrace embed --scheme zero-width --capacity
pytest tests/test_cli.py -k capacity -v
# F3/F4 attack harness
pytest tests/test_attacks.py tests/test_robustness.py -v
wmtrace attack steal --scheme unigram-toy --docs 20 --length 120
wmtrace attack robustness --length 200
# Task 7 contracts + refactor safety
pytest tests/test_contracts.py -v
pytest tests/test_golden_examples.py -v # no published number moved
git diff --stat HEAD # src/ line count flat or down
Every task ends with a full pytest run. A task is not complete while the
suite is red.
Stated plainly, because a testing document that only lists strengths is a marketing document.
TestClient. The JavaScript, the rendering, the theme toggle, and the
heatmap are unverified by automated tests; they were checked by hand and by
headless screenshot.findings/FINDINGS.md — the external review these tests were hardened againstfindings/IMPLEMENTATION_PLAN.md — task-by-task remediation, test-firstdocs/RESEARCH.md — the research spec whose examples the golden tests pinCLAUDE.md — project invariants that the contract tests enforce