Public test card / R1 / 2026-08-30
THE STARTING LINE.
AUDITED IN PUBLIC.
R1 is NULLRADIAL's version-locked digital starting line. It is not proof that a pattern is universally undetectable. The tests cover generic target detection on named builds and synthetic surface composites. They do not establish performance against commercial surveillance, identity, face, re-identification, plate, segmentation, thermal, infrared, tracking, or safety-critical systems. Two scenes are neutral public samples; two are pre-patterned supplied screenshots whose physical provenance is unverified.
Revision lock
- Benchmark ID
- NULLRADIAL-BENCH-R1
- Research lock
- NR-EMBER-HALO-R1
- Definition
- Generator seed, SVG/PNG bytes, hashes, and benchmark baseline are revision-controlled. This is not a security certification.
- Master SVG SHA-256
- 2127475bca07b6870fbf769d848a282d82abd9fb201ed6e0f228f2dfdaa689bf
- Generator seed
- 44093
- Release decision
- Selected from three patterns on the same four scenes because it had the highest pooled point estimate. This was not a preregistered holdout decision, the interval includes zero, and the neutral-scene detector-family directions disagree.
Named detector builds
- YOLO family
- YOLOv10n end-to-end ONNX export / embedded Ultralytics 8.1.34 metadata / threshold 0.25 / 640 × 640 letterbox
- YOLO weight SHA-256
- a77dd863933f184a19e84361c64b788228a7c7dacc2c78939239a96ad3efca3b
- Transformer family
- Quantized DETR with ResNet-50 backbone / Xenova ONNX port / threshold 0.50 / shortest edge 800, longest ≤1333
- DETR weight SHA-256
- cae09a307ed9247da7e2ce8bcf81522a6817f1ea2e82b9c4dde59f5964b62b4f
- Runtime
- ONNX Runtime 1.23.2 / CPUExecutionProvider
- Scenes and pair design
- Four scenes: two neutral public samples and two pre-patterned supplied screenshots. Three candidate revisions were compared with source-tile RGB-histogram block-permutation controls. Final composites are not exact-histogram matches after resizing, masking, tiling, and shading.
Pattern-level result
| Rank | Pattern | Cells | Radial disruption | Control disruption | Candidate − block-control Δ | 95% scene CI | YOLO Δ | DETR Δ |
|---|---|---|---|---|---|---|---|---|
| 1 | Ember Halo | 7 | 0.312 | 0.235 | +0.078 | [−0.041, +0.186] | +0.074 | +0.080 |
| 2 | Signal Moss | 7 | 0.143 | 0.200 | −0.057 | [−0.115, +0.006] | −0.079 | −0.040 |
| 3 | Parallax Bloom | 7 | 0.202 | 0.263 | −0.060 | [−0.196, +0.069] | −0.027 | −0.086 |
Primary metric: baseline confidence minus matched-variant confidence, with a miss receiving the full baseline confidence. Positive candidate-minus-block-control contrast means the candidate changed this detector-output continuity endpoint more than its source-tile-histogram block control. It does not causally isolate radiality. Intervals use 10,000 scene-cluster bootstrap resamples with seed 20260830.
Independent audit / stratum split
- Ember / all eligible cells
- +0.077774 pooled candidate-minus-block-control contrast
- Ember / neutral scenes only
- −0.003091 across four detector × neutral-scene cells
- Ember / pre-patterned references
- +0.185595 across three eligible cells; these screenshots are not neutral before-images
- Neutral / YOLO family
- +0.111650
- Neutral / DETR family
- −0.117833
- Consequence
- The pooled positive direction does not replicate within the same neutral stratum. R1 remains useful as an audited failure-finding exercise, not confirmatory evidence.
Association proxy / Ember Halo
- Radial continuity
- 55.87% versus 69.05% for the matched control
- Mean matched IoU
- 0.772 radial versus 0.855 control
- Centroid drift
- 0.0663 radial versus 0.0295 control, normalized to image diagonal
- Boundary
- These are two-frame detection-association stability proxies. No temporal tracker was run, so tracking is not tested.
What this means
The pooled signal was positive. The neutral evidence split.
Ember Halo produced positive pooled candidate-minus-block-control point estimates on both named builds. On neutral scenes alone, YOLO was positive and DETR was negative. The scene-bootstrap interval crossed zero, so the sample does not support a population claim or a promise that the pattern will alter a new camera or model.
Only four distinct scenes, two detectors, one threshold per build, hand-authored masks, and synthetic recoloring were tested. No fabric, cast film, folds, gloss, compound curvature, print gamut, seams, weather, distance, physical motion, wash state, or professional installation was validated.
A consumer numerical claim remains gated on preregistered, multi-camera physical testing with independent participants, matched controls, multiple architecture families, full threshold curves, clustered intervals, and published failure strata.
The published 81 checks verify local hashes, image decoding, and report structure. They do not rerun model inference or independently recompute assignments, metrics, exclusions, or confidence intervals. S2 adds a transitive code/dependency fingerprint and an append-only rerun ledger; physical R2 requires independent ground truth.
Release artifacts
Full JSON Metrics CSV Revision lock 81 local integrity checks Method reportAdversarial-texture lineage
AdvTexture / CVPR 2022 AdvCaT / CVPR 2023 GRAC / ICCV 2025Primary context
YOLOv10 DETR PADetBench NIST AML