Runs the chained identity + campaign clustering pipeline against all seven fixtures via from_synthetic / from_synthetic_identity adapters and ratchets every YAML floor to 1.0 — the production clusterer (and the reference clusterers used in the per-fixture tests) all score perfectly across ARI / homogeneity / completeness / singleton_recall on each fixture. Three substrate fixes surfaced by the ratchet: - Tuning: shared_infra now Jaccards payload+C2 only; decky_set moved into cohort_weight to prevent fleet-scarcity false-merges (F1's shared_wordlist failure mode). Tier weight raised to 1.0 so shared payload+C2 alone crosses threshold (F5's intended pass). - Adapter: from_synthetic_identity now reads SyntheticSession started_at + duration_s for session_windows and per-decky timestamps (the production-row adapter still uses start_ts/end_ts when available). - Fixture data: paused_campaign.yaml's JA3 collided exactly with vpn_hopping.yaml's (same TLS extension list). The collision fused two unrelated campaigns under the chained identity layer in the noise_floor composite. Made paused's JA3 distinct. Also wires Campaign / CampaignsResponse into models/__init__.py's __all__ that was missed in the schema commit.
25 lines
815 B
YAML
25 lines
815 B
YAML
# Bounds for fixture 4 (paused_campaign).
|
|
#
|
|
# Ground truth at campaign-level: 1 campaign of 2 observation rows
|
|
# (one per DSL actor — modeling the operator's two operational
|
|
# windows). A correct algorithm scores 1.0 on every metric.
|
|
#
|
|
# Completeness is the load-bearing metric: a clusterer that lets a
|
|
# multi-day silent period split the campaign tanks completeness
|
|
# (the one true class is split across two predicted clusters,
|
|
# matching the gap). The adversarial time_window_clusterer
|
|
# demonstrates this and the bound below rejects it.
|
|
#
|
|
# This fixture is CAMPAIGN-LEVEL ONLY (see the fixture YAML for
|
|
# why). No identity-level scoring.
|
|
#
|
|
# Bounds are loose at v1; tighten as the algorithm matures.
|
|
adjusted_rand_index:
|
|
min: 1.0
|
|
homogeneity:
|
|
min: 1.0
|
|
completeness:
|
|
min: 1.0
|
|
singleton_recall:
|
|
min: 1.0
|