Skip to content

Evidence and limits

Evidence layers

  1. 47 real hook runs establish Claude Code WebSearch integration feasibility.
  2. E2 pilot uses real Qwen3-14B inference and the production hook on ten controlled synthetic retrieval bundles. With all poison sources covered by known IOCs, fake-brand recommendation fell from 10/10 to 0/10.
  3. Offline tests establish URL normalization, matching, failure semantics and API constraints.

The pilot is not a live-search experiment. All fake-brand evidence resides on matched sources, so it establishes semantic impact only under complete known-source coverage.

Four evaluation layers, not one total score

LayerCore questionPrimary tests
MechanismIs a covered source denied before retrieval?normalization, evasions, network zero-arrival, redirects, end-to-end latency
Shared networkCan one verdict reliably serve multiple clients?propagation, revocation, cache expiry, multi-client consistency
Account governanceAre enumeration, brigading and privilege abuse constrained?concurrent quotas, immediate bans, RBAC, Sybil accounts, malicious reports and appeals
Semantic outcomeDoes removing polluted evidence change recommendations?E2 and FORGE under known, partial, unknown and legitimate-host parasite conditions

FORGE is naturally a content-pollution and recommendation-semantics benchmark. It does not directly test URL matching, pre-request zero arrival, verdict propagation, strongly consistent quotas or governance. It is therefore used only in the fourth layer, never as a mechanism test or system-wide score. Fake-brand mention, first recommendation, rank, recommendation strength, poison citation, benign-brand suppression and document counts are reported; model and hook latency are separated.

IOC lifetime

Private operational IOCs, authentication and quotas are defense in depth against efficient enumeration. Their actual effect on GEO domain rotation remains a longitudinal research question.

Research prototype · Operational IOCs remain private