Evidence and limits
Evidence layers
- 47 real hook runs establish Claude Code WebSearch integration feasibility.
- E2 pilot uses real Qwen3-14B inference and the production hook on ten controlled synthetic retrieval bundles. With all poison sources covered by known IOCs, fake-brand recommendation fell from 10/10 to 0/10.
- Offline tests establish URL normalization, matching, failure semantics and API constraints.
The pilot is not a live-search experiment. All fake-brand evidence resides on matched sources, so it establishes semantic impact only under complete known-source coverage.
Four evaluation layers, not one total score
| Layer | Core question | Primary tests |
|---|---|---|
| Mechanism | Is a covered source denied before retrieval? | normalization, evasions, network zero-arrival, redirects, end-to-end latency |
| Shared network | Can one verdict reliably serve multiple clients? | propagation, revocation, cache expiry, multi-client consistency |
| Account governance | Are enumeration, brigading and privilege abuse constrained? | concurrent quotas, immediate bans, RBAC, Sybil accounts, malicious reports and appeals |
| Semantic outcome | Does removing polluted evidence change recommendations? | E2 and FORGE under known, partial, unknown and legitimate-host parasite conditions |
FORGE is naturally a content-pollution and recommendation-semantics benchmark. It does not directly test URL matching, pre-request zero arrival, verdict propagation, strongly consistent quotas or governance. It is therefore used only in the fourth layer, never as a mechanism test or system-wide score. Fake-brand mention, first recommendation, rank, recommendation strength, poison citation, benign-brand suppression and document counts are reported; model and hook latency are separated.
IOC lifetime
Private operational IOCs, authentication and quotas are defense in depth against efficient enumeration. Their actual effect on GEO domain rotation remains a longitudinal research question.