Bookmark Context
Type to search documentation.

Prompt-injection scanner

How Bookmark Context flags web pages that try to manipulate an AI reader.

Saved pages are third-party text that an AI assistant will later read. The scanner looks for content written to manipulate that assistant — “ignore previous instructions”, fake system tokens, requests to read local files — and flags it. It never edits or blocks text on its own; it produces a score and the caller decides.

Categories and weights

Each chunk is matched against regex categories. The weight is how strongly a match implies a real injection attempt:

CategoryWeightExample signal
system_token0.95<|im_start|>system, [INST], ### System:
exfiltration0.90~/.ssh/id_rsa, curl https://evil…, password: …
direct_override0.85“ignore all previous instructions”
llm_targeted_instruction0.82“if you are an AI, you must…”
concealment0.80“do not mention this to the user”
role_reassignment0.75“you are now DAN”, “enable developer mode”
persistent_instruction0.70“for all future requests, always…”

Score

score = max(category weight) + 0.05 × (number of extra categories that fired)

capped at 1.0. A clean chunk scores 0.0.

Save-time gate

POST …/bookmarks returns a ScanWarning (and saves nothing) when any chunk scores above 0, or when the page looks like an unrendered SPA shell (content_unscannable). Override with ?force=true — see Overriding a scan warning.

After indexing

Every stored chunk keeps its injection_risk and injection_signals as Chroma metadata, and both appear in search_collection / ask_collection results so a client can down-weight or ignore a suspicious passage.

What it deliberately ignores

Credential-shaped placeholders that appear in every API doc — api_key="YOUR_API_KEY", <your-key>, os.getenv("API_KEY"), empty strings — are not flagged. A warning that is always wrong teaches you to dismiss the one that is right.

Last updated Sep 9, 2026