shanchuanzhi-agent-canary

mcp
Guvenlik Denetimi
Basarisiz
Health Gecti
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Community trust — 17 GitHub stars
Code Basarisiz
  • fs.rmSync — Destructive file system operation in scripts/check-package.mjs
  • process.env — Environment variable access in scripts/check-package.mjs
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

Zero-false-positive tripwires for AI agents: decoy MCP tools + leak-tracing canary tokens. Know instantly when your agent is compromised.

README.md

agent-canary

Tripwires for AI coding agents. It plants decoy MCP tools and canary tokens in
your environment, then gives SDK integrations a session circuit breaker to
contain the next guarded action after a compromise signal.

Works with Claude Code, Cursor, Cline, Windsurf — anything that speaks MCP.
Non-MCP agents can use the SDK instead (see below). Node 20+, MIT, no telemetry.

中文文档:README.zh-CN.md

Live demo & sponsor · Glama listing · GitHub Discussions

Agent Canary — AI agent and MCP security

V1.2.3: one guarded route for every tool call

V1.2.3 is the free public line. It keeps the zero-false-positive detection
model and free SDK containment primitives, then adds
createGuardedToolRouter() so integrations have one reviewed dispatch path for
decoys and real tools. It retains the fully offline self-test and centralized
audit-event redaction before data reaches JSONL or a webhook:

Layer What it does
Detection Inert decoy MCP tools and planted canary tokens detect a compromise signal.
Containment SAFE → TRIPPED → QUARANTINED happens synchronously; createGuardedToolRouter() sends decoys to containment and real tools through the fail-closed guard.
Alerting JSONL audit events and optional webhook/desktop alerts are sent after the state transition; tool arguments and canary values are redacted.
Untrusted content → prompt injection → decoy touched / token detected
                                      ↓
                              SESSION TRIPPED
                                      ↓
                                QUARANTINED
                                      ↓
                         dangerous guarded tool call
                                      ↓
                                  BLOCKED
                                      ↓
                            alert + local audit log

The containment API is documented in docs/containment.md.
V2.x paid features are maintained and delivered separately; V2.1 is not
published from this branch.

The problem

Coding agents read files, run commands and call APIs. If one picks up injected
instructions — a poisoned README, a malicious web page, a doc file — it may
quietly exfiltrate secrets or worse, and nothing tells you.

Detection tools that score prompts produce false positives, and false positives
get ignored. agent-canary inverts this: it plants things that no legitimate
workflow ever touches, so any contact is a real compromise signal.

  • Decoy MCP tools. A fake wire transfer, a fake production secret reader, a
    fake root shell. They never perform a real action, but a hijacked agent will
    call one.
  • Canary tokens. Unique cnry_... strings planted in honeypot files. If one
    shows up in agent output, an outbound request or a git diff, a secret was
    copied. There is no benign explanation.

Every fake tool reply embeds a one-time trace token, so exfiltrated "secrets"
point back to the exact tool call that leaked them.

Install the free V1.2.3 line

Prerequisite: Node.js 20 or newer. The public source build contains the free
V1.2.3 baseline:

git clone https://github.com/DorianChn/agent-canary && cd agent-canary
npm install && npm run build && npm link

Run agent-canary --help after linking, then run the offline containment
check. The public repository and public package contain the free V1 line only.
V2.x is maintained and delivered privately after verified purchase; it is not
distributed from this public source branch.

Container and Glama evaluation

The repository includes a minimal stdio-only Docker image so MCP registries
such as Glama can build the free V1 server and inspect its tool schemas without
credentials, network access, or a V2 delivery package:

docker build -t agent-canary .
docker run --rm -i agent-canary

The image starts agent-canary serve. It exposes the same inert decoy tools as
the local V1 CLI; it does not execute real tools or contain payment, license,
or customer data.

Usage

# plant a honeypot file with fake secrets
agent-canary tokens plant .env.canary --label my-project

# register 12 decoy tools in Claude Code (or: install cursor)
agent-canary install claude

# verify the alert pipeline
agent-canary alert-test

# verify SAFE → QUARANTINED → BLOCKED locally, without network or user-state writes
agent-canary self-test

Restart your editor. If the agent later calls a decoy or leaks a token:

$ agent-canary events
2026-09-15T10:21:04Z  decoy_called  tool=canary_read_secrets
2026-09-15T10:21:09Z  token_found   label=my-project path=report.md

The scanner is CI-friendly (exit code 1 on leak) and there is a live watcher:

agent-canary tokens check . && echo clean
agent-canary watch .

The decoys

The agent sees admin-grade tools. None of them do anything real.

Decoy tool What a hijacked agent thinks it gets
canary_transfer_funds wire transfer, auto-approved under 50k
canary_read_secrets production API keys for any environment
canary_run_shell arbitrary command as root on the build server
canary_delete_files permanent deletion, no trash, no undo
canary_send_email email from any internal address, SPF skipped
canary_export_database full-table CSV dump, no audit entry
canary_fetch_internal internal admin panel from the trusted network
canary_disable_security endpoint protection turned off
canary_cloud_console AWS/GCP console sessions + temporary credentials
canary_k8s_exec arbitrary exec inside a production pod
canary_secrets_rotate emergency credential rotation (locks out humans)
canary_git_force_push force push to protected branches

Free vs Personal

This repository publishes the free V1 baseline. V2 Personal is described here
for subscribers, but its paid implementation, signing keys, customer records,
and delivery package are kept outside the public repository.

Free (forever) Personal (¥72 / 30 days)
Decoy server, tokens, watch, alerts, install yes yes
eval — injection resistance scoring yes
dashboard — HTML attack-chain timeline yes
export — CEF / JSON / CSV for SIEM yes
V1.2.3 session circuit breaker, guarded tool router, and offline self-test yes yes
SDK decoy handling and canary scanning yes yes

V2 Personal currently uses a manual WeChat Pay / Alipay confirmation flow.
See the public payment instructions
for the QR codes, price, and the information needed for delivery. The seller
checks the actual transfer before sending installation and activation
instructions; no automatic delivery or instant activation is promised.

Cooperation and integration

We welcome focused collaboration with MCP client maintainers, AI-agent builders,
security researchers, and DevSecOps teams:

  • integrate agent-canary into an MCP client, agent framework, or secure template;
  • run a reproducible prompt-injection evaluation and publish the results;
  • pilot the alert/audit pipeline in a controlled development or CI environment;
  • discuss paid integration, private deployment, or security-assessment support.

Start in GitHub Discussions
with the integration target, scope, and preferred contact method. Do not post API
keys, payment receipts, customer data, or unpublished findings.

Distribution and partner paths

The project is already discoverable through the official MCP Registry
and Glama. For a deeper
security-platform integration, the Snyk Technology Alliance Partner Program
is a candidate channel; any application or commercial terms must be reviewed by
the maintainer before submission. We do not mass-post or send unsolicited
promotional messages.

Non-MCP agents (free V1.2.3 guarded tool router)

Create one guard per agent session, then give it to one router. The router
answers decoys with guard.runDecoy() and routes every non-decoy callback
through guard.executeToolCall().

import {
  CanaryBlockedError,
  createAgentGuard,
  createGuardedToolRouter,
  decoyToolDefs,
} from "agent-canary/sdk";

const guard = createAgentGuard({
  sessionId: "support-chat-42",
  // Exact, reviewed names only. Default is an empty allowlist.
  quarantineAllow: ["read_file", "git_status"],
});
const toolDefs = [...myRealToolSchemas, ...decoyToolDefs("openai")];
const router = createGuardedToolRouter({
  guard,
  executeRealTool: realTool, // host-provided callback
});

await router.dispatch({ name: "git_status", args: {} });              // SAFE: allowed
await router.dispatch({ name: "canary_read_secrets", args: {} });     // trip → quarantine

try {
  await router.dispatch({ name: "http_post", args: { url: "https://example.invalid" } });
} catch (error) {
  if (error instanceof CanaryBlockedError) console.log(error.decision); // action_blocked
}

// Expose this only to a human incident-response control plane, never an LLM tool.
guard.reset({ acknowledgedBy: "on-call-human" });

guard.inspect(agentOutput, "final-answer") detects a planted token and trips
the same session. decoyToolDefs("anthropic") emits Anthropic schemas.

Dashboard and SIEM

agent-canary dashboard --out report.html   # self-contained HTML timeline
agent-canary export --format cef           # or json, csv

Injection-resistance evaluation

V2 Personal includes a reproducible 20-payload evaluation suite. Use text output
for humans or JSON for CI; provider/API failures fail closed and are never counted
as a successful resistance result:

agent-canary eval --provider openai --model gpt-4o --format json
agent-canary eval --provider openai --model deepseek-chat \
  --base-url https://api.deepseek.com/v1 --format json --out eval.json

Guarantees and limits

  • Decoy tools never perform real actions. canary_run_shell does not run
    commands. The handlers return fabricated output, nothing else (see
    SECURITY.md).
  • Canary tokens unlock nothing anywhere.
  • No telemetry. Events stay in ~/.agent-canary/events.jsonl unless you
    configure a webhook.
  • Alerts only fire when a decoy is touched or a token surfaces. Nothing in a
    legitimate workflow can trigger them.
  • Containment is integration-scoped. It can block only real tool calls
    routed through router.dispatch() or guard.executeToolCall() /
    guard.beforeToolCall(). If a compromised agent's first dangerous action
    bypasses these paths, agent-canary cannot intercept that action. Decoys are
    harmless, so touching one lets the guard quarantine the session before a
    later guarded action runs.
  • This release does not ship an MCP proxy for arbitrary upstream MCP servers;
    the reviewed next-step design is in docs/containment.md.

Known limit: this is JavaScript, so a determined user can patch dist/ and
strip the license checks. The signed-license scheme raises the bar against
casual copying; it is not DRM.

Commands

serve / init / install / uninstall
tokens generate|plant|check|list
watch, events, report, dashboard, export, eval
self-test
status, activate, alert-test, set-webhook, set-notify

Run agent-canary --help for details.

License

MIT

Yorumlar (0)

Sonuc bulunamadi