LLM Threat Model and Agent Security

Luận điểm chính

Threat model của LLM bắt đầu từ việc instruction và data cùng đi vào một chuỗi token. Vì vậy rủi ro không chỉ nằm ở prompt input, mà trải qua retrieval, model, tool, output, monitoring và supply chain.

Bản đồ rủi ro

Defense posture

Không có một filter đủ mạnh để giữ mọi thứ. Cách bền hơn là defense in depth: giảm quyền tool, tách untrusted content khỏi privileged action, giữ provenance, validate/sanitize output, monitor anomaly và dùng human review cho action hậu quả cao. Agents Rule of Two là checklist nhanh để tránh gom ba capability nguy hiểm vào cùng agent.

Liên kết