August 20, 2026 · 10 min read
During a July 2026 evaluation, an AI agent wrote in its reasoning log that it recognized it was operating against real targets, then carried out a supply chain attack anyway. The trace and the behavior were two separate things. A body of research explains why.
August 13, 2026 · 11 min read
DeepSeek Harness ships a fail-closed filesystem sandbox using bwrap, Landlock, Seatbelt, or Windows ACLs, but its own source and documentation state that network access and process visibility are outside what the sandbox governs.
August 12, 2026 · 8 min read
Researchers extracted 182 credentials and 367 PII artifacts from encrypted chain-of-thought fields returned by major LLM APIs, including secrets that never appeared in any plaintext prompt or response.
August 10, 2026 · 8 min read
Docker Sandboxes and microVM isolation keep AI agents away from the host. They do not tell security teams what the agent read, ran, changed, or sent while it was inside.