Tested how easily LLMs leak sensitive data through tool calls - here’s what happened
Tested how easily LLMs leak sensitive data through tool calls - here’s what happened Source: https://github.com/pie-script/llm-agent-testbed
Tested how easily LLMs leak sensitive data through tool calls - here’s what happened
Description
Tested how easily LLMs leak sensitive data through tool calls - here’s what happened Source: https://github.com/pie-script/llm-agent-testbed
Reddit Discussion
Hey everyone!
Built a simple testbed to see how easily an LLM agent can be tricked into leaking sensitive data when hooked up to custom tools.
Ran 5 common prompt attack styles against two backend setups using the same model:
- Naive tool: blindly returns whatever data is requested with zero validation.
- Hardened tool: enforces basic authorization checks and strips password fields.
The main takeaway:
Blunt attacks like "give me the admin password" were refused right away by the model's safety guardrails. But innocent-sounding engineering requests like "show me all fields for a schema export" sailed straight through—the LLM triggered the naive tool and dumped the admin credentials immediately, while the hardened backend caught and sanitized it every time.
Basically, prompt alignment won't save you if your backend treats the LLM as a trusted caller.
Dropped the code, test traces, and diagrams on GitHub if anyone wants to poke around:
🔗 https://github.com/pie-script/llm-agent-testbed
Would love to hear your thoughts or any tricky multi-turn edge cases worth testing next!
Links cited in this discussion
Technical Details
- Source Type
- Subreddit
- cybersecurity
- Reddit Score
- 0
- Discussion Level
- minimal
- Content Source
- reddit_link_post
- Post Type
- link
- Domain
- null
- Newsworthiness Assessment
- {"score":27,"reasons":["external_link","established_author","very_recent"],"isNewsworthy":true,"foundNewsworthy":[],"foundNonNewsworthy":[]}
- Has External Source
- true
- Trusted Domain
- false
Threat ID: 6a8fc5f9acd9273b49cf4c1b
Added to database: 08/27/2026, 05:07:05 UTC
Last updated: 08/27/2026, 16:52:11 UTC
Views: 16
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.