RedCell
RedCell — Security Evaluation for Tool-Using LLM Agents
Creator and Researcher
An open-source defensive security evaluation platform for tool-using LLM agents. It runs budgeted attack searches, distinguishes unsafe intent from attempted and realized tool actions, and preserves every result as auditable evidence.
View source codeWhat did the agent intend, attempt, and actually do?
Record unsafe intent separately from execution.
Challenge
Tool-using LLM agents call databases, files, and business APIs, so a manipulated agent can leak data or take actions it shouldn't. A fixed list of jailbreak prompts doesn't cover this, because the attack surface comes from natural-language decisions and non-deterministic models.
Approach
I treated red teaming as an evidence-producing experiment system, not a prompt generator. Deterministic instrumentation scores canaries, permission checks, and cross-user access; fixed token and cost budgets plus fingerprinted conditions make Static, Random, Thompson, and LLM-driven controllers comparable.
Implementation
Delivered an end-to-end arena, multi-turn execution, trace persistence, crash-safe resume, and a fail-closed six-condition experiment runner. A 776-test offline suite validates the engineering path; an 18-run, 1,080-attempt online pilot did not support adaptive superiority, and the next formal provider matrix remains pending.
View technical detailsHide technical detailsTech stack and implementation notes04
Tech stack
Python 3.11, asyncio, Pydantic, SQLAlchemy, Typer, Jinja2, SQLite, pytest, Ruff, and Black. The automated suite runs fully offline without provider credentials.
- Built an open-source defensive security evaluation platform for tool-using LLM agents, covering prompt injection, sensitive-data disclosure, and unauthorized tool use with deterministic Intent / Attempt / Impact evidence.
- Designed a six-condition, token-budgeted experiment framework comparing Static, Random, Thompson, and LLM-driven controllers, with fingerprinted configurations, role-level token accounting, SQLite traces, and crash-safe resume.
- Hardened live-model execution with fail-closed condition validation, cross-process SQLite rate limiting, and shared 429 cooldowns; validated the system with a 776-test offline suite plus Ruff and Black.
- Ran 18 online Phase 0 pilot runs covering 1,080 attack attempts with zero abandoned attempts; reported the adaptive hypothesis as NOT SUPPORTED and retained Static/Random baselines.

