BoundBench

Generative Agents (Smallville)

Research code for 'Generative Agents: Interactive Simulacra of Human Behavior': an LLM-driven multi-agent town simulation with a Django visual frontend.

github.com/joonspk-research/generative_agents · 2026-10-04 · fe05a71

Defense-in-depth score

4.4 / 10

Minimal

Generative Agents is a research simulation, not a tool-using agent: its LLM-driven characters can only walk around a fixed map, talk to each other and remember things, so most of its safety comes from having almost nothing dangerous to do. It ships no controls of its own. The main risk is in the local web frontend, whose handling of model-written text and of requests is not hardened. Memory is written by the model without checks and carries into every forked run.

Criteria

C1 Identity & least privilege

Moderate 0.55 / 1.00

The simulation holds one credential: the operator's OpenAI API key, read from a hand-written utils.py and used only for model and embedding calls. The model has no tools, so it cannot use that key or any other operator credential to act on outside systems. Nothing narrows the key itself (it is a long-lived account key), and request handling on the local Django server that the simulation talks to is not locked down. A hijack therefore reaches the OpenAI quota and the local simulation files, not the operator's wider accounts.

C2 Approval gates

N/A · full credit 1.00 / 1.00

The agents in this project take no consequential actions. Their model outputs only pick a destination on a fixed game map, an emoji, a short action description, chat lines and memory entries, all written to the simulation's own JSON files. There is no shell, HTTP client, email or other side-effecting tool for an approval gate to guard. The operator commands that delete data (exit) are typed by the human, not chosen by the model.

C3 Tool & action scoping

Strong 0.72 / 1.00

The model's only real 'action' is choosing where a character walks, and that choice is checked against the map's fixed set of addresses before any path is computed. That is a narrow, allow-listed action space. The weak spot is how the text the model produces (chat lines and action descriptions) is handled in the browser, which is not sanitized. There are no tools to enable or disable and the model cannot add any.

C4 Code-execution isolation

Minimal 0.05 / 1.00

The project never runs model-generated code on purpose: there is no shell, eval or code tool. The report rates this criterion on how the simulator's browser frontend handles model-written text, which is not sanitized or isolated. The browser's own sandbox is the only boundary, and nothing in the project adds isolation.

C5 Untrusted input blast radius

Minimal 0.05 / 1.00

The agents read content their operator did not write: other agents' chat lines, memories carried over in forked simulation folders, and history CSVs loaded on request. All of it goes into prompts with the same standing as the system's own instructions, and nothing structurally limits what a manipulated agent can then do. In practice the damage is small because agents have no tools, no access to secrets and no outbound channel from the backend. The exception is the browser frontend, where chat and description text is not sanitized.

C6 Memory, context & configuration integrity

Minimal 0.15 / 1.00

Each agent keeps a long-term memory of events, chats and model-written reflections that is saved to JSON and reloaded whenever a simulation is forked. Anything the model writes, including text from other agents, is stored without validation and later retrieved into prompts as trusted context. Each agent's memory lives in its own folder and every fork is a copy, so earlier simulation states remain intact, but there is no review, provenance tag or expiry enforcement on what gets remembered. A poisoned memory persists across the operator's sessions and can keep steering agents' output, including the chat shown in the frontend.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The project loads no third-party code at runtime: no plugins, MCP servers, downloaded tools, model files or package installs. Models are reached only through the OpenAI API. pickle and selenium are imported but never used.

C8 Secrets & sensitive-data protection

Minimal 0.17 / 1.00

The OpenAI key lives in a plaintext Python file the operator creates (git-ignored), and secret handling in the committed Django settings is not locked down. The key is never put into prompts or logs and no telemetry is sent anywhere, but this follows from the design, not from any masking or secret-handling mechanism. The key is a long-lived account key with no scoping.

C9 Audit & traceability

Minimal 0.40 / 1.00

The simulation writes a JSON file for every step with each agent's movement, action description and chat, and saves memory on request. That is a structured, timestamped trail of what agents did, but it omits the prompts and model responses behind them. The files sit in the simulation folder that the local web server can also write to, and the 'exit' command deletes the whole folder. Errors are swallowed, so gaps go unnoticed.

C10 Limits & kill switch

Minimal 0.30 / 1.00

Each run is bounded by the step count the operator types (run N), and per-call retries and conversation turns are capped in code. There is no wall-clock limit, no API timeout and no token or cost cap, so a large step count spends without limit. Stopping is by Ctrl-C on a single process, but bare except blocks around the API calls can swallow the interrupt and let the loop continue.