Agent Safety Codex Note
AI Implementation Operations
Observation LogOBSERVATION RECORD / B59BFD69RECORDED : 2026-05-10DOMAIN : IMPLEMENTATIONSTATUS : ARCHIVED

Keep AI agents out of production.
Build a development sandbox where it can break things and learn

The most dangerous assumption when delegating work to an AI agent is that it will be fine because the AI is smart. AI works quickly, which also means the damage from a mistake can spread quickly. Safe delegation does not come from trusting AI; it comes from giving it an environment that can be restored after something breaks.

A workflow that keeps AI agents out of production, tests changes in a sandbox, and applies them only after diff review
Framework: Test in a sandbox → limit the scope of changes → review the diff before applying them

The confidence to delegate to AI comes from recoverability, not trust

Given a task, an AI agent can read and modify many files and even create a commit. That is powerful, but also risky. It may make sweeping changes to areas a person would handle carefully while misunderstanding the context.

The question is not whether to trust AI. The workflow should distinguish what AI can handle from what still needs human judgment. AI brings speed; people bring judgment and verification. With that division in mind, let the agent work somewhere that can be restored if it breaks something.

In other words, safety comes from the environment, not the AI's disposition. A sandbox repository, a working branch, a pull-request checkpoint, diff review, and clear prohibitions make it easier to delegate work to AI.

Before

Make changes directly in production

  • Change a broad area all at once
  • Proceed while misunderstanding the context
  • Let the impact of a mistake spread
After

Test in a sandbox

  • Work in a sandbox repository
  • Define the delegated scope in advance
  • Build safety into the environment

The confidence to delegate to AI comes from recoverability, not from trusting AI.
It comes from a setup that can be restored after a failure.

Set up a sandbox where things can break before letting AI touch the production branch

If you delegate site changes or article publishing to AI, keep it away from production at first. CSS, JavaScript, shared templates, category structure, navigation, and the sitemap can all have site-wide effects when changed in one place.

A sandbox repository or backup environment makes it easier to delegate broader tasks. A failed attempt does not affect production; if the sandbox breaks, discard it. If the work succeeds, apply only the reviewed changes to production. This sequence greatly reduces the risk of using AI.

Branches are useful too, but beginners may find it clearer to separate work into different repositories. It is immediately obvious which one is production and which one is the sandbox. In AI-era development, the structure should be technically sound and easy for its operators to understand.

Before
  • Often stops at simply forbidding production access
  • The scope can still be unclear around shared CSS and the sitemap
  • The broader the task, the harder it is to predict its impact
With a sandbox
  • The sandbox does not affect production
  • Beginners can separate work by repository
  • Use a pull-request checkpoint as a safety valve

AI needs guardrails more than a lecture

Telling AI to “be careful” does not make it safe. Specify what it may and may not touch, what counts as complete, and how the work will be reviewed. These are less like prompt instructions and more like guardrails around a worksite.

For example, when adding an article, limit changes to the article directory, index pages, and sitemap. Do not let the agent modify rules files such as AGENTS.md or DESIGN.md. Allow changes to shared CSS only when explicitly requested in advance. Treat source images as read-only and have the agent copy them into the article directory.

These rules are not meant to restrict AI for their own sake. They allow you to delegate more safely. Guardrails let AI move quickly within a defined area; accidents happen when it is allowed to run fast with no boundaries.

Step 01
The confidence to delegate to AI comes from recoverability, not trustGiven a task, an AI agent can read and modify many files and even create a commit.
Step 02
Set up a sandbox where things can break before letting AI touch the production branchWhen delegating site changes or article publishing to AI, keep it away from production at first.
Step 03
AI needs guardrails more than a lectureTelling AI to “be careful” does not make it safe.
Step 04
Use automated pull requests as the default, not automatic publishingThe most reliable way to operate an AI agent is to let it create the work automatically, then stop at a pull request—not to publish everything automatically.

Use automated pull requests as the default, not automatic publishing

The most reliable approach to AI agents is to let them create the work automatically and stop at a pull request, rather than publish automatically. Let AI write an article, add files, and update index pages and the sitemap. Before anything goes live, a person reviews the diff.

This extra review step does not significantly reduce the speed of AI-assisted work. It lets you delegate more with confidence. Full automation may look appealing, but a final approval checkpoint helps protect the quality and brand of published work.

Safe AI operations are not about constantly doubting the AI. Define where it can work at full speed and keep the final decision with a person. This lets AI take on more of the work while people retain responsibility for the outcome and its quality.

Starting point
Keep AI agents out of production; start in a sandbox…
Design
Treat AI not as a “great employee” but as an unstable, high-speed actor…
Asset
A pull-request checkpoint is a better safeguard than automatic publishing…

Distinguish weak practices from strong ones

The important question is not the surface-level task name, but what to preserve and what to turn into something reusable.

Weak Pattern

Unbounded delegation

  • Delegate a broad scope to AI
  • Experiment in an environment close to production
  • Rely on reminders to keep the work safe
Strong Pattern

Isolated operation

  • Test in a sandbox repository
  • Treat AI as a fast-moving execution agent
  • Define what must remain untouched in advance

An AI agent is not a capable employee; it is an unstable actor that moves at extreme speed.

An AI agent is not a capable employee; it is an unstable actor that moves at extreme speed. What it needs is not trust, but careful design around isolation, permissions, diffs, and approval.

An AI agent is not a capable employee; it is an unstable actor that moves at extreme speed.
What it needs is not trust, but careful design around isolation, permissions, diffs, and approval.

Source Notes
  • AI is much safer when it starts in a sandbox repository instead of production
  • Treat AI not as a “great employee,” but as an unstable actor that moves at extreme speed
  • A pull-request checkpoint is a better safety valve than automatic publishing
  • For a working system, instructions that prevent damage matter before expanding its scope