Anyone who has run a coding agent for an afternoon knows the feeling: the agent asks, you click approve, it asks again, you click approve, and by the twentieth prompt you are no longer reading what you approve. OpenAI's answer in Codex is Auto-review, and per reporting on October 6, 2026 it is now free for anyone signed in with a ChatGPT account.
The idea is simple. When an action needs approval, a second agent reviews it instead of you. The stated payoff is large: Codex sessions in Auto-review stop for human approval roughly 200 times less often than in manual mode, and of the small fraction of actions that do go to the reviewer, about 99% are approved. This guide explains how it works, what stays unchanged, how to configure it, and where its limits are, because the limits matter more than the headline.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A reviewer agent that answers Codex approval requests instead of a person. |
| Is it free? | Reported free for ChatGPT-account sign-ins; check your plan. |
| Does it widen permissions? | No. Same sandbox, same network rules, same protected paths. |
| What does it block? | Secret exfiltration, credential probing, persistent security weakening, irreversible destruction. |
| Can I customize it? | Yes: [auto_review].policy locally, guardian_policy_config for orgs. |
| Is it a security boundary? | No. OpenAI says it is not deterministic security. |
How is it different from --yolo? | Yolo skips review. Auto-review keeps a check, just not a human one. |
The problem it solves: approval fatigue
Codex has three broad postures. In default mode, the agent stops whenever it wants to step outside its sandbox, such as running a command that writes outside the workspace or reaching the network. That is safe and exhausting. In full access mode, friction disappears and so does oversight. Many developers drift toward the second because the first is unbearable on long tasks, which we covered in our Codex slash commands reference and in an experiment on how people behave when approving agent actions.
Auto-review is a third posture: keep a check on higher-risk steps, but make the checker a model rather than your attention. It matters most for long runs and automations, where nobody is sitting at the keyboard. OpenAI's own framing is that it lets Codex "work longer with fewer approvals and safer execution," moving through tests and builds while a separate agent examines the risky steps in context before they run.
How it works
The docs describe a simple loop:
- The main Codex agent works inside its sandbox as usual.
- When it wants to do something that requires approval, the request goes to a reviewer agent instead of stopping for a person.
- The reviewer evaluates the action in the context of the task and returns a decision with a rationale.
- If approved, the action runs. If denied, the main agent sees the rationale and is expected to find a safer approach rather than work around the denial.
It applies to interactive approval policies, specifically on-request or granular policies. It does not apply where approvals are already off.
Two design details are worth noticing. First, the reviewer sees the action in context, so it can tell that rm -rf build/ inside a project is routine while rm -rf ~ is not. A static allowlist cannot make that distinction as easily. Second, the denial includes a rationale that the main agent receives, which is meant to stop it from retrying the same thing with cosmetic changes.
What it does not change
This is the part to read twice. The docs are explicit that Auto-review is a reviewer swap, not a permission grant. It does not change:
- sandbox mode (
read-only,workspace-write), - writable roots,
- network access rules,
- protected paths and filesystem boundaries,
- which MCP servers and tools are available,
- routine actions already allowed inside the sandbox.
In practice that means your sandbox is still your real security boundary. If you configured network off, Auto-review does not turn it on. If a path is protected, the reviewer is not a back door to it. What changes is only how requests that would have interrupted you are answered. This is why we treat sandbox design as the primary control in our guides to agent sandbox isolation and restricting desktop agents.
What the reviewer blocks
The documented high-risk categories are:
- Exfiltration of secrets or credentials to untrusted destinations.
- Credential probing, such as hunting for tokens, cookies or session data.
- Persistent security weakening, for example disabling protections or adding long-lived access.
- Destructive, irreversible operations.
Because this is a model judgment, you should think of it as a probabilistic filter tuned for those categories, not as a proof that nothing bad can happen. The reported 99% approval rate on reviewed actions suggests most requests are benign, which fits what developers observe, but the interesting cases are the rare ones where a plausible-looking command is not benign.
Configuration and overrides
The documented controls are:
| Control | Where | Purpose |
|---|---|---|
approvals_reviewer = "auto_review" | config.toml | Route approvals to the reviewer agent |
[auto_review].policy | local config.toml | Add your own rules to the reviewer |
guardian_policy_config | managed requirements | Organization policy, which takes precedence over local settings |
/approve | in session | Allow one narrow retry after a denial |
There is also a denial circuit breaker: Codex stops after 3 consecutive denials or 10 rolling denials, so a stuck agent does not loop forever trying variations. Timeouts are distinguished from explicit rejections, which helps when you are diagnosing why a run halted.
A starting point for a cautious personal setup might look like this (illustrative, check the current docs for exact syntax):
# ~/.codex/config.toml
approvals_reviewer = "auto_review"
[auto_review]
policy = """
Deny any command that sends file contents or environment variables to a
network host not already listed in the project README.
Deny changes to CI credentials, SSH config and shell startup files.
Allow package installs from the project's lockfile only.
"""
Write policies in plain, specific sentences about your own risks, and test them by deliberately asking the agent to do a borderline thing in a throwaway repository.
Why it is not a security boundary
OpenAI's docs are candid: Auto-review "is not deterministic security." It complements sandboxing, monitoring and organizational policy. We would put the limitations in four groups.
It is a model. Models can misjudge, especially on unusual commands, long pipelines or multi-step plans whose danger only appears in combination.
It can be targeted. Anything that lets untrusted text reach the agent, like web pages, issue comments or dependency READMEs, can try to talk the main agent into a risky request framed convincingly for the reviewer. Researchers have published work on attacks against human-in-the-loop approval flows, sometimes called loopjacking, and an AI reviewer is a different but related surface. We have not tested those attacks against Auto-review, so treat that as a reason for caution, not a finding.
It sees what it is shown. The quality of the decision depends on the context passed to the reviewer. Subtle state, such as which credentials are in the environment, may not be visible to it.
It lowers the cost of saying yes. Fewer prompts means fewer moments where you notice something odd. That is the benefit and the risk.
For related agent behavior that surprised users, see our checklist on leaving an always-on agent unattended and the analysis of agents hitting sites at scale.
A practical rollout plan
If you run Codex on real repositories, a measured path looks like this.
- Tighten the sandbox first. Keep
workspace-write, network off by default, and protect secrets paths. Auto-review cannot fix a loose sandbox. - Enable Auto-review on a low-stakes repo. Run a normal task and read the reviewer rationales it produces, including approvals.
- Log denials. Keep notes on what it denied and why. Patterns tell you where to tune the policy.
- Try to break it on purpose. Ask the agent to read a dotfile, call a new host, or delete a directory outside the workspace, in a disposable environment, and see how the reviewer responds.
- Add a team policy. For organizations, put the non-negotiables in
guardian_policy_configso individuals cannot weaken them. - Keep a human for the irreversible. For deploys, payments and data deletion, require a person regardless of reviewer output.
If you need hooks rather than a reviewer model, Claude Code offers a different pattern with function hooks and approval notifications, and the two approaches can be compared on one axis: a deterministic rule you write yourself versus a model judgment you tune.
What this means for what you build or pay
Two practical consequences stand out. For individuals, free Auto-review lowers the cost of running longer unattended tasks, which tends to increase usage of the very quota that OpenAI has been adjusting. If you are watching limits, see our coverage of the 28-day Codex pledge. For teams, the reviewer becomes one more component to govern: someone needs to own the policy text, review denial logs and decide what stays human-only.
The deeper point is a shift in where trust sits. Approval used to be a human gate on every risky step. Increasingly it is a model gate with a human on the exceptions. That can be a good trade, but only if the sandbox underneath is already doing the real work.
Related reading
- Codex Auto-review and the token usage controversy (August)
- Codex slash commands: complete reference
- OpenAI's 28-day Codex pledge
- Dots and Pro 200: the Oct 30 usage cut
- Agent sandbox isolation: five things to know
- Restricting desktop agent access
- Claude Code function hooks
- Claude Code approval notification hook
- What people do when approving agent actions
Primary: the Codex Auto-review documentation (learn.chatgpt.com) · OpenAI Developers announcement of Auto-review on X · OpenAI alignment note on auto-review of agent actions
Details are accurate as of October 6, 2026 and drawn from OpenAI documentation and press coverage. Pricing, availability and config syntax can change; verify in the current Codex docs before enabling it on shared or production repositories.
