vibebuilt
/ Vibe Coding / Autonomous AI Agent: What It Can Do Unsupervised
Vibe Coding • • 14 min read

Autonomous AI Agent: What It Can Do Unsupervised

What an autonomous AI agent can do without a human, mapped to the permission modes, sandboxes and hard caps that Claude Code, Codex and Copilot document.

Autonomous AI Agent: What It Can Do Unsupervised

An autonomous AI agent is a program that takes a goal, decides its own next steps, and executes them through tools without waiting for a person to approve each one. What it can do unsupervised is not a property of the model. It's a list of actions somebody allowed, and for the three coding agents that publish their permission systems in detail, that list is knowable to the line. Claude Code will edit files and run commands on its own in auto mode but refuses to force push or merge an unapproved pull request. Codex runs inside a sandbox with the network off by default. GitHub's Copilot cloud agent gets one branch, one pull request, and 59 minutes.

The pages ranking for this query define autonomy as "acts without human input" and then list industries. I'd rather define it as a permission ledger, because that is the only version you can configure, audit, or explain after something goes wrong.

Autonomy Is A Ledger, Not A Personality

Snowflake's page has the cleanest definition on the first page of results. An autonomous agent is a system that can perceive its environment, reason about a defined objective and take action with limited ongoing supervision. The loop is perceive, reason, act, check the result, repeat.

That loop says nothing about what "act" is allowed to include. Reading a file and running terraform destroy are both actions. So the practical question behind "autonomous" is which action classes run without a prompt, which prompt, which are denied outright, and what the agent physically cannot reach because of a sandbox or a firewall. Those four answers are the ledger.

The rest of this article fills that ledger in for three shipping agents, then shows what an unattended run looks like with each, then maps the ways unattended runs go wrong to the specific control that catches each one. The broader AI coding agents guide covers what these agents are; this post is about what they're allowed to do.

The Ledger For Three Real Agents

Everything in this table comes from each vendor's own permission documentation as fetched on 2026-09-15. Defaults change, so I'd treat the links as the source of truth and the table as a snapshot.

Action Class Claude Code (auto mode) Codex (workspace-write, on-request) Copilot Cloud Agent
Read repository files Runs Runs Runs
Edit files in the workspace Runs, classifier watches Runs Runs, on its own copilot/ branch only
Run tests and build commands Runs, classifier watches Runs inside the sandbox Runs in an ephemeral environment
Reach the network Runs unless it looks like exfiltration Off by default, asks Restricted by a firewall
Edit files outside the workspace Classifier decides; protected paths never auto-approve Asks Cannot, single repository
Delete pre-existing files irreversibly Blocked by default Within sandbox rules Within its branch
Push to a branch Runs; force push blocked Asks if network is off Runs, single branch
Trigger CI workflows Runs, by pushing Asks if network is off Waits for a human to click Approve and run workflows
Merge a pull request Blocked unless a human approved it Asks if network is off Cannot; a human must review and merge
Deploy, migrate, or destroy infrastructure Blocked by default Asks if network is off Cannot reach it
Time limit None documented None documented 59 minutes, hard

Three things stand out to me. First, only Copilot's agent documents a hard wall clock. The permission pages for the other two say nothing about one, so as far as those docs go they run until the task ends or you stop them. Second, the three products enforce the ledger in different places. Claude Code's auto mode uses a second model, a classifier, to review each action against the request. Codex uses an operating-system sandbox and a network switch. Copilot uses repository plumbing, branches, workflow approval and review rules. Third, none of the three lets the agent approve its own work. GitHub prevents the requester from approving the PR the agent opened, and Claude Code's classifier blocks approving Claude's own pull request.

Claude Code: Modes And A Classifier

Claude Code's permission modes page lists six modes. Manual, whose config value is default, runs reads only without asking. acceptEdits adds file edits and common filesystem commands. plan reads and proposes. dontAsk runs reads and pre-approved tools and denies anything that would have prompted. bypassPermissions runs everything and the docs reserve it for isolated containers and VMs. auto runs everything with background safety checks, and on Pro, Max, and Team plans it is now the built-in starting mode for terminal sessions.

Auto mode is the interesting one for this article, because it is the closest thing on the market to "unsupervised by default". The mechanism is a separate classifier model that reviews actions before they run and blocks anything that escalates beyond the request, targets infrastructure it doesn't recognize, or looks driven by hostile content the agent read. The blocked-by-default list on the page is longer and more specific than I expected. Download-and-execute patterns like curl | bash. Sending sensitive data to external endpoints. Production deploys and migrations. Force push. git reset --hard and other discards of uncommitted work. terraform destroy. Merging a pull request no human approved. Printing a live credential into the transcript. Launching another agent loop with --dangerously-skip-permissions.

There's a smaller list of things no mode auto-approves, including rm on critical paths and any tool matched by an explicit ask rule. And deny rules block in every mode, including bypassPermissions.

The docs also carry a warning I'd quote to anyone who reads "auto" as "safe". Auto mode reduces permission prompts but does not guarantee safety, and Anthropic says to use it for tasks where you trust the general direction, not as a replacement for review on sensitive operations. The classifier trusts your working directory and the git remotes that existed when the session started. A remote added mid-session is treated as external, which is a nice detail, and also a reminder that the trust boundary is something you set up before the run rather than something the model figures out.

Codex: A Sandbox And A Network Switch

OpenAI's approvals and security page describes three sandbox modes. read-only lets Codex read and run within a constrained sandbox but make no changes. workspace-write is the default, labeled Auto, where Codex can read files, make edits, and run commands in the workspace. danger-full-access is no sandbox and no approvals, and the page marks it not recommended.

Approval policy is separate from the sandbox. Under on-request, commands the sandbox allows run without a prompt, and Codex asks when it wants to edit outside the workspace or run something that needs the network. Under never, it doesn't ask and stays inside whatever the sandbox permits. The older untrusted policy has been retired and the docs say to remove it from configs.

The default that matters most is the network. By default, the agent runs with network access turned off, and you enable it explicitly with network_access = true. That single switch is why the Codex column above says "asks" for push, merge, and deploy. Without the network, those actions are impossible rather than merely discouraged. My read is that this is the most honest autonomy control of the three, because it's enforced below the model. A prompt injection can talk a model into wanting to push. It can't talk a closed socket open.

For unattended runs the page is explicit. Use codex exec --sandbox workspace-write. The old --full-auto flag is kept as a deprecated compatibility path and prints a warning.

Copilot Cloud Agent: Plumbing As Policy

GitHub's agent takes a different approach. Instead of a classifier or a sandbox flag, autonomy is bounded by the repository itself. The risks and mitigations page says only users with write access can trigger it. It pushes to a single branch, created for it with a copilot/ prefix. Workflows aren't triggered until a person with write access reviews the code and clicks Approve and run workflows. The draft pull requests it opens must be reviewed and merged by a human, and the person who asked for the PR can't be the one who approves it. Internet access goes through a firewall you can customize.

Then the wall clock. Each session has a maximum execution time of 59 minutes, which GitHub calls a hard limit that cannot be extended or bypassed. Combined with one repository, one branch, and exactly one pull request per task, the worst case is bounded in a way the other two can't match without you adding a container and a timeout yourself.

The tradeoff is capability. This agent can't deploy, can't touch a second repository, and can't merge. For a lot of teams, and for the kind of nightly cleanup task I'd hand off first, that is precisely the point.

An Unattended Run Contract

If you're going to let any of these run without watching, I'd write the run down before starting it. Not a prompt. A contract, with the flags.

Task: [one sentence, with the acceptance check named]
Agent and mode: [product, permission mode or sandbox, approval policy]
Isolation: [container or VM, non-root user, or "none" written honestly]
Network: [off, allowlist, or on, and why]
Allowed actions: [exact tool or command allowlist]
Denied actions: [deny rules, protected paths]
Stop conditions: [time cap, cost cap, "ask if the test count drops"]
Evidence to keep: [transcript, diff, command output, exit statuses]
Who reviews before merge: [a named person who is not the requester]

Filled in for the three products, the "agent and mode" and "isolation" lines look like this.

For Claude Code in CI with an exact allowlist, the docs give claude -p "run the test suite" --permission-mode dontAsk --allowedTools "Bash(npm test)" "Read". Nothing outside that list runs, and anything that would have prompted is denied. For a fully unattended run the docs give claude -p "<prompt>" --dangerously-skip-permissions and then require a container, VM, or the sandbox runtime, run as a non-root user. That "required" is doing a lot of work. The mode that removes every prompt is only documented for environments where the blast radius is a disposable machine.

For Codex, codex exec --sandbox workspace-write with the network left off is the documented shape. You get edits and test runs and nothing that leaves the machine.

For Copilot, you don't write flags at all. You assign an issue, and the platform limits described above are the contract. The stop condition is 59 minutes and the reviewer is whoever picks up the draft PR, who by rule is not you if you assigned it.

What I would not do is put bypassPermissions or danger-full-access on a machine that holds real credentials and call it autonomy. It's the absence of a ledger, and in my own monorepo the rule is the same one I apply to cluster manifests, which is that anything not written down gets reverted the next time someone deploys. Repository rules in an AGENTS.md file can tell the agent what you'd like, but they are guidance in its context window, not enforcement, and an unattended run needs enforcement.

What Goes Wrong Unsupervised, And What Catches It

Snowflake's page lists the failure modes honestly, which is rare on this results page. Cascading flawed reasoning, state corruption over long workflows, prompt injection through malformed data, runaway compute cost, and debugging difficulty. Those are real. What's missing is the mapping to a control, so here it is.

Failure What It Looks Like Control That Catches It
Scope creep The agent "fixes" three things you didn't ask for and one of them is a config file Allowlists (dontAsk, --allowedTools), single-branch and single-PR limits, diff review before merge
Test weakening The suite passes because an assertion was loosened or a test deleted A stop condition on test count, human review of the test diff, the classifier's check for changes concealed relative to the request
Secret leakage A key ends up in a commit, a PR body, or an external request Claude Code's classifier blocks printing live credentials and pushes that would send secrets outside the repo; Codex's network-off default; Copilot's firewall
Prompt injection via tool output Fetched content tells the agent to run something Classifier review of actions "driven by hostile content Claude read"; no network in the Codex sandbox; the firewall on Copilot
Runaway time or cost The loop retries for hours Copilot's 59-minute cap; for the other two, a timeout you add in the container or CI job, since neither permissions page documents one
Irreversible destruction git reset --hard, deleted files, terraform destroy Blocked by default in auto mode; outside the Codex workspace; unreachable from Copilot's branch

The row I'd worry about most is time and cost, because two of the three products leave it to you. An agent that can't push anything harmful can still burn a night's budget arguing with a flaky test. Put the timeout in the job runner, not in the prompt.

Is ChatGPT An Autonomous Agent?

ChatGPT the chat app is not. It answers the message in front of it and stops. The agent that ships with it is Codex, which OpenAI's pricing documentation says is included on every ChatGPT tier, and Codex is what runs in the sandbox described above. So the honest answer is that ChatGPT contains an agent and is not one, in the same way that Claude's chat interface is not Claude Code.

When someone says "ChatGPT did X to my repo", they mean Codex, and the next question should be which sandbox mode it was in.

What Are The 7 Types Of AI Agents?

The number seven isn't standard. Salesforce's page lists six types and Snowflake's lists five. Between the two lists the recurring names are simple reflex, model-based, goal-based, utility-based, and learning agents, and the count moves depending on whether a vendor adds multi-agent systems or hybrid setups to reach a rounder number. The taxonomy is about how an agent decides, not about what it's permitted to do, and it doesn't help you answer the unsupervised question. The scheduled piece on types of AI agent takes that list apart with working examples, so I'll leave it there.

Which Autonomous AI Agent Is The Best?

For unattended coding work, the best one is whichever gives you a ledger you can read and a wall you can trust, and the three above are the only ones on this results page whose docs let you answer "what can it do unsupervised" without guessing. Ranking them beyond that depends on where you want the enforcement to live.

If you want the enforcement in the platform, Copilot's cloud agent has the tightest built-in box and the smallest capability. If you want it below the model, Codex's sandbox with the network off is the hardest boundary to argue past. If you want the agent to do more and you're willing to trust a classifier plus a container, Claude Code's auto mode is the most capable of the three and its block list is the most detailed thing any of these vendors has published. My concern with all three is the same. The defaults are only as good as the machine they run on, and a permissive mode on a laptop full of credentials is not an autonomous agent, it's an unlocked door.

Start With A Bounded Task

I'd pick one task with a named acceptance check, fill in the contract above, and run it once while I watch. Then run it again without watching and compare the two transcripts. The difference between them is your actual autonomy budget, and it's usually smaller than the marketing and larger than the fear.

When that works, the next piece in this cluster, on how to build an AI agent, covers building one that does a single job with the same kind of ledger designed in from the start.