How far should we let AI agents act and write code unsupervised?

Set an agent's authority by the consequence and reversibility of what it can do. Separate proposing a change from approving and releasing it, keep a working stop mechanism, and name who resolves uncertain outcomes. Judge coding agents by accepted, maintainable work on your hardest code.

Draft for Tamir's review. Not published.

Tamir Khason · Updated · Decision guides

Set authority by consequence

List the actions an agent can take that spend money, create external commitments or change customer records. Mark the ones that are irreversible or hard to unwind. Those need limits, an approval gate and a stop mechanism that works. If an agent repeated a completed action after a retry, suspend that action until you can confirm completion.

Give coding agents task-specific access, and keep release authority with a person. Check whether the agent can read production secrets, and test branch protections against its actual credentials. Treat documents an agent reads as untrusted evidence, never as instructions.

Measure accepted work

Compare accepted changes with the review and correction effort they create. Code volume can rise without useful progress, so it is a poor measure and a worse incentive.

When a vendor claims AI tools make a rewrite several times faster, test the claim on the hard part of your system. Fund the module with the most undocumented business logic first, with acceptance tests you write. Measure the rate including review and defect fixing. The tools cannot recover rules nobody wrote down.

Depending on your seat

If you're on the board, ask management to define the agents' authority before widening it, with legal and finance confirming the delegation. Ask that engineering incentives rest on accepted capability rather than generated code volume.

If you're the CEO or CTO, grant access around useful work and observable outcomes. Before you shrink a team because a partner uses AI, list the responsibilities the smaller team covers, who reviews generated code and what knowledge you must keep.

What to check before you decide

  • List agent actions that spend, commit the company or change customer records, and mark the irreversible ones.
  • Separate code generation from approval and release, and test branch protections against the agent's real credentials.
  • Check whether retrieved documents can trigger actions or expose another customer's data.
  • Test interrupted and repeated actions, and name who resolves uncertain outcomes.
  • Compare accepted changes with the review and correction effort they created.
  • Test an AI rewrite claim on your hardest module, with acceptance tests you write.
  • Confirm ownership of generated code and the tool terms on your source code.

Questions people ask

What limits should boards set before agents can commit company resources?

Define authority by consequence and reversibility before widening autonomy. Approval depends on commitment limits, attributable actions, exception handling, and legal review where actions can bind the company.

Should boards reward engineering output measured mainly by AI-generated code volume?

Use measures connected to accepted capability and sustainable delivery. A suitable approach depends on product goals, quality, review effort, and the behaviors the incentive will encourage.

Should we give every developer unrestricted access to paid coding agents?

Set access around useful work and observable outcomes. Expansion depends on accepted changes, review burden, security boundaries, and whether agent use improves the team's actual delivery constraints.

Should we hire fewer engineers because an outsourcing partner uses AI?

Compare required capabilities and accountable outputs rather than headcount claims. The choice depends on review capacity, product knowledge, security responsibilities, and how exceptions are handled.

Vendor claims AI coding tools make our legacy rewrite several times faster, how do I test that claim before funding it?

AI tools speed up the parts of a rewrite that were never the bottleneck, so test the claim on the hard part of your system, with your data and your acceptance tests. Fund a bounded first module with defined outputs and judge the rate from that. It depends on how well your legacy behavior is documented and tested, because the tools cannot recover rules nobody wrote down.

Is our financial document assistant ready for untrusted customer files?

Treat retrieved documents as untrusted evidence rather than operational instructions. Viability depends on permission boundaries, constrained tools, and observable behavior under adversarial inputs.

Should coding agents be allowed to approve or release production changes?

Give agents task-specific access and separate proposal from release authority. The boundary depends on repository controls, secret exposure, and which actions require accountable approval.

Should we suspend autonomous customer actions after an AI agent repeats transactions?

Suspend the affected autonomous action until its outcome and safe operating boundary can be established. It depends on the consequences of repetition, the ability to confirm completion, and an approved supervised alternative.

How I can help with this decision

Ask or talk (Free)
I give my view on which agent actions need limits first, and the bounded test that shows what the tools really deliver on your system.
Review (Pay if it was worth it)
I write an independent assessment of the agents' authority, permissions and test evidence, or of a vendor's speed claim. I recommend what to allow, restrict, gate or defer.
Retain (When it makes sense)
I stay close as use grows, to review agent permissions, material exceptions and each module's acceptance.