Replace Agent Adjectives With Verifiable Verbs

“Be careful” may be the least testable instruction in a coding repository.

It sounds responsible. So do “make a high-quality change,” “be thorough,” and “use good judgment.” These phrases give an instruction file a reassuring tone and give the agent almost no observable boundary.

When the result was wrong, the only diagnosis available was that it had not been careful enough.

A useful operating contract needs a different grammar: replace adjectives about desired character with verbs that produce evidence.

“Careful” has no completion state

Suppose an agent edits six files while fixing a one-line bug. Was that careless?

Maybe. The extra files could contain required tests and generated outputs. Or they could be unrelated cleanup. The adjective cannot decide.

An operational rule can:

Before editing, identify the smallest set of files required for the request.
After editing, inspect the diff and remove unrelated changes.

Now there are actions to observe: identify scope, inspect diff, remove unrelated changes.

The same conversion works across common instructions:

  • “Be accurate” becomes “trace factual claims to an approved source; mark unresolved claims unknown.”
  • “Be safe” becomes “do not mutate external state without an explicit apply or approval boundary.”
  • “Test thoroughly” becomes “run the repository’s relevant test command and report its exact result.”
  • “Avoid scope creep” becomes “name the requested behavior and do not modify files unrelated to it.”

The rewritten versions are longer. They are also debuggable.

Separate policy domains

One paragraph often tries to make an agent safe in every possible way:

Make focused, accurate, secure changes and verify them carefully.

That sentence mixes at least four policy domains.

Scope asks what may change. Provenance asks what may be asserted. Authority asks which side effects are permitted. Completion asks what evidence is required before the task is done.

Keeping those domains separate matters because one can pass while another fails. A two-line patch may be perfectly scoped and based on an invented API. A well-sourced implementation may still publish or delete something without permission. A safe local edit may be declared complete without running its test.

Use headings that match the failure you want to diagnose:

## Scope

## Evidence

## Mutation authority

## Verification

## Completion

When a rule is violated, the review can name the boundary instead of grading the assistant’s personality.

Give uncertainty an output

Instructions often demand certainty indirectly. They define what a complete answer looks like but provide no valid form for missing evidence.

The model then faces a bad choice: stop the entire task or fill the gap with something plausible.

An operating contract should make uncertainty representable:

If a required fact cannot be verified, label it [UNKNOWN], state what source is missing, and continue with claims that are supported.

That rule does not reward ignorance. It prevents unknown information from being disguised as completion.

The same pattern applies to code. If a platform path cannot be tested in the environment, report the unrun check and the limitation. Do not quietly promote a static inspection into runtime proof.

A visible gap is actionable. A smooth guess is not.

Completion needs an evidence-bearing closeout

Agents can perform the right work and still summarize it vaguely.

“Tests pass” omits which tests. “Updated the implementation” omits where. “No other changes” may be written without inspecting the diff.

Require a small evidence-bearing closeout:

  • files changed and why;
  • exact verification commands and outcomes;
  • warnings or checks not run;
  • remaining assumptions or risks;
  • a final diff or status inspection.

This is not ceremonial reporting. It is the last point where an accidental file, skipped command, or unresolved assumption can become visible before review.

The closeout should stay proportional. A documentation typo does not need a compliance dossier. It may need a link check and a clean diff. A release automation change needs more.

Prose should point toward enforcement

Instruction files are not sandboxes. A sentence cannot physically prevent a push, a deletion, or an invented claim.

But observable rules reveal which failures recur. A repeated “inspect the diff” violation can become a hook. A recurring unverified metadata problem can become a schema check. A mutation boundary can move into tool permissions.

Vague adjectives cannot make that transition because there is nothing precise to enforce.

Open one agent instruction file and circle every word such as careful, quality, appropriate, or thorough. For each, ask what a reviewer should be able to see in the transcript, filesystem, or test output.

Write that action instead.

The goal is not to make the agent sound less human. It is to make success less dependent on whether “careful” happened to mean the same thing to both of you today.