A verifiable coding task tells Autohand Code what behavior should change, where to look, what must remain compatible, and how to check the result. Use examples of inputs and outputs to resolve ambiguity. Treat the final diff and validation results as the evidence of completion.

The task template

Outcome: [observable behavior to implement or repair]
Context: [entry point, file, issue description, or failing command]
Constraints: [public APIs, dependencies, files, and data to preserve]
Acceptance: [input -> expected output, plus error and boundary cases]
Validation: [repository-defined command and any manual checks]
Deliverable: explain the diff, tests run, results, and remaining limitations.

Use the actual repository test command. A placeholder such as npm test is only appropriate when that script exists. If you do not know the correct command, ask the agent to inspect the project instructions first.

Debug a failing behavior

The task list marks every task complete when I complete task a.
Inspect src/tasks.js and test/tasks.test.js. Reproduce the failure first.
Complete only the matching ID and preserve the input array.
Add unknown-ID and empty-array cases. Keep the exported API unchanged.
Run npm test and explain why the new assertions catch the original bug.

The strongest bug prompt includes the observed failure and expected behavior. Avoid prescribing an implementation until you understand the cause; an instruction to add a delay or suppress an exception may hide the failure.

Add a feature with explicit boundaries

Add filterTasks(tasks, status), where status is all, open, or done.
Return matching tasks in their original order without mutating the input.
Reject any other status with TypeError. Do not add dependencies.
Write tests for each status, invalid input, and an empty list.
Run npm test and npm run check. Report any unchecked assumptions.

Choose the right amount of context

Task Include Avoid
Bug fix Reproduction, error, expected behavior, related test Entire unrelated logs
Refactor Compatibility contract and observable baseline A broad request to clean everything
UI change Target screen, interaction, viewport, acceptance criteria A style adjective without behavior
Review Diff boundary, risk areas, severity definition A request for findings with no evidence threshold

For UI tasks, tests of component state do not prove rendering or keyboard behavior. Ask for the specific browser interaction you need and review the result at the relevant viewport.

Course-correct early

If the agent expands scope, restate the outcome and exclusions in a short follow-up. If it is stuck, ask for the failing command, output, current hypothesis, and smallest next experiment. If it reports success without execution, ask it to distinguish a proposed check from a completed check.

Completion checkpoint

A reviewer can determine whether the change meets the task without reading the entire conversation. For a worked example, follow fix a bug with a regression test. For planning a larger change, see plan mode.

Frequently asked question

What should a coding prompt include?

Include the observable outcome, relevant source context, compatibility constraints, acceptance examples, validation commands, and the evidence you expect in the final report.