Delivering Angular with AI coding agents: clear roles, checks and human ownership

AI-assisted development, Coding agents, Angular, CI · Published 1 October 2026 · Updated 2 October 2026

The useful part

Agree the outcome and the acceptance check before delegating. A passing test informs a human decision; it cannot own product quality.

An AI-generated change can compile, pass its own tests and still be the wrong change. For a client, the useful question is who owns the result: who chooses the scope, checks the design and accepts the release. This is the process used for GeoAtlas, clearcraft’s own product, and the delivery structure I bring to an Angular engagement.

Delivery flow: the human defines the brief, an agent implements within scope, and the human checks the evidence.
In practice / Process diagram. A suggested structure for bounded tasks, not measured output from a client engagement. Open the image for a larger view.

One human, bounded agent work

The human owns product goals, design decisions, scope and final acceptance. A main working session can help reason through those decisions. Delegated agents receive mechanical work whose decisions have already been made: gather specified evidence, run a named check or implement a fully specified change.

Read-only work is not automatically mechanical. Listing imports is evidence collection; deciding whether the architecture is good is a review. That distinction matters more than the model’s name or price. A stronger model does not expand a delegated agent’s authority.

  • Evidence collection: return the requested files, counts or command output with their revision and scope.
  • Implementation: change the assigned files against explicit behaviour and acceptance criteria; report a missing decision instead of inventing one.
  • Verification: execute named checks and return the raw result. The main session reviews the diff and evidence; the human remains accountable for acceptance.

A brief is a contract

Give each task one observable outcome, explicit file ownership, constraints, a proving command and a stopping condition. Supply the context it needs even when the tool can inherit a conversation. Two agents should not edit the same file concurrently; separate checkouts help, but shared configuration and CI still need coordination.

Example brief: preserve selection when filtering a table

Illustrative task, not a client result. The decision is already made: selected record IDs survive a filter change, including IDs temporarily hidden from the view.

Outcome: changing a filter does not clear selected record IDs.
Own: selection-state.ts and selection-state.spec.ts only.
Keep: public API, row IDs and permissions unchanged.
Prove: select A; filter A out; remove the filter; A is selected.
Check: run the selection-state spec and paste its output.
Stop: report NEEDS_CONTEXT if the behaviour is ambiguous.

The brief names a user-visible result. “Improve selection handling” leaves the most important decision to the implementer.

Tests first, with a bite

For a behaviour change, first run the existing relevant tests. Add a regression test and see it fail for the intended reason, then implement the change. A targeted mutation — for example clearing the selection on every filter update — should make the new test fail again. Restore the implementation and verify it passes.

This checks that a test protects the behaviour. It does not prove that the requirement itself is correct or that all edge cases are covered. That still needs review. Cosmetic changes need proportionate visual and accessibility checks, rather than tests that merely repeat the markup.

The gates provide evidence

In GeoAtlas, cheap checks can run locally. Browser suites and visual baselines run on a controlled CI runner. A fast gate, a full browser gate and a nightly deployment check answer different questions; the CI article explains the split and its release trade-off.

Passing tests are necessary evidence for the behaviours they cover. They cannot decide whether the interface is useful, the article is accurate or the architecture fits the product. The person reviewing the change makes those decisions.

Evidence rules for reports

  • Report the exact command, revision, exit code and relevant output. Distinguish a test declaration count from executed passing tests.
  • Name skipped and unexecuted checks. A passing unit spec is not a passing end-to-end suite.
  • Calling a failure pre-existing, flaky or unrelated needs evidence from a baseline or a controlled rerun. Without it, the cause is unresolved.
  • Keep the hand-back short and link to durable logs. The main session reads the diff and independently checks the evidence before accepting it.

Failure modes worth designing for

Shared-page locators. A new demo can make an older page-wide locator ambiguous. Scope locators to the relevant region and run the affected page’s tests, including its existing scenarios.

Concurrent edits. Restoring a whole file during a mutation check can erase another task’s work. Use isolated checkouts or targeted edits, and coordinate changes to shared types and configuration.

Competing CI work. A gate and an interactive test can starve each other on one runner. An explicit queue or shared admission rule makes the test environment part of the evidence.

Unbounded context. Large transcripts and vague briefs increase repeated exploration. Smaller tasks and durable notes make resuming work cheaper to reason about. Model cost and elapsed time depend on the task and provider; an anecdotal day of usage is not a transferable estimate.

What the client receives

A useful handover contains the change, its reason, the checks for that revision, known limits and a rollback path. AI is part of how the work gets done. The deliverable is a maintainable application and evidence a team can inspect.

The delivery process shows the wider sequence, from a user task to release. Start with one bounded stage and agree what would demonstrate that it works before implementation begins.

Sources and scope

This is a description of clearcraft’s working rules and own-product experience. The brief is illustrative; there is no claim of a controlled productivity or cost comparison.