Skip to content

Project management with AI agents: from objective to verifiable tasks

Series II: From infrastructure to agentic projects in production. Episode 06 of 20.

In the previous episode, we turned the infrastructure of the fictional B2B portal into verifiable configurations. We must now organize the work so people and AI agents know what outcome they are pursuing, what they may change and how they demonstrate completion. The main decision is simple: do not delegate a vague intention. Delegate a bounded task with an acceptance owner and observable evidence.

An agent can propose a plan, work in the repository and report blockers. It does not independently set the commercial priority, the deadline promised to a customer or the acceptable trade-off between speed, cost and risk. Those decisions remain with the people accountable for the product and project.

From objective to verifiable outcome

The objective describes the useful change, not the technical activity. For the portal, “implement document upload” is too broad. A better version would be: “an authorized supplier can attach a PDF document to a request, and the operator can see the file and the audit event.”

We also establish boundaries from the objective. In our hypothetical scenario, the first version does not extract the document's content automatically, accept other formats or migrate the old archive. These exclusions reduce divergent interpretations and prevent the agent from expanding the work without a decision.

A backlog is the ordered list of known work. It is not a collection in which every item has the same priority. The order should reflect value, risk and dependencies. The Scrum Guide defines the Product Goal as the target against which planning takes place, while the backlog describes what might fulfil that target. A company does not have to adopt Scrum in full to use the helpful idea: clarify the outcome first, then order the work.

A small backlog ordered by dependencies

For the upload capability, the team can start with these work packages:

Order

Task

Depends on

Primary evidence

1

Confirm access and retention rules

product decision

recorded decision and approved cases

2

Define the upload API contract

task 1

reviewed schema and error examples

3

Store the file and create the audit event

task 2

automated tests and verified log

4

Build the upload and display interface

tasks 2 and 3

passed end-to-end scenario

5

Enable the capability in a controlled manner

tasks 3 and 4

rollback plan and target-environment verification

A dependency shows that one task is blocked by another. GitHub Issues, for example, supports sub-issues and “blocking” relationships, but the tool is secondary. A board with many columns cannot repair an incorrect order. If the access rule has not been decided, the agent should not invent roles and continue with implementation.

The task brief for a person or an agent

The same brief should be readable by a colleague and a development agent. A minimal version contains:

  • desired outcome: the observable behaviour after completion;
  • authorized context: documents, files and decisions that are sources of truth;
  • in scope and out of scope: the limits of the change;
  • dependencies: decisions, services or tasks that must be completed first;
  • acceptance criteria: concrete conditions the result must meet;
  • mandatory verification: tests, static analysis, manual scenarios or relevant captures;
  • acceptance owner: the person who compares the result with the need;
  • stop conditions: situations in which the agent asks for clarification instead of making assumptions.

For its coding agent, GitHub recommends clearly described problems, complete acceptance criteria and guidance about the relevant code area. The recommendation is product-specific, but the principle is general: an executable task reduces room for interpretation. This does not mean prescribing every line of code. Within the architectural boundaries, the agent can propose the technical solution and explain alternatives.

For the storage task, a verifiable criterion would be: “a user without permission is denied, and the file is not stored.” “Make it secure” is insufficient because it does not identify behaviour that can be tested.

Acceptance criteria and Definition of Done

Acceptance criteria belong to a particular requirement. They describe the examples and conditions that must be demonstrated. The Definition of Done, meaning the shared definition of completed work, applies to all relevant changes: reviewed code, passing tests, updated documentation, no secrets in the commit and a release procedure, for example.

The Scrum Guide defines the Definition of Done as the formal description of the state in which the result meets the product's quality measures. For the portal team, “the agent has finished” is only a progress statement. The status becomes “ready for acceptance” when the diff, verification results and explanation of decisions are available. It becomes “done” only after the owner confirms both the criteria and the shared definition.

The SWE-bench study shows why the problem is not reduced to code generation: real tasks can require coordinated changes across multiple functions, classes and files, within an execution environment. The benchmark evaluates resolutions to problems associated with real issues and changes, but it is not a guarantee for a company's project. The practical lesson is that evaluation must be tied to the repository, tests and required behaviour, not to how persuasive the agent's report sounds.

What the agent may update

The development agent can propose the decomposition of work, link its change to the task, mark completed checks and flag a blocker. If it discovers an API incompatibility, it updates the status with evidence and requests a decision. It does not independently change the scope, accept its own work or move a commercial deadline.

We separate this role from the agent integrated into the portal. In a later episode, the in-application agent might classify requests or extract information from documents. It is a product feature, with its own requirements and risks. The development agent works on the project and delivers changes for review. Confusing the two leads to incorrect permissions, metrics and accountabilities.

A control rhythm for a small team

A small company does not need heavy ceremonies. It can use a short rhythm:

  1. the product owner confirms the objective, boundaries and order;
  2. the agent or developer proposes the plan and identifies dependencies;
  3. the team accepts the plan only as an execution hypothesis, not a guaranteed estimate;
  4. the work produces artifacts and evidence on a separate branch;
  5. a person verifies the result and decides to accept, correct or stop it.

Deadlines are estimated separately, using the team's actual capacity and newly discovered unknowns. An agent that executes one stage quickly does not eliminate the time needed for clarification, integration, review or remediation of side effects.

The piece added to the project is the reusable task brief and the backlog ordered by dependencies. They turn delegation into a verifiable working agreement. In the next episode, we will prepare the repository so agents can find the instructions, architecture and test environments required for these tasks.

Sources

Next step

Want to turn an idea into a backlog that people and agents can execute and verify? Talk with the i8 team.

Recommended for you

IDE, terminal, or cloud? How to choose AI agent development tools

Preparing code for AI agents: context, rules and test environments

Who has the project keys? Identities and secrets for applications and agents

Cookies

We use cookies required for the site to work. With your consent we also enable additional features (videos, maps) or anonymous statistics. You can change your choice at any time from the site footer.

Cookie policyPrivacy noticeTerms and conditionsCookie preferences