AI Agent Task Management: Build a Durable Task Board
Build an AI agent task board with stable IDs, explicit status transitions, bounded permissions, completion evidence, and human review.
AI agent task management works best when the task board is durable operational state, not a list buried in one agent conversation. Give every task a stable ID, a small status vocabulary, an explicit owner, a bounded transition policy, and evidence that proves completion. The agent can then resume work across sessions without guessing what "done" meant.
The practical workflow is:
- Define which work belongs on the board.
- Create a task schema with stable identity.
- Write the allowed status transitions beside the data.
- Connect the agent with the smallest useful permission.
- Claim one task by exact ID.
- Record blockers and handoffs as structured state.
- Attach completion evidence and move the task to review.
- Verify the board before reporting progress.
This guide uses a claim -> work -> prove -> review contract. Claiming establishes who is responsible. Work records the current state and blockers. Proof links the task to a concrete artifact or result. Review keeps the agent that performed the work from silently declaring its own success.
In this guide
- What AI agent task management is
- Choose the board boundary
- Design the task schema
- Define status transitions
- Connect the agent
- Run the task lifecycle
- Handle concurrency and retries
- Verify the board
- AI agent task management FAQ
What is AI agent task management?
AI agent task management is the practice of storing delegated work as structured records that an authorized agent can inspect and update. A useful task record preserves identity, ownership, workflow state, acceptance criteria, blockers, and completion evidence across model runs, chat sessions, and handoffs.
The board is not the agent runtime. It does not decide when a model runs, execute tools, or prove that an external action succeeded. It is the shared source of truth for what work exists and what state each task is in. The runtime still needs to schedule the agent, enforce tool permissions, and pause for approval when an action crosses a risk boundary.
OpenAI's current guide describes agents as systems that combine models, tools, and guardrails, with tools providing the ability to act and guardrails defining oversight (OpenAI, checked July 2026). A task board gives those runs durable coordination state. It should not be confused with the agent's conversational memory; the memory-versus-structured-state guide explains that boundary.
1. Choose the task-board boundary
Put work on the board when another session, person, or agent needs to inspect, claim, review, or resume it. Keep transient reasoning and scratch notes in the runtime or working files.
Good task rows describe addressable outcomes:
- verify the export flow against named acceptance criteria
- draft onboarding copy and attach the pull request
- classify a bounded batch of feedback and flag uncertain rows
- investigate a failed job and record the evidence
Avoid task rows such as "work on marketing" or "improve the app." They have no finish line, so an agent cannot prove completion or decide when to stop. Split broad goals into tasks that can produce one inspectable result.
Also separate coordination from external authority. A task that says send customer update does
not grant permission to send it. The board can hold a draft, approval state, and evidence reference,
but the messaging tool and its approval policy remain separate. The
human-in-the-loop workflow shows how to place a real approval
boundary before consequential actions.
2. Design an AI agent task schema
Create an agent_tasks dataset indexed by task_id. A stable index lets every agent address the
same task directly instead of searching by a title that may change.
| Column | Type | Purpose |
|---|---|---|
task_id |
text, index | Stable identity for reads, updates, retries, and handoffs |
title |
text | Short outcome-oriented label |
description |
text | Scope and relevant context |
status |
choice | todo, doing, blocked, review, done, or cancelled |
owner |
text | Person or agent responsible for the next action |
priority |
choice | A small documented scale such as P0 through P3 |
acceptance_criteria |
text | Conditions that must be true before review |
blocker |
text | Exact missing decision, dependency, permission, or input |
claimed_at |
datetime | When the current owner started work |
updated_at |
datetime | Last meaningful task-state change |
evidence_ref |
text | URL, dataset key, ticket ID, test-run ID, or other evidence locator |
reviewed_by |
text | Person or independent agent that checked the result |
run_id |
text | Runtime execution associated with the latest change |
This is intentionally smaller than a full project-management schema. Add fields only when a
workflow rule or useful query needs them. For example, add due_at when missed deadlines matter,
or parent_task_id when the board needs explicit decomposition. Do not add a percentage-complete
field unless the workflow can define and verify the number.
Use choice columns for bounded fields and descriptions for ambiguous ones. Store durable operating rules in dataset instructions. Rowset's schema-design guide documents semantic types, choice values, column descriptions, metadata, and instructions.
3. Define the status transition contract
Status is a workflow decision, not a decorative label. Define who may move each task and what must be present after the move.
| Current state | Next state | Required condition |
|---|---|---|
todo |
doing |
Set owner, claimed_at, and run_id |
doing |
blocked |
Set a specific blocker; keep ownership explicit |
blocked |
doing |
Record what resolved the blocker and start a new run if needed |
doing |
review |
Acceptance criteria checked; evidence_ref present |
review |
doing |
Reviewer records what failed and returns ownership |
review |
done |
Reviewer is named and evidence still resolves |
| any active state | cancelled |
Record the decision and preserve existing evidence |
Put the contract in dataset instructions so a later agent session receives the same rules:
Use task_id for every lookup and update.
Let one scheduler or an external transactional claim service assign todo tasks.
Mirror the assigned owner, claimed_at, and run_id together; do not treat the Rowset patch as a lock.
Set status=blocked when progress requires a missing input, permission, or decision.
Move a task to review only when acceptance_criteria are checked and evidence_ref is present.
Do not move your own task from review to done.
Do not delete tasks. Use cancelled and preserve the evidence.
Read the task back after every uncertain write.
These instructions guide the agent, but they do not create a transactional state machine. If a workflow needs hard transition enforcement, simultaneous worker claims, or strict reviewer separation, implement those controls in the application or orchestration layer.
4. Connect the agent through MCP or REST
Use hosted MCP when the agent benefits from discovering tools and their input schemas at runtime. The MCP architecture lets clients list tools before calling them, and each tool can expose a JSON Schema for its inputs (Model Context Protocol, checked July 2026). Use the Dataset API when an application or worker already speaks HTTP.
For Rowset:
- Create a private dataset and keep public preview disabled.
- Give the agent a Read + write key only if it must change task rows.
- Connect with the hosted MCP guide or the Dataset API.
- Call
get_datasetbefore row work to load the index, schema, and instructions. - Read and update tasks by
task_id.
OWASP recommends minimum tool access, separate permission sets for different trust levels, and explicit authorization for sensitive operations (OWASP, checked July 2026). Apply that outside the prompt. Rowset's current Read + write role is account-wide and also permits destructive MCP operations such as deleting rows, dropping columns, removing relationships, and archiving datasets. It cannot isolate a task agent to safe task-row patches. Restrict the tools exposed by the agent runtime, require confirmation around destructive calls, and use a separate application proxy or narrower integration when a hard server-side task-only boundary is required.
5. Run the claim -> work -> prove -> review lifecycle
The lifecycle turns a task card into a resumable operating contract.
Claim one exact task
Do not use a Rowset read followed by a patch as a multi-worker claim. Two agents can both read
status=todo, both patch the row, and both start work even if each reads back its own claim before
the other overwrite arrives.
Choose one safe assignment path:
- Single scheduler: one runtime selects and assigns tasks, then workers receive exact task IDs.
- Serialized worker: only one agent can claim from this board at a time.
- External transactional claim: a queue, lock, or lease service atomically assigns the task,
then the worker mirrors
status=doing,owner,claimed_at, andrun_idinto Rowset.
After assignment, read the row by task_id and confirm the mirrored owner and run details before
work. That read-back detects drift; it does not make the preceding patch atomic.
Do not claim "the first open task" from an old in-memory list. Another worker may have changed the board. Let the scheduler or transactional claim service choose one task and pass its exact ID to the worker.
Work against acceptance criteria
Treat the acceptance criteria as the exit conditions for the run. OpenAI's agent-building guidance describes runs as loops that continue until a defined exit condition, structured output, error, or turn limit (OpenAI, checked July 2026). The board should make the task-specific exit condition visible before the agent starts.
When the agent discovers new scope, create another task or propose a change. Do not silently widen the current task and then claim the original criteria passed.
Record blockers as actionable state
Use blocked only when the agent cannot make meaningful progress without a named dependency:
Needs approval for the production migration planWaiting for API credentials from the account ownerTest service returns 503; retry after incident INC-204
"Stuck" is not enough. Keep the owner, describe the next action, and attach a source reference when one exists. A later session should be able to decide whether the blocker still applies without reconstructing the whole conversation.
Prove completion before review
Move the task to review, not directly to done. Attach evidence that another person or agent can
inspect: a pull request, deployed URL, test output, exported report, dataset key, or ticket.
Completion evidence should prove the acceptance criteria, not merely prove that the agent produced something. A file path proves a file exists. It does not prove the file is correct. For important work, record the validation command or review artifact alongside the output.
Review independently
The reviewer checks the evidence and either returns the task to doing with a concrete reason or
moves it to done with reviewed_by set. The MCP tools specification recommends keeping a human
able to deny tool invocations and presenting confirmation for operations
(Model Context Protocol, checked July 2026).
Use risk-based human review for consequential actions; independent agent review can help with
bounded, reversible work.
6. Handle concurrent agents and retry failures
A Rowset task board is useful coordination state, but a read followed by a patch is not an atomic
compare-and-set claim. Two agents can read the same todo row before either update arrives, and
read-back cannot prevent both from starting. Serialize assignment even on a small board when
duplicate work would be expensive or external. For a multi-worker queue, place an external
transactional claim or lease service in front of execution.
Use stable IDs and absolute values for recoverable updates. If a write times out:
- Read the task by
task_id. - Compare the stored fields with the intended final state.
- Continue if the desired values are already present.
- Retry only the missing update.
- Stop and report a conflict if another owner or run changed the task.
The idempotent AI-agent update guide covers this
identity -> desired state -> confirmation pattern in detail. Include a run_id so an operator can
connect a task mutation to the runtime attempt that made it. For stronger reconstruction, join the
board with the separate AI agent audit-trail pattern.
7. Verify the task board before reporting progress
Do not accept "tasks updated" as the final result. Reconcile the board with deterministic checks:
- every
doingtask has an owner, claim time, and run ID - every
blockedtask has a specific blocker and next responsible owner - every
reviewtask has acceptance criteria and a resolvable evidence reference - every
donetask has evidence and a named reviewer - no active task has two owners encoded in free text
- cancelled tasks preserve their prior evidence and decision context
- uncertain writes were read back before retry
Return a short run report with task IDs: claimed, moved to review, completed, blocked, conflicted, and unchanged. The IDs make the report verifiable against the board.
Rowset's existing agent task-board starter gives you a compact schema. Use this guide when you need the fuller lifecycle, transition rules, retry behavior, and review contract. You can test the pattern with Rowset's 7-day hosted trial.
AI agent task management FAQ
Should an AI agent use a Kanban board or its memory?
Use a task board for durable operational state: task identity, owner, status, blockers, acceptance criteria, and evidence. Use agent memory for retrieved context and learned preferences. A task that must survive a new session or handoff should not depend on the agent remembering a conversation.
Can multiple AI agents claim tasks from one Rowset dataset?
They can read and update the same private dataset, but Rowset row updates do not provide an atomic queue claim, and read-back cannot prevent both workers from starting. Use one serialized scheduler or an external lock, transactional queue, or lease service before assigning work. Mirror the result to Rowset for coordination and review.
What counts as completion evidence?
Use an artifact another reviewer can inspect against the acceptance criteria: a pull request, test run, deployed page, report, dataset key, ticket, or source record. A status change or agent message alone is not evidence that the requested outcome is correct.
Should agents mark their own tasks done?
For low-risk personal work, that may be acceptable if the evidence is deterministic and easy to
recheck. For shared, customer-facing, destructive, or costly work, move the task to review and
require a separate person or trusted reviewer to approve completion.
Is Rowset a full project-management system?
No. Rowset provides private structured datasets with stable row identity, semantic schema, instructions, MCP and REST access, exports, and optional read-only previews. It does not provide full project planning, atomic worker queues, schedules, notifications, or workflow enforcement.