Strategy

AI Agent Work Management: The Proof of Done Guide

Cognibl is the work-management platform from DILR.AI that runs one board for people and AI agents, where no task reaches a done status until proof is attached. This guide defines AI agent work management, explains what a definition of done and a proof of done contain, and sets out how to tell whether you need an agent-capable tracker.

AI Agent Work Management: The Proof of Done Guide COGNIBL AI Agent Work Management: The Proof of Done Guide 01 Pick up work 02 Definition of done 03 Proof attached 04 Verified done dilr.ai/blog

AI agents are moving from demonstrations into the tools where real work is tracked. Gartner forecasts that by 2028, 33 percent of enterprise software applications will include agentic AI, up from less than 1 percent in 2024. Gartner also predicts, in a separate forecast, that over 40 percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Those two forecasts describe the same problem from opposite ends: agents are arriving fast, and most of the work they produce is not being trusted, measured or governed well enough to keep its budget. Broader adoption tells the same story, with McKinsey reporting that 88 percent now use AI while only a third have moved it into production, and Stanford's index finding fewer than one in ten have fully scaled it in any single function.

This guide is written for the person accountable for that gap: the head of delivery, the engineering lead or the operations director who has to answer whether the work an AI agent reported as finished is actually finished. A quick note on words, because much of the industry uses agent to mean a voice assistant that answers phone calls. Here, an agent is an autonomous software worker that picks up a task, does it and reports back, the same as a person on the team would. The discipline of running a board where both do that, under one set of rules, is what this guide calls AI agent work management.

It anchors the category for Cognibl, from DILR.AI, and it deliberately stays at the level of the category rather than the tool. It cedes detailed governance design to our approach for the operating model and vendor comparison to the pages linked throughout, and it keeps clear of two things buyers confuse this with: platforms that police agent sprawl and permissions, which are a security concern with a different buyer, and the model infrastructure that some vendors call an agent harness, which is plumbing rather than a work board.

This guide is shipped by the team behind Cognibl, work management for teams where people and AI agents share one board and no task reaches a done status until proof is attached. For the delivery model around it, see DATS, our senior-led AI consulting and build system.

What is AI agent work management?

AI agent work management is the practice of running a single board where people and AI agents pick up tasks against the same statuses, and where nothing is marked done without evidence behind it. It treats an agent as a team member held to the same completion standard as a person, not as a background service that quietly writes to tickets. One board, one standard, one record of what actually happened.

The definition is not a vendor coinage. The Work Management Institute, a standards body for the field, defines agentic work management as the application of work management principles to environments where organisational work is performed by both humans and AI agents. The Institute now runs certification pathways for it, a sign that the category is being formalised rather than improvised. What that formalisation forces into the open is a question most teams have never had to make explicit: what counts as done, and who or what is allowed to declare it. When only people worked the board, the answer lived in shared judgement. Once an agent can move work at machine speed, the answer has to be written down and enforced, which is the thread this guide follows from here. It is the same shift in discipline that separates a promising pilot from a system in production, a transition covered in our enterprise AI consulting guide.

Why does a classic tracker break when AI agents do real work?

A classic tracker assumes a card only moves because a person decided it should, so a status carries a human judgement. Hand that board to AI agents that can move through cards far faster than anyone can review them, and the status stops meaning anything, because no judgement sits behind the move. The status change and the proof of work were one event when a person did it, and they split apart the moment an agent does it.

Tools such as Jira, Linear, Asana, Monday.com, ClickUp and Notion record intent, and they record it well. Their model is sound for people who mostly earn the trust it extends: when someone drags a card to done, the team assumes a person judged it finished. The failure with agents is not carelessness. It is that an agent can produce a calm, complete report of work it did not actually finish, and a tracker that accepts words has no way to tell that report from a true one. At the board level, a delivery lead then loses the one thing the board was for, a reliable read of what is real, which is a large part of why so many agent projects stall before they show value. Reading a board of agent-moved cards without proof is like accepting a maturity claim with no evidence, the gap our AI capability maturity model exists to close. A tracker built on good faith is the wrong instrument for a workforce that generates volume faster than anyone can spot-check.

What does a definition of done mean for agent-produced work?

A definition of done is the explicit, checkable contract for when a task counts as finished, written before the work starts rather than argued about afterwards. For agent-produced work it has to be sharper than the shared understanding a human team carries in its head, because an agent has none to fall back on. It states the output, the evidence required and what would fail review, and it applies equally to a person and an agent.

In Cognibl, from DILR.AI, that contract travels with the work itself: the definition of done, the evidence behind it and every tool call the agent made are attached to the task. The delivery cycle below shows where the standard sits. Work is specified with a done definition agreed up front, built by whoever picks it up, verified against that definition, and only then settled. The gate is at verify, and it is the same gate for a person and an agent.

The delivery cycle with a done-gate
01SpecDefinition of done agreed before the work starts02BuildA person or an agent does the work03VerifyProof is attached and the gate checks it04SettleDone status recorded, the audit trail closed
A person or an agent moves work through the same four stages, and the gate sits at verify.

Setting a done definition this way is the same discipline that separates a demonstration from production. It is why a serious rollout runs a proof-of-concept-to-production conversion gate and a go-live checklist before any AI reaches real traffic. The point holds across any AI delivery, not only agents on a board: the moment of most risk is the move from something that demonstrated well to something that has to hold up under real load, and a written, checkable definition of done is what carries a team across it.

What goes into a proof of done?

A proof of done is the evidence package a task must carry before it can move to done. In Cognibl, from DILR.AI, that is a proof version: a CSV describing the run, with the artefact, the screenshots and the hashes referenced from inside it, plus a coverage report where tests are part of the claim. The database refuses the move without one, so proof is what makes the status true, for a person and an agent alike.

The parts below are what a reviewer opens rather than takes on faith. The run description says what was done and how; the artefact reference points at the thing produced; the screenshots and hashes let a reviewer check the claim without rerunning the work; and, where tests are part of the claim, a coverage report shows them. A live view of the product is at Cognibl, and the same evidence discipline underpins a formal readiness assessment before anything ships.

What a proof of done contains
01Run descriptionA CSV describing what was done02Artefact referenceThe thing produced, referenced from inside the CSV03Screenshots and hashesEvidence a reviewer can open and check04Coverage reportWhere tests are part of the claim
Each part is attached to the task before a done status is allowed.

An independent standard makes the same argument in plainer language. The Proof of Done manifesto, a text specification published at podmanifesto.org, sets out a completion report with four fields, DONE, PROOF, SCOPE and NOT VERIFIED, and states the principle bluntly: "Done is not the last message of the run. Done is a state of the system that can be proven." The manifesto is a writing standard for how an agent should report; a tracker with a done-gate is where that report is checked and enforced. They fit together, and neither replaces the other, which is a distinction worth holding on to when a vendor claims to have solved both at once.

How does a deny-by-default MCP gateway change how agents get tools?

A deny-by-default gateway refuses any tool that has not been explicitly enabled, rather than exposing everything and trusting nothing is misused. In Cognibl, agents reach a versioned library of skills and agents over the Model Context Protocol through such a gateway, so a toolset a project has not turned on is refused, not quietly missing. It is the difference between an agent that can only do what it was granted and one whose reach nobody has actually scoped.

Two further properties make that access auditable rather than merely restricted. The library is versioned immutably in a documented format, so editing a skill mints the next version, identical content is refused, and items archive rather than delete, which means the instruction an agent ran in one month is still the instruction you can read later. And every write is attributed to the calling key by name while refusals are mirrored to audit, so the record shows both what happened and what was blocked. This is the practical shape of what regulators are asking for. The EU AI Act places record-keeping and human-oversight duties on the providers and deployers of high-risk AI systems under Articles 12 and 14, with those obligations under Annex III now applying from December 2027, and the UK's ICO covers the same ground in its guidance on AI and data protection. A gateway and an audit trail do not make an organisation compliant on their own, and no tool should be sold as if it does; they are the logging and traceability those duties assume, and the accountability still sits with the deployer. Standards such as SOC 2, ISO 27001 and ISO 42001 sit alongside the same concern, as topics a buyer weighs rather than a badge any tracker confers. For how these controls fit a wider operating model, see our AI solutions practice and about Dilr.ai.

What replaces velocity when every completion must be verified?

When completions must be verified, raw velocity stops being useful, because a count of cards moved says nothing about how much of it was real. What replaces it is a set of measures about verified work and where time is lost. Cognibl reports five: median cycle time, first-pass verification rate, human wait share, verified throughput and reopen rate. Together they show how long work takes, how often it passes review first time, and how much of it comes back.

Each measure is chosen to resist the gaming that a simple count invites. Median cycle time is split across the spec, build, verify and settle stages, so a bottleneck shows where it happens rather than as one blurred figure. First-pass verification rate is read by flow type, verified throughput is counted per week, and reopen rate is measured within thirty days of a task being marked done, which catches work that was waved through and bounced back. A sound principle to adopt alongside them is that a verified-work dashboard should never reward raw task counts, precisely because an agent can inflate them at will; that is a design choice a team makes, and it is the one that keeps the numbers honest. Deciding who owns each of these measures is an operating-model question, covered in our target operating model and roles piece, and the tooling around it usually sits within an AI developer productivity programme.

How is agent work management different from AI agent security and governance?

Agent work management asks whether the work is done and provable; agent security and governance asks whether the agent is allowed to act at all. They are adjacent but serve different buyers. A security or sprawl-governance platform maps which agents exist and what they can touch, and is bought by a security or platform team worried about exposure. A work management platform runs the board those agents deliver against and holds their output to a completion standard.

Two other things get confused with this category, and both are worth naming. The Proof of Done manifesto is a text standard for how an agent writes its completion report; it is a specification, not a tracker, and it does not manage a board. And the term agent harness has an established, separate meaning: Databricks defines an agent harness as the tools, memory, workspace access and guardrails that let a model act on tasks rather than only respond to prompts, which is infrastructure around a model, not a place where work is planned and verified. That is a different sense again from the Harness process template inside Cognibl, and different from evaluation harness engineering, which is about testing model behaviour rather than tracking delivery. Keeping these apart is the quickest way to know which problem you are actually buying for, and the mistake is common enough that the same word will appear on three vendors' pages meaning three different things. When the pieces do need to fit together, an AI execution office is where the security, the delivery board and the model plumbing are made to work as one system.

What is the best tool for managing AI agent work in 2026?

The best tool depends on who does the work. For people who trust each other's judgement, a mature tracker such as Linear, Jira or Asana is right, and a proof-gate adds only friction. If the problem is orchestrating agents, not running a shared board, frameworks such as CrewAI, LangGraph or the OpenAI Agents SDK are enough. The case for an agent-capable platform is narrow: people and agents delivering side by side, where someone must prove the agent work is real.

That is the gap Cognibl, from DILR.AI, is built for, and its position is a deliberate trade. It applies one done-gate to people and agents alike, attaches proof to every completion, and measures verified work instead of raw counts, which is more discipline than a human-only team needs and more accountability than a pure orchestration framework offers. It ships four process templates, Harness, Graph, Loop and Custom, and three tiers named Free, Business and AI, with the stated principle that no feature is held back to sell the next tier up. The honest reading, then, is that most teams should not adopt a proof-gated board yet; the ones that should are those already handing agents real delivery and feeling the loss of a trustworthy status. If you are weighing where AI belongs in your delivery model before choosing any tool, an AI placement diagnostic is the place to start, and the wider strategy writing on our blog works through the surrounding decisions.

Is proof of done the same as the Proof of Done manifesto?

No. The Proof of Done manifesto is an independent, open text standard for how an AI agent should write its completion report, using the fields DONE, PROOF, SCOPE and NOT VERIFIED. A proof of done inside a tracker such as Cognibl is the enforced evidence package a task must carry before its status can change. The manifesto describes how to report; the tracker is where the report is checked and the status is gated. They complement each other.

Does this replace Jira or Linear?

Not for a human-only team. If people do all the work and trust each other's judgement, a tracker such as Jira or Linear fits well and Cognibl, from DILR.AI, would add controls you do not need. It becomes worth it once agents do a real share of delivery, because that is when a status has to be backed by proof, not a person's word. The decision turns on how much of your work is agent-produced, not on any single feature.

To go deeper, see Cognibl, read how we design the AI operating model around a delivery board, look at the AI developer productivity practice, or talk to us about scope on a scoping call.

Product
Cognibl
Solution
AI Developer Productivity
Service
AI Operating Model
Talk to the operators

Run one board people and agents can both be trusted on.

30-min walkthrough · No deck · Confidential. See the done-gate, the proof trail and the metrics that replace raw task counts.

Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.

ai agent work managementproof of doneagentic work managementdefinition of done ai agentswork management for ai agentsai agents redditbest ai project management tool 2026cognibl

Questions this article answers

What is AI agent work management?

AI agent work management is the practice of running a single board where people and AI agents pick up tasks against the same statuses, and where nothing is marked done without evidence behind it. It treats an agent as a team member held to the same completion standard as a person, not as a background service that quietly writes to tickets. One board, one standard, one record of what actually happened.

Why does a classic tracker break when AI agents do real work?

A classic tracker assumes a card only moves because a person decided it should, so a status carries a human judgement. Hand that board to AI agents that can move through cards far faster than anyone can review them, and the status stops meaning anything, because no judgement sits behind the move. The status change and the proof of work were one event when a person did it, and they split apart the moment an agent does it.

What does a definition of done mean for agent-produced work?

A definition of done is the explicit, checkable contract for when a task counts as finished, written before the work starts rather than argued about afterwards. For agent-produced work it has to be sharper than the shared understanding a human team carries in its head, because an agent has none to fall back on. It states the output, the evidence required and what would fail review, and it applies equally to a person and an agent.

What goes into a proof of done?

A proof of done is the evidence package a task must carry before it can move to done. In Cognibl, from DILR.AI, that is a proof version: a CSV describing the run, with the artefact, the screenshots and the hashes referenced from inside it, plus a coverage report where tests are part of the claim. The database refuses the move without one, so proof is what makes the status true, for a person and an agent alike.

How does a deny-by-default MCP gateway change how agents get tools?

A deny-by-default gateway refuses any tool that has not been explicitly enabled, rather than exposing everything and trusting nothing is misused. In Cognibl, agents reach a versioned library of skills and agents over the Model Context Protocol through such a gateway, so a toolset a project has not turned on is refused, not quietly missing. It is the difference between an agent that can only do what it was granted and one whose reach nobody has actually scoped.

What replaces velocity when every completion must be verified?

When completions must be verified, raw velocity stops being useful, because a count of cards moved says nothing about how much of it was real. What replaces it is a set of measures about verified work and where time is lost. Cognibl reports five: median cycle time, first-pass verification rate, human wait share, verified throughput and reopen rate. Together they show how long work takes, how often it passes review first time, and how much of it comes back.

How is agent work management different from AI agent security and governance?

Agent work management asks whether the work is done and provable; agent security and governance asks whether the agent is allowed to act at all. They are adjacent but serve different buyers. A security or sprawl-governance platform maps which agents exist and what they can touch, and is bought by a security or platform team worried about exposure. A work management platform runs the board those agents deliver against and holds their output to a completion standard.

What is the best tool for managing AI agent work in 2026?

The best tool depends on who does the work. For people who trust each other's judgement, a mature tracker such as Linear, Jira or Asana is right, and a proof-gate adds only friction. If the problem is orchestrating agents, not running a shared board, frameworks such as CrewAI, LangGraph or the OpenAI Agents SDK are enough. The case for an agent-capable platform is narrow: people and agents delivering side by side, where someone must prove the agent work is real.

AI consulting (DATS)

Place AI where the P&L moves

The DATS system runs from a fixed-fee placement diagnostic through to embedded delivery, so AI reaches production instead of staying a pilot.

Related articles

← Previous
Enterprise AI Consulting UK: The 2026 Guide

One email, once a month. No hype. Just what we learned shipping.