Strategy

LangGraph Work Management Layer: Adding a Done Gate

Cognibl is the work-management platform from DILR.AI that puts a done-gate on work, so a task reaches done only once proof is attached. This guide explains where the LangGraph runtime ends, where a shared work board begins, and when a LangGraph team does not need that second layer at all.

LangGraph Work Management Layer: Adding a Done Gate COGNIBL LangGraph Work Management Layer: Adding a Done Gate 01 Graph run 02 Interrupt pause 03 Proof attached 04 Verified done dilr.ai/blog

A team that builds agents on LangGraph has made a sound engineering choice. LangChain describes LangGraph as "an agent runtime and low-level orchestration framework", states that it is MIT-licensed and free to use, and its open-source repository showed 42.7k stars on GitHub on 5 October 2026. The runtime question is settled for most of these teams. The question that arrives next is different: when a LangGraph agent says a task is finished, who checks, where is the evidence kept, and what does the operations manager read on Monday morning?

That gap is not a LangGraph defect. LangGraph's own documentation says the library "is very low-level, and focused entirely on agent orchestration", and it does that job well. The gap sits one layer up, where a business tracks work, agrees what done means and decides whether to trust a completion claim. McKinsey's State of AI 2025 puts the size of the problem plainly: 88% of organisations use AI in at least one function, yet only about a third have it in production and roughly 6% capture material EBIT impact. That gap has many causes, and this guide deals with one operational part of it: knowing which agent work actually counts as done.

This guide is written for the engineering lead or head of delivery who already runs LangGraph and wants to know what a work management layer adds on top of it, what it duplicates, and when it is not worth adding at all. It does not compare trackers in general; that head-to-head lives in our AI agent work management guide, and the product definition of Cognibl, from DILR.AI, lives in our Cognibl explainer. Here the scope is one stack: LangGraph for execution, a done-gated board for the work.

This guide is shipped by the team behind Cognibl, work management where people and AI agents share one board and a task reaches done only once proof is attached. Or see DATS, our five-stage system for placing AI where it pays.

What is a work management layer for LangGraph?

A work management layer for LangGraph is the system that records what an agent was asked to do, what done means for that task, and the evidence that it was achieved. LangGraph runs the agent: state, branching, memory and pauses. The work layer sits above the run and holds the task record that a business owner and an engineer both read. It answers whether the work is finished, not how the graph executed.

The distinction matters because the two layers have different readers. An engineer debugging a graph wants node-level state, retries and model calls. A practice manager or COO wants a list of tasks, their owners, the agreed definition of done and a yes or no on whether each one shipped. Putting both audiences on the execution trace forces one of them to read the wrong record.

Cognibl, from DILR.AI, is one example of a work layer built for this split. On Cognibl's product page, agents pick up work under their own name, against the same statuses the team uses, and "the definition of done, the evidence behind it and every tool call are attached to the work". A classic tracker such as Jira or Linear can also serve as the work layer; what changes is whether completion is gated on evidence or on someone dragging a card.

If you are still deciding whether agents belong in the workflow at all, that is a placement question before it is a tooling question, and an AI placement diagnostic is the cheaper place to answer it.

What does LangGraph already give you for oversight and audit?

LangGraph already gives a team real oversight tools. Its documentation lists persistence through checkpointers, human-in-the-loop interrupts that pause a run for approval, time travel to replay or fork from a saved checkpoint, and debugging through LangSmith traces, which LangChain calls the record of what agents did in production. Any honest comparison starts from that baseline, because a LangGraph team is not working blind.

The LangGraph overview sets out the core benefits: mixing deterministic and agentic steps in one graph, persistence so an agent can "persist through failures and can run for extended periods", and human-in-the-loop oversight by inspecting and modifying agent state at any point. The persistence documentation separates checkpointers, which hold a thread's graph state, from stores, which hold application data across threads.

On the observability side, the LangSmith documentation is direct: "Traces are the record of what your agents did in production." LangSmith also works across frameworks, so a team running more than one agent stack can trace them in one place. The time travel documentation adds a useful caution: replaying from a checkpoint re-executes later nodes, including model calls and API requests, which may return different results.

None of this is a gap to sell against. A LangGraph team with LangSmith tracing and interrupts on its risky nodes already has strong run-level control. The useful question is what those records are for, and who is meant to read them. Our note on evaluation harness engineering covers the testing side of the same stack.

Where does the LangGraph runtime end and the work board begin?

The LangGraph runtime ends when a graph run finishes, pauses or fails. The work board begins with the task that caused the run: who asked for it, what the acceptance criteria were, which run or runs addressed it, and whether the result was accepted. One task can span several runs, a human edit and a review, so the task record has to outlive any single thread.

A simple test separates the layers. Ask what happens to each record when the run is deleted or replayed. A checkpoint is tied to a thread and exists to resume or replay execution. A trace exists to explain execution. A task record exists to state an outcome, and it should survive the run being re-executed with different results, because the outcome claim is what a business acts on.

The table below sets out where each record usually lives in a LangGraph stack with a separate work layer. It describes a common division, not a rule; some teams keep more in LangSmith, others more on the board.

RecordUsually lives inPrimary reader
Graph state and checkpointsLangGraph persistenceEngineer resuming or replaying a run
Execution trace of model and tool callsLangSmith or another tracerEngineer debugging behaviour
Approval at a paused nodeLangGraph interrupt, resumed with CommandReviewer approving one action
Task, owner and definition of doneWork boardDelivery lead and business owner
Proof that the outcome was metWork board, attached to the taskBusiness owner accepting the work
Delivery metrics across many tasksWork boardCOO or head of delivery

The first three rows are execution records and LangGraph already handles them well. The last three are work records. They are where a team decides whether agent output counts, and they are where a governed AI operating model assigns accountability, because a RACI needs a task to hang on, not a thread ID.

Two layers around one agent task
01Task createdOwner and definition of done set on the board02Graph runLangGraph executes, checkpoints state03Interrupt pauseA reviewer approves a risky action04Trace recordedModel and tool calls kept for debugging05Proof attachedEvidence of the outcome linked to the task06Verified doneStatus moves only once proof exists
Execution records sit in the runtime and tracer; the outcome record sits on the work board, attached to the task.

What should count as done for a LangGraph agent's task?

Done for a LangGraph agent's task should mean the outcome named in the task's definition of done has been met and evidenced, not that the graph reached its end node. A run can complete cleanly and still produce the wrong artefact. The definition of done should name the artefact, the check that proves it, and who accepts it, before the agent starts work on the task.

In practice a good definition of done for agent work has three parts. First, the artefact: a merged pull request, a populated record, a generated report. Second, the check: tests passing, a reconciliation matching, a reviewer signing off. Third, the evidence format: what gets attached so that someone who was not present can verify the claim later. Our proof of done guide goes deeper on writing these for mixed teams of people and agents.

Cognibl makes the evidence step mechanical. On the same Cognibl page, a task reaches a done status only once proof is attached, "a CSV describing the run, with the artefact, screenshots and hashes referenced from inside it", and without a proof version the database itself refuses the move. The same rule holds for a person and for an agent, which removes the temptation to hold agents to a stricter or looser standard than the people beside them.

For a LangGraph team, the practical consequence is that the graph's final node is not the place to declare success. Whichever board the team uses, the run should produce the evidence the task's definition of done asks for, and acceptance should rest on that evidence rather than on the run ending. A worked example in a regulated setting is our post on evidencing agent work for Making Tax Digital.

How do LangGraph interrupts differ from a done-gate?

LangGraph interrupts and a done-gate both stop work until a person acts, but they guard different things. An interrupt pauses a graph mid-run so a reviewer can approve, edit or reject one action before execution continues. A done-gate sits on the task record and blocks the status from moving to done until evidence of the finished outcome is attached. One governs an action; the other governs a completion claim.

LangChain's interrupts documentation describes the mechanism precisely: "Interrupts allow you to pause graph execution at specific points and wait for external input before continuing." When an interrupt fires, LangGraph saves the graph state through its persistence layer and waits until the run is resumed with a Command. That is the right tool for a node that sends an email, writes to a production table or spends money.

A done-gate answers a later question. After every interrupt has been approved and the graph has finished, is the task actually complete? An approval on step four says nothing about whether the report produced at step nine matches what the business asked for. A run can pass every approval and still deliver the wrong artefact, because no approval was ever asked about the final outcome.

The two work best together. Put interrupts on irreversible actions inside the graph, and gate completion on the task record outside it, on whatever board the team uses. Our design notes on human approval gates were written for voice agents, but the reasoning about where to place a human checkpoint carries over to any agent stack.

Does a LangGraph team need a separate work layer at all?

A LangGraph team does not always need a separate work layer. If engineers are the only consumers of agent output, LangSmith traces plus an existing Jira or Linear board may be enough, and a proof gate would add friction without changing a decision. The case for a separate, evidence-gated layer appears when people and agents deliver side by side and someone outside engineering must trust the agent's completions.

Three signals suggest the extra layer will pay for itself. The first is mixed ownership: agent tasks and human tasks feed the same deliverable, so one board has to show both under the same statuses. The second is an outside reader: a client, an auditor or a finance lead needs to accept work without reading traces; for one sector view of where that applies, see our SaaS and technology guide. The third is rework: tasks marked complete are coming back, which means the status is recording activity rather than outcomes.

If none of those apply, keep the stack lean. If one or more do, the cheapest next step is usually to write definitions of done for the agent tasks you already run, before choosing any tool, and our DATS consulting team can help scope them. Once agent work is in production, our AI execution office provides embedded delivery with production placements the client owns, and teams that want help with the engineering side can look at our AI developer productivity work.

Cognibl measures whether the layer is working. Its delivery metrics include cycle time, first-pass rate, human wait share, verified throughput, meaning completions per week that cleared the gate, and reopen rate, meaning tasks reopened within 30 days. The Cognibl page describes reopen rate as the number that tells you whether done really meant done.

How do EU AI Act record-keeping duties bear on a LangGraph stack?

The EU AI Act sets record-keeping and oversight duties for high-risk AI systems, which fall mainly on their providers, and whether an internal delivery agent is high-risk depends on its use case. For a high-risk system, the Act sets Article 12 automatic event recording over its lifetime and Article 14 design for effective human oversight. Those are duties on the system's design and its provider, once the high-risk regime applies.

The text of Article 12(1) reads: "High-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system." Article 14 sets the oversight requirement in similar design terms. Neither article says a work board must exist, and neither can be met by a tracker alone; the logging requirement is about the system's own capability.

For a LangGraph team, the practical reading is modest. Checkpoints and traces are execution logs, and they are the natural place to start if a system is ever classified as high-risk. A task-level evidence record is a separate, complementary record of outcomes and approvals. Do not treat either as a compliance programme; classification is a legal and design question, not a tooling one, and our AI solutions work starts from that design question. Classification, deployer duties and timing need advice specific to your use case; our AI tool inventory guide covers the first step, which is knowing which systems you run.

Where Cognibl fits here is narrow and factual. As the Cognibl page sets out, every call through its gateway is traced, every write is attributed by key name, and records are append-only and hash-chained. That gives the work record an append-only history. It does not make a system compliant.

What is the best way to manage LangGraph agent work in 2026?

The best way to manage LangGraph agent work in 2026 depends on who reads the output. For engineer-only workflows, LangGraph with LangSmith tracing and an existing tracker such as Jira or Linear is often the right answer. For teams where people and agents share delivery and outsiders must trust completions, a done-gated work layer on top of LangGraph earns its place.

Score the options on four criteria. Does the record survive a run being replayed? Can a non-engineer accept a task without reading a trace? Is completion gated on evidence or on a status change? Do the metrics count verified outcomes or activity? LangSmith is built for tracing, evaluating and debugging runs, which a work board is not. A classic tracker has the advantage that your team already uses it. Other agent frameworks, such as CrewAI or the OpenAI Agents SDK, raise the same question about where outcome evidence lives.

Cognibl, from DILR.AI, is built for the case where all four criteria matter at once. It is a keyboard-first tracker with backlog, sprint board, roadmap and version-control links, plus the done-gate, a versioned skills and agents library in SKILL.md format reached through a deny-by-default Model Context Protocol gateway, and the five delivery metrics. It ships Harness, Graph and Loop process templates, and two AI flows that summarise and flag problems but never decide. It is built and run by Dilr.ai Ltd in London.

The concession is real. If your agents only open pull requests that engineers review, a mature tracker plus LangSmith will serve you with less change, and you should keep it. Choose a done-gated layer when the reader of the result is not the person who wrote the graph, and if you are unsure which side of that line you sit on, talk to our team before buying anything. Our enterprise AI consulting guide covers how to make that kind of placement decision across a whole portfolio.

Does Cognibl replace LangGraph or LangSmith?

Cognibl does not replace LangGraph or LangSmith. LangGraph remains the runtime that executes the agent, and LangSmith remains the place engineers trace and evaluate runs. Cognibl is a work management platform: it holds the task, the definition of done, the attached proof and the delivery metrics. The three answer different questions, and a team can keep its runtime and tracer unchanged while adding a work layer.

That separation is deliberate. The Cognibl product page describes a trust layer that "sits between the agent and the business, so work is finished when it is proved, not when it is claimed". The principle is about where the outcome record lives, not about how a graph is built, and how any particular stack connects to a given board is an engineering question to confirm during evaluation. For the broader product picture, see what Cognibl is or the main Cognibl page.

Is a LangGraph checkpoint the same as proof of done?

A LangGraph checkpoint is not the same as proof of done. A checkpoint saves a thread's graph state so execution can resume, replay or fork, and replaying can produce different results. Proof of done is evidence that a task's agreed outcome was met, attached to the task so someone else can verify it later. A checkpoint describes where a run was; proof describes what the work achieved.

The two can sit side by side. A team may choose to note in its evidence which run produced an artefact, and a reviewer may still open the trace to see how it was made. What matters is that the acceptance decision rests on the outcome evidence, not on the run state.

Want to see this in production? Try Cognibl on your own board, book an AI placement diagnostic, see our DATS methodology, or browse more AI strategy guides on placing agents inside real delivery teams.

Product
Cognibl
Solution
AI Developer Productivity
Service
AI Operating Model
Talk to the operators

Make agent work count as done only when it is.

30-min scoping call · No deck · Confidential. We will tell you whether a work layer fits your LangGraph stack, and where agent output should be gated.

Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.

LangGraph work management layerLangGraph human in the loopLangGraph task trackingLangGraph audit trailproof of done for ai agentsai agents redditbest ai agent work management tool 2026cognibl

Questions this article answers

What is a work management layer for LangGraph?

A work management layer for LangGraph is the system that records what an agent was asked to do, what done means for that task, and the evidence that it was achieved. LangGraph runs the agent: state, branching, memory and pauses. The work layer sits above the run and holds the task record that a business owner and an engineer both read. It answers whether the work is finished, not how the graph executed.

What does LangGraph already give you for oversight and audit?

LangGraph already gives a team real oversight tools. Its documentation lists persistence through checkpointers, human-in-the-loop interrupts that pause a run for approval, time travel to replay or fork from a saved checkpoint, and debugging through LangSmith traces, which LangChain calls the record of what agents did in production. Any honest comparison starts from that baseline, because a LangGraph team is not working blind.

Where does the LangGraph runtime end and the work board begin?

The LangGraph runtime ends when a graph run finishes, pauses or fails. The work board begins with the task that caused the run: who asked for it, what the acceptance criteria were, which run or runs addressed it, and whether the result was accepted. One task can span several runs, a human edit and a review, so the task record has to outlive any single thread.

What should count as done for a LangGraph agent's task?

Done for a LangGraph agent's task should mean the outcome named in the task's definition of done has been met and evidenced, not that the graph reached its end node. A run can complete cleanly and still produce the wrong artefact. The definition of done should name the artefact, the check that proves it, and who accepts it, before the agent starts work on the task.

How do LangGraph interrupts differ from a done-gate?

LangGraph interrupts and a done-gate both stop work until a person acts, but they guard different things. An interrupt pauses a graph mid-run so a reviewer can approve, edit or reject one action before execution continues. A done-gate sits on the task record and blocks the status from moving to done until evidence of the finished outcome is attached. One governs an action; the other governs a completion claim.

Does a LangGraph team need a separate work layer at all?

A LangGraph team does not always need a separate work layer. If engineers are the only consumers of agent output, LangSmith traces plus an existing Jira or Linear board may be enough, and a proof gate would add friction without changing a decision. The case for a separate, evidence-gated layer appears when people and agents deliver side by side and someone outside engineering must trust the agent's completions.

How do EU AI Act record-keeping duties bear on a LangGraph stack?

The EU AI Act sets record-keeping and oversight duties for high-risk AI systems, which fall mainly on their providers, and whether an internal delivery agent is high-risk depends on its use case. For a high-risk system, the Act sets Article 12 automatic event recording over its lifetime and Article 14 design for effective human oversight. Those are duties on the system's design and its provider, once the high-risk regime applies.

What is the best way to manage LangGraph agent work in 2026?

The best way to manage LangGraph agent work in 2026 depends on who reads the output. For engineer-only workflows, LangGraph with LangSmith tracing and an existing tracker such as Jira or Linear is often the right answer. For teams where people and agents share delivery and outsiders must trust completions, a done-gated work layer on top of LangGraph earns its place.

AI consulting (DATS)

Place AI where the P&L moves

The DATS system runs from a fixed-fee placement diagnostic through to embedded delivery, so AI reaches production instead of staying a pilot.

Related articles

← Previous
Corporate KYC Review Cost in UK Banks: A Diagnostic Guide

One email, once a month. No hype. Just what we learned shipping.