Industries

AI Agent Governance for Law Firms: The Supervision Record

Cognibl is the work-management platform from DILR.AI that refuses to mark a task done until proof is attached. This guide shows UK law firms what Ayinde, the SRA Code and Law Society guidance expect of AI agent work, what a supervision record should contain, and why a named solicitor still approves the output.

AI Agent Governance for Law Firms: The Supervision Record COGNIBL · LAW AI Agent Governance for Law Firms: The Supervision Record 18 of 45 cited cases in one High Court referral did not exist Source: High Court, Al-Haroun, 2025 dilr.ai/blog

A law firm can delegate a task to an AI agent. It cannot delegate the accountability for that task. Since the Divisional Court's judgment in Ayinde on 6 June 2025, a risk partner has good reason to ask not only "is this tool any good?" but "show me how this output was produced, who checked it, and against what". An AI workflow cannot answer that question if the record of the work lives in a chat window, an inbox or someone's memory.

This guide is written for the COO and the risk partner of a UK firm who have been asked to approve AI agents on matter operations work: collating disclosure, tracking outstanding intake documents, drafting client updates, assembling compliance evidence. It sets out what the courts, the Solicitors Regulation Authority's Code of Conduct for Solicitors and the Law Society now expect, and what a supervision record for agent work should contain before anyone signs it off. It deliberately does not cover where AI pays in a law firm or the adoption and complaints economics, which our AI for law firms in the UK guide already sets out, nor voice-led client intake, which sits in our AI voice guide to law firm client intake.

One framing note before we start. Law is an inferred fit for Cognibl, from DILR.AI, rather than a vertical we have deck evidence for: our own industry research names a governed agent desk for matter operations as a candidate use, not a proven one. We say so plainly, and the controls in this guide apply whichever product a firm chooses.

This guide is shipped by the team behind Cognibl, the work-management platform where a task reaches done only once proof is attached. Or see DATS, our consulting system for placing AI inside regulated operations.

What does AI agent governance mean for a UK law firm?

AI agent governance in a UK law firm means the firm can show, for every task an agent touched, what the agent was asked to do, what it was allowed to access, what it produced, and which named solicitor checked the output before it reached a client or a court. The rules of professional conduct already bind the firm and its solicitors; governance is the evidence that those duties were met.

That definition is narrower than most vendor marketing, and it is meant to be. Governance here is not a policy document or an ethics statement. It is a record that survives an investigation. If the SRA, a court or a client asks how a piece of work was produced in March, the firm should be able to open the record in September and read the same answer.

The distinction matters because the Law Society's own research says the profession is still sorting out its vocabulary. Its April 2026 report on the future of agentic AI in legal practice found that many solicitors confuse agentic AI with AI agents and with generative AI, and warned that the change may arrive incrementally, through solicitors delegating "just one more task". A firm that governs agent work task by task is protected against exactly that drift, because each delegation leaves its own record. For the underlying model of tasks, proofs and done-gates, our AI agent work management guide is the long version, and the AI operating model service is where we design the RACI around it.

What did the High Court say in Ayinde about checking AI work?

In Ayinde v Haringey and Al-Haroun v Qatar National Bank, decided on 6 June 2025, the Divisional Court held that lawyers who use artificial intelligence for legal research have a professional duty to check its accuracy against authoritative sources before relying on it. The court said that duty also covers lawyers who rely on work others produced with AI, and called for action from those with leadership responsibilities, such as managing partners, and from regulators.

The judgment, published on Find Case Law, dealt with two referrals under the court's Hamid jurisdiction. In Ayinde, the defendant's wasted costs application rested in part on five fake cases cited in the grounds. In Al-Haroun, the schedule of references prepared for the court was far larger, and the numbers are the reason this case belongs in every risk partner's file.

Citations put before the court in Al-Haroun
45Citations listed18Cases did not exist27Cases existed
Of 45 citations listed in the court's schedule of references, 18 named cases that did not exist; many of the remaining 27 did not support the propositions cited. Source: High Court, Ayinde and Al-Haroun [2025] EWHC 1383 (Admin), para 74

Two passages carry the weight for anyone designing agent workflows. The first extends the duty past the person at the keyboard:

"This duty rests on lawyers who use artificial intelligence to conduct research themselves or rely on the work of others who have done so." (Ayinde, para 8)

The court compared that reliance to a lawyer relying on a trainee solicitor or a pupil barrister. The second passage, at paragraph 9, says practical and effective measures must now be taken by those with individual leadership responsibilities, naming heads of chambers and managing partners, and by those who regulate legal services. In Al-Haroun the court added that a lawyer is not entitled to rely on a lay client for the accuracy of citations: it is the lawyer's own professional responsibility.

Read as an operating requirement, Ayinde does not ban AI. It says the check is not optional, it must be done against authoritative sources, and the people running the firm are expected to make sure it happens. An agent workflow that cannot show the check happened fails that test, however good the agent.

Who is accountable when an AI agent does matter work?

The solicitor who supervises the work stays accountable when an AI agent does matter work, and the firm stays bound by its own regulatory duties. The SRA Code of Conduct for Solicitors says a solicitor who supervises others remains accountable for the work carried out through them, and the SRA made the same point about accountability when it authorised the first AI-driven law firm.

Paragraph 3.5 of the SRA Code of Conduct for Solicitors says that where you supervise or manage others providing legal services, you remain accountable for the work carried out through them and you effectively supervise work being done for clients. Note who the duty binds: the solicitor, and the regulated firm through its own SRA obligations. It never binds a software supplier, which is why no vendor can sell a firm out of it.

The clearest worked example is the SRA's own. In its news release of 6 May 2025, SRA approves first AI-driven law firm, the regulator explained what it checked before authorising Garfield.Law Ltd to provide regulated legal services in England and Wales:

  • processes to quality-check work, keep client information confidential and safeguard against conflicts of interests;
  • how the firm manages the risk of AI hallucinations, including that the system will not propose relevant case law;
  • that the system is not autonomous and only takes a step where the client has approved it, with supervision and monitoring in place;
  • that named regulated solicitors remain ultimately accountable, and responsible for all the system outputs and for anything that goes wrong.

That list is a practical governance template, because it is what the regulator actually asked for. The Law Society reaches the same place from the profession's side. Its guidance, Generative AI: the essentials, updated in June 2026, says a solicitor's duties apply regardless of whether AI was used, and whether it was used by the solicitor personally or by anyone under that solicitor's supervision.

The same accountability logic underpins our AI placement diagnostic, which maps every proposed AI task to the named person who will answer for it before any tool is chosen.

What should a supervision record for agent work contain?

A supervision record for AI agent work should contain the instruction the agent ran, the data and tools it was permitted to use, the output it produced, the evidence a person used to check that output, the name of the supervising solicitor who approved it, and the time of each step. Each field answers a question the SRA, a court or a client could ask.

The fields below are our synthesis of the three sources above, not a published regulatory template. Each row names the source that prompts the field, or says where it is our own inference.

Field in the recordThe question it answersWhere the expectation comes from
The instruction, versionedWhat was the agent told to do, exactly, on that date?Our inference from Ayinde para 8: the duty extends to reliance on others' work
Permitted data and toolsCould the agent see privileged material it should not have?SRA Code para 6.3, confidentiality
The output itselfWhat did the agent actually produce?Law Society essentials: document inputs, outputs and errors
The check and its sourcesWas it verified against authoritative sources?Ayinde para 7
The approving solicitorWho is accountable for this output?SRA Code para 3.5; SRA Garfield.Law release
A timestamped, tamper-evident historyWhen did each step happen, and can the record be trusted months later?Our design inference from the above

The Law Society checklist is specific on the third row. It tells firms to document inputs, outputs and any errors of the generative AI tool if this is not automatically collected and stored, and to implement a consistent process for checking and verifying outputs. "If this is not automatically collected" is the operative phrase: a firm that relies on people remembering to save the evidence will have gaps, and the gaps will be in exactly the busy weeks that produce the errors.

The final row is the one most tools skip. A record that anyone with access can edit quietly is weak evidence. A record that is append-only, where each entry is chained to the last, lets the firm show nothing was rewritten after the event. That is the design point behind how Cognibl records work, which dilr.ai/cognibl describes as the same discipline behind our harness engineering practice, turned into a product.

How would a governed Matter Operations Desk work?

A governed Matter Operations Desk is a set of AI agents, each with one narrow matter-operations job, whose every output lands on a shared board as a task that cannot be marked done until its proof is attached, with a supervising solicitor's approval built into the firm's own workflow. The agents prepare work; solicitors decide. Our law research names six candidate roles for such a desk.

The six role names come from the law industry research behind this blog, where the desk appears as a candidate design for the profession rather than a deployed product. The one-line job descriptions are our own illustration of how each role could work under supervision:

  • a disclosure analyst that prepares a first-pass index and review set for a solicitor to check;
  • a client intake coordinator that tracks which onboarding documents are still outstanding, for a fee earner to chase;
  • a time capture assistant that drafts time entries for the fee earner to confirm;
  • a precedent librarian that retrieves the firm's own approved precedents for a drafting task;
  • a matter status communicator that drafts client updates for approval before they are sent;
  • a compliance evidence clerk that assembles the file a risk partner needs for an audit, for the risk partner to review.

Every role on that list ends with a person. That is deliberate, and it echoes the SRA's description of Garfield.Law: a system that is not autonomous, only takes a step once the client has approved it, and leaves named solicitors accountable for all its outputs. A desk that sends client updates or files time entries on its own is a different proposition, and not one this guide recommends.

One agent task, from instruction to approved output
01Scoped instructionVersioned, per role02Permitted accessOnly enabled tools03Proof attachedOutput plus evidence04Solicitor approvesThe firm's sign-off st…
Each step adds to the task's record. Cognibl refuses a done status until proof is attached; the named solicitor's approval is a step the firm builds into its own workflow.

Here is how those steps map onto Cognibl as the live product stands today. On dilr.ai/cognibl, Cognibl is described as a tracker for teams where people and AI agents work side by side, with agents picking up work under their own name against the same statuses the team uses. The definition of done, the evidence behind it and every tool call are attached to the work. A task reaches a done status only once proof is attached, a CSV describing the run that references the artefact, screenshots and hashes, and without a proof version the database itself refuses the move. The same rule holds for a person and for an agent, which matters in a firm where a trainee and an agent may do the same job.

Be clear about the current limits, because a risk partner will ask. The agents that Cognibl runs at cognibl.com today are coding agents on Claude Code and Codex, each of which opens a pull request that only a person can merge, and connected tools are read-only for now. So for matter operations the honest model is agent-prepared work products with a person approving each one, which keeps a named person accountable for every output, as the SRA required of Garfield.Law. Firms that want the board without agents first can use it as a plain tracker; our what is Cognibl explainer covers that ground.

How does client confidentiality change the design?

Client confidentiality changes the design by making access, not output quality, the first control. Under the SRA Code, a solicitor keeps client affairs confidential unless disclosure is required or permitted by law or the client consents, so an agent workflow must prove which matter data each agent could reach, and refuse everything else by default rather than by policy.

Paragraph 6.3 of the SRA Code is short, and it binds the solicitor rather than the vendor: you keep the affairs of current and former clients confidential unless disclosure is required or permitted by law or the client consents. Applied to agents, that turns into three design questions. Which matters can this agent see? Which tools can it call? And can the firm prove both answers after the event?

The strongest answer is an access model that fails closed. On dilr.ai/cognibl, agents reach the skills and agents library over the Model Context Protocol through a gateway that is deny by default: a toolset that has not been enabled is refused, not quietly missing. Every call is traced, every write is attributed by key name, and records are append-only and hash-chained. In matter terms, an agent can only call the toolsets its project has enabled, and anything else is refused. Which matters' data each connected tool exposes remains a design decision for the firm's own systems, and it belongs in the risk partner's sign-off.

Two further points sit outside the work-management layer. Where the work involves personal data, the ICO guidance on AI and data protection sets out the accountability questions a firm should be able to answer. And where privileged documents must never leave the firm's perimeter, the model question is separate from the governance question. Our private clinical small language models are one example from another regulated sector, and the open models guide to document extraction discusses running extraction on models an organisation controls.

If you want this designed around your firm rather than read about, our AI execution office embeds alongside your team to run the first agent workflows under these controls.

How should a firm roll this out without losing supervision?

A firm should roll out agent work one narrow, reversible task at a time, starting where an error is cheap to catch, checking every output in full until the record shows the agent is reliable, and only then moving to sampled review with the sampling rule written down. Supervision should loosen because evidence justifies it, never because a busy week demands it.

A practical sequence, in the order we would run it:

  1. Start with internal, low-consequence work. Compliance evidence assembly and precedent retrieval are good first tasks: the output goes to a colleague, not a client or a court.
  2. Write the definition of done before the agent runs. For a precedent retrieval, done might mean "returns only firm-approved precedents, each with its source location". A vague definition makes every check a judgement call.
  3. Check every output in full at first. Ayinde makes the check non-negotiable for research, and full review is the only way to learn what the agent gets wrong.
  4. Track the reopen rate. A task marked done and later reopened is a useful signal of an unreliable agent. Cognibl's delivery metrics include a reopen rate within 30 days and the human wait share, the time a task spends waiting on a person.
  5. Move to sampled review only with a written rule. The Law Society report notes that some legal teams already assess large-scale document review models by statistical sampling instead of manual review. That can be defensible, but only if the sampling rule is recorded and the risk partner has approved it.
  6. Keep client-facing and court-facing outputs under full review. Client updates and anything filed with a court stay with a named solicitor, whatever the reopen rate says.

The Law Society report names the risk this sequence guards against: if left unaddressed, incremental delegation could normalise an invisible substitution of professional judgement. Each step above keeps the substitution visible, because each one leaves a record a person signed. The same staged discipline runs through how we approach AI placement, and our Making Tax Digital evidence guide applies it to an accounting practice under HMRC's rules.

What is the best AI agent governance approach for a law firm in 2026?

The best AI agent governance approach for a law firm in 2026 is the one that produces a supervision record automatically, fails closed on access, and keeps a named solicitor as the approver of every client-facing or court-facing output. Which product delivers that depends on what the firm already runs, and for some firms the right answer is not a new tool at all.

Judge any option, including ours, against five criteria:

  • Proof before done. Can a task be closed without evidence attached?
  • Access that fails closed. Is an unapproved tool refused, or merely not configured?
  • A record that cannot be quietly rewritten. Append-only and chained, or editable by anyone with admin rights?
  • Versioned instructions. Can you read the exact instruction an agent ran months ago?
  • People keep the decision. Does any AI feature approve work, or only summarise and flag it?

On those criteria, a tracker such as Jira, Asana or Linear is strong on planning and visibility; Cognibl's own comparison on dilr.ai/cognibl frames the difference as classic trackers assuming a human moved the card, so check whether your tracker can require evidence before done. Agent frameworks such as LangGraph orchestrate an agent's steps; our LangGraph work management layer guide argues the record of done is a separate layer. Cognibl is designed around several of these criteria: proof before done, a deny-by-default gateway, an append-only record and versioned SKILL.md instructions, and on dilr.ai/cognibl its AI flows flag problems and do not decide, so people keep the decision.

The concession is real. A firm whose practice management and document systems already enforce matter walls, log every access and require approval before anything leaves the building may be better served by tightening those systems than by adding a new board. And a firm not yet using agents at all needs a placement diagnostic more than a platform. Our AI consulting services start there for that reason.

No AI agent can sign off legal work in a way that moves accountability. Under the SRA Code, a solicitor who supervises others remains accountable for the work carried out through them, and when the SRA authorised Garfield.Law in May 2025 it said named regulated solicitors stay responsible for all the system outputs. An agent can prepare and evidence work; a solicitor approves it.

The practical consequence is that every agent workflow needs an approval step with a named person attached, and the record should show who approved what and when. Cognibl applies the same proof rule to a person and an agent, but approval of legal work stays a human act.

According to the Law Society, not yet in any evidenced way. Its Generative AI essentials guidance says the Law Society's April 2026 report on the future of agentic AI in the legal profession found no evidence that agentic AI is currently being used in legal practice, though the report expects adoption to arrive incrementally, one delegated task at a time.

That gap is an advantage for a firm that wants to set its governance before the tools arrive rather than after. For the wider picture of how the DILR.AI lines map to a UK law firm, the industry guide covers each in turn, and our industries blog category collects the sector guides.

Want to see this in practice? Open Cognibl at cognibl.com, read our enterprise AI consulting guide, see AI developer productivity for governed coding agents, or talk to our team about a supervised pilot.

Product
Cognibl
Solution
AI Developer Productivity
Service
AI Operating Model
Talk to the operators

Put agent work on a record your risk partner can sign.

30-min scoping call · No deck · Confidential. We will tell you which matter tasks are safe to delegate first, and what the supervision record needs to show.

Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.

ai agent governance law firmuk law firms aiai agents law firm uksra ai supervisionayinde ai fake citationsai agents redditbest ai agent governance tool 2026cognibl

Questions this article answers

What does AI agent governance mean for a UK law firm?

AI agent governance in a UK law firm means the firm can show, for every task an agent touched, what the agent was asked to do, what it was allowed to access, what it produced, and which named solicitor checked the output before it reached a client or a court. The rules of professional conduct already bind the firm and its solicitors; governance is the evidence that those duties were met.

What did the High Court say in Ayinde about checking AI work?

In Ayinde v Haringey and Al-Haroun v Qatar National Bank, decided on 6 June 2025, the Divisional Court held that lawyers who use artificial intelligence for legal research have a professional duty to check its accuracy against authoritative sources before relying on it. The court said that duty also covers lawyers who rely on work others produced with AI, and called for action from those with leadership responsibilities, such as managing partners, and from regulators.

Who is accountable when an AI agent does matter work?

The solicitor who supervises the work stays accountable when an AI agent does matter work, and the firm stays bound by its own regulatory duties. The SRA Code of Conduct for Solicitors says a solicitor who supervises others remains accountable for the work carried out through them, and the SRA made the same point about accountability when it authorised the first AI-driven law firm.

What should a supervision record for agent work contain?

A supervision record for AI agent work should contain the instruction the agent ran, the data and tools it was permitted to use, the output it produced, the evidence a person used to check that output, the name of the supervising solicitor who approved it, and the time of each step. Each field answers a question the SRA, a court or a client could ask.

How would a governed Matter Operations Desk work?

A governed Matter Operations Desk is a set of AI agents, each with one narrow matter-operations job, whose every output lands on a shared board as a task that cannot be marked done until its proof is attached, with a supervising solicitor's approval built into the firm's own workflow. The agents prepare work; solicitors decide. Our law research names six candidate roles for such a desk.

How does client confidentiality change the design?

Client confidentiality changes the design by making access, not output quality, the first control. Under the SRA Code, a solicitor keeps client affairs confidential unless disclosure is required or permitted by law or the client consents, so an agent workflow must prove which matter data each agent could reach, and refuse everything else by default rather than by policy.

How should a firm roll this out without losing supervision?

A firm should roll out agent work one narrow, reversible task at a time, starting where an error is cheap to catch, checking every output in full until the record shows the agent is reliable, and only then moving to sampled review with the sampling rule written down. Supervision should loosen because evidence justifies it, never because a busy week demands it.

What is the best AI agent governance approach for a law firm in 2026?

The best AI agent governance approach for a law firm in 2026 is the one that produces a supervision record automatically, fails closed on access, and keeps a named solicitor as the approver of every client-facing or court-facing output. Which product delivers that depends on what the firm already runs, and for some firms the right answer is not a new tool at all.

Dilr Voice

Voice AI built for your sector

Dilr Voice answers and places calls 24/7 with compliance rules for regulated industries, from clinics and estate agents to financial services.

Related articles

← Previous
Embedded AI Delivery Team vs Project Consulting: UK Guide

One email, once a month. No hype. Just what we learned shipping.