BLOG

What Does It Mean to Delegate Work to AI? What We Learned Running Mr.AI In-House

What comes to mind when you hear "delegating work to AI"? You ask AI a question in a chat and get an answer back. For many people, we suspect, "using AI" still looks like this.

Since July 2026, we have been delegating the handling of change requests from our clients to "Mr.AI," an AI agent built into our own business tool, projectAI. Mr.AI carries the work forward from investigating the cause, writing up the specification, fixing the program and testing it, all the way to deploying it to the production environment.

This article is not an introduction to what we built with Mr.AI. The full picture of how it works is in our case study on running Mr.AI in-house. Here, we write about what it means to delegate work to AI, what we delegated and what we did not, and how we thought things through and fixed them when they did not go well.

Key points of this article: delegating means changing people's role from those who carry the work to those who decide it. Four questions for drawing the line, three principles for checking results, and five things to decide first

How "Using" AI Differs from "Delegating" to It

When you "use" AI, a person gives instructions each time, receives the result, and carries it to the next step. However capable the AI is, it is the person who moves the work forward.

When you "delegate" to AI, the AI carries the work from one step to the next. People make the decisions at key points.

Researchers at the University of Washington divide the autonomy of AI agents into five levels according to the role the person plays: an operator who gives instructions and controls it, a collaborator who works alongside it, a consultant who gives advice, an approver, and an observer who watches over it. They state that the level at which an agent operates does not follow naturally from the AI's capabilities, but is "a design decision" that its builders make deliberately[1].

In this framework, the level we chose for Mr.AI is the one where the person is the "approver." The AI moves the work forward, and people give approval at four gates: the specification, the merge, the check in the development environment, and the deployment to production.

In other words, delegating work to AI does not mean handing decisions over to AI. As we see it, it means changing people's role from those who "carry" the work to those who "decide" it.

Drawing the Line Between Work You Can Delegate and Work You Cannot

Four Questions We Use to Draw the Line

When deciding what to delegate to AI, we ask the following four questions.

1. Can the result be checked outside the AI? A change to a program can be checked with automated tests. A translation can be compared with the original. Work that has no way of being checked other than the AI's own "done" is hard to delegate.

2. Can it be undone if it goes wrong? If a draft or an investigation is wrong, you can simply throw it away. A deployment to the production environment or a promise to a client is hard to take back. Before any operation that is hard to undo, we always place a human gate.

3. Does the decision involve responsibility or values? Prices, priorities and commitments to clients have no single right answer. Decisions that raise the question of who takes responsibility stay with people.

4. How long is the piece of work delegated at one time? METR, a research organization that evaluates AI, measures how long a task AI can complete by the time the same task takes a person. For publicly available models as of February to March 2026, the length of task they could complete with a 50% probability was about 12 hours, but the length they could complete with an 80% probability was about 1.5 hours[2]. The more of a long piece of work you hand over in one go, the higher the chance it fails somewhere. That is why we divide work into short steps and check the result at each step.

What We Delegated and What We Did Not

Dividing Mr.AI's work according to the four questions gives the following.

What we delegated to AIWhat we kept with people
Investigating the cause of bugs and posting its assessmentJudging whether the specification is right
Drafting the specification (what to change, acceptance criteria, estimated effort)Judging whether a change can be merged
Asking questions about unclear pointsJudging whether it can go to production
Changing the program, and running and fixing automated testsPrices and effort
The work of deploying to the development and production environmentsPriorities
Monitoring for stalled work, and notifications in set situationsCommitments to clients, and contacting people on its own judgment

The left column is the work of "carrying" the work forward. The right column is the judgment that "decides" where the work goes.

In one survey, a research team at Stanford University asked 1,500 workers in 104 occupations how much of 844 tasks they would want to delegate to AI agents[3]. Workers were positive about automation for 46.1% of the tasks overall. The reason chosen most often was "it would free up time for high-value work" (69.38%). And in 45.2% of occupations, the most desired form of involvement was "people and AI working together as equal partners."

People do not want to give up all of their work. They want to let go of the carrying and spend their time on the deciding. Our own line ended up in the same place.

How We Check the Results of Delegated Work

We Do Not Take the AI's "Done" at Face Value

AI agent failures come in one particularly troublesome form: confidently reporting "Completed" when the work is not actually finished.

In a 2026 study, in a test environment simulating customer service, 45–48% of AI agent failures were of this "said it was complete, but it was not actually finished" kind[4]. The study also reports that having another AI judge the AI's responses caught almost none of these failures (it is a workshop paper that has not yet been peer-reviewed).

So with Mr.AI, whether to move on to the next step is never decided by the AI's self-report. Whether the tests passed is read directly from the results of the testing system. Whether a deployment to production has finished is also checked directly against the deployment's result before the work is marked complete.

Giving Approvers Material They Can Check

Having a person approve does not count as checking if all they do is read the AI's explanation and click because it "looks fine."

In an experiment published in 2025 by researchers at Princeton University and Microsoft Research, when the AI's answer came with an explanation, people were more likely to trust the AI, whether the answer was right or wrong[5]. On the other hand, when supporting sources were shown, or inconsistencies in the explanation were visible, over-reliance on wrong answers decreased.

For this reason, the approval screen shows not the AI's explanation but material people can check for themselves. For approving a specification, that is the summary of the specification, the acceptance criteria and the estimated effort. For the check in the development environment, it is the verification steps and the actual screens. Sending something back requires entering a reason, and if the content changes after approval, the approval has to be given again.

Do Not Add Too Many Approvals

If checking is the goal, it might seem best to have a person approve every step. But as the number of approvals grows, each check gets shallower and approval turns into a formality.

We placed approvals at four points: the specification (what to build), the merge (whether to put it into the main code), the check in the development environment (whether it works as expected), and the deployment to production (whether to release it into the client's operations). Each is either a point where the direction is set or the moment just before an operation that is hard to undo. The other steps are checked by test results and the monitoring system.

Cases That Did Not Go Well, and How We Fixed Them

Most Failures Were Not About How Smart the AI Was

Since we began delegating change requests to Mr.AI, a number of things have not gone well. Looking back, most of them were not a matter of "the AI got it wrong" but of "the work was not handed off properly."

Researchers at the University of California, Berkeley analyzed more than 1,600 failures in systems where multiple AI agents work together, and sorted 14 failure types into three groups: problems in system design, misalignment between agents, and insufficient verification of results[6]. Almost all of our own experience fits one of these three.

What Actually Happened

A "wait" arose that nobody noticed (late July 2026). A change that was about to be merged conflicted with another change, so the automated tests never started, and Mr.AI kept waiting for the results. Because the AI was only "waiting," no error was raised either. → We made it check whether the change can be merged before waiting for tests. In addition, a monitoring system that runs every 10 minutes compares everything against the actual state of the external systems.

Work on the same request ran twice (late July). When a person manually triggered a retry, the work already running and the new work proceeded at the same time. → We made sure that only one piece of work is ever running for a single request.

A late notification rolled back a step that had already advanced (early August). An old notification arrived late and overwrote a state that had already moved ahead. → We made it re-read the latest state just before writing a new one.

The AI kept working on a request a person had closed (late August). Work did not stop even on requests the person in charge had set to "Completed" or "Rejected," and unnecessary changes were being created. → We made work stop when a request is closed, and never resume from there.

The same investigation was done twice (late August). Investigating the cause and working out the fix started at the same time, duplicating the same research. → We added a step that waits for the investigation to finish before moving on, and hands the work to a person after a set time. Mr.AI itself has carried out part of this fix as a change request.

Progress was invisible to the person in charge (late September). Even when tests dragged on, the person in charge had no way of knowing. → We made it notify the person in charge when waiting for test results passes 60 minutes, and stop and hand over to a person at 3 hours.

What the Fixes Had in Common

Lined up side by side, most of the fixes were not about "making the AI smarter." They were about using the design of the system to prevent handoff problems: stalling, overlapping and rolling back.

The other thing that mattered was having a setup that lets us keep fixing things. A 2025 report compiled by a Massachusetts Institute of Technology project found that only a small fraction of organizations were getting measurable returns from their investment in generative AI[7]. The report names as a major cause that AI systems do not retain feedback from the people using them and do not adapt to how they are used (the results are described as preliminary, based on interviews and surveys). Work delegated to AI is not finished on the day you delegate it. Recording what did not go well and continuing to fix the system is the precondition for widening what you delegate.

What Companies Starting Out Should Decide First

When considering an AI agent, it is tempting to start by deciding "which AI to use." In our experience, though, these are the five things to decide first.

  1. Split the work into "carrying" work and "deciding" work. Delegate the carrying work first, and keep the deciding work with people.
  2. Set criteria outside the AI for judging whether work may proceed. Tests, checklists, a person's review: criteria other than the AI's self-report.
  3. Decide who approves and what material they look at. Limit approvals to points where the direction is set and to the moments just before operations that are hard to undo.
  4. Decide whom to notify, and when, if work stalls. Without monitoring and someone to hand over to, the AI's work stops without anyone noticing.
  5. Decide what information and permissions the AI must not have. For example, financial information, other clients' information, and the authority to contact people on its own judgment.

With these five decided, you can start by delegating up to investigation and drafting, then widen the scope little by little while looking at the records. Our overall approach to AI adoption is set out in "Our Philosophy on AI Adoption, and How We Put It into Practice," and common failures in "Common AI Adoption Failures and the Consultations We Actually Receive."

Frequently Asked Questions

How much can we delegate to an AI agent?

The basic rule is to start with work whose results can be checked outside the AI and that can be undone if it goes wrong. Investigation, drafting, and work that can be verified by tests fall into this category. Before decisions that set the direction, and before operations that are hard to undo, put a person's approval in place.

With so many approvals, doesn't the work for people stay the same in the end?

It does if you put approvals at every step. We limit them to four points, where the direction is set and just before operations that are hard to undo, and check everything else with test results and the monitoring system. Approvers receive only the material they need to decide, gathered in one place.

If the AI reports "Completed," can we trust it?

Do not take it on trust; check it against results from outside the AI. We recommend having a system that checks test results, and whether a deployment actually succeeded, separately from the AI's report.

Can a small company delegate work to an AI agent too?

Yes. You can start by delegating investigation and drafting to AI in one area of work. Through our AI Implementation Support, we help you from the very first step: running your first agent in your own work.

Summary: Delegating Means Becoming the One Who Decides

Delegating work to AI does not mean handing decisions over to AI. It means handing the role of carrying the work to AI, so that people can focus on the role of deciding.

To do that, you need to draw a line around what you delegate, check results against criteria outside the AI, limit where approvals happen, and decide in advance who takes over when work stalls. And by continuing to fix the system based on what did not go well, the range of work you can delegate gradually widens.

The system we built with Mr.AI is described in our case study on running Mr.AI in-house, and the use of AI across projectAI as a whole in our case study on running projectAI company-wide. If you would like to work out together where to draw the line on what to delegate to AI, please consult us through AI Implementation Support.

References

  1. K. J. Kevin Feng, David W. McDonald, Amy X. Zhang, "Levels of Autonomy for AI Agents," 2025
  2. METR, "Frontier Risk Report (Feb–Mar 2026)," May 2026
  3. Yijia Shao et al. (Stanford University), "Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce," 2025 (revised 2026)
  4. Laksh Advani, "From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents," FAgen@ICML 2026
  5. Sunnie S. Y. Kim et al., "Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies," CHI 2025
  6. Mert Cemri et al. (University of California, Berkeley), "Why Do Multi-Agent LLM Systems Fail?," 2025
  7. MIT NANDA, "The GenAI Divide: State of AI in Business 2025," July 2025