Most pilots never reach daily work. We look at why - on implementations we took all the way into shifts and still support.
Because a pilot usually answers the question "can the model do this", while the work starts with the question "who reads this and when". The demo goes well, then the output lands somewhere nobody opens, and two weeks later it is forgotten.
The second common case is an agent built into a process that does not exist. If nobody did that reconciliation before, its absence does not hurt, and the agent becomes optional.
The ones where the answer can be checked. A parsed invoice is checked against the document, a price gap against purchase history, an unwritten-off balance against stock. Checkability matters more than difficulty: a hard but checkable task holds, an easy and uncheckable one does not.
In our implementations that means document work and reconciliation between systems above all: there the agent removes a manual step entirely rather than speeding it up.
Where a mistake is paid for with money, with people or with a supplier relationship. Refusing a payment, changing a price, letting somebody go, stopping a delivery - these are decisions with consequences, and a human takes them.
The split we settled on: the agent brings the task to the point where all that is left is to agree or refuse. The person spends a minute instead of an hour, but the decision stays theirs.
Not the model, the surroundings. A supplier changes the layout of a delivery note, a new category appears in the system, the person who acted on the answers goes on leave. Any of these quietly switches the agent off.
That is why support in our projects is not "call us if it falls over" but a regular check that the agent still answers the question it was given.
Record two figures before launch: how long the task takes now and how often a mistake is found in it. Afterwards compare against those, not against a feeling.
We calculate the effect on the client side and on their data. That is why this note carries no aggregate percentage: averaging savings across different companies and presenting it as a general result is not something we will do, because such a figure belongs to nobody.
How long the effect holds. Our oldest agents have run for less time than it takes to speak about years, and our honest answer to "what happens in three years" is one line: we do not know yet, we keep watching.
These observations come from implementations ITHS ran itself and still supports: 150+ projects per ITHS internal records as of 7 October 2026. Client data is not disclosed, only anonymised observations are used. We publish no aggregate figures on this topic until the sample is closed - this note holds only what repeats from project to project.