The problem
A log that records a violation has already permitted it.
Ask an enterprise how it governs AI and you will usually hear about traces, dashboards, audit logs,
and evaluation suites. Every one of those instruments acts after the model has acted. They are
excellent at telling you that something went wrong. None of them stopped it.
That gap is tolerable when AI drafts a summary someone reads before sending. It stops being tolerable
the moment an agent writes to a repository, moves money, touches a member record, or opens a pull
request against a regulated system. At that point the question is not what happened. It is
what was allowed to happen, and what could have prevented it.
The models are extraordinary at producing output that sounds correct, and plausibility is not
correctness. Professional-services firms have retracted AI-assisted reports over fabrications. Law
firms have been sanctioned over invented citations. The failure mode is not that the model refuses to
work — it is that it works confidently and wrongly, and nothing in the path objected.
There is a second, slower failure that matters more to a CFO than to an engineer. A Carnegie Mellon
study of AI coding agents found the initial three-to-five-times velocity gain
dissipates within roughly three months, as security, maintainability, reliability and
complexity defects accumulate faster than the features do. Unverified speed does not compound. It
converts into debt, and the debt eats the gain.
That finding is third-party and industry-reported. It is cited here as the argument for verification
discipline, not as a measured result of any one programme.