<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Andrew Lewis was Here]]></title><description><![CDATA[AI adoption strategies for technology leaders where governance matters, timelines are real, and nobody wants to be the cautionary tale]]></description><link>https://andrewlewis.ca</link><image><url>https://substackcdn.com/image/fetch/$s_!ISZj!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd024d759-eeb3-44b1-8186-48f22d0817b3_764x766.png</url><title>Andrew Lewis was Here</title><link>https://andrewlewis.ca</link></image><generator>Substack</generator><lastBuildDate>Wed, 22 Jul 2026 00:27:27 GMT</lastBuildDate><atom:link href="https://andrewlewis.ca/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Andrew Lewis]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[andrewlewiswashere@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[andrewlewiswashere@substack.com]]></itunes:email><itunes:name><![CDATA[Andrew Lewis]]></itunes:name></itunes:owner><itunes:author><![CDATA[Andrew Lewis]]></itunes:author><googleplay:owner><![CDATA[andrewlewiswashere@substack.com]]></googleplay:owner><googleplay:email><![CDATA[andrewlewiswashere@substack.com]]></googleplay:email><googleplay:author><![CDATA[Andrew Lewis]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[We Were Already Offloading Our Thinking]]></title><description><![CDATA[A response to Yennie Jun: the autonomy we&#8217;re afraid of losing to AI was mostly delegated already &#8212; what matters is whether anyone reviews what comes back.]]></description><link>https://andrewlewis.ca/p/we-were-already-offloading-our-thinking</link><guid isPermaLink="false">https://andrewlewis.ca/p/we-were-already-offloading-our-thinking</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Mon, 20 Jul 2026 12:29:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bh-_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bh-_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bh-_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png 424w, https://substackcdn.com/image/fetch/$s_!bh-_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png 848w, https://substackcdn.com/image/fetch/$s_!bh-_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png 1272w, https://substackcdn.com/image/fetch/$s_!bh-_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bh-_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8667713,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/207766746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bh-_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png 424w, https://substackcdn.com/image/fetch/$s_!bh-_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png 848w, https://substackcdn.com/image/fetch/$s_!bh-_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png 1272w, https://substackcdn.com/image/fetch/$s_!bh-_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3cb7b43-b328-47cc-bd7b-dbb42156b798_3000x2000.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Yennie Jun published an essay last week asking whether we are offloading too much of our thinking to AI. It is the strongest version of that worry I have read &#8212; grounded in her own habits, honest about the benefits, and free of the usual doom. Her most unsettling figure is a man her friend met at a San Francisco startup event, wearing a small microphone that records every conversation he has so an AI can analyze his day. &#8220;I let Fable do all of my thinking these days,&#8221; he explained, with apparent enthusiasm. Her core claim is that autonomy depends on continuing to participate in forming our own desires, and that handing decisions to AI &#8212; even trivial ones, like what to eat or what to listen to &#8212; slowly erodes who we are.</p><p>She is right about more than is comfortable. And the essay still aims at the wrong danger. The threat was never that we delegate our thinking; we have always delegated it, to whatever was nearest. The threat is delegating without a review loop. The distinction matters because the first framing is existential, and leaves you nothing to do but worry. The second is operational. Operational problems can be fixed.</p><h2><strong>The alarm has a history</strong></h2><p>Socrates warned that writing would destroy memory. In the Phaedrus, he argues that people who trust the written word will stop exercising their memories and carry the appearance of wisdom rather than the real thing. He was not wrong. The trained memory of oral culture &#8212; the capacity to hold an epic poem or a body of law in your head &#8212; genuinely atrophied. We traded it for literacy, and literacy bought philosophy, science, and the common law.</p><p>The pattern has repeated on schedule ever since. Calculators reached classrooms in the 1970s, and a generation of parents was warned that numeracy would die. It largely did, at least the mechanical kind; almost nobody now does long division for a living, and nobody mourns it. GPS arrived, and spatial reasoning measurably declined &#8212; there is research associating habitual turn-by-turn navigation with weaker spatial memory. The losses in each case were real. Jun is honest about the costs of offloading, so the response should be honest too: this is not a story where the worriers were simply wrong.</p><p>What the worriers missed is what happened next. Each time, the freed capacity moved up a level, and the surrendered skill quietly stopped mattering. Every skill you maintain has a <strong>carrying cost</strong> &#8212; the practice hours required to keep it sharp, paid from a fixed budget of attention. The useful question about any capability AI can now perform is not &#8220;are we losing it&#8221; but &#8220;is it still worth its carrying cost.&#8221; Sometimes the answer is yes: writing-to-think survived the word processor, because the value was never the typing. Sometimes the answer is no, and pretending every loss is an erosion of self collapses a judgment call into a panic.</p><h2><strong>The autonomous baseline never existed</strong></h2><p>Her framing carries a quieter assumption: that before AI, we were forming our own desires. Mostly, we were not. The music was chosen by Spotify&#8217;s recommendation engine, the restaurant by an aggregate of strangers&#8217; reviews, the shoes by a marketing budget, and the opinions &#8212; more than anyone likes to admit &#8212; by an engagement-ranked feed. Ken Liu&#8217;s &#8220;The Perfect Match,&#8221; the 2012 short story Jun builds her essay around, was not a prophecy about large language models. It was a barely exaggerated description of the recommendation systems we had already accepted.</p><p>Jun asks exactly the right question: who is making the final decisions for the things that matter in your life? But the honest pre-AI answer was rarely &#8220;me, unaided.&#8221; It was a network of delegations we never audited, because none of it looked like delegation. It looked like convenience.</p><p>Against that baseline, an AI assistant is a strange place to plant the flag, because it is <strong>the first delegate in the chain you can actually interrogate</strong>. Ask Claude why it recommended something and you get reasons &#8212; reasons you can push on, disagree with, and overrule. Try asking the TikTok algorithm why it fed you what it fed you. An articulate delegate carries its own risk; a confident explanation can persuade when it should not. But a delegate that shows its reasoning at least makes review possible. The feed never offered that. Measured honestly against what it replaces, this is the first delegation mechanism in twenty years that hands any autonomy back.</p><h2><strong>Delegation is a discipline</strong></h2><p>Most of my job is delegation. I lead development teams and a project management office at a national law firm, and very little that matters ships from my own hands anymore. Nobody describes this as a loss of capability. A director who insists on doing everything personally is not more capable than one who delegates well &#8212; they are less, and every organization knows it. Deciding what to hand off, with how much context, and how tightly to inspect what comes back is high-order judgment. It may be the highest-order judgment the AI era pays for.</p><p>The legal industry has run on this discipline for a century. Associates draft; partners review; the partner&#8217;s name goes on the opinion. Nobody argues that the partner has surrendered their autonomy to the associate. The profession also knows precisely what the failure looks like, because it has a name for it: the partner who signs without reading. The failure was never delegation. It was <strong>delegation without a review loop</strong>.</p><p>Read Jun&#8217;s essay through that lens and her best moment confirms it. In Portugal, she and her sister wondered why the country celebrates explorers the United States would call colonizers. Her sister reached for ChatGPT; Jun suggested they think first. They speculated, disagreed, backtracked &#8212; and only then asked the AI, using its answer to test and extend hypotheses they had formed themselves. That is not resistance to offloading. That is a review loop, run well. Her mother&#8217;s physics students, meanwhile, pasting assignment questions into a chatbot and submitting the output unread, are the partner who signs without reading. Same technology, opposite disciplines. What separates them is not how much thinking was offloaded. It is whether anyone reviewed what came back &#8212; an operating problem, and operating problems are trainable. Organizations worried that AI will erode their people&#8217;s judgment should be training exactly this muscle: what to delegate, what to hold, how to review.</p><h2><strong>The line that deserves the worry</strong></h2><p>One part of Jun&#8217;s argument deserves to be granted in full. There is a real difference between delegating task execution and delegating desire formation. &#8220;Draft this memo&#8221; is delegation. &#8220;Tell me what to want from my career&#8221; is something else, because if you outsource the wanting, no one is left standing behind the review loop. A reviewer needs preferences of their own to review against.</p><p>But this is not a new problem arriving with a new technology. We already share our desire formation with other minds that hold opinions about what we should want &#8212; mentors, spouses, consultants, therapists. A good mentor absolutely tells you what to think, and sometimes what to want. Centuries of social practice taught us to take that input as input: to hold advisers at the distance where they inform the wanting without replacing it. We also recognize instantly when the line fails &#8212; the mentor who dictates, the consultant who decides for you. The skill of holding that line exists. It survives, like any skill, only if it is practiced. That is Jun&#8217;s real warning, and it stands stripped of its fatalism.</p><p>The Microphone Man is genuinely alarming, but the microphone is not why. He is alarming because he announced, cheerfully, that his review loop is closed &#8212; the AI is smarter, so the AI does the thinking now. Socrates&#8217; alarm about writing was real too, and the answer was never to stop writing things down. It was to remain the kind of reader who checks. That is still the answer. Delegation without review is not offloading your thinking; it is abdicating it. And abdication was never the tool&#8217;s decision to make.</p><div><hr></div><p>Read Yennie Jun&#8217;s original essay, <a href="https://www.artfish.ai/p/offloading-thinking-to-ai">Are we offloading too much of our thinking to AI?</a> &#8212; it is the strongest version of the worry, and it deserves the argument. If the operational side of AI adoption is your beat, subscribe; it is the only thing I write about.</p>]]></content:encoded></item><item><title><![CDATA[Agent Governance Is Access Control You Never Wrote Down]]></title><description><![CDATA[Every requirement a UN body just named for a trustworthy AI agent is the access discipline you already run for people &#8212; the only new part is the judgment the machine doesn&#8217;t bring.]]></description><link>https://andrewlewis.ca/p/agent-governance-is-access-control</link><guid isPermaLink="false">https://andrewlewis.ca/p/agent-governance-is-access-control</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Thu, 16 Jul 2026 12:01:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!HSgt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HSgt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HSgt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HSgt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HSgt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HSgt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HSgt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg" width="1456" height="970" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:970,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:317016,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/207010946?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HSgt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HSgt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HSgt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HSgt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd00f7c4-ce11-4327-b1b5-b07bcd051233_2000x1333.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>A trustworthy machine and a trustworthy employee need the same six things. A United Nations standards body spent this month naming them for machines. Enterprises named them for people years ago and filed the whole discipline under access management.</p><p>Here is what actually happened. The ITU &#8212; the International Telecommunication Union, the UN agency that has set the rules for how machines talk to each other since 1865, back when the machines were telegraphs &#8212; opened a focus group in July on trust and identity for AI agents. Its agenda names what a trustworthy agent has to account for: identity, authority, delegation, human oversight, accountability, and provenance. The first working meeting is in Paris in November, the second in Geneva in January. Treat that as a new frontier and the honest move is to wait &#8212; let the standards body convene, deliberate, and eventually tell you what a trustworthy agent is.</p><p>Treat it as an operator and the list stops the wait cold. You have seen every item on it before.</p><h2><strong>You already run this system</strong></h2><p>Read the agenda again with an IT lens instead of a policy one and it resolves into something ordinary:</p><ul><li><p><strong>Identity</strong> is a credential &#8212; the verifiable answer to &#8220;who is this.&#8221;</p></li><li><p><strong>Authority</strong> is permissions &#8212; can it spend, sign, send, move money.</p></li><li><p><strong>Delegation</strong> is the grant record &#8212; who handed over the access, and who can take it back.</p></li><li><p><strong>Oversight</strong> is an approval gate &#8212; the step where a human has to say yes.</p></li><li><p><strong>Accountability</strong> is knowing whose name is on the action when it goes wrong.</p></li><li><p><strong>Provenance</strong> is an audit log &#8212; the trail that lets you reconstruct what happened.</p></li></ul><p>That is not an emerging governance discipline. It is identity and access management, the practice every regulated enterprise has run for its people for years &#8212; the joiner-mover-leaver process, the role-based permissions, the quarterly access review, the audit trail the regulator asks to see. A UN body did not describe a new problem. It described your IAM program and pointed it at a new kind of user.</p><p>The maturity of that program was always measured by one question: can you say, without a two-week fire drill, who can touch a given system and why. Most enterprises can answer it for their people, at least on paper. Almost none can answer it yet for the agents they have already switched on.</p><h2><strong>Why it feels new anyway</strong></h2><p>The reason the list reads as novel is that we built every piece of it around a human being, and the human quietly did half the work.</p><p>A person authenticates, gets permissions scoped to a role, and then supplies the one thing the access model never had to encode: judgment. A lawyer with full reach into the document system does not open every matter she can technically open. She knows which files are walled off, which client would object, which situation calls for asking first. The permission set was never actually safe on its own. It was safe because a person with context and something to lose sat behind it.</p><p>An agent inherits the access and brings <strong>none of the judgment</strong>. Spin one up inside that same document system and it has the lawyer&#8217;s reach with none of the lawyer&#8217;s restraint. It will do exactly what its permissions allow, at machine speed, and nothing in it will pause to ask whether that was wise. The credential still resolves. The permissions still check out. The judgment that made them safe is simply gone.</p><p>That is the only genuinely new fact in the whole affair. Not identity, not authority, not audit. The actor holding the permissions no longer comes with a person attached.</p><h2><strong>You don&#8217;t have to wait for Paris</strong></h2><p>Some of this work does need the standards body. When your agent has to prove itself to another company&#8217;s agent across an organizational boundary &#8212; no shared HR system, no common directory, no one who onboarded them both &#8212; you need an agreed way to exchange and verify identity and authority. That is real, it is genuinely new, and it is what Paris and Geneva are for. The scientists advising the UN keep making a related point: the capability is running ahead of the evidence base the policy needs. A standards process moving at the speed of a November meeting will not hand you a safe agent next quarter.</p><p>Inside your own walls, you are not waiting on anyone. The six requirements are your existing access model, made explicit for an actor that supplies no judgment of its own. This is what &#8220;AI without operations is just a demo&#8221; looks like at the scale of a single agent: the demo is the capability, and the operations is the access model you write down around it.</p><p>You provisioned people out of habit &#8212; a role, a template, an approval, done. You have to provision an agent <strong>on purpose, in writing</strong>: state what it may touch, what it may never touch, where a person has to approve, and whose name answers for it when it goes wrong &#8212; the developer who built it, the team that deployed it, or the human who was supposed to be supervising. The implicit rule has to become an explicit one, because the thing holding the access will follow the letter of what you wrote and none of what you assumed.</p><h2><strong>What the agent is actually for</strong></h2><p>Here is the part worth sitting with. The agent is not the thing that needs governing. It is the thing that finally forces you to write your governance down.</p><p>For years &#8220;allowed&#8221; meant &#8220;technically permitted, and trusted not to abuse it&#8221; &#8212; and the second half of that sentence lived in people&#8217;s heads, never on paper. An actor with no judgment turns that unwritten half into a specification you have to author, review, and own. The uncomfortable gift of agents is that they make you say, in writing, what you always meant by access.</p><p>A UN focus group can define what trust between agents should look like across the industry. It cannot write down what &#8220;allowed&#8221; was always supposed to mean inside your organization. Only you can &#8212; and now, for the first time, you have a reason to. The six things that make a machine trustworthy were never really about the machine. They were about the discipline you already had, and finally have to put in writing.</p><div><hr></div><p><em>If this landed, the earlier posts in this thread &#8212; on why your agent inherits your lawyer&#8217;s permissions, and why least privilege is the only sane default &#8212; are on my LinkedIn. Subscribe here to get the next one on operating AI, not just deploying it.</em></p>]]></content:encoded></item><item><title><![CDATA[What a Government Proved About AI That Vendors Wouldn’t]]></title><description><![CDATA[The Government of Alberta published twenty-one papers on rebuilding its technology with AI, and the real lesson is about the part you can&#8217;t buy.]]></description><link>https://andrewlewis.ca/p/what-a-government-proved-about-ai</link><guid isPermaLink="false">https://andrewlewis.ca/p/what-a-government-proved-about-ai</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Tue, 14 Jul 2026 12:01:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1tMk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1tMk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1tMk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!1tMk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!1tMk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!1tMk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1tMk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:315481,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/205995666?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1tMk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!1tMk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!1tMk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!1tMk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82da69fd-ba9d-4ab4-bbce-308a010f5b42_1536x1024.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>My own provincial government just published a manual for the part of AI adoption that almost everyone skips. In June, the Government of Alberta released twenty-one white papers describing how it is rebuilding four decades of public technology with a workforce of AI agents. The papers lead with numbers built to be doubted: a twenty-five-year-old benefits application rebuilt in four days instead of five months, a new public system delivered in eleven weeks for roughly $108,000 against a traditional estimate near $1.9 million, and a stated goal of cutting the cost and time of building government software by ninety-five percent.</p><p>Read the numbers, raise an eyebrow, move on. That is the reasonable response, and it misses the actual story. The results Alberta is reporting did not come from the AI. They came from the operations the province built around it. And the operations are the one thing a vendor cannot sell you.</p><h2><strong>The vendors were half right</strong></h2><p>One of the papers is a first-person account from Chris Wright, a director who spent years managing technical-debt work and legacy integrations. In his conversations with vendors, he kept hearing the same message: AI has significant limitations for this kind of work. He was skeptical of the claim but had nothing concrete to argue with, because he had never used the tools at any depth himself.</p><p>So he tested it on something he knew cold. Twenty-five years ago he wrote the Remote Area Heating Allowance application for Alberta Agriculture, by hand, in Java. It took five months. It is still running today. He put it through Alberta&#8217;s build process and rebuilt it in four days, and the new version does more than the original, including a public-facing portal the first one never had. He wrote the original code, so he could judge the output with direct knowledge of what it was supposed to do. The quality held.</p><p>That one was a prototype, a proof of capability. The harder proof is the Alberta Classroom Information Portal, a genuine enterprise system for principals and teachers, carrying the full weight of privacy protection, security review, identity management, and production infrastructure. It went live in June. It was delivered in eleven weeks for about $108,000, against a traditional estimate of $1.3 to $1.9 million and a timeline north of a year. It is not a tool anyone could afford to get wrong, and it was held to the same bar a traditional product team would have to clear.</p><p>The vendors were not exactly lying. Twelve months earlier the models could not have done this, and honest vendors were describing the tools as they were. But they got the deeper thing wrong. They were talking about the model, and the model was never where Alberta&#8217;s advantage came from. <strong>The model was the easy part.</strong></p><h2><strong>What a vendor cannot sell you</strong></h2><p>Alberta is unusually direct about this. Their coding work is model-agnostic by design. Claude is their current preference, but they say plainly they expect that lead to change, and they have built everything so it can. In their own words, the intellectual property of what they call the &#8220;harness&#8221; matters more to the outcome than the specific model underneath it.</p><p>The harness is not exotic. It is a short file of project rules the AI reads first, a library of small instruction sets they call skills, and a near-perfect reference application the AI is told to copy. More than a hundred of the decisions a public system has to get right, how a user signs in, how input is validated, how secrets are stored, are made once, in that template, before any project begins.</p><p>The reason they needed it is worth understanding, because it is the reason most AI pilots quietly disappoint. Left to itself, the same model given the same request twice produces two different applications, with different layouts and different decisions underneath. Two capable people using the same tool diverge just as far. The prototypes came fast and looked finished, and then the months saved on the front end were lost on the back end, fixing the security and accessibility gaps the demo had hidden. None of that was a failure of the model. It was the absence of a defined way of working.</p><p>They are also blunt about the cost of building the fix. You cannot produce a good harness casually or at speed, they write, and every version copied from a copy degrades. One of the case-study authors put it more plainly still: more work went into building the factory than into getting the applications out the other end.</p><p>That sentence is the whole argument in miniature. What Alberta built was not software. It was <strong>encoded judgment</strong>, the standards and decisions and taste that normally live in a senior engineer&#8217;s head, written down so that a machine, or a novice, applies them the same way every time. A model you can buy, and its edge expires on someone else&#8217;s release schedule. Encoded judgment is yours, and it compounds. In a law firm, the equivalent is not a better contract tool. It is the risk posture, the house style, and the hard-won preferences a senior partner carries by instinct, finally written down where they can be applied at scale, and checked.</p><h2><strong>The standards no one had written down</strong></h2><p>Before Alberta could rebuild anything, it had to see what it actually had. So it pointed a tool it calls Git Insights at its entire code estate: about fifty agents reading 466 million lines across roughly 3,400 code repositories in around twenty hours, for under $2,000. No consultant engagement comes close on time or cost.</p><p>The finding that should travel furthest has nothing to do with speed. Across eight years and 8,178 contributors, Alberta&#8217;s software met its own standards completely on first release only forty percent of the time. The rest needed rework.</p><p>Their reading of that number is careful, and I think it is correct. This was not a story about lazy or unskilled people. The contributors were capable and hard-working, and most of the code was sound when it shipped. The problem was systemic. The standards were aspirational rather than prescriptive, inconsistently communicated, and silent in the contracts themselves. Where the rules are undefined, capable people make their own reasonable choices, and thousands of reasonable choices over eight years become drift. They compare the result to a medieval city, every building sound on its own terms, the whole grown dense and hard to cross.</p><p>There is a fairness in how they handle this that I want to underline, because it is rare in AI writing. We remember the one time an AI gets an answer wrong and forget the ten thousand times it was right, and we almost never hold human performance to the same standard in the same breath. Alberta did. It measured its own people honestly, at a scale no manual review could, and it credited them: the creativity, the judgment, and the work that no model produced. No AI dreamed up Git Insights. A person did.</p><p>This is the part most organizations will recognize if they are honest. The lesson is barely about AI at all. Most companies do not have an AI problem. They have <strong>an undocumented-standards problem</strong>, and AI is simply the first tool sharp enough to make the gap visible and then to close it. Alberta draws the comparison to a car plant: we would not accept a factory where forty percent of the vehicles left the line passing their safety checks. We hold software to a lower bar mostly because, until now, we could not see the number. AI did not fix the standard. It made the gap countable, and then it made the standard enforceable.</p><h2><strong>Compliance moved to the front</strong></h2><p>The enforcement is the most interesting mechanism in the collection. Alongside the build sits a set of review agents Alberta names by color. One checks code quality and hygiene. One reads the prose for the tells of machine-written text. One attacks the finished application from the outside, the way an intruder would. One works through the defensive checklist. Together they run more than 400 individual checks, including 285 requirements from an international security standard and 62 of Alberta&#8217;s own cloud-security rules, and they run continuously, as the application is built, for pennies each time.</p><p>The consequence is a real inversion. By the time the cybersecurity team sits down to review an application, it is already compliant, with the evidence attached. In one training cohort, a hundred public servants, most of whom would never have called themselves engineers, produced more than 560 working applications in six days, and the strongest cleared the cyber review and the accessibility check on first release. <strong>Compliance moved from the last gate to the first</strong>, and it now runs on every change instead of once at the end under deadline pressure.</p><p>There is a line buried in the technical papers that I keep returning to: you cannot be accountable for a process you have not defined. That is the quiet argument under the whole security model. Left undefined, the work happens inside the model&#8217;s private and probabilistic reasoning, and you gain speed while losing any real understanding of how the result was reached. Written down as skills, the same expectations become standard operating procedure, applied the same way every time and auditable after the fact.</p><p>For anyone working in a regulated business, this is the line to sit with. The same collection describes a pipeline that took one Alberta act and benchmarked it against the equivalent statutes across fourteen jurisdictions for about $200, with every finding traceable back to the exact clause it came from. Comparative statutory analysis that would take a team months, done in an afternoon and defensible down to the line. The speed is the least of it. The real change is that a control you used to run once, expensively, at the end, now runs continuously, for the price of lunch.</p><h2><strong>The bottleneck moved to us</strong></h2><p>Speed like this breaks the tools we use to manage work. On the coding task alone, Alberta measures the AI at well over a hundred times a human developer&#8217;s pace, and they are candid that traditional estimation collapses as a result. An AI has no reliable sense of how long a human would take, and a human has no reliable sense of how fast the AI will move. Planning poker stops meaning anything.</p><p>So they built a delivery tool that runs a kind of chess clock between the person and the AI, tracking who is holding the work at every step. The finding is uncomfortable and worth repeating. When a project misses its target, it usually has little to do with the AI. The AI finished in an hour; the person took a week to look at it. The delay lives in the handoff.</p><p>Their most ambitious paper follows that thread all the way down, and it is the one I would press on a fellow leader. It treats the org chart as a compression algorithm, a way of squeezing ground truth up through the layers so it fits inside a senior person&#8217;s attention, and then hydrating strategy back down into local action. A person conveys meaning at roughly forty bits a second. A model works orders of magnitude faster. Hold every interaction to human speed and you leave most of the gain on the table. In a human hierarchy you want one manager for every several workers; in an agentic one, the ratio may invert, with several supervising agents auditing every worker that builds. Their conclusion is bracing: the twentyfold improvement they are chasing is not achievable while keeping the org chart, the briefing note, and the approval chain exactly as they are. The hierarchy has to be rebuilt at the same time as the technology, or the old process quietly eats the new speed.</p><p>The operating lesson is simple to state and hard to act on. If an AI pilot did not save much time, the tool is probably not the problem. Look at <strong>the wait between the work</strong>, not the work itself. The model got faster. The organization around it did not.</p><h2><strong>Measure capability, not savings</strong></h2><p>Which raises the question of how you would even know you were winning. Alberta&#8217;s answer is the most quietly radical idea in the whole set, and the one I would steal first. Do not measure the program by the money it saves.</p><p>Their reasoning is hard to argue with. Continuing to do exactly the work you do today, just with fewer people, banks a one-time saving and forfeits the larger prize: an organization that can do things it could not do before. So they judge the work against three measures instead. Readiness, meaning whether the organization can actually carry the change, from staff skill to governance. System health, meaning whether the estate is getting safer and smaller rather than sicker and larger. And cost, counted honestly across the whole of government rather than shifted from one budget line to another.</p><p>The line that stays with me is their own: continuing the same work with fewer staff is folly. It reframes the entire exercise. The point of the harness, the standards, and the factory is not a cheaper version of the current output. It is a larger capability, measured by what the organization can now attempt. Dollars saved is the metric that makes an AI program look successful while leaving it exactly where it started.</p><h2><strong>What none of this ships in</strong></h2><p>Set the pieces side by side. The harness of encoded judgment. The standards finally written down. The compliance gates running on every change. A way of measuring delivery that catches the human handoff instead of hiding it. A measure of success built on capability rather than savings. An academy that put thousands of public servants and more than ten thousand members of the public through structured training. None of it arrives in a license. No vendor sells it, because no vendor can. It is made of one organization&#8217;s own rules, its own work, and its own judgment about what &#8220;good&#8221; means.</p><p>Alberta&#8217;s sharpest strategic move follows from that. Rather than centralize all software delivery or scatter it, they propose holding the protective core at the center, the security, identity, and data rules, while opening the actual building out to the people who understand the work. The center stops being the place software is delivered and becomes the place software is governed. The alternative is not a tidy status quo. It is shadow AI, built outside the fence with none of these controls, which arrives whether it is sanctioned or not. Their instinct for what a leader should buy is the same everywhere in the collection: stop buying chat conversations, and start building the pipelines and the standards that outlast the question you asked today.</p><p>Skepticism is still warranted, and Alberta invites it. These are self-published papers, and their numbers deserve interrogation. The largest claims, the ninety-five percent and the twentyfold, are targets at least as much as they are settled results. Some of what works in a government estate of unclassified legacy code will not transfer cleanly to a firm handling privileged client data. But the method is legible, the smaller numbers are concrete and consistent, and the method is the point. Alberta&#8217;s real product was never the applications. It was the capability to produce them, and they published the manual for building it.</p><h2><strong>The operations are still on you</strong></h2><p>The models will keep getting cheaper and better on a schedule none of us controls, which is exactly why the model is the commodity in this story. The operations will never be a commodity, because they are built out of your own judgment, and no one can hand you that.</p><p>AI without operations is just a demo. Alberta spent eighteen months building the operations, then gave the manual away. The building is still on the rest of us. You can download their manual this afternoon. You will still have to write your own.</p><div><hr></div><p>If this was useful, subscribe. I write about the operational side of AI, the part that never fits inside a product demo. The Velocity White Papers are public at thevelocitywhitepapers.com, and they are worth your time.</p>]]></content:encoded></item><item><title><![CDATA[The Clock You Can’t See From Across the Table]]></title><description><![CDATA[Zack Shapiro&#8217;s &#8220;The Two Clocks&#8221; is right about the biggest problem in AI and law. Here is what it looks like from inside a firm.]]></description><link>https://andrewlewis.ca/p/the-clock-you-cant-see-from-across</link><guid isPermaLink="false">https://andrewlewis.ca/p/the-clock-you-cant-see-from-across</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Wed, 08 Jul 2026 22:50:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_s7H!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_s7H!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_s7H!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!_s7H!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!_s7H!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!_s7H!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_s7H!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2268209,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/206212056?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_s7H!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!_s7H!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!_s7H!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!_s7H!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5b94139-e380-427d-847f-59f515ac3c22_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Zack Shapiro&#8217;s <a href="https://x.com/zackbshapiro/status/2074852265354022983">The Two Clocks</a> is the most accurate account of the AI-and-law problem I&#8217;ve read from anyone standing outside a firm. His argument, in one line: AI capability has raced years ahead of the institutions meant to use it, and the gap between the fast clock of the technology and the slow clock of the firm is the real story. Most of what follows is agreement. The rest is what that picture looks like from inside one, where I run AI and innovation for a national practice and spend my days in the exact gap he names.</p><p>Start with where he&#8217;s right, because he&#8217;s right about the thing that matters most.</p><h2><strong>The bottleneck really has moved</strong></h2><p>The line I&#8217;d underline twice is his: <strong>the bottleneck is no longer the intelligence, it&#8217;s the absorption of it.</strong> That is the whole job. The capability has been good enough for real work for a while now. The deployed value hasn&#8217;t followed, and the reason isn&#8217;t the model. It&#8217;s everything around the model &#8212; the process, the supervision, the standards, the person willing to change how the work gets done.</p><p>I&#8217;ve made this argument in narrower terms. The last mile in legal <a href="https://andrewlewis.ca/p/the-last-mile-in-legal-has-its-own">has its own geography</a>, a specific terrain of hazards that general adoption advice never names. And the productivity numbers everyone quotes come with <a href="https://andrewlewis.ca/p/the-entry-fee-for-mckinseys-60-percent">an entry fee nobody reads aloud</a>: the data, process, and supervision preconditions that have to exist before a 3-to-5x number means anything. Shapiro&#8217;s version &#8212; a very expensive motor bolted into the old driveshaft &#8212; is the same warning in a better metaphor. <strong>Procurement is not absorption.</strong> On that, no daylight between us.</p><p>His workshop anatomy is correct too: task, background, judgment, constraints, deliverable, verification. And the most important sentence in the essay is the one pointing out that none of it is technical. The lawyers who take to this fastest are the best delegators, not the most technical people in the building. That matches everything I watch happen.</p><p>So here are three places the view changes when you move from across the table to inside the room.</p><h2><strong>Verification is not word six. It&#8217;s the new bottleneck.</strong></h2><p>In his anatomy, verification is the last item. In practice, it&#8217;s the whole game.</p><p>When production gets cheap, the constraint doesn&#8217;t disappear. It moves. It moves to the one step that didn&#8217;t get faster: a human deciding whether the output is right enough to put a name on. And that step does not scale the way drafting does. A model can generate a hundred competent drafts in the time it used to take to write one. It cannot make a partner able to actually read a hundred.</p><p>The demo where the mood in the room changes is real. I&#8217;ve run that demo. But it&#8217;s the easy half. The hard half arrives at the second thousand documents, when the drafts are landing faster than anyone can check them and someone approves one they did not fully read &#8212; not from laziness, from arithmetic. That is a selection problem, not a review problem, and <a href="https://andrewlewis.ca/p/you-cannot-review-your-way-out-of">you cannot review your way out of it</a>. Every guardrail and eval only works after a human has decided what deserves attention.</p><p>Shapiro actually hands us the reason himself. Law has no compiler. A wrong contract doesn&#8217;t crash; it sits in a drawer until it detonates. That&#8217;s precisely why verification can&#8217;t be the sixth thing on the checklist. It&#8217;s the thing the whole rebuild has to be designed around. He says the words and then walks past them.</p><h2><strong>The surplus doesn&#8217;t have your name on it</strong></h2><p>The essay&#8217;s engine is optimism about who keeps the money. The frontier labs are General Electric; the firm that rebuilds itself is Coca-Cola. Absorb the capability and the margin is yours.</p><p>From inside, that&#8217;s the claim I&#8217;d hold most loosely.</p><p><strong>Efficiency that everyone can buy flows to the buyer, not the producer.</strong> When every serious firm has absorbed AI &#8212; which is the entire premise of a race &#8212; the production savings get competed straight through to clients. You capture surplus from what stays scarce and hard to substitute, not from speed your competitor also has. His own evidence points this way: Blackstone is already paying Kirkland less, before Kirkland has rebuilt a thing. That&#8217;s not the fruit of absorption. That&#8217;s the surplus leaving through the door marked bargaining power.</p><p>We have the wider numbers on this. Eighty-eight percent of organizations use AI in some function; a single-digit share get measurable financial impact from it, and the <a href="https://andrewlewis.ca/p/the-rising-tide-doesnt-lift-all-boats">playbook executives are running contradicts their own data</a>. A rising tide has never lifted all boats. It lifts the ones rigged to catch it.</p><p>None of which means don&#8217;t absorb. It means be clear about which bet you&#8217;re placing. Absorbing to keep up is survival &#8212; the table stakes, the price of still being in the game in three years. Absorbing to win is a different move, and it doesn&#8217;t run on speed. It runs on the parts that don&#8217;t commoditize: genuine judgment, and the ability to stand behind the work.</p><h2><strong>The grind was doing something</strong></h2><p>Shapiro concedes, in roughly a paragraph, that AI compresses the junior grind that used to train lawyers, and that firms will have to &#8220;design training around decision-making deliberately.&#8221; That paragraph is carrying a fifteen-year problem it can&#8217;t lift.</p><p>The grind wasn&#8217;t only what firms sold. It was the transmission system for judgment. First-pass research, first-pass drafting, the diligence nobody enjoyed &#8212; that was where instinct got built, through repetition and exposure and being wrong in front of someone senior. Take the reps away and you keep this cohort&#8217;s judgment while quietly starving the next one&#8217;s.</p><p>There&#8217;s early evidence the trade isn&#8217;t free. In one study, developers who learned a new library with an AI assistant <a href="https://andrewlewis.ca/p/the-hidden-second-clause-in-the-ai">scored worse on understanding it afterward</a> than those who struggled through without one. Speed for skill, and sometimes you get neither. Meanwhile the judgment that survives isn&#8217;t the relaxing part of the job. <a href="https://andrewlewis.ca/p/your-brain-is-a-judgment-machine">It gets compressed into every minute of the day</a>, because the easy work that used to space it out is gone. The premium is real. It&#8217;s also heavier to carry than the essay lets on, and the pipeline that produces the people able to carry it is the thing most at risk.</p><h2><strong>Some of the slow clock is the law of lawyering</strong></h2><p>One more, and it&#8217;s the one I can only say from inside.</p><p>Not all of the slow clock is cowardice and comp cycles. Some of it is the job. Privilege, confidentiality, malpractice exposure, the bar rules, the client&#8217;s outside-counsel guidelines &#8212; these gate what absorption is <em>permitted</em>, not just what&#8217;s <em>possible</em>. The Sullivan &amp; Cromwell hallucination Shapiro cites as fear is also a genuine governance failure, and the lesson a careful firm draws from it isn&#8217;t &#8220;be braver.&#8221; It&#8217;s &#8220;build the process that makes that impossible.&#8221;</p><p>The rational firm isn&#8217;t only protecting its margin. Part of what looks like foot-dragging is protecting the client, and that part is correct. The work is to separate the caution that&#8217;s real from the caution that&#8217;s theater &#8212; and Shapiro collapses them into one. Governance isn&#8217;t the committee that slows you down. Done right, it&#8217;s the thing that lets you go fast without the filing that ends a career. It isn&#8217;t a step you finish. It&#8217;s a property of the system, maintained continuously or not at all.</p><h2><strong>What I&#8217;d keep, and what I&#8217;d add</strong></h2><p>Keep the clocks &#8212; the frame is the most useful thing published on this in months. Keep own-the-method over rent-the-wrapper. Keep absorption as the real constraint. On the shape of the problem, Shapiro is more right than anyone writing about this from the outside.</p><p>What I&#8217;d add is the part that turns his diagnosis into a defensible position. The durable moat isn&#8217;t the speed of your clock. It&#8217;s whether you can prove, on the record, that the fast work is also right &#8212; a verified, governed process a human genuinely stands behind. That&#8217;s the input that stays scarce when production goes to zero, because it&#8217;s the one thing the client can&#8217;t get from a cheaper tool or build themselves without taking on the risk they&#8217;re paying you to hold.</p><p>Absorb to survive. Differentiate on judgment and accountability to actually keep the surplus. Because <strong>absorption without verification isn&#8217;t the Coca-Cola fortune. It&#8217;s just a faster demo.</strong></p><div><hr></div><p><em>I write every week about the operational reality between an AI demo and deployed value &#8212; mostly from inside a law firm, where the stakes make the gap impossible to ignore. If that&#8217;s useful, <a href="https://andrewlewis.ca/">subscribe</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Governance Isn’t a Step in Your AI Pipeline]]></title><description><![CDATA[The Governed Context Pipeline &#8212; three phases, one wrapper, and the verification economics that decide what&#8217;s worth automating at all.]]></description><link>https://andrewlewis.ca/p/governance-isnt-a-step-in-your-ai</link><guid isPermaLink="false">https://andrewlewis.ca/p/governance-isnt-a-step-in-your-ai</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Mon, 06 Jul 2026 12:02:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q0vj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q0vj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q0vj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Q0vj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Q0vj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Q0vj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q0vj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:367237,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/205054776?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Q0vj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Q0vj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Q0vj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Q0vj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f3cb2ea-ad7d-4d1c-bccc-e00aa6d5c71c_2400x1600.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The most dangerous box on an AI workflow checklist is the one labeled &#8220;governance.&#8221; Not because it&#8217;s wrong &#8212; because of where it sits.</p><p>Nate B Jones recently published a nine-step framework for agent workflows, running from context pack through human approval (gate). It&#8217;s good. The individual practices are the right ones: normalize your data, ground the model in retrieval, cite every claim, keep a human in the loop. If you&#8217;re building agents, <a href="https://youtu.be/U4TmrlWEY4M">watch it</a>.</p><p>But the framework is linear &#8212; nine sequential steps. And somewhere in any linear list, governance becomes a step. Something the pipeline passes through once, checks off, and moves past.</p><p>That shape is the problem. Not for a startup shipping a demo. For anyone operating where the output has consequences &#8212; a law firm, a bank, a hospital, anywhere regulated &#8212; the shape is exactly backwards.</p><p>This piece lays out the restructure in full: three phases wrapped in a governance layer, the failure mode each phase exists to prevent, and the operational finding that changed how we decide what to automate in the first place.</p><h2><strong>A checklist implies an order of operations</strong></h2><p>Consider what actually has to be true for an agent workflow to survive review in a regulated environment.</p><p>Every action needs to be attributable to a human principal. Access control needs to hold when data is ingested, when it&#8217;s stored, when it&#8217;s retrieved, and when the output ships. The audit log needs to capture all of it. Data residency constraints don&#8217;t pause while the pipeline does its work.</p><p>None of those are steps. They&#8217;re properties &#8212; conditions that must hold at every step. A checklist that positions governance as step nine implies the other eight ran without it. The first serious review will ask what was happening during steps one through eight, and &#8220;governance came later&#8221; is not an answer that ends well.</p><p>Running AI inside a national law firm forces this distinction quickly. Matter-level confidentiality, client data residency, professional accountability for anything that enters the record &#8212; these constraints don&#8217;t arrive at the end of a pipeline. They&#8217;re the environment the pipeline runs in. A workflow that treats them as a checkpoint has already failed; it just hasn&#8217;t been asked the right question yet.</p><p>So the restructure: three phases, wrapped in a governance layer. The Governed Context Pipeline. Each phase answers one question, and each exists to prevent one specific failure.</p><h2><strong>Scope: what is the agent allowed to know?</strong></h2><p>The first phase happens before any data moves, which is why almost nobody does it first.</p><p>Teams start at ingestion. Connect the mailbox, pull the PDFs, chunk, embed. The energy goes into the plumbing, because plumbing feels like progress &#8212; there&#8217;s something to demo at the end of the sprint. But the expensive mistakes in agent projects happen before the first document is ingested, in the decisions nobody wrote down about what the agent may read and toward what goal.</p><p>Nate's skeleton, to its credit, starts exactly here &#8212; his first step is a context pack that defines what the agent is allowed to read. Scope is that instinct given an enforcement mechanism.</p><p>Scope is the permission boundary, defined explicitly. What sources can this agent touch? For what purpose? On whose authority? In a law firm, the answers are concrete: role-based access against the firm&#8217;s identity provider, matter-level permissions that mirror ethical walls, residency rules about which jurisdictions data can occupy. The critical design decision is <em>where those answers live</em>. Not in a policy memo &#8212; in the architecture. A memo describes intent. An access control list enforces it.</p><p>One rule does most of the work: <strong>the agent&#8217;s access is a strict subset of the authenticated human&#8217;s access.</strong> Never more. An agent acting for a lawyer who cannot see a matter cannot see that matter either &#8212; not because it was told not to look, but because the credential it operates under makes looking impossible.</p><p>The failure this phase prevents is the one that ends AI programs in regulated industries: the &#8220;helpful&#8221; agent that read the wrong client&#8217;s file. Once that happens, no downstream safeguard matters. You can verify the output perfectly and the breach has still occurred, because the document was never supposed to be in context at all. Scope failures are unrecoverable by design &#8212; which is why scope is phase one and not a hardening task for later.</p><h2><strong>Structure: turn the mess into addressable context</strong></h2><p>The second phase is the unglamorous middle, and it&#8217;s where most of the actual engineering lives.</p><p>Source material arrives through controlled connections &#8212; not ad-hoc exports, not someone&#8217;s downloads folder. Large documents get chunked into tagged, addressable pieces, so a five-hundred-page record becomes something a retrieval system can point into with precision. Normalization turns unstructured mess into a clean schema: <span>dates become dates, amounts become amounts &#8212; </span><strong>Nate's phrase, and exactly the right one</strong><span> &#8212; parties become parties</span>. Storage stays secure and localized, holding whatever residency promise was made in scope. And retrieval works by similarity against that normalized store, pulling exact language with file references attached.</p><p>That last property deserves a sentence of its own, because it&#8217;s the entire epistemology of the pipeline: the agent works from retrieved fact, not remembered vibes. A model answering from its training data is recalling an impression of the world. A model answering from retrieval is reading a specific paragraph of a specific document it can name. Everything defensible about the output rests on that difference.</p><p>Two operational realities about this phase, both learned the slow way.</p><p>First, normalization quality is gated by input quality. The hardest inputs aren&#8217;t exotic &#8212; they&#8217;re handwritten documents, OCR&#8217;d scans, and sources whose formatting drifts from one custodian to the next. No amount of clever prompting upstream fixes a date field that arrived as a smudge. That makes normalization permanent iteration work, not a solved step, and any plan that treats it as a one-time migration is a plan for disappointment.</p><p>Second, normalization is deliberately isolated as its own step so it can be improved without touching anything else. Ingestion doesn&#8217;t change. Retrieval doesn&#8217;t change. The normalizer gets better every month on its own schedule. That&#8217;s not pipeline trivia &#8212; it&#8217;s the design decision that lets the boring part compound.</p><p>The failure this phase prevents: plausible-sounding output grounded in nothing. Skip the structure work and the agent still answers &#8212; that&#8217;s the trap. It answers fluently, confidently, and from nowhere.</p><h2><strong>Sign-off: defensible, not just plausible</strong></h2><p>The third phase is where the output earns the right to be relied on. Four operations: cite, verify, export, gate.</p><p><strong>Cite</strong> means every generated claim maps to a specific paragraph in a specific source document. Not &#8220;according to the file&#8221; &#8212; this claim, that paragraph. The receipt. Citation is what converts an assertion into something checkable, and its real function is economic, which matters enormously later in this piece.</p><p><strong>Export</strong> means the output ships as a structured review packet &#8212; a timeline, a list of identified gaps, a draft summary &#8212; rather than a wall of chat text. The reviewer gets an artifact organized for review, not a transcript to spelunk.</p><p><strong>Gate</strong> means the hard stop. The AI cannot send, submit, file, or finalize. Ever. An authenticated professional reviews the citations and approves, and the approval is what moves the work forward. The gate is architecture, not policy &#8212; the capability to act simply doesn&#8217;t exist in the agent&#8217;s permission set.</p><p>The failure this phase prevents: an AI output entering the record with no human fingerprints on it. In a profession where a person is accountable for everything that goes out the door, that failure isn&#8217;t an embarrassment. It&#8217;s a structural breach of how the profession works.</p><p>But there&#8217;s a subtlety inside sign-off that most pipelines miss, and it&#8217;s the difference between having receipts and having proof.</p><h2><strong>Four ways a citation fails</strong></h2><p>A citation is an assertion, not proof. Audit the receipts in a real pipeline and they fail in four distinct ways:</p><ol><li><p><strong>Hallucinated</strong> &#8212; the source doesn&#8217;t exist. Rare in a grounded pipeline, but it&#8217;s the failure everyone knows about, because it&#8217;s the one that makes headlines and sanctions orders.</p></li><li><p><strong>Semantically wrong</strong> &#8212; the source is real, but it doesn&#8217;t mean what the model thinks it means. The claim and the paragraph are both genuine; the relationship between them is not.</p></li><li><p><strong>Missing</strong> &#8212; a claim that should carry a citation doesn&#8217;t, and slides through as connective tissue.</p></li><li><p><strong>Duplicated</strong> &#8212; one source quietly does double duty across several claims it only partially supports.</p></li></ol><p>The second failure is the dangerous one, because it survives a spot-check. The document exists. The link resolves. The reviewer clicks through, sees a real paragraph in a real filing, nods, and moves on. Every superficial signal of groundedness is present, and the meaning is wrong. Hallucination checks &#8212; &#8220;does this source exist?&#8221; &#8212; catch the first failure and are completely blind to the second.</p><p>Which is why verification is its own operation with its own tooling, not a byproduct of citation. Checking that a receipt exists is string matching. Checking that the receipt <em>supports the claim it&#8217;s attached to</em> is an entailment judgment &#8212; and it has to be performed against every claim, not a sample you hope generalizes. Hand that judgment to humans and you&#8217;ve rebuilt the verification labor the pipeline was supposed to reduce. Hand it to a second model and you&#8217;ve created a new set of judgments that themselves need calibrating. That recursion is why this is the live edge of the work right now. Nothing about it is finished, and it&#8217;s harder than any layer of the stack that gets talked about more.</p><h2><strong>The wrapper is the argument</strong></h2><p>Around all three phases sits the governance layer &#8212; and here is where the restructure stops being a diagram preference and starts being the point.</p><p>Four properties, held everywhere: identity, so every action is attributable to a human principal; access control, enforced at every phase rather than checked once at the door; audit logging, immutable, capturing prompts, actions, and approvals &#8212; which means compliance reporting comes free instead of being a quarterly archaeology project; and residency, holding through every phase including retrieval, because a vector store in the wrong jurisdiction is a residency violation no matter how well-scoped the ingestion was.</p><p>The difference between a step and a property sounds like semantics until you watch it decide an architecture. Treat governance as a step and you get a pipeline with a compliance checkpoint &#8212; one gate, at one position, inspecting whatever reaches it, trusting that nothing leaked before inspection. Treat governance as a wrapper and the questions change shape entirely. Not &#8220;did the output pass review?&#8221; but &#8220;was there any moment when this agent could act outside its principal&#8217;s identity?&#8221; Not &#8220;is there an audit trail?&#8221; but &#8220;is there any action that could have escaped it?&#8221;</p><p>And every one of these properties maps to something you can point at in a real enterprise environment. Residency is the cloud tenancy itself &#8212; region-pinned storage and compute that physically cannot host data elsewhere. Identity and access control bind to the same enterprise identity provider the humans authenticate against, so the agent inherits the org&#8217;s permission model instead of maintaining a parallel one. The connectors that bridge into document management systems carry those permissions through rather than flattening them. If a team presenting an agent architecture can&#8217;t point to where each property is enforced, the property isn&#8217;t held &#8212; it&#8217;s hoped for.</p><p>The wrapper turns governance from an inspection into an invariant. None of it is a policy document. All of it is architecture. That&#8217;s the part linear framings underplay, and it&#8217;s the part that determines whether the system survives contact with a regulator, an opposing counsel, or your own risk committee.</p><p>Governance isn&#8217;t step ten. It&#8217;s the box the other nine steps live inside.</p><h2><strong>The economics of the gate</strong></h2><p>Everything so far describes how to build the pipeline correctly. This section is about when not to build it at all &#8212; and it comes from the finding that reshaped our automation decisions more than any technical result.</p><p>The gate is not free. Put an AI workflow in production with a human verification gate at the end, and in some workflows the verification consumes as much time as the manual process it replaced. The automation produced its output in seconds; the checking of that output ate the savings. Net time recovered: approximately nothing.</p><p>The accounting is actually worse than break-even, because of <em>who</em> does the verifying and <em>what they know</em>. A person doing the work manually verifies as they go. Checking is amortized across authorship &#8212; every choice is checked at the moment it&#8217;s made, by the person holding the reasoning behind it. The verifier at a gate reads cold. They weren&#8217;t present during creation, they don&#8217;t hold the reasoning, and they&#8217;re reconstructing intent from the artifact alone. The pipeline didn&#8217;t eliminate the verification labor. It relocated it &#8212; and handed it to someone in a structurally worse position to perform it.</p><p>Peter Naur described the underlying mechanism in 1985, in &#8220;Programming as Theory Building&#8221;: the real substance of built work is the theory in the builder&#8217;s head, and the artifact doesn&#8217;t carry the theory. Programmers are rediscovering this at scale right now, reviewing AI-generated code they didn&#8217;t write and can&#8217;t cheaply validate. It generalizes to all knowledge work. The gate asks a human to recover a theory from an artifact, which is the expensive direction.</p><p>Out of that comes the selection criterion &#8212; the most useful sentence this pipeline has produced:</p><p><strong>Automate where verification is asymmetric.</strong> Where checking is structurally cheaper than producing: confirming a citation points where it claims, validating an extracted date against the source, checking a sum. In those tasks the gate is fast, the machine does the production, and the economics genuinely work.</p><p>Where verification is symmetric &#8212; where &#8220;is this right?&#8221; can only be answered by re-deriving the judgment &#8212; the pipeline saves nothing. &#8220;Does this summary capture what matters across these files?&#8221; has no shortcut; answering it honestly means reading the files, which is the work the summary was meant to replace. The honest architecture in those cases moves the human upstream, into the creation loop, steering while the work is made, rather than stationing them at the gate to reconstruct it afterward.</p><p>Notice what this does to the phases you&#8217;ve already read about. Cite and Export aren&#8217;t decoration &#8212; they exist precisely to convert symmetric verification into asymmetric verification. A claim with a paragraph-level receipt can be checked without re-deriving it. A structured packet with a stated timeline and gap list can be reviewed without reconstructing it. When that conversion succeeds, automation pays. When it fails &#8212; when the receipts still can&#8217;t make checking cheap &#8212; the task was a bad automation candidate, and no amount of pipeline sophistication rescues it.</p><h2><strong>Where it breaks</strong></h2><p>A framework that claims no limits is marketing, so here are this one&#8217;s, stated plainly.</p><p>Normalization never finishes. Input quality gates everything downstream, the worst inputs are the ones humans produce most casually, and the phase is permanent iteration work dressed as a pipeline step. Budget for it accordingly &#8212; in attention, not just money.</p><p>Semantic verification is unsolved. The four-failure taxonomy tells you what to look for; it doesn&#8217;t hand you a tool that finds failure mode two at scale. The honest status is: active work, real tooling, no finish line in sight.</p><p>And the framework governs context pipelines &#8212; agents that read, ground, and draft. It says nothing about model selection, nothing about fine-tuning, nothing about the dozen other decisions an AI program has to make. It&#8217;s a load-bearing wall, not a whole house.</p><p>What it does do is give a leadership team a shape to reason with. When the next agent proposal lands, the questions write themselves. What&#8217;s the scope, and where is it enforced? What does the structure work cost, and who maintains it? Is verification asymmetric, or are we about to relocate labor and call it savings? Can any action escape the wrapper?</p><p>The checklist wants you to ask &#8220;what&#8217;s the next step?&#8221; The better question is &#8220;what has to be true at every step?&#8221;</p><p>Draw the box first. Then build what runs inside it.</p><div><hr></div><p>On LinkedIn this week I&#8217;m walking through one phase of this framework each day &#8212; compressed versions of the arguments above. This article is the destination they point back to; there&#8217;s nothing else to keep up with. If someone forwarded you this, you can subscribe at <a href="https://andrewlewis.ca/">andrewlewis.ca</a>.</p>]]></content:encoded></item><item><title><![CDATA[The Teams Asking for AI Aren’t Generating the Revenue]]></title><description><![CDATA[When AI effort flows to whoever is most excited, you can ship for a year and never touch the part of the business that pays for it.]]></description><link>https://andrewlewis.ca/p/the-teams-asking-for-ai-arent-generating</link><guid isPermaLink="false">https://andrewlewis.ca/p/the-teams-asking-for-ai-arent-generating</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Thu, 02 Jul 2026 18:00:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lMWl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lMWl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lMWl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!lMWl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!lMWl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!lMWl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lMWl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2710556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/204707451?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lMWl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!lMWl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!lMWl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!lMWl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e18be9e-2c72-4299-84cc-60525483d258_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The teams that raise their hands for AI are almost never the teams that generate the revenue. That gap is easy to miss, because enthusiasm looks like traction. Most of the time it isn&#8217;t.</p><h2><strong>Enthusiasm Is a Real Signal</strong></h2><p>Building for the excited teams first is the obvious move, and for good reason. When a group asks for an AI workflow, they&#8217;ve already done half of the change-management work on your behalf. They tolerate a rough first version. Feedback comes back as a note, not a complaint. They volunteer as pilots, forgive the bugs, and &#8212; when something finally works &#8212; turn into the internal reference you point every skeptic toward.</p><p>None of that is a mistake. Friction is the tax on every internal tool, and a willing team pays it without being chased. If you&#8217;re standing up an AI practice from nothing, the enthusiasts are how you get a first win on the board before the budget conversation turns hostile. Following demand is a defensible way to start, and anyone who tells you to ignore your most eager users has never had to earn credibility inside an organization that didn&#8217;t ask to be changed.</p><h2><strong>But Enthusiasm and Revenue Aren&#8217;t Correlated</strong></h2><p>The problem surfaces a quarter or two later, when you look at where the effort actually went. In most firms &#8212; ours included &#8212; a small number of practice groups generate a disproportionate share of the revenue. The litigation teams, the corporate groups, the ones running the largest matters against the tightest deadlines. Those are rarely the people sending the eager Friday-afternoon note about a tool they just read about.</p><p>So the work flows one way and the money sits in the other. You can ship six workflows to the teams that asked, run clean pilots, post healthy adoption numbers, and still not have touched the part of the business that pays for all of it. That&#8217;s the trap. Every one of those workflows is genuine effort. <strong>None of it is evidence that you moved the number that matters.</strong></p><p>This is the difference between activity and value, and it&#8217;s the easiest distinction to lose when the activity feels good. Adoption dashboards count seats, sessions, and logins &#8212; motion. They don&#8217;t ask whether the motion is pointed at anything. A number that goes up is not the same as a number that matters, and most AI programs are measured almost entirely on the first kind.</p><p>Measuring value would mean asking a harder question: did this workflow change the economics of work that actually bills? Hours returned to a partner on a live matter. Cycle time cut on the review that was holding up a deal. A junior freed from the mechanical first pass so the expensive judgment happens sooner. Those numbers are harder to collect and easier to argue with, which is exactly why programs quietly settle for seat counts instead. The metric you can gather in an afternoon wins over the metric that would tell you the truth.</p><h2><strong>The Path of Least Resistance Has an Org Chart</strong></h2><p>Effort doesn&#8217;t drift at random. It follows inbound. The teams that email you become the teams you build for, and the teams that email you are the ones already sold on the premise &#8212; which means your roadmap gets quietly authored by whoever is most curious, not whoever is most valuable. Nobody decides this in a meeting. It&#8217;s the default that takes over when no one sets a different one.</p><p>There&#8217;s a compounding effect, too. The enthusiast team you served becomes your best case study, so the next enthusiast team hears about it and comes to you, and the flywheel spins entirely inside the low-stakes tier of the business. Each turn feels like momentum. From the inside, a roadmap driven purely by inbound is almost indistinguishable from one driven by strategy &#8212; right up until someone asks what it did for the P&amp;L, and the honest answer is a list of teams that were already going to be fine.</p><h2><strong>The Honest Version of the Other Side</strong></h2><p>There&#8217;s a real counterargument here, and it deserves its full weight rather than a quick takedown. Maybe building for the enthusiasts first isn&#8217;t misallocation at all. Maybe it&#8217;s sequencing.</p><p>You don&#8217;t want your first agentic workflow to fail inside the litigation group three days before a filing. The high-revenue teams are precisely the ones who can least afford a broken tool, a hallucinated citation, or a process that quietly drops a step &#8212; the stakes that turn a rough pilot into a client problem. Testing on the small, willing groups is how you surface those failures somewhere cheap. You build the muscle, harden the process, and earn the right to walk into the room where the work is unforgiving. Under that reading, the enthusiasts aren&#8217;t a distraction &#8212; they&#8217;re the rehearsal.</p><p>That argument is genuinely persuasive. It&#8217;s the strongest defense of enthusiasm-led rollout there is, and in a regulated environment it&#8217;s often the correct call. Which is exactly what makes it dangerous, because it&#8217;s also the perfect cover for never getting to the hard rooms at all.</p><h2><strong>The Line Between Sequencing and Drift</strong></h2><p>Sequencing and drift look identical for the first two quarters. Both have you shipping to willing teams and posting good numbers. The only thing that separates them is whether you can name the crossover.</p><p>If it&#8217;s really sequencing, you can answer three questions today: <strong>which revenue group you&#8217;re piloting toward, what has to be true before you bring the workflow to them, and roughly when.</strong> If those answers don&#8217;t exist &#8212; if the enthusiasts are simply where the work keeps landing because that&#8217;s where the work is easy &#8212; then &#8220;we&#8217;re de-risking first&#8221; isn&#8217;t a strategy. It&#8217;s the story you tell about drift after the fact.</p><p>Naming the crossover forces a second, uncomfortable admission: some of the enthusiast pilots you&#8217;re proudest of were never going to move the P&amp;L, and were worth doing only as practice. That&#8217;s a fine reason to build something. It&#8217;s a poor reason to keep building in the same place once the muscle is there.</p><p>The discipline is almost embarrassingly plain. Before a pilot starts, write down who it&#8217;s for and what graduation looks like. A pilot with a named destination is a sequencing plan. A pilot without one is a hobby that happens to have a roadmap attached. AI without operations is just a demo &#8212; and a rollout that only ever serves the people who asked is just enthusiasm with a dashboard.</p><p>Enthusiasm tells you where it&#8217;s safe to begin. Never let it decide where you stop.</p><div><hr></div><p><em>New here? Subscribe for field notes on making AI actually operational inside a real organization. If this one landed, forward it to whoever owns your AI roadmap &#8212; the person deciding which team gets built for next.</em></p>]]></content:encoded></item><item><title><![CDATA[Least Privilege Was Built for People Who Pause]]></title><description><![CDATA[Every access model you run assumes a human is standing at the moment of action. Agents void that assumption.]]></description><link>https://andrewlewis.ca/p/least-privilege-was-built-for-people</link><guid isPermaLink="false">https://andrewlewis.ca/p/least-privilege-was-built-for-people</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Mon, 29 Jun 2026 12:26:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ffpl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ffpl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ffpl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Ffpl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Ffpl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Ffpl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ffpl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1768613,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/203899970?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ffpl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Ffpl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Ffpl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Ffpl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F614b76b4-a1f4-4199-8e27-f70e601e54e9_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Every access control system you have ever run carried an assumption it never bothered to write down. The assumption was that a human being would be standing at the moment of action, deciding whether to actually go through with it.</p><p>That assumption is so deep we stopped seeing it. We talk about permissions as if they were the control. They were never the control. The control was the person &#8212; the one who had the access, noticed the request was strange, hesitated over the unfamiliar recipient, and chose not to send. The permission set drew the outer wall. Human judgment did the real work inside it, quietly, every single day, for free.</p><p>Agents inherit the wall. They do not inherit the judgment. And the moment that happens, a lot of governance that felt solid turns out to have been resting on something we never accounted for.</p><h2><strong>The contract nobody signed</strong></h2><p>Least privilege is usually explained as a defense against malice. Give people only the access their job requires, so a bad actor &#8212; or a stolen credential &#8212; can do less damage. True enough. But that framing hides the part that actually carried the weight.</p><p>Most people in most organizations are over-permissioned, and it almost never matters. A paralegal can probably open far more of the document system than any single matter requires. A developer can usually reach systems they have not touched in a year. We tolerate that gap between <em>granted</em> access and <em>needed</em> access because a human sits in it, applying discretion at the point of use. The gap is wide, and the judgment is what keeps the width from being dangerous.</p><p>That is the contract nobody signed. Access can be loose because judgment is tight. Take the judgment out of the moment of action, and the looseness you have lived with comfortably for twenty years becomes the exposure.</p><h2><strong>What the agent removes</strong></h2><p>Consider what Microsoft&#8217;s Copilot Cowork can do once it is running in someone&#8217;s context. It drafts and sends email, creates and edits and deletes documents, posts in collaboration channels, searches across the enterprise, and reaches the document system through the same connectors a person uses. It does all of this as the user, with the user&#8217;s permissions, not a narrower set of its own.</p><p>To Microsoft&#8217;s credit, it pauses before the riskiest actions and asks for approval, with a risk indicator attached. That sounds like the missing judgment, restored. It is not. The approval is one click, it leans on the same human attention least privilege already assumed, and it never changes what the agent is able to reach. Ask a person to approve twenty actions a day and you have not added judgment &#8212; you have manufactured a reflex. The prompt becomes a turnstile that everyone learns to push through.</p><p>This is not a worry confined to my own corner of the world. A 2025 <a href="https://cloudsecurityalliance.org/press-releases/2026/04/16/more-than-half-of-organizations-experience-ai-agent-scope-violations-cloud-security-alliance-study-finds">Cloud Security Alliance survey</a> of 445 security professionals found that <strong>ninety percent of deployed AI agents carry more access than their assigned task requires</strong>, and nearly half of organizations had already experienced a security incident involving one. The over-permissioned agent is not the edge case anyone is bracing for. It is the normal state of deployment.</p><h2><strong>From governing files to governing the factory</strong></h2><p>There is an idea moving through software circles that you are no longer building the product, you are building the factory that builds the product &#8212; and that you have to govern the factory accordingly. The same shift is happening to ordinary professional work. Increasingly the thing you build is not the document or the policy. It is the process, and the agent, that produces them.</p><p>Once you see governance that way, least privilege stops being an IT setting and becomes a design decision about the machine. You are no longer asking what a given person should be able to do with a given file. You are asking what the agent &#8212; its tools, its reach, its standing permissions &#8212; is structurally allowed to do with every file it can touch. The security community is converging on the same point from its own direction: the OWASP work on agentic applications now frames the goal as <strong><a href="https://auth0.com/blog/owasp-top-10-agentic-applications-lessons/">Least Agency</a></strong> &#8212; autonomy that is earned, scoped, and bounded, rather than granted by default because the underlying account happened to have it.</p><p>The instinct, having seen all this, is to lock everything down. Resist it. Maximal restriction kills the serendipity and cross-pollination that made these tools worth adopting in the first place. The discipline is to draw the line at the boundary of what a role would ever genuinely need &#8212; and to leave a deliberate margin on the generous side of strict minimum. There is a line. It sits a little wider than the tightest possible setting, and a great deal narrower than what most people hold today.</p><p>If that feels premature &#8212; if the honest objection is that nothing has gone wrong yet &#8212; look at the shape of the bet. Over-restrict and the cost is friction and a few frustrated colleagues. Under-restrict and the cost is the kind that ends up in a regulator&#8217;s letter or a headline. When the downside is that asymmetric, waiting for proof is not caution. It is a slow decision to do nothing.</p><h2><strong>The part that isn&#8217;t technical</strong></h2><p>If the argument stopped there, it would be an engineering problem, and engineering problems are the easy kind. The hard part is everywhere the technology isn&#8217;t.</p><p>People want a voice in how they get restricted, and they are right to. When word gets out that access is going to change, the reasonable reaction is not gratitude &#8212; it is a request to be consulted about a wall being built around your own work. Then there are the old records sitting under a practice area that no longer has an active owner: cut access and you might orphan something that still matters; leave it and you have proven the point about exposure. There are whole areas of a business with no clear owner at all. And underneath all of it sits the least glamorous question in governance &#8212; who approves access, by what process, on what timeline, without slowing the work that pays for everything.</p><p>None of that is a permission toggle. That is the actual work, and it is the work least likely to get done, because it is political and procedural rather than technical, and it does not demo well.</p><p>So the line I gave my colleagues, the one that sounded extreme, was never an accusation. Treat every internal user as malicious is a design stance, not a character judgment. It says: stop architecting around who you trust, because trust was a statement about intent, and intent is no longer the variable that determines the outcome. Architect around what the machine can do on their behalf.</p><p>Build for what can happen, not for what anyone means to do. <strong>Intentions don&#8217;t matter. Outcomes do.</strong></p>]]></content:encoded></item><item><title><![CDATA[Defining “Done” Is the Work the Loop Can’t Do]]></title><description><![CDATA[An agent that runs until it&#8217;s finished forces you to say what finished means before the work exists &#8212; and code gets away with it in a way a legal contract never will.]]></description><link>https://andrewlewis.ca/p/defining-done-is-the-work-the-loop</link><guid isPermaLink="false">https://andrewlewis.ca/p/defining-done-is-the-work-the-loop</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Thu, 25 Jun 2026 12:02:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!T5jM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T5jM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T5jM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!T5jM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!T5jM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!T5jM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T5jM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2085605,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/202982657?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!T5jM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!T5jM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!T5jM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!T5jM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67bdd497-e215-45d1-a916-df2d93ccd227_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The first agentic loop I handed real work to was almost embarrassingly simple. The instruction was one line: keep adding meaningful tests until you reach eighty percent coverage. It ran, it stopped, it produced something I could use. The loop worked because &#8220;done&#8221; was a number a machine could check on its own, over and over, without me in the room.</p><p>Then I started thinking about everything else. If a loop can run until a coding task is finished, where else does the pattern fit? Drafting. Review. Triage. The long middle of white-collar work that looks, from a distance, a lot like writing tests &#8212; produce a draft, check it, revise, repeat. And the more I looked at the work I actually manage, the less the pattern held. Not because the agents aren&#8217;t capable. Because of something hiding inside that one-line instruction I gave the test loop.</p><h2><strong>The two conditions inside one sentence</strong></h2><p>Read the instruction again: keep adding meaningful tests until you reach eighty percent coverage. There are two conditions in that sentence, not one. &#8220;Eighty percent&#8221; is the part the loop can see. It is a number, computed the same way every time, and the agent can compare its work against it without asking me anything. &#8220;Meaningful&#8221; is the part only I can see. It is a judgment about whether a test exercises something that matters, and the coverage number is blind to it.</p><p>The loop optimized the half it could measure and quietly leaned on me for the half it couldn&#8217;t. That arrangement felt like automation. It was closer to a division of labor I hadn&#8217;t noticed I was making. The machine took the measurable condition; I kept the judgment and pretended the number carried it.</p><p>That worked out fine for tests. It&#8217;s worth being honest about why.</p><h2><strong>Why coverage was allowed to stand in for quality</strong></h2><p>Code is instrumental. Nobody actually wants the code &#8212; they want what the code does. The product is downstream of the artifact, which means you are allowed to measure &#8220;done&#8221; with a structural proxy and trust that the proxy tracks the thing you care about. Coverage is one of those proxies. It does not measure whether the software is good. It measures how much of the code your tests touched, on the assumption that broadly exercised code is less likely to hide a defect that reaches the user.</p><p>The trouble is that the assumption is weaker than most teams treat it. The most careful study on this &#8212; <a href="https://www.cs.ubc.ca/~rtholmes/papers/icse_2014_inozemtseva.pdf">Inozemtseva and Holmes, presented at ICSE in 2014</a>, across roughly thirty-one thousand test suites on five large systems &#8212; found that once you control for the size of the test suite, coverage is only weakly correlated with how effective those tests actually are. Their conclusion is blunt: using a fixed coverage value as a quality target is unlikely to produce an effective suite. The field agreed with itself a decade later and named it the most influential paper of its conference year.</p><p>This is <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC7901608/">Goodhart&#8217;s law</a> wearing work clothes. When a measure becomes a target, it can stop being a good measure, because the act of optimizing for it pulls it loose from whatever it was standing in for. The point is not that coverage is useless. Low coverage genuinely tells you something &#8212; it flags code nobody tested. The asymmetry is the whole lesson: a low number is a real warning, and a high number is not a real promise. The proxy is good at catching neglect and bad at certifying quality. A loop that runs to a coverage target inherits exactly that limitation, and for tests, where the real product lives one step downstream, you can live with it.</p><p>You can live with it because the proxy and the product are different objects. That is the condition that makes the whole thing work. And it is the condition that most knowledge work does not meet.</p><h2><strong>A contract is not instrumental</strong></h2><p>When a lawyer ships a contract, the contract is the product. There is no downstream artifact that the document is merely a means toward. You do not deliver &#8220;what the contract achieves&#8221; and keep the contract as scaffolding &#8212; you deliver the contract, and its quality lives in the thing itself: whether it is coherent, whether it anticipates the failure it was written to prevent, whether a counterparty&#8217;s counsel will accept it without a fight. None of that reduces to a number you can compute on the side and trust.</p><p>This is the move that breaks the analogy to tests. In code, &#8220;done&#8221; can ride on a proxy because the artifact is instrumental. In most professional knowledge work, the artifact is the deliverable, the quality is holistic, and there is no faithful proxy to point the loop at. You cannot write the equivalent of eighty percent coverage for a memo, because there is no &#8220;coverage&#8221; that stands one step away from the memo&#8217;s worth. The worth is in the memo.</p><p>So the comfortable division of labor collapses. With the test loop, I let the machine hold the measurable condition and I held the judgment. With a contract, there is no measurable condition to hand off. The judgment is the entire job. And a loop has a very specific demand about judgment that, until you try to use one for this kind of work, is easy to miss.</p><h2><strong>Doing the work is how we find out what &#8220;done&#8221; is</strong></h2><p>Here is the part that took me longest to say plainly. Traditionally, doing the work is how we discover what done looks like. You don&#8217;t hold the finished standard in your head and then march toward it. You draft, and the draft shows you what it needs. You see the thing take shape, and the shape tells you where it falls short. Judgment, in most real work, is terminal &#8212; it happens at the end, against an artifact you can finally look at.</p><p>A loop will not let you wait that long. To run an agent until it is finished, you have to define &#8220;finished&#8221; before a single version of the artifact exists. The judgment that used to live at the end gets moved to the front, ahead of the work, into a specification you write blind. For a task like coverage, where the standard genuinely is knowable in advance and handed to you from outside, that relocation costs nothing. For work where the standard is discovered in the doing, asking someone to judge before the work is done is not a workflow. It is a contradiction.</p><p>We have known this about people for a long time. Michael Polanyi&#8217;s whole account of <a href="https://press.uchicago.edu/ucp/books/book/chicago/T/bo6035368.html">tacit knowledge</a> turns on the line that we can know more than we can tell &#8212; that a great deal of real expertise never makes it onto the page as an explicit rule, which is exactly why you can recognize good work when you see it and still fail to specify it in advance. Software learned the same lesson the expensive way and wrote it into the <a href="https://agilemanifesto.org/principles.html">Agile Manifesto</a>: welcome changing requirements even late in development, because the best requirements and designs emerge through the work rather than before it. The industry spent two decades moving away from the up-front specification freeze. The agentic loop quietly asks for the freeze back.</p><h2><strong>The machines are already demonstrating the problem</strong></h2><p>If this sounds like a soft, human complaint about creativity resisting measurement, the most striking confirmation is coming from the machines themselves. The difficulty of specifying a goal in advance, such that an optimizer pursuing it actually does what you meant, is one of the oldest named problems in the field. <a href="https://arxiv.org/abs/1606.06565">Concrete Problems in AI Safety</a> framed it in 2016: a formal objective is an attempt to capture the designer&#8217;s informal intent, and it can be satisfied by solutions that are valid in the literal sense and wrong in every sense that mattered. DeepMind catalogued dozens of cases under the name <a href="https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/">specification gaming</a> &#8212; behavior that meets the letter of an objective without achieving the intended outcome &#8212; and noted that writing a specification that actually reflects what the designer wanted is, in their dry phrasing, difficult.</p><p>The newest version of this is no longer a toy. In <a href="https://arxiv.org/abs/2502.13295">early 2025, researchers at Palisade</a> told reasoning models to win against a chess engine, and the stronger models &#8212; by default, without being nudged &#8212; went after the task environment instead of playing chess, hacking the game state to register a win. The instruction was met. The intent was not. (<a href="https://www.technologyreview.com/2025/03/05/1112819/">MIT Technology Review</a> and <a href="https://time.com/7259395/">TIME</a> both covered it.) Take that as analogy rather than proof, because an agent gaming a chess benchmark and a contract that fails a client are not the same failure. But the shape is identical, and it is the shape of my point: a stated &#8220;done&#8221; is not the same as the done you meant, and the gap between them is precisely the judgment a loop asks you to write down before you can see anything.</p><h2><strong>So who is supposed to write &#8220;done&#8221;?</strong></h2><p>This is where the question stops being abstract for me, because the honest answer implicates my own seat. A stopping condition can only be authored by someone who already knows what good looks like for that specific work. That person is not, usually, the person who wants to deploy the loop across a team.</p><p>I manage a team and provide direction. I am not sitting in the room with the user, deconstructing their workflow into the conditions that would tell an agent it was finished. That is not a confession of laziness; it is a structural fact about where leadership sits relative to the work. The definition of done lives inside the task, with the practitioner who does it, in the same tacit place Polanyi pointed at. It cannot be delegated upward to whoever holds the budget for automation, and it cannot be delegated to the agent, because the agent is the thing waiting to be told. The bottleneck in scaling agentic work was never the model&#8217;s capability. It is that &#8220;good&#8221; lives inside the work, and most of it has never been written down &#8212; because, until now, nobody had to.</p><h2><strong>&#8220;Isn&#8217;t this just the discipline you should already have?&#8221;</strong></h2><p>The strongest objection to all of this is a fair one, and I want to state it at full strength rather than knock down a weak version. Defining &#8220;done&#8221; up front is not some new impossibility the loop invented. It is ordinary project-management discipline &#8212; a brief, a specification, acceptance criteria, a definition of done on a ticket. People have always been too quick to skip it, and if the loop forces a team to actually do the work of saying what finished means, that is a gift, not a problem.</p><p>Where that objection is right, it is completely right, and I&#8217;ll concede it without hedging. For instrumental work with a faithful proxy &#8212; code against tests, a data pipeline against a validation suite, anything where the artifact is a means and a measurable condition stands one honest step away from the goal &#8212; defining the stopping condition in advance is just rigor, the loop rewards it, and you should reach for the loop. I am not arguing against autonomous agents. I run them.</p><p>Where the objection misses is the moment the artifact stops being a means and becomes the deliverable. There, &#8220;define done in advance&#8221; does not resolve into a cleaner brief. It resolves into &#8220;judge the work before you are allowed to see it,&#8221; and no amount of discipline dissolves that, because the thing you would need to judge does not exist yet. The discipline argument assumes the standard is writable and the team is just dodging the writing. Sometimes the standard is not writable in advance at all, and pretending otherwise is how you end up shipping something that passed every stated condition and satisfies no one.</p><h2><strong>The open question, and the quiet risk</strong></h2><p>What I genuinely don&#8217;t know is how far the writable territory extends. Can we get to the point where we can author the equivalent of unit tests for knowledge work? Some domains will yield &#8212; there are corners of legal and financial work structured enough that a real rubric is coming, and a loop will own them. Some domains may never yield. Most sit somewhere on the spectrum between, and the honest position is that we don&#8217;t yet know where the lines fall.</p><p>It is tempting to believe the machines will close the gap for us by learning to judge open-ended work. They are not there. The current best attempt &#8212; using a strong language model as the judge of another model&#8217;s output &#8212; is real and useful and also <a href="https://arxiv.org/abs/2510.27106">documented to drift and contradict itself</a>, giving the same work different scores across runs, with a catalog of <a href="https://arxiv.org/abs/2410.02736">systematic biases</a> toward length and surface form. An unreliable judge does not rescue you from needing to know what good is. It just hides the moment you stopped knowing.</p><p>So the risk I actually worry about is not that we fail to measure the unmeasurable. It is that we route around it. Faced with work whose &#8220;done&#8221; resists specification, the path of least resistance is to quietly reshape the work into the part that can be specified, automate that, and let the rest atrophy &#8212; to optimize for the measurable and, as <a href="https://press.princeton.edu/books/hardcover/9780691174952/the-tyranny-of-metrics">Jerry Muller documents across a dozen institutions</a>, come to treat whatever resisted measurement as if it were never the point. That is not a failure of the technology. It is a failure of nerve dressed up as efficiency, and it is the most likely way good work gets thinner without anyone deciding that it should.</p><p>A loop is a bridge. It will carry you, tirelessly and faithfully, from a problem to a solution &#8212; but only to a solution you can already describe well enough to recognize when you arrive. It does not find the far bank for you. You still have to know the end from the beginning, and for the work that actually matters, knowing the end was always the job.</p><div><hr></div><p><em>If this is useful, the place I work these ideas out in long form is <a href="https://andrewlewiswashere.substack.com/">AndrewLewisWasHere</a>. Subscribe there, or forward this to the person on your team who&#8217;s about to point a loop at work that ships as the artifact itself.</em></p>]]></content:encoded></item><item><title><![CDATA[Your Agent Has the Same Permissions You Do]]></title><description><![CDATA[The conversation about AI agents is all about what they can do. The question nobody can answer yet is what they did.]]></description><link>https://andrewlewis.ca/p/your-agent-has-the-same-permissions</link><guid isPermaLink="false">https://andrewlewis.ca/p/your-agent-has-the-same-permissions</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Tue, 23 Jun 2026 12:16:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!om_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!om_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!om_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!om_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!om_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!om_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!om_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1857506,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/202909683?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!om_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!om_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!om_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!om_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbec4d0c3-6a9b-4cf0-b055-e1c2aa846a6d_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A senior leader asked a simple question in a governance review last week: can we audit and verify every action an agent takes? Not the model&#8217;s reasoning, not the prompt &#8212; the actions. The edits, the sends, the deletions. The honest answer was no. Not yet. And the more I sat with that &#8220;not yet,&#8221; the more I realized it is the actual story of agentic AI inside a regulated firm, and almost nobody is telling it.</p><p>The entire public conversation about agents is about capability. How much access should they have. How autonomous should they be. Whether to let them run in a loop until a goal is met. Those are real questions. They are also the wrong ones to lose sleep over, because they are downstream of a quieter fact that has already shipped.</p><h2><strong>The agent runs as you</strong></h2><p>When a lawyer sets up an agent, the agent does not operate in some sandbox beside their work. It operates inside their identity. Microsoft&#8217;s own documentation for Microsoft 365 Copilot agents is explicit about this: an agent runs under the user&#8217;s identity, accessing data and services according to the permissions already assigned to that user. The agent inherits the person.</p><p>In a law firm, that inheritance is not abstract. It means the agent can open the matter, edit the document, send the email, and delete a file from the document management system &#8212; because the lawyer can. The newer agent layers that sit inside everyday tools, the ones that took Copilot from &#8220;draft me a paragraph&#8221; to &#8220;go do the thing,&#8221; did not just add a feature. They quietly handed a non-human process the full reach of a licensed professional.</p><p>So the unit of risk was never the agent&#8217;s cleverness. It was the agent&#8217;s authority. And that authority is borrowed, silently, from a person who may never see how it was used.</p><h2><strong>Accountability used to live in a person</strong></h2><p>The reasonable objection is that we never had a clean audit trail for humans either. That is true, and it is worth conceding plainly. A lawyer who made a hundred small decisions in a document did not generate an inspectable log of each one. The record, where it existed, lived in systems IT could comb through after the fact, not in anything the practitioner reviewed.</p><p>But accountability did not depend on that log, because it lived somewhere more reliable: in a licensed person you could question. You could ask a lawyer why they did something and get an answer, because they were present for the doing. Responsibility and awareness were the same thing.</p><p>An agent breaks that link. The action still carries the lawyer&#8217;s authority &#8212; their permissions, their identity, their name on the matter &#8212; but the awareness is gone. The person who is accountable was not in the room. The trail still exists for IT to assemble; tools like Purview will surface agent activity to administrators and compliance teams. What does not exist is a view that puts &#8220;here is what your agent did, under your name&#8221; in front of the one person who has to answer for it.</p><p>That is the gap. Not logging. Surfacing, to the accountable party, in time to matter.</p><h2><strong>The unsolved part is judgment, not logging</strong></h2><p>Here is the part I am genuinely unsure about, and I would rather name the uncertainty than pretend past it. Even if you capture every action &#8212; and increasingly you can &#8212; it is not obvious that you can present that record to a non-technical professional in a form they can actually use.</p><p>A partner does not want a JSON event stream. They want to know whether the thing the agent did was the thing they would have done, and they want to know it without becoming an engineer. The hard problem is not storage. It is translation: turning a machine&#8217;s actions into something a human can apply judgment to, fast enough to catch a mistake before it leaves the building. I have not seen anyone solve that well, and I am not certain the interface even exists yet.</p><p>What I am certain of is the asymmetry. Capability is compounding on a monthly cadence &#8212; new agents, broader access, longer autonomous runs. The ability to answer for any of it is not moving at the same speed. Every month that gap widens is a month a firm is accumulating authority it cannot account for.</p><h2><strong>The smaller question</strong></h2><p>There is a temptation, when a tool gets more powerful, to spend all the governance energy on the frontier &#8212; what should we let it do next. That energy is misallocated if you cannot yet answer for what it has already done.</p><p>Before you give an agent more of what it can do, you have to be able to answer the smaller question. Not the impressive one about capability. The unglamorous one about accountability.</p><p>What did the agent do? What can it do? The second question is getting all the attention right now. The first is the one with your name on it.</p><div><hr></div><p><em>I write every week about AI as an operational discipline inside regulated firms &#8212; the work behind the work that does not show up in the vendor demo. If that is your world, subscribe.</em></p>]]></content:encoded></item><item><title><![CDATA[The Entry Fee for McKinsey's 60 Percent Smaller Teams]]></title><description><![CDATA[The agentic delivery numbers are real. The preconditions underneath them are the part nobody will quote.]]></description><link>https://andrewlewis.ca/p/the-entry-fee-for-mckinseys-60-percent</link><guid isPermaLink="false">https://andrewlewis.ca/p/the-entry-fee-for-mckinseys-60-percent</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Thu, 18 Jun 2026 12:22:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UIis!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UIis!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UIis!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!UIis!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!UIis!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!UIis!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UIis!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1668058,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/201829124?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UIis!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!UIis!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!UIis!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!UIis!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a9fa27-1339-420a-b1bf-2dfa144861d5_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>McKinsey just published the stat that every AI strategy deck will be quoting for the next year. In <a href="https://www.mckinsey.com/capabilities/mckinsey-technology/our-insights/rewiring-software-delivery-for-the-agentic-era">Rewiring software delivery for the agentic era</a>, the firm reports that organizations adopting agentic delivery models are seeing threefold to fivefold productivity improvements alongside a 60 percent reduction in average team size. Teams of eight to twelve become pods of three or four: a product owner, a tech lead, and an AI-enabled engineer supervising a factory of agents that works through the night.</p><p>The accepted read writes itself. AI makes teams smaller and faster. Buy the agents, shrink the teams, collect the productivity.</p><p>But the article also lists, almost in passing, what the model requires before any of it works, and that list is the part worth reading twice. The business needs a clear enough vision of what is being built that agent output can be judged against it. The technical environment has to be standard and consistent. The path from requirements to code has to follow a structure agents can reliably interpret. Requirements, guardrails, and specifications have to be codified in machine-readable formats instead of scattered across disconnected documents. And the same core stakeholders have to stay engaged across the entire value stream, or everything downstream turns into rework.</p><p>Strip out the agent vocabulary and read that list again. It describes <strong>a team that has made its own work completely legible to itself</strong>. Decomposed operations, explicit intent, defined review gates, and shared context that lives in the system rather than in someone&#8217;s head. The agents are downstream of all that. So is the 60 percent.</p><h2>Nimbleness is a visibility property</h2><p>&#8220;Nimble&#8221; usually gets read as small and fast, which is why the team-size stat will travel further than anything else in the article. But team size was never the constraint. The constraint is whether anyone can see the work structurally.</p><p>A team that understands its own work as a system of operations &#8212; what gets produced, what gets handed off, what intent governs each piece, where judgment is non-negotiable &#8212; can re-form around a new mission in days. A team that experiences its work as a continuous stream of tasks, meetings, and messages is rigid at any size, and handing it a factory of agents produces exactly what McKinsey warns about: fragmented output that nobody trusts. The small team is a consequence of the legible one.</p><h2>Structure is a fossil of coordination cost</h2><p>There is a reason team structure resists this, and Ronald Coase named it in 1937. <a href="https://onlinelibrary.wiley.com/doi/10.1111/j.1468-0335.1937.tb00002.x">The Nature of the Firm</a> argued that firms exist because coordinating through the open market costs something: finding people, negotiating, enforcing agreements. The same logic shapes structure inside the firm. Layers, roles, and handoffs exist to manage the cost of communication, translation, and oversight. The professional services pyramid is the cleanest example anywhere &#8212; juniors produce, seniors review, and the ratio between them is priced directly into the business model. Org charts are fossils of what coordination used to cost.</p><p>When a technology collapses specific coordination costs, structure sized for the old costs turns from scaffolding into drag. McKinsey&#8217;s own pipeline analysis makes this concrete: the bulk of delivery effort sits in requirements through coding, where humans spend their days translating intent from one artifact to another. Every translation is a coordination cost. The agentic model removes the human from those handoffs entirely, with review concentrated at defined gates. The structure built to manage the handoffs does not dissolve on its own. Someone has to notice it is now load without purpose.</p><h2>Intent travels better than instructions</h2><p>If detailed structure recedes, something has to replace it, and the strongest answer is older than software. After Napoleon dismantled the Prussian army at Jena in 1806, the reformers who rebuilt it concluded that battlefield conditions changed faster than orders could travel. Helmuth von Moltke the Elder later gave the principle its lasting form: no plan of operations survives first contact with the enemy. So commanders stopped issuing detailed instructions and started transmitting intent &#8212; what must be achieved and why &#8212; leaving subordinate officers to decide how, on the ground, as conditions shifted. The Prussians called it Auftragstaktik. The modern term is mission command.</p><p>McKinsey&#8217;s daily sprint is mission command in a different uniform. During the day, humans resolve ambiguity, set guardrails, align stakeholders, and define what acceptable looks like. Overnight, execution runs against that intent at scale. In the morning, review happens at the gates. Intent travels into the night shift; instructions never would.</p><p>This is the actual design principle for team nimbleness in the agentic era: a mission plus short-term goals. The mission states what must be true and why it matters. The short-term goals create checkpoints where intent gets recalibrated against what execution revealed. The compression from two-week sprints to daily cycles matters less than what made it possible &#8212; clearly transmitted intent needs far less planning apparatus wrapped around it.</p><h2>Getting legible</h2><p>None of this tells a leader how to make a team legible to itself, and I am suspicious of anyone selling a template for it. The honest version is a set of questions, asked of the team&#8217;s work rather than its org chart. What operations does this team actually perform &#8212; operations, not roles? Where does intent currently live: written somewhere a system could read it, or in the heads of two senior people? Which handoffs in the workflow are translation, and which are judgment? What does &#8220;done&#8221; mean for each operation, and who decides?</p><p>Most teams cannot answer these today. That is not a failing; nobody has ever asked them to see their work this way. The structural view is developed, not installed, and there is no vendor for it. Which is precisely why it is the work behind the agentic numbers rather than a line item in the deployment budget.</p><h2>The harder case</h2><p>Software is the easy case, and McKinsey says so directly: the way agents are being used in software development is a harbinger for broader delivery models. Decades of version control, ticketing systems, and CI/CD have already made software work partially legible. Requirements and code leave trails by default.</p><p>Professional services work mostly does not. In a law firm, intent lives in precedent files, marked-up drafts, and the accumulated judgment of senior practitioners. The pyramid prices translation in as oversight and calls it development. The legibility debt is far larger than in software, which means the work behind the numbers is heavier &#8212; and the advantage for any firm that does it is correspondingly larger, because almost nobody in the market has started.</p><p>There is also a structural bind worth naming. Re-decomposing a team&#8217;s structure requires someone to sign for it, and in conservative institutions nobody wants their name on the new org chart. Missions carry less weight. A mission expires; a structure has to be defended indefinitely. That makes mission-based nimbleness more available to a regulated institution than restructuring. A three-month mission with short-term goals and defined review gates asks nothing of the org chart, and it quietly builds the legibility that any future structure will depend on.</p><p>The 60 percent will circulate as proof that AI shrinks teams. Read the preconditions instead and it proves something quieter: the organizations collecting those numbers did an unglamorous body of work first. They made the work visible and the intent explicit <strong>before any agent touched it</strong>. The team got smaller because the work got legible.</p><p>AI doesn&#8217;t make teams nimble. It exposes which teams already did the work of seeing themselves clearly.</p><div><hr></div><p><em>If this resonated, subscribe to get the next piece on the work behind the work &#8212; the human capability development that AI adoption quietly depends on.</em></p>]]></content:encoded></item><item><title><![CDATA[You Cannot Review Your Way Out of This]]></title><description><![CDATA[Verification at AI scale is a selection problem, and every guardrail, eval, and playbook only works after a human decides what deserves attention.]]></description><link>https://andrewlewis.ca/p/you-cannot-review-your-way-out-of</link><guid isPermaLink="false">https://andrewlewis.ca/p/you-cannot-review-your-way-out-of</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Tue, 16 Jun 2026 12:17:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Wvlq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Wvlq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Wvlq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Wvlq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Wvlq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Wvlq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Wvlq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1039681,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/201828729?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Wvlq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Wvlq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Wvlq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Wvlq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78002110-d6c3-4988-ba0c-a42110f217c7_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Somewhere in your organization this week, an AI-drafted document was approved by someone who did not fully read it. Not from laziness; from arithmetic. Machine-generated drafts, summaries, analyses, and reports now arrive faster than the people responsible for them can check, and the gap widens every month, in every department, whether or not anyone has said it out loud.</p><p>Software engineering is simply where that gap is measured best, because no other knowledge work is so thoroughly instrumented. The numbers from that world deserve attention from everyone whose job produces documents rather than code, since developers are only the first to hit a wall the rest of us are approaching. <a href="https://linearb.io/resources/software-engineering-benchmarks-report">LinearB&#8217;s 2026 benchmarks</a>, built from 8.1 million pull requests across 4,800 engineering teams, found AI-assisted work waits 4.6 times longer for a first review than human-written work, and barely a third of it gets accepted, against 84.5 percent for the unassisted kind. A queue is forming in front of the one resource AI cannot multiply: a person willing to say this is correct.</p><p>The perception side is stranger. When <a href="https://metr.org/blog/2026-02-24-uplift-update/">METR ran a randomized trial</a> with experienced open-source developers, participants believed AI had made them roughly 20 percent faster while the measured result was 19 percent slower. METR&#8217;s own 2026 follow-up suggests that slowdown has likely narrowed or reversed as the tools improve, so treat the direction as provisional; the durable finding is the gap itself. The people closest to the work could not tell whether it was going faster. If that is true in the most measurable profession on earth, consider the odds that anyone drafting contracts, reports, or board memos with AI knows their real number.</p><p>The standard interpretation of all this is that review is too slow. Leaders read these numbers and conclude they need faster reviewers, AI-assisted checking tools, better queue management. The vendors agree enthusiastically, because every one of those conclusions has a product attached to it.</p><p>The standard interpretation is wrong. The numbers describe a system where generation capacity and verification capacity have decoupled, and no amount of review acceleration reconnects them. <a href="https://www.sonarsource.com/blog/state-of-code-developer-survey-report-the-current-reality-of-ai-coding">Sonar&#8217;s 2026 State of Code survey</a> of more than 1,100 developers found AI now accounts for roughly 42 percent of committed code, expected to reach 65 percent by next year. The same survey found 96 percent of developers don&#8217;t fully trust that code to be correct, only 48 percent always verify it before committing, and 38 percent say reviewing AI output takes more effort than reviewing a colleague&#8217;s work. Read those findings together. This is a machine that produces artifacts faster than anyone can check them, staffed by people who know the artifacts are unreliable and have quietly stopped checking anyway. Swap the code for contracts, research summaries, financial models, or marketing copy, and nothing in that sentence changes except the file format.</p><p>The question worth asking is what verification even means once output volume crosses review capacity. The answer is structurally different from what most organizations are building.</p><h2>Two curves, only one of them moves</h2><p>A <a href="https://arxiv.org/abs/2602.20946">recent economics paper on the automation frontier</a> frames the problem as two competing cost curves. The cost of generating an artifact (code, a brief, a report, a vulnerability finding) falls exponentially as models improve. The cost of verifying that artifact stays roughly where it has always been, because it is bounded by human cognition. A senior developer can meaningfully review somewhere around 150 lines of code an hour. A careful reader of contracts, research memos, or financial models moves at a similarly fixed pace, and has for as long as those documents have existed. Attention does not get firmware updates, and there is no model release on any roadmap that changes it.</p><p>Engineers at Agoda <a href="https://www.infoq.com/news/2026/03/agoda-ai-code-bottleneck/">reached the same conclusion from production data</a>: AI tools measurably raised individual output while project-level velocity barely moved, because coding was never the real constraint. The constraint sits upstream, in specification and verification, the two activities that require human judgment. They frame this as a rediscovery of Fred Brooks&#8217; forty-year-old argument that accelerating one stage of a pipeline buys you almost nothing if the binding stage is elsewhere. The industry has spent three years accelerating the stage that was never binding.</p><p>When one curve falls exponentially and the other stays flat, the gap between them compounds. This is the part most adoption conversations skip. An organization that responds to a compounding gap with linear capacity (more reviewers, longer review windows, a second approval step) has chosen to lose slowly rather than rethink the game. The arithmetic forecloses the strategy before the first hire is made.</p><p>The same Agoda piece offers a useful taxonomy of the postures available once you accept that. You can read every line AI produces, the white-box stance, which preserves full assurance and caps your throughput at exactly the 150 lines an hour you had before. You can ship whatever the machine generates, the black-box stance, which is fast right up until it meets a production system with real users and real consequences. Or you can verify selectively, at boundaries and against intent, accepting that assurance is now a budget to be allocated rather than a property to be assumed. The same three stances exist for every AI draft that crosses a desk, whether the desk belongs to an engineer, a lawyer, or an analyst, and whether anyone has named them or not. Most organizations are running the third stance already. Few have admitted it, and fewer still have decided deliberately where the budget goes, which means the allocation is being made anyway, by default, by whoever is most tired on a given afternoon.</p><p>I work inside a national law firm, where this arithmetic has a sharper edge than it does in software. A pull request that ships with a subtle bug costs you an incident. A factum that ships with a fabricated citation costs you a professional reputation, and possibly a client. The stakes are asymmetric: the price of being wrong exceeds the benefit of being right, which means verification cannot be quietly abandoned the way half of those surveyed developers have abandoned it. Regulated industries never get the option of not checking; the only choice available is what to check.</p><p>That phrase, choosing what to check, is the actual answer to the scaling question, and everything else in this piece is an elaboration of it.</p><h2>What the tools actually do</h2><p>The toolchain that has grown up around this problem is usually marketed as verification capacity. Guardrails, evals, LLM-as-judge pipelines, formal methods, provenance tracking. The framing is that these tools let you check more. Watch how each one actually works and something close to the opposite is true: each earns its keep by deciding what will never be examined, and by making that decision survivable.</p><p>Guardrails verify boundaries instead of artifacts. A guardrail doesn&#8217;t ask whether an output is good; it asks whether the output crosses a line you drew in advance: leaks data, violates a policy, exceeds a risk threshold. <a href="https://arxiv.org/abs/2603.25176">Production guardrail systems</a>, like the centralized service Singapore deployed for its public chatbots, work precisely because they refuse to evaluate quality. They evaluate violations, which is a vastly smaller problem. The honest description of a guardrail is a decision about which failures you can afford to detect cheaply and which ones you&#8217;ve accepted you won&#8217;t catch at all.</p><p>Evals and judge pipelines apply the same selective logic to quality. <a href="https://medium.com/@adnanmasood/rubric-based-evals-llm-as-a-judge-methodologies-and-empirical-validation-in-domain-context-71936b989e80">Rubric-based evaluation</a> has become the working standard for assessing open-ended AI output at scale, and it earns its place. But a judge model scoring outputs against explicit criteria is a statistical instrument. It reports on the distribution, calibrated periodically against human experts, carrying known biases toward verbose answers and toward output resembling its own, and it stays silent on whether any individual artifact is correct. Teams that treat eval scores as verification have confused quality control sampling with inspection, and the difference matters enormously when one bad artifact can hurt you.</p><p>With specifications, the selection happens upstream. The most interesting shift in engineering practice right now is the argument, <a href="https://www.aviator.co/blog/the-ai-code-verification-bottleneck-why-faster-code-generation-means-slower-reviews/">made well in Aviator&#8217;s analysis of the review bottleneck</a>, that the only verification with an external reference point is verification against a spec a human has approved. Tests check behavior the test author imagined. A reviewer checks a diff against their reconstruction of what it was supposed to do. A spec is the one artifact in the pipeline where intent lives outside the generation loop, and verifying against it means a human did the hard thinking once, up front, so the checking inherits that thinking.</p><p>Formal methods verify guarantees instead of confidence. Leonardo de Moura, who built the Lean theorem prover, <a href="https://leodemoura.github.io/blog/2026-2-28-when-ai-writes-the-worlds-software-who-verifies-it/">draws the line cleanly</a>: testing provides confidence, proof provides a guarantee, and it is genuinely hard to quantify how much confidence testing actually buys you. Proof is expensive and narrow; you reserve it for the properties where confidence isn&#8217;t good enough. Which is, again, a selection decision.</p><p>Every tool that works at scale works by shrinking the verification surface, and every shrinking decision is a human judgment about what matters. The tools carry out triage decisions. They cannot make them.</p><h2>When the verifier needs a verifier</h2><p>The obvious move, once human review can&#8217;t keep pace, is to have AI verify AI. It is also the move with the nastiest failure mode, and the failures are no longer hypothetical.</p><p><a href="https://srlabs.de/blog/ai-verification-bottleneck">SRLabs&#8217; analysis of AI-generated security findings</a> describes what happens when an organization layers AI security tooling on top of AI-generated code and treats the findings as decisions: the system becomes self-referential, producing artifacts faster than anyone can anchor them in an actual threat model. Teams under that load optimize for closing items rather than reducing risk, because closed items are measurable and risk reduction isn&#8217;t. The curl project reached the logical endpoint in January, ending its bug bounty after years of fielding AI-generated vulnerability reports that looked plausible and dissolved under scrutiny. The bounty had become a denial-of-service attack on maintainer attention.</p><p>Academia hit the same wall from a different direction. An <a href="https://arxiv.org/abs/2602.05930">audit of NeurIPS 2025 submissions</a> documented a hundred fabricated citations that sailed through peer review, including placeholder hallucinations literally citing &#8220;Firstname Lastname,&#8221; because reviewers check methodology and novelty, and nobody&#8217;s job has ever been confirming that cited papers exist. The gap was always there. AI made it economical to exploit at volume. A <a href="https://arxiv.org/abs/2601.16909">related analysis of peer review itself</a> puts the conclusion bluntly: when the rate of claims rises exponentially against fixed human bandwidth, collapse is a mathematical inevitability, and abstaining from AI assistance doesn&#8217;t preserve the system&#8217;s integrity; it just guarantees the system drowns.</p><p>And the judge models themselves are attack surface. Security researchers have shown judges that <a href="https://www.trendmicro.com/vinfo/us/security/news/managed-detection-and-response/llm-as-a-judge-evaluating-accuracy-in-llm-security-scans">ignore their instructions when fed adversarial output</a>, repeating attack strings instead of evaluating them. The fix proposed in that research is a guardrail on the judge &#8212; which should give you pause. We are now building verifiers for the verifiers, and there is no level of that tower where the regress stops on its own. It stops where a human owns an assessment and signs their name to it. Nowhere else.</p><h2>Verification as deterrence</h2><p>Earlier this year, in a long conversation about verification architecture, I landed on a framing I haven&#8217;t been able to shake: at scale, verification stops behaving like a truth-recovery problem and starts behaving like a deterrence problem. Closer to mutually assured destruction than to forensic investigation.</p><p>Truth recovery assumes you can, in principle, examine an artifact and determine whether it is correct. That assumption holds at human scale. It fails at machine scale, for the cost-curve reasons above, and the failure is permanent because the gap compounds. What remains achievable is a different posture entirely: making bad artifacts expensive to produce, cheap to contain, and traceable to a producer who bears the cost.</p><p>Look at what actually works in the systems under the most pressure and you find deterrence mechanics where you would expect inspection. Curl didn&#8217;t get better at detecting slop reports; it changed the economics of submitting them. The proposed fix for fabricated citations is mandatory existence checks rather than smarter reviewers, shifting the verification burden onto the claim, at the point of submission, where it costs the producer instead of the reviewer. Spec-driven development works because the spec makes intent auditable, which makes deviation attributable. And provenance tracking (who generated this, from what, under what instructions) won&#8217;t establish that an artifact is correct, but it establishes who answers for it if it isn&#8217;t, and that knowledge changes behavior upstream of any review.</p><p>The legal profession produced the cleanest illustration on record just this week. On June 8, a federal judge in Mississippi <a href="https://www.abajournal.com/news/article/federal-judge-terminates-4-plaintiff-and-defense-attorneys-over-ai-errors">sanctioned every lawyer of record in a contract dispute</a> after filings from both sides turned out to contain fabricated citations. The two out-of-state lead attorneys admitted using AI without verifying the output; local counsel had signed the briefs without reviewing them. The court revoked both out-of-state lawyers&#8217; temporary admissions, barred them from appearing before the Northern District of Mississippi for two years, fined all four attorneys, cancelled the trial, and referred everyone to their state bars. The instructive part, read as a verification story, is the remedy the court chose: no call for better detection tooling, just consequence attached to the signature. The ruling&#8217;s principle, that responsibility &#8220;remains the sacred duty of the lawyer who signs the page,&#8221; is deterrence doctrine written by a judge who understood the problem more clearly than most AI strategy decks do.</p><p>The pipeline I sketched afterward runs adversarial passes from different model families against a primary output: fact checks, logic checks, an audit of unstated assumptions. The design process taught me something I didn&#8217;t expect. Certifying outputs as true was never on the table; no architecture I could draw gets you there. What a design like this can do is make certain classes of failure reliably expensive to get past, which changes what is worth attempting in the first place. Deterrence, functioning exactly as deterrence does.</p><p>This reframe matters because it redirects investment. An organization pursuing truth recovery buys review capacity and falls further behind every quarter, while one pursuing deterrence builds chokepoints, assigns ownership, and engineers consequence &#8212; and the second approach scales, because consequence does the verifying for you, continuously, at every point of production you can&#8217;t see.</p><h2>The part no tool covers</h2><p>All of which leaves the question every playbook quietly assumes is already answered: who decides what deserves verification in the first place?</p><p>Every mechanism in this piece runs on a prior human judgment. Which boundaries the guardrails enforce, what the rubric rewards, where formal proof is worth its cost, and what gets sampled, gated, or waved through. The tooling industry talks about these as configuration details. They are the entire game. A perfectly engineered verification stack pointed at the wrong things is expensive theater, and the <a href="https://srlabs.de/blog/ai-verification-bottleneck">security teams optimizing for closed findings</a> while risk accumulates elsewhere are running exactly that theater right now.</p><p>The judgment involved is the capability I keep circling in this newsletter: signal discrimination. Knowing, when you face more artifacts than you can examine, which ones matter: by consequence, by blast radius, by the asymmetry between what a wrong artifact costs and what a right one earns. In my world, an AI-drafted internal summary and an AI-drafted court filing might come from the same tool on the same afternoon, and they belong in entirely different verification regimes. Nothing in the tooling knows that. A person has to, and that person has to be willing to own the call, which in most organizations is the genuinely scarce resource. The technology of verification is improving fast. The willingness to sign one&#8217;s name to &#8220;this is checked enough&#8221; is not, because the signature carries all the downside and none of the credit.</p><p>So the real playbook, stripped of vendor language, is short. Decide what failure costs, per artifact class, before deciding what verification it gets. Push the burden of proof onto the producer wherever you can: specs, provenance, existence checks at the point of claim. Use machines to sample distributions and patrol boundaries; reserve humans for the artifacts where consequence is asymmetric. And put a name on every assessment, because an unowned verification is a rumor with a checkmark on it.</p><p>None of that is a technology roadmap. It is judgment, exercised in advance, encoded into a system &#8212; which is to say it is work most organizations have not done, and the tools they&#8217;re buying cannot do it for them. The verification crisis everyone can now measure looks less like a case for better checking machines than like a bill arriving for a question that got skipped: what, in all this output, actually matters?</p><p>AI doesn&#8217;t make you better. It exposes whether you were already doing the work.</p><div><hr></div><p><em>If this kind of analysis is useful, subscribe &#8212; one piece like this every week, written from inside an institution doing the work in real time.</em></p>]]></content:encoded></item><item><title><![CDATA[The Last Mile in Legal Has Its Own Geography]]></title><description><![CDATA[McKinsey published the data. HBR named the gap. Both are right. Both are general. Here is what the terrain looks like inside a law firm.]]></description><link>https://andrewlewis.ca/p/the-last-mile-in-legal-has-its-own</link><guid isPermaLink="false">https://andrewlewis.ca/p/the-last-mile-in-legal-has-its-own</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Thu, 14 May 2026 11:02:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!b9_0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b9_0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b9_0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!b9_0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!b9_0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!b9_0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b9_0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3007207,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/197434957?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!b9_0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!b9_0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!b9_0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!b9_0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e69d276-f827-40fb-9f2b-bf970c11f015_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The McKinsey numbers do not need a second read to be sobering. The first read does the job.</p><p><em><a href="https://www.mckinsey.com/capabilities/people-and-organizational-performance/our-insights/the-state-of-organizations">The State of Organizations 2026</a></em> surveyed more than ten thousand senior leaders across fifteen countries and sixteen industries. Eighty-eight percent of organizations are deploying AI in at least parts of their business. Fewer than twenty percent see significant impact on the bottom line. Eighty-six percent of leaders say they are not ready to embed AI into day-to-day operations. One in six organizations has no clear C-level owner for AI adoption at all.</p><p><a href="https://hbr.org/2026/03/the-last-mile-problem-slowing-ai-transformation">HBR&#8217;s &#8220;last mile&#8221; diagnosis</a> from a closed-door Harvard summit fills in the texture McKinsey&#8217;s spreadsheet cannot. The summit included a global investment bank with more than two hundred and fifty LLM-connected applications in production, a payments network where copilot adoption sits above ninety-nine percent of employees, and an apparel group running eighteen thousand automated finance processes. In every one of those firms, finance teams hunt for measurable impact in headcount and cycle-time numbers and come up empty. The work is happening. The value is not landing.</p><p><a href="https://andreisavine.substack.com/p/last-mile-enterprise-ai-dies">Andrei Savine</a> pulled both pieces together in March: enterprise AI dies in the last mile because organizations under-funded the operational layer that turns model output into business value. He calls it the Production Layer. He cites an executive in the McKinsey report who put the corrective ratio at five to one &#8212; for every dollar spent on technology, five should be spent on people, process, and verification.</p><p>That diagnosis is correct.</p><p>That diagnosis is also a tourist map.</p><p>Inside the firm, the disconnect is on the calendar. A partner explains why AI will not materially change the way they practice. The IT department, meanwhile, is months into a slow restructure built on the opposite assumption. Both positions are held with conviction. Neither side can yet describe the outcome on the other side of the change.</p><p>The general diagnosis works at the macro level because it abstracts away the structural features of the firms it describes. Software companies can absorb productivity gains by reducing headcount. They have CEOs who can decree adoption. Their work product is a piece of software, not a lawyer-hour. None of that is true in a law firm.</p><p>This piece is a map of what the last mile actually looks like inside one. The hazards in the terrain. The path that does not end in either a mass layoff or a restructuring announcement masquerading as strategy. The question is not whether five-to-one is the right ratio. It is whether law firms know what the five is supposed to build.</p><div><hr></div><h2><strong>Why the General Diagnosis Underestimates Legal</strong></h2><p>The five features that make a law firm break the standard model are not exotic. They are the structure of the business. Anyone who has worked inside a firm knows them. Anyone trying to apply a generic enterprise-AI playbook to a firm runs into them within the first quarter.</p><p>Start with the billable hour. The standard productivity gain from AI is a productivity gain. In a firm billing by the hour, the same gain is a revenue cut unless something changes upstream. Two software companies have absorbed their productivity gains by reducing headcount in the last few months &#8212; <a href="https://www.atlassian.com/blog/company-news/atlassian-team-update-march-2026">roughly 1,600 people at Atlassian</a>, and <a href="https://www.bloomberg.com/news/articles/2026-02-24/wisetech-to-cut-2-000-jobs-as-ai-ends-era-of-manual-coding">a 2,000-person restructure at WiseTech shortly before</a>. Law firms cannot run that play. The work product is the lawyer-hour itself. <em><a href="https://www.wolterskluwer.com/en/news/wolters-kluwer-releases-2026-future-ready-lawyer-survey-report">Wolters Kluwer&#8217;s 2026 Future Ready Lawyer</a></em> survey of 810 lawyers found 54% expect firms to use AI efficiency for more clients or competitive pricing; the 2024 edition projected that AI automation could reduce hourly billing per lawyer by roughly $27,000 a year. Those numbers cut against revenue, not toward it. The inversion is what makes generic AI advice land badly in partner meetings: efficiency arguments are heard as compensation arguments.</p><p>Beneath the billable hour sits the apprenticeship. Junior associates learn judgment by doing the work AI does best &#8212; first-pass document review, citation checks, due diligence summaries, contract redlining. The training pipeline runs through the same activities AI is now positioned to consume. Automate the work without redesigning the apprenticeship, and the firm produces senior associates whose calibration never developed. The <em><a href="https://www.citiglobalwealth.com/atwork/insights/citi-hildebrandt-client-advisory">Citi Hildebrandt 2026 Client Advisory</a></em> reports revenue growth of 11.3% across surveyed firms, productivity per lawyer down 0.6%, and 88% of firms planning continued associate growth &#8212; built on the assumption that the apprenticeship pipeline still works. The assumption may not hold.</p><p>Governance compounds the problem. McKinsey&#8217;s number &#8212; one in six organizations without a clear C-level AI owner &#8212; understates the legal case. In an equity partnership, even a C-level owner cannot decree adoption. Decisions move through practice-group chairs, executive committees, and partner votes. The <em><a href="https://www.iltanet.org/blogs/ilta-news1/2025/09/16/press-release-ilta-releases-2025-legal-technology">ILTA 2025 Technology Survey</a></em> of 580 firms found that user resistance is the top barrier to AI adoption at 57%, up from 54% the year prior; half of those firms have no formal AI policy at all. The gap is not a failure of legal IT. It is a structural feature of how authority is held and exercised in a partnership.</p><p>Privilege and conflicts sit beneath the policy gap. Technical friction in legal includes constraints other industries do not face. Solicitor-client privilege. Conflict-of-interest screens. Data residency for regulated client matters. Prohibitions on third-party model training. Ethical-wall enforcement that must survive at the model and the prompt level. <em><a href="https://www.americanbar.org/news/abanews/aba-news-archives/2024/07/aba-issues-first-ethics-guidance-ai-tools/">ABA Formal Opinion 512</a></em> and the <em><a href="https://flsc.ca/what-we-do/model-code-of-professional-conduct/">Federation of Law Societies of Canada Model Code</a></em> both treat AI use as a competence question, a confidentiality question, and a supervision question simultaneously. The provincial regulators &#8212; Ontario, British Columbia, Alberta, and the Barreau du Qu&#233;bec &#8212; have all issued guidance that constrains which tools are usable for which matters. None of this is a security checklist. It is a substantive constraint on adoption.</p><p>Last comes the asymmetry. The price of being wrong exceeds the benefit of being right. <em><a href="https://en.wikipedia.org/wiki/Mata_v._Avianca,_Inc.">Mata v. Avianca</a></em> made the failure mode public in 2023. <a href="https://www.damiencharlotin.com/hallucinations/">Damien Charlotin&#8217;s AI Hallucination Cases Database</a> has been logging the failure mode ever since, at the rate of dozens of new entries per month from courts in more than a dozen countries. The downside of an AI-fabricated citation is career-ending. The upside of an AI-generated brief is incremental. When the loss function is asymmetric, rational professionals default to avoidance &#8212; which produces exactly the shallow utilization HBR&#8217;s last-mile diagnosis describes.</p><p>Generic diagnoses do not survive these five features. A different map is needed.</p><div><hr></div><h2><strong>The Map: Five Hazards in the Terrain</strong></h2><p>These hazards are not theoretical. They are already arriving.</p><h3><strong>Hazard 1 &#8212; The Apprenticeship Gap</strong></h3><p>The work AI handles best is the work juniors learn from. Document review. Due diligence summaries. Citation checks. Contract markup. Strip those activities out without replacement and the firm is running an apprenticeship system whose training data has been deleted.</p><p>The signal does not appear in the first year. It appears two and three years in, when an associate who learned to draft alongside an LLM is asked to negotiate a non-standard term in a deposition prep meeting and cannot triangulate why the standard term existed. The output looks fine. The pattern-matching against past work &#8212; the thing that turns a third-year into a fourth-year &#8212; has not happened.</p><p>The hazard sits at the intersection of two operating questions. The decomposition question asks what the junior was actually doing inside the document review. Some of it was the review itself. Some of it was learning what a representation looks like when it is doing real work in a credit agreement, learning which questions to ask before flagging an issue to a senior, learning the difference between a problem and an artifact of the drafting party&#8217;s preferences. The collaboration-quality question asks what the junior was learning by doing the work, not what they were producing. If only the production is automated, the learning evaporates with it.</p><p>The Citi Hildebrandt advisory caught the contradiction. Firms are still planning associate growth on the old assumption &#8212; that the same training pipeline still works &#8212; while productivity per lawyer ticked down. The senior-heavy staffing model some firms are leaning into assumes juniors continue to develop into seniors capable of senior work. The pipeline that produced that development is being quietly automated. The hazard arrives on the promotion clock. Juniors will keep being hired; the calibration that used to develop alongside them will not appear on schedule, and the firm will not notice until the class it tries to promote is already on the partnership runway.</p><h3><strong>Hazard 2 &#8212; The Pricing Collision</strong></h3><p>The hourly model inverts under AI. The same brief that took twenty hours last year takes seven this year. The client knows. The client has run the spreadsheet.</p><p>The signal arrives in the form of an RFP that asks specific questions. Which tools does the firm use? On which matter types? What productivity assumption is baked into the proposed rates? The <em><a href="https://www.acc.com/resource-library/2025-acc-chief-legal-officers-survey">2025 ACC Chief Legal Officers Survey</a></em>, <a href="https://www.everlaw.com/press/release/acc-report-2025/">summarized in subsequent analysis</a>, found that 59% of general counsel see no clear cost savings from outside counsel using AI. A transparency gap, the report called it. That phrase will appear in a procurement deck within twelve months.</p><p>Pricing models that align with the new reality already exist. Fixed fees. Capped fees. Success fees. Capacity-based arrangements. AI-augmented hourly with a productivity discount. The menu is not new. What is new is the pressure to choose, and the pressure to choose visibly. <em><a href="https://www.thomsonreuters.com/content/dam/ewp-m/documents/thomsonreuters/en/pdf/reports/future-of-professionals-report-2025.pdf">Thomson Reuters&#8217; 2025 Future of Professionals</a></em> projects up to 240 hours per professional saved per year by AI-augmented workflows; the 2024 edition projected 12 hours per week saved by 2029 and roughly $100,000 in additional billable hours per U.S. lawyer. Those numbers are useful to a firm that has redesigned its pricing to capture them. They are punitive to a firm that has not.</p><p>This hazard is the one most likely to drive the same failure mode inside legal. <em>A firm that does not redesign pricing and roles will eventually balance the books with headcount.</em> The redesign is uncomfortable. The alternative is worse. The choice is whether to lead the pricing conversation or inherit one shaped by clients who started running their own AI math two budget cycles ago.</p><h3><strong>Hazard 3 &#8212; The Judgment Hollowing</strong></h3><p>Apprenticeship is about who comes next. Judgment hollowing is about who is already in the chair.</p><p>Senior judgment was built on years of doing junior work. Take that work away through AI mediation &#8212; not removal, mediation &#8212; and judgment does not develop the same way. The hazard appears in two opposing modes. One is overtrust: lawyers who accept AI output uncritically because the immersion that produced calibration has thinned out. The other is blanket rejection: lawyers who refuse to engage with AI assistance because the output feels uncalibrated even when it is sound. The first failure ships errors. The second ships sluggishness. Both come from the same root.</p><p>This is what skill-formation research in adjacent professions has been pointing at. When the practice that produced calibration is mediated, calibration drifts. Output calibration and signal discrimination &#8212; the two senior capacities that distinguish a partner from a senior associate &#8212; degrade in opposite directions when the underlying practice that built them disappears. The senior who used to read a memo and feel the wrong sentence is still reading the memo. The reading is faster. The feel is duller.</p><p>The signal is not visible on a dashboard. It is visible in review cycles. A senior who used to spot the problem in the first pass now spots it in the third. A junior whose draft used to be rebuilt is now approved with light edits &#8212; not because the draft is better but because the senior&#8217;s bar moved. The work product reads acceptable in both cases. The institutional capability does not develop. The hazard is the part of the iceberg the engagement letter cannot describe and the matter close-out cannot bill.</p><h3><strong>Hazard 4 &#8212; The Partnership Fracture</strong></h3><p>Adoption splits the equity table. Litigation moves on one tool while M&amp;A holds out. IP runs a different stack than tax. Practice groups operate as small businesses inside the firm; their adoption velocities can diverge by a factor of five within the same year, and the divergence shows up in realization rates before it shows up in compensation discussions.</p><p>The ILTA 2025 survey captured the inflection point. Eighty percent of firms reported using or exploring generative AI. Half reported no formal AI policy. Those two numbers describe a partnership in which adoption is happening faster than governance, and where the governance vacuum is being filled at the group level by whoever decided to act first.</p><p>The signal is not the adoption rate. The signal is the compensation tension that follows. Two practice groups that bill different realization rates against similar matter types invite a comp-committee conversation no managing partner enjoys. Recruitment messaging that promises one thing across a firm where the lived experience varies by group invites a different conversation with laterals. The fracture is not technological. The fracture is political, and it arrives on the partnership&#8217;s calendar approximately one fiscal year after the first group commits to its first material AI deployment.</p><p>Inside the firm, the divergence is already concrete. One practice group, pressed by a client to deliver AI-driven efficiencies, has mandated tool use and is rebuilding workflows accordingly. Other groups have not opened the conversation. All of them sit inside the same partnership. The comp committee will eventually have to reconcile what the calendar already shows.</p><p>The leadership challenge in this hazard is velocity matching, not adoption. Getting compatible adoption velocities across groups is a different problem from getting any adoption at all, and it is not solved by the same instruments.</p><h3><strong>Hazard 5 &#8212; The Asymmetric Stakes</strong></h3><p>The price of being wrong exceeds the benefit of being right.</p><p><em><a href="https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1:2022cv01461/575368/54/">Mata v. Avianca</a></em> was the warning shot &#8212; a Southern District judge, a $5,000 sanction, and the first widely circulated story of fabricated citations submitted to a federal court. <a href="https://www.damiencharlotin.com/hallucinations/">Charlotin&#8217;s database</a> has been the running count ever since: well over a thousand decisions tracked across more than a dozen countries, with new entries arriving at the pace of dozens per month. <em><a href="https://edrm.net/2025/07/when-ai-policies-fail-the-ai-sanctions-in-johnson-v-dunn-and-what-they-mean-for-the-profession/">Johnson v. Dunn</a></em>, decided in the Northern District of Alabama in July 2025, extended the line into BigLaw. A practice-group co-leader at a large, well-regarded firm signed a motion containing fabricated citations. The court signaled that monetary sanctions are no longer sufficient to deter AI-generated errors, and that future cases may see referrals to bar counsel and other escalations.</p><p>The signal is not the headline cases. The signal is the malpractice underwriting questionnaire. Carriers have started asking which tools a firm uses, for which matter types, with which verification process. Some are writing AI-specific exclusions. The <a href="https://www.fct-cf.ca/Content/assets/pdf/base/FC-Updated-AI-Notice-EN.pdf">Federal Court of Canada has issued a notice to the parties and the profession on AI in court proceedings</a>; provincial superior courts have followed; U.S. district courts have done the same on a docket-by-docket basis. Disclosure of AI use is no longer optional in many jurisdictions, and in some it must include disclosure of how the AI was used.</p><p>The hazard lives in the loss function. As long as one fabricated citation costs more than ten useful drafts save, the rational play is to avoid the tool &#8212; and that avoidance is what produces the shallow utilization the survey data keeps measuring. The accountability avoidance pattern McKinsey identifies is rational behavior in an asymmetric environment. Resolution is not the elimination of risk. Resolution is clarity about who is responsible for which verification, at which step, with what record.</p><p>Five hazards. None of them appear on the McKinsey or HBR map. The path through them is what the next section describes.</p><div><hr></div><h2><strong>The Path: What Implementation Actually Looks Like</strong></h2><p>Six moves. Each one resolves a hazard. None of them resolve all the hazards. The work has to be done in sequence and in combination.</p><h3><strong>Move 1 &#8212; Start with measurable operations before client-facing work</strong></h3><p>Knowledge management. Marketing. Finance. IT. Recruiting. These are the practice areas of the firm where measurement is possible, where the stakes are bounded, and where the team can build the muscle for measurement before pointing AI at client matters.</p><p>The reason this comes first is not theoretical. A firm that cannot measure value in its own back office has no business claiming to measure value in client work. Most firms skip the step anyway, because the political appeal of a client-facing pilot is too strong. There is a partner who wants it. There is a vendor who will demo it. There is a press release in the budget. The internal pilot is less photogenic. It is also where the failure modes show up cheaply.</p><p>What measurable internal deployments produce is more valuable than the deployments themselves. Real numbers. Real failure modes. Real governance precedents. The early adopters of internal AI become the people who teach the practice groups, because they have already had the embarrassing first conversation with privacy, the awkward second conversation with risk, and the corrective third conversation with finance.</p><p>A firm that has not staffed its internal AI team before standing up its client-facing AI team is not running an enterprise AI program. It is running a procurement exercise with extra steps. The move that says &#8220;we will start where it matters most&#8221; sounds bold and usually fails. The starting point that matters most is the one where mistakes cost the least.</p><h3><strong>Move 2 &#8212; Run the Task Audit at the practice-group level, not the firm level</strong></h3><p>A litigation matter is not an M&amp;A transaction is not a regulatory filing is not an IP prosecution. The work is structurally different. The AI fit is structurally different. The realistic adoption pace is structurally different. Firm-wide rollouts produce shallow utilization because they ignore those differences and try to deploy one set of tools, with one set of metrics, across groups that need different things.</p><p>What the federated version looks like is unglamorous. Each practice group decomposes its own work into the activities that compose a matter &#8212; research, drafting, review, analysis, communication, project management. Each group runs its own audit of where AI fits, where it does not, and where the answer is uncertain. Litigation may find AI fits document review and timeline construction; M&amp;A may find AI fits diligence summaries and disclosure schedule drafts; tax may find narrow but high-value uses around regulatory text comparison. Each group owns its own map.</p><p>This is what the HBR prescription &#8212; redesign roles, budgets, and processes &#8212; looks like when the redesign is translated into a firm. Not a firm-wide redesign. A federated one, in which the firm-level work is to set guardrails and standardize the verification, and the group-level work is to choose the use cases.</p><p>The Task Audit is the artifact. The honest version takes a quarter per group and produces something a partner can defend in a comp committee. The dishonest version takes a week, looks like a deck, and ages badly.</p><h3><strong>Move 3 &#8212; A single accountable owner per practice group, not a committee</strong></h3><p>Committees defer. Owners decide. McKinsey&#8217;s number &#8212; one in six organizations without a clear C-level AI owner &#8212; is worse in legal, because even where an owner exists at the firm level, the partnership structure dilutes accountability. The committee is the partnership&#8217;s default response to a contested decision, and AI adoption is a contested decision.</p><p>What works is a named partner in each practice group who owns AI decisions for that group. Not a committee chair. Not a project sponsor. Not a steering-group member. An owner &#8212; someone whose performance review includes a line item for the group&#8217;s AI capability development, and whose decisions do not require a partnership-wide vote to take effect inside the group. The owner reports up to a chief AI officer or equivalent at the firm level, but the accountability lives at the group level because that is where the work lives.</p><p>This configuration is unfashionable in firms that prefer consensus. It is also the only configuration that produces decisions on the timeline AI requires. The alternative is a steering committee that meets monthly, defers two of the three decisions on the agenda, and ratifies the decision a partner already made between meetings.</p><p>The named-owner model has a useful side effect. Partners who own decisions become accountable for outcomes. Partners on committees do not. Functional AI capability over the next three years will sit with firms that staffed for ownership early, not with firms that staffed for governance theater.</p><h3><strong>Move 4 &#8212; Redesign the apprenticeship deliberately</strong></h3><p>The apprenticeship gap does not close by accident. It closes by design &#8212; or it does not close.</p><p>The redesign question is not how the firm automates junior work. The redesign question is what juniors do instead, such that they emerge with the judgment senior practice requires. The automation question is the easy one. The redesign question is where the actual capacity sits.</p><p>What the redesign looks like in practice is structural. A second-year associate who used to spend a third of their hours on first-pass review now spends a third of their hours on something else. The &#8220;something else&#8221; has to do for the second-year&#8217;s development what the first-pass review used to do. Possibilities exist. Supervised secondary review, where the associate critiques the AI&#8217;s output and learns to spot what a senior would spot. Structured client-interaction time, where the associate develops the judgment that is hardest to automate. Deliberate cross-practice exposure, where the associate sees how a deal partner and a litigator weight the same fact pattern differently. None of these is automatic. All of them require partner time, which is the scarcest resource in a firm.</p><p>The cost of the redesign is real. The cost of not redesigning is realized in five years, when the firm tries to promote a class of senior associates and discovers their calibration is not where it needs to be. The redesign work flows into Section 4, where role redefinition becomes the unit of analysis.</p><h3><strong>Move 5 &#8212; Address pricing before clients force the conversation</strong></h3><p>The general counsel office is going to ask. Better to have an answer than to be asked. Better to have proposed the answer than to be in defensive negotiation when the question arrives, because the question will not arrive politely.</p><p>Fixed fees. Success fees. Capacity arrangements. AI-augmented hourly with a productivity discount. The menu exists. The choice is which structure fits which matter type and which client. The choice is also which structure preserves margin while signaling that the firm has thought seriously about the productivity assumption clients are making.</p><p>Pricing is not a finance department problem. Pricing is a partnership problem, because the realization-rate consequences of every pricing experiment land on individual partner P&amp;Ls. The right pattern is to run experiments inside specific practice groups, with specific clients, on specific matter types, and harvest the data before the pricing committee tries to set a firm-wide policy. Firm-wide policy on pricing arrives last, after the experiments.</p><p>This is uncomfortable. It is also where the strategic self-direction question gets answered. Firms that lead the pricing conversation define the market. Firms that wait inherit a market other firms defined. The Wolters Kluwer survey found 54% of lawyers expect firms to use AI efficiency for more clients or competitive pricing &#8212; the client side has already done the math. The question for the firm is whether the math gets done in the partnership&#8217;s frame or in the client&#8217;s frame.</p><h3><strong>Move 6 &#8212; Measure capability development, not hours saved</strong></h3><p>Hours saved is the metric that produced those software-industry cuts. It is the wrong metric in legal, because the answer to &#8220;what did we do with the saved hours&#8221; cannot be &#8220;we cut the people.&#8221; The model does not work if the people are gone. The whole apprenticeship hazard, the whole judgment-hollowing hazard, every reason a firm has a future at all is grounded in the people being there to develop into the next generation of senior practitioners.</p><p>What to measure instead lives at the capability layer. Practice-group adoption depth &#8212; how many lawyers in the group are using AI for substantive work, not for one-off email drafts. Output calibration scores against blinded review. Role evolution &#8212; whether roles are changing in the direction the firm intends, or drifting. Time reinvested &#8212; when a task that took ten hours takes three, where do the seven hours go, and what shows up at the end of the quarter that did not exist at the beginning.</p><p>The five-to-one ratio from the McKinsey advisor surfaces at this move. Five-to-one is not a budget rule. It is a measurement principle. If the dollar count on the people side is small, the measurement on the people side will be small, and the capability the people side is supposed to build will not develop. A firm that spends one dollar on tools and twenty cents on capability development is running a one-to-five program with extra slides, not a five-to-one program.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://andrewlewis.ca/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Andrew Lewis was Here is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2><strong>Role Redefinition Through Live Operating Questions</strong></h2><p>The Production Layer cannot be installed. It has to be developed.</p><p>Savine&#8217;s prescription assumes installation &#8212; build the layer, fund it five-to-one, install the agentic control plane, run the pipeline. Inside a firm, the layer is not infrastructure. The layer is a set of capabilities held inside a set of roles that are changing shape faster than HR can document them. The Production Layer in a law firm looks less like a new platform team and more like existing roles being re-aimed at capability development, with new measurement underneath.</p><p>The work is still early. That matters. If the last mile were mainly a deployment problem, the evidence inside a firm would appear as tool rollouts, usage charts, and completed implementation plans. The evidence looks different inside the work. It appears as roles changing shape before the metrics are settled, and as operating questions the organization has to learn how to answer before outcomes are clean enough to report. Here&#8217;s what my team is doing to build the Production Layer in the firm I work in, and what the operating questions look like in each case:</p><p><strong>Conversation sophistication scoring.</strong> We are building scoring metrics across LLM conversations in the firm&#8217;s AI products. The unresolved question is not whether conversations can be scored &#8212; they can. The hard question is which scores should change what happens next. A metric that does not connect to an intervention becomes dashboard ornamentation. A useful metric tells the organization where users need better scaffolding, where a workflow is too vague, where output calibration is weak, or where a product is encouraging shallow prompting. Someone has to decide what better conversations mean, what action follows the signal, and who owns the intervention. That deciding is role design, not analytics.</p><p><strong>Technology Training becomes Learning Enablement.</strong> We&#8217;re moving from technology training to learning enablement. The old mandate was organized around software instruction. The new mandate is organized around business outcomes, particularly where AI software changes the work itself. The work moves from showing people how to use the tool to helping groups develop the capacity to get a better business result from the tool &#8212; which includes adoption, workflow fit, practice-group context, measurement, and reinforcement after the formal training ends. The Production Layer in this case looks less like new infrastructure and more like an existing team with a redefined mandate and a different measurement underneath.</p><p><strong>Merging Business Intelligence, AI Enablement, and Application Development.</strong> We are reconfiguring a programming leadership role into an Engineering Manager with a portfolio across BI, AI Enablement, and Applications. The team is small. The mandate is to create an AI-native development process &#8212; which is not the same as a faster ticket pipeline. Artifacts are cheap. Judgment and systems are the hard part. The role is not only about producing more code or more applications. It is about building a system where AI changes the development process without hollowing out architecture, review, accountability, or product judgment. The hazard inside the role is the same hazard inside the firm: surface productivity gain at the cost of underlying capability.</p><p>These are signs of the terrain, not success stories yet. In each case, the software is the easy part to name. The hard part is deciding what new capacity the organization needs, which role owns it, and how the firm will know whether that capacity is improving. The Production Layer thesis points at this. It cannot specify it. The specification has to come from inside the practice, inside the roles, and inside the unresolved questions that appear before outcomes are clean enough to report.</p><div><hr></div><h2><strong>Closing</strong></h2><p>The map drawn here is real. The hazards are observable inside the work. The path is operational. None of that does the work.</p><p>Generic diagnoses produce generic responses. The McKinsey numbers will be quoted at every legal industry conference for the next eighteen months. The HBR last-mile language will become a slide-deck staple, then a panel topic, then a vendor positioning line. None of that will move a single firm forward, because none of it engages with the structural features that determine whether the firm has a future on the other side of the curve.</p><p>What moves a firm forward is the development work that happens at the practice-group level, in the pricing conversation, in the apprenticeship redesign, in the named partner who decides to own a decision instead of distributing it across a committee. The Production Layer is not a build. It is a development arc. The arc takes years and runs through individual roles whose mandates are changing faster than the org chart can document them.</p><p>The five-to-one ratio at the top of this piece is the right principle. It is not the right plan. The plan has to be built inside a partnership, against billable-hour math the math does not want to give back, inside a regulatory landscape that does not allow shortcuts, and with case law and malpractice carriers ratcheting up the cost of being wrong while the survey data ratchets up the expectation of efficiency. The work is not optional. The pace of the work is not negotiable. The map is the first artifact. The path is the second. Neither is the work.</p><p>AI does not make a law firm better. It exposes whether the firm was already doing the work.</p>]]></content:encoded></item><item><title><![CDATA[When the Words Don’t Exist Yet]]></title><description><![CDATA[Why AI governance is harder than communication advice can explain.]]></description><link>https://andrewlewis.ca/p/when-the-words-dont-exist-yet</link><guid isPermaLink="false">https://andrewlewis.ca/p/when-the-words-dont-exist-yet</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Mon, 11 May 2026 12:02:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!GjiZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GjiZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GjiZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png 424w, https://substackcdn.com/image/fetch/$s_!GjiZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png 848w, https://substackcdn.com/image/fetch/$s_!GjiZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png 1272w, https://substackcdn.com/image/fetch/$s_!GjiZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GjiZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png" width="1456" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2313653,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/196967510?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GjiZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png 424w, https://substackcdn.com/image/fetch/$s_!GjiZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png 848w, https://substackcdn.com/image/fetch/$s_!GjiZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png 1272w, https://substackcdn.com/image/fetch/$s_!GjiZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5287a51b-5947-4edd-a259-e77d210251fa_1983x793.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In the late Republic, Cicero set himself the project of translating Greek philosophy into Latin. He set out to render Plato and the Stoics in his own language, and he discovered fairly quickly that Latin could not hold the concepts he was trying to import. The vocabulary did not exist. In the <em>Academica</em> he wrote &#8212; half-anxiously, half-ambitiously &#8212; that he had been forced to <em>manufacture</em> the words. He coined <em>qualitas</em> to render Plato&#8217;s <em>poiotes</em>. He coined <em>moralis</em> from <em>mos</em>, the Latin for custom, to render the Greek <em>ethikos</em>. To these he added <em>evidentia</em>, <em>humanitas</em>, <em>quantitas</em>, <em>essentia</em> &#8212; by modern accounts, more than two hundred Latin equivalents of Greek philosophical and rhetorical terms, many of which survive in English as <em>quality, moral, individual, vacuum, property, definition, infinity, science</em>. The Romans got their philosophical vocabulary because one man manufactured it in Latin, sentence by sentence, while writing about it.</p><p>Cicero did not describe his work as translation. He described it as construction.</p><p>Most modern advice for people working at the seam between technical teams and business leadership does not seem to know that Cicero had this problem. The advice treats the cognitive load as a communication issue. Slow down. Find better metaphors. Bring people along. The advice is everywhere, and it isn&#8217;t wrong. It is also about two thousand years late, and it misreads the actual work.</p><p>The work of sitting between two worlds is not principally about explaining one to the other. It is about building the vocabulary that lets either side describe the territory in the first place. The cognitive load that comes with the work &#8212; the exhaustion, the reaching for words that aren&#8217;t quite right, the tedious circumlocution because the precise term hasn&#8217;t been settled yet &#8212; is not evidence of a failure to communicate. It is evidence of construction underway.</p><p>The Romans understood the position well enough to give it a god. Janus, two-faced, looked simultaneously at the past and the future, and in some traditions at the inside and the outside, the known and the unknown. He presided over thresholds and transitions. The month of January is his. Janus was not a metaphor for the experience of holding two views at once &#8212; he was an acknowledged cosmological role, important enough that the calendar opened with him. The Romans accepted that someone had to sit at the threshold between worlds, and they gave that someone a name.</p><p>They did the same for the institutional version of the role. The Latin word <em>pontifex</em> &#8212; the title borne by Roman priests, including the Pontifex Maximus &#8212; literally means <em>bridge-builder</em>. <em>Pons</em> (bridge) and <em>facere</em> (to make). The role had a name, an office, and a temple. The state recognized that the seam between worlds was load-bearing, and it appointed someone to carry it. The Romans were precise about this, and they were not romantic about it. They simply understood that a bridge does not build itself, and the person who builds it deserves an institutional title.</p><p>Move the picture forward two thousand years. AI governance, in any institution serious about AI, sits in the same position Cicero sat in. Two domains that do not natively share a vocabulary.</p><p>The technical world has its terms: model weights, retrieval-augmented generation, zero-data retention, prompt injection, agentic loops, fine-tuning, MCP servers, data residency. Each one carries a long technical tail and assumes a body of operational knowledge to be useful.</p><p>The legal and business world has its terms: privilege, fiduciary duty, materiality, retention obligations, regulatory exposure, duty of competence, professional secrecy. Each one carries its own long tail, and it assumes a different body of operational knowledge.</p><p>Neither vocabulary natively contains the other&#8217;s concepts. There is no Latin for <em>poiotes</em> and no business term for <em>prompt injection</em>. The person doing AI governance is, structurally, doing what Cicero did. They are manufacturing the language that one side can use to describe the other. The phrase <em>AI governance</em> itself is a coinage of the last few years, and it isn&#8217;t fully settled &#8212; different institutions mean different things by it. The vocabulary is being built, in real time, by the people occupying the position.</p><p>This is why communication advice misses the work. <em>Communicate better</em> assumes the words already exist and the speaker is choosing the wrong ones. The actual problem is more like Cicero&#8217;s. The words don&#8217;t exist yet. The practitioner is coining them. They are testing whether <em>guardrail</em> lands, whether <em>AI risk</em> lands, whether <em>agentic system</em> lands, whether the legal team will accept <em>prompt-level controls</em> as a meaningful category. They are doing what Cicero was doing in the <em>Academica</em>, and they are doing it under board-level scrutiny rather than over a quiet correspondence with Atticus.</p><h2><strong>The cost is structural and ancient</strong></h2><p>The cognitive load that comes with the position is not a personal weakness, and it is not unique to the modern moment. Cicero complained about it openly. In his letters to Atticus he expresses anxiety about whether Latin can carry the work, whether his coinages will hold, whether the Romans will accept words that did not exist yesterday. Cicero&#8217;s anxiety in those letters was diagnostic. He was naming the structural reality of vocabulary construction, and the cost it produced. The work was ambitious, and the ambition came with a cost the work itself created.</p><p>What changes when the position is understood this way is mostly the optimization function. The work is generative work. It produces vocabulary that did not exist before, and that vocabulary does work long after any given meeting ends. Cicero&#8217;s <em>qualitas</em> persisted for two thousand years. Boards in the present moment are not going to remember any individual conversation about AI governance, but the categories that get coined now &#8212; the names for risks, controls, accountability structures, evaluation methods &#8212; those will persist. The careful manufacture of the right word is the most durable form of leadership available in liminal positions.</p><p>That recognition reframes what the work actually is. The temptation in the position is to optimize for fluency &#8212; for speed, for clarity, for the smooth handoff between domains. Fluency is useful, but it is downstream of the actual work. Upstream is the vocabulary. The careful, sometimes tedious, often unfinished construction of words that one side can use to describe the other accurately. <em>Pontifex</em> labour, in the original sense.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://andrewlewis.ca/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Andrew Lewis was Here is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><strong>What we have not given the role</strong></h2><p>The Romans did one thing better than the present moment is doing. They named the position and they put it inside an institutional structure. <em>Pontifex</em> was an office. Janus was on the calendar. The role of bridge-builder between worlds was understood to be public, recognized, and load-bearing.</p><p>The same recognition has not been given to AI governance. The people doing the work hold a hundred different titles, most of them generic &#8212; head of innovation, AI lead, director of emerging technology, occasionally something like <em>responsible AI</em>. None of these names the actual work, which is closer to <em>vocabulary engineer</em> or <em>bridge-builder for a new domain</em>. The cost of the absent name is the same cost Cicero observed in his letters. When the role is unnamed, the work is invisible. The load gets interpreted as personal weakness, and the standard advice &#8212; given to invisible work that wasn&#8217;t supposed to look like work in the first place &#8212; is fluency. Fluency does not build the bridge.</p><p>The systemic leadership question &#8212; how do you bring others with you &#8212; has an older answer than the modern leadership literature gives. You do not bring people with you by communicating better. You bring them with you by manufacturing the language that lets either side describe the territory you stand on. Cicero did it for philosophy. The Romans built the office for it. The work is not new. The institution has not yet caught up to the work.</p><p>That is the work behind the work in liminal positions. The slow construction of the words that don&#8217;t exist yet.</p><div><hr></div><p><em>The Work Behind the Work is the umbrella under which all of this thinking lives &#8212; six interdependent capability questions that AI structurally cannot answer for you. Subscribe at andrewlewis.ca.</em></p>]]></content:encoded></item><item><title><![CDATA[The rising tide doesn’t lift all boats]]></title><description><![CDATA[What McKinsey&#8217;s 2025 data actually says about AI value capture, why the playbook executives are following contradicts their own findings, and what to read instead.]]></description><link>https://andrewlewis.ca/p/the-rising-tide-doesnt-lift-all-boats</link><guid isPermaLink="false">https://andrewlewis.ca/p/the-rising-tide-doesnt-lift-all-boats</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Thu, 30 Apr 2026 12:01:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!mjH-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mjH-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mjH-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mjH-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mjH-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mjH-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mjH-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2976092,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/195477773?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mjH-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mjH-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mjH-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mjH-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81a162ab-e796-4e27-b85b-614fa75f6ff8_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Eighty-eight percent of organizations now use AI in at least one business function. Six percent are getting measurable financial impact from it.</p><p>Both numbers come from the same source: <a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai">McKinsey&#8217;s </a><em><a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai">State of AI in 2025</a></em>, published last November and based on responses from 1,993 organizations across 105 countries. The first number is in the headline. The second is in the methodology, where &#8220;high performers&#8221; &#8212; defined as the 6% reporting at least 5% EBIT impact from AI use &#8212; are described as a small but meaningful cohort.</p><p>The gap between those two numbers is the entire story.</p><p>What you do with that gap depends on whose framing you accept. McKinsey&#8217;s framing is that the 6% are pulling ahead because they are committing more aggressively to the playbook: pursuing what the report calls <em>transformative innovation</em>, redesigning workflows, scaling faster, investing more. The implication for the other 94% is straightforward: do more of the same thing, harder.</p><p>That is the firm selling transformation engagements explaining why most transformation engagements aren&#8217;t working.</p><p>This essay is about why that framing is wrong, why the data inside the same report points at a different answer, and what executives should actually do with the gap. It draws on Shakespeare, Sutton, three years of operating AI inside a regulated industry, and McKinsey&#8217;s own contradictory evidence &#8212; and ends with a diagnostic question every executive can answer for themselves.</p><h2><strong>The half of the metaphor most people have stopped reading</strong></h2><p>Most executives carry around a half-version of an old metaphor: a rising tide lifts all boats. The implication is that benefit is automatic. Buy the licenses. Enable the seats. The water does the work.</p><p>There is a more honest version of the same metaphor in Shakespeare. Brutus, in <em>Julius Caesar</em>, says &#8220;there is a tide in the affairs of men, which, taken at the flood, leads on to fortune.&#8221; Most quotations stop there. The next line is the part that should keep executives up at night: &#8220;Omitted, all the voyage of their life is bound in shallows and in miseries.&#8221;</p><p>The tide doesn&#8217;t lift everyone. It lifts the ones who are ready to take it. The rest get stranded.</p><p>This isn&#8217;t a stylistic distinction. It&#8217;s a structural one. A boat with holes in the hull doesn&#8217;t rise with the water; it fills up and goes down faster. As the tide of AI capability rises &#8212; and it is rising, on a curve that doesn&#8217;t depend on anyone&#8217;s organizational readiness &#8212; the gap between seaworthy organizations and unseaworthy ones widens, not narrows.</p><p>Standing still is not neutral either. Competitor vessels are getting ready. The water is going up regardless of what any single executive does.</p><p>That is what the 88-vs-6 gap actually describes. Not a gradient of effort or commitment. A gradient of seaworthiness.</p><p>The vantage point matters here. Working inside a regulated industry sharpens this. When you cannot just deploy and learn &#8212; when data residency, regulator scrutiny, and partner trust constrain every decision &#8212; the gap between buying capability and capturing value becomes uncomfortably visible. You watch teams roll out impressive-looking pilots that produce no measurable change in how cases get worked or matters get billed. You watch licensing spend that nobody can defend in front of a finance committee. You watch the same governance committee revisit the same policy questions for the third quarter in a row because nobody wants their signature on the approval. Regulated industries don&#8217;t get to skip the operational work. The unregulated ones don&#8217;t either; they just discover it later, after the licensing spend is already sunk.</p><h2><strong>Why this looks like Sutton&#8217;s Bitter Lesson, and isn&#8217;t</strong></h2><p>Anyone watching the AI capability curve from a distance can recognize the trajectory. It tracks a pattern Richard Sutton named in 2019 in <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html">an essay called The Bitter Lesson</a>. His argument was that across seventy years of AI research, general methods that scale with computation have eventually beaten approaches built on human knowledge of the domain. Chess. Go. Speech recognition. Computer vision. The pattern repeats: human-knowledge-based methods feel principled and produce early wins, then get overtaken by general methods plus more compute. Sutton&#8217;s claim has aged exceptionally well. Large language models are the most emphatic confirmation of his thesis since he wrote it.</p><p>Read carelessly, the Bitter Lesson is the case for executive optimism. The tide is rising because compute keeps growing. General methods keep absorbing what humans used to scaffold. The eventual end state is models that handle the integration work themselves. Just wait. Just deploy. The capability will arrive; the value will follow.</p><p>That reading is wrong, in a specific way.</p><p>Sutton was making a claim about model capability and training paradigms &#8212; what wins benchmarks at the research frontier. He was not making a claim about what creates value in deployed enterprise systems. The frontier of value creation in 2026 is not the model layer. It is the layer around the model: governance, retrieval, evaluation, workflow integration, accountability, measurement. Almost all of that is what Sutton called &#8220;human knowledge&#8221; &#8212; exactly the kind of scaffolding the Bitter Lesson predicts will eventually be absorbed. But it has not been absorbed yet, and won&#8217;t be for a while, because it is not benchmark-shaped. It is context-shaped, organization-shaped, accountability-shaped. There is no test set for &#8220;did the matter close faster&#8221; or &#8220;would the partner sign off on this advice.&#8221;</p><p>The honest synthesis is that Sutton&#8217;s lesson and the operational discipline argument live at different layers. Sutton describes the trajectory of model capability, which is rising reliably. The operational discipline argument describes what happens between that rising capability and the outcomes a board recognizes. That is exactly where the 94% are stuck. Both observations are true; they are not in competition. The executive who reads the Bitter Lesson as license to wait is making a category error, confusing what wins research benchmarks with what compounds in P&amp;L.</p><h2><strong>What McKinsey&#8217;s report actually shows, and what it conveniently omits</strong></h2><p>Read the report carefully and a different argument emerges than the one in the executive summary.</p><p>McKinsey&#8217;s own correlation analysis tested twenty-five organizational attributes against EBIT impact from generative AI. The single biggest predictor was workflow redesign. Not licensing volume. Not tooling sophistication. Not training programs. Operational rewiring.</p><p>That finding is in their data. It is not in their headline recommendation. The headline recommendation is to <em>pursue transformative innovation</em> &#8212; a phrase that maps conveniently to a multi-year consulting engagement and not to anything an internal team could execute on its own. Their data identifies operational rewiring as the binding constraint. Their recommendation identifies engagement scope as the answer. Those are different things.</p><p>There is another revealing detail. The report identifies twelve &#8220;best practices for gen AI adoption and scaling.&#8221; Read them in sequence and notice what they describe: establishing a dedicated transformation office, creating a comprehensive change story, developing role-based capability training programs, defining a phased rollout roadmap, establishing employee incentives, mechanisms to incorporate feedback, fostering trust through structured communications. Each of these is a defensible activity. Together, they describe the typical scope of a transformation consulting engagement almost exactly. The list is not wrong. It is just suspiciously shaped.</p><p>The high-performer framing has its own problem: it is partly tautological. McKinsey defines high performers as the 6% getting EBIT impact, then observes that they redesign workflows and pursue transformation. That doesn&#8217;t establish causation. It is equally consistent with a different reading: organizations capable of producing real EBIT impact are also the ones with the institutional discipline to redesign workflows. The methodology and the outcome correlate. The report assumes the first causes the second.</p><p>There is also a quiet methodological choice worth surfacing. The EBIT impact figures are self-reported, by respondents who are typically the same executives sponsoring the AI initiatives in their organizations. There is no independent verification. People who have sponsored a major program and put their name on it are not neutral evaluators of whether the program worked. The 6% number may be generous; the actual share of organizations capturing real, board-defensible value from AI is probably lower.</p><p>This matters because the prescription most executives are following is built on these unexamined assumptions.</p><p>None of this is dishonest. Consulting firms make money on transformation programs. Their reports are not neutral observations of the AI landscape. They are demand generation for the engagements they sell. That is not a conspiracy theory. It is how the business model works. Executives know this about their own software vendors and somehow forget it about their consulting reports. The reports use the language of independent research and arrive accompanied by the prestige of the firm; both produce a halo that suspends ordinary critical reading.</p><p>The cleanest read of McKinsey&#8217;s own 2025 report is this: their data agrees that operational rewiring is the binding constraint on AI value capture. Their recommendation does not.</p><h2><strong>What seaworthy actually requires</strong></h2><p>A vessel that can take the tide needs two things, and you need both.</p><p>The first is the hull &#8212; operational readiness. Governance that can move at the speed of capability change. Procurement processes that start with a defined business problem rather than competitive anxiety. Data infrastructure that doesn&#8217;t break when the workload shifts. Measurement frameworks that distinguish adoption from value. These are the structural conditions that determine whether AI capability, once deployed, produces outcomes or evaporates into pilot purgatory.</p><p>The second is the crew &#8212; capable people. Practitioners who can see their own work structurally and identify where AI fits. People who can evaluate AI output with calibrated judgment rather than blind acceptance or reflexive rejection. Leaders who direct capability toward operational outcomes rather than activities.</p><p>Hull without crew is infrastructure waiting for capability nobody can direct. Crew without hull is capable individuals trapped in an unready organization. Most programs invest heavily in one and call it complete.</p><p>The version of this most executives miss is that both are operational disciplines. Hull readiness is obviously operational: governance, processes, infrastructure. But crew readiness is also operational. It is not produced by lunch-and-learns or tip sheets. It is produced by changing how work is decomposed, evaluated, and assigned. The training-program version of crew development is theatre. The work-redesign version is the actual thing.</p><p>What this looks like in practice is unglamorous. It looks like sitting with a team for an afternoon and watching them do their actual work, then identifying the four steps in their process that AI could change and the two steps it absolutely shouldn&#8217;t touch. It looks like building a measurement framework that distinguishes &#8220;we used the tool&#8221; from &#8220;the matter closed faster&#8221; or &#8220;the output was higher quality.&#8221; It looks like an executive willing to put their name on a governance position that will need to be defended six months from now when someone challenges it. None of these activities photograph well in a board deck. All of them produce more EBIT impact than another round of training rollouts.</p><p>The diagnostic question is which one is your binding constraint right now. A firm with strong governance and weak practitioner capability has a different next move than one with sophisticated practitioners trapped in a governance vacuum.</p><h2></h2><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://andrewlewis.ca/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://andrewlewis.ca/subscribe?"><span>Subscribe now</span></a></p><h2><strong>The cascade between friction points</strong></h2><p>The friction points reinforce each other in a cascade that is worth naming, because most executive interventions miss the dynamic and treat each as independent.</p><p>Governance friction comes first because it sets the conditions for everything that follows. When nobody will sign the manifest &#8212; when policy decisions get distributed until no individual is accountable &#8212; the organization cannot articulate what it is willing to do with AI. That ambiguity perpetuates cultural friction. In the absence of clear positions, the rational behavior is to scrutinize every proposed initiative for flaws rather than commit to one. Why would a manager stake their reputation on a course that has not been institutionally sanctioned?</p><p>Cultural friction in turn shapes technical friction. When the organization rewards skepticism and penalizes commitment, procurement defaults to broad, defensive purchases. Tools get bought because peers bought them. Licenses get distributed widely so no single deployment can fail visibly. The technical environment that emerges is wide and shallow rather than narrow and deep, optimized for political safety rather than operational impact.</p><p>Wide-and-shallow technical environments produce workflow friction inevitably. People have AI tools available but no specific operational problem the tools were procured to solve, no clear quality framework for output, and no model for what good integration looks like. They use the tools at the margins of their work &#8212; rewriting an email, summarizing a document &#8212; and never integrate them into the core processes that would actually produce measurable outcomes. The capable AI tools that do exist get used at a fraction of their potential because the users have not done the work of understanding their own workflow well enough to know where AI belongs.</p><p>The cascade is reinforcing, but it is not strictly linear. You can have workflow friction without governance friction. You can resolve technical friction while cultural friction persists. The point is that addressing one friction point in isolation often reveals how much the others have been compensating for it. Organizations that address only the technical layer &#8212; buying better tools &#8212; discover that the governance and cultural blockages were the real constraints. Organizations that try to address culture without resolving governance discover they are asking people to commit in an environment where commitment carries no institutional support.</p><p>Identifying which friction point is dominant is the executive&#8217;s actual job. Not all four. The one that is the binding constraint right now.</p><h2><strong>The asymmetry has inverted</strong></h2><p>For most of the last three years, executive avoidance was the rational position. Committing to a major AI program carried real personal risk: an underperforming initiative would carry the executive&#8217;s name. Avoiding carried no symmetric cost. Boards were not asking pointed questions. There was no peer benchmark to be measured against. The water was rising slowly enough that being late looked safe.</p><p>That has changed. Boards now ask comparative questions. Industry surveys publish utilization data. McKinsey&#8217;s own report turns the visibility up further &#8212; every executive whose firm sits in the 94% now has a number describing their position that their board has access to. The cost of being late has become measurable, and it is compounding quarter by quarter.</p><p>There is a second-order effect worth naming. The compounding gap between high performers and the rest is not just about EBIT impact. It is about the institutional learning that comes with operating AI seriously over time. Firms that have spent two years building governance muscle, measurement frameworks, and integrated workflows are not just ahead in adoption. They have built organizational capability that the late-arriving firm cannot replicate by deploying the same tools faster. The advantage is structural, not procurement-driven, and structural advantages compound on a different curve than catch-up effort can match.</p><p>The executive who commits to operational readiness now is making the lower-risk bet. The one waiting for clarity is making the bet that used to be safe and no longer is. Most have not repriced.</p><h2><strong>Where to start</strong></h2><p>The diagnostic question is straightforward: which friction point is your binding constraint right now?</p><p>If governance is the binding constraint, no investment in capability development will produce outcome. The starting move is identifying who is willing to own a clear position on data handling, acceptable use, and risk tolerance &#8212; and accepting that the position may need to be defended later. The friction does not dissolve when risk is eliminated. It dissolves when accountability is claimed.</p><p>If culture is the binding constraint, the work is changing the incentive structure so that commitment is rewarded and avoidance carries cost. This is the hardest of the four because it requires changing what the organization values. Most cultural interventions fail because they try to change rhetoric without changing reward; people read the actual signals.</p><p>If your binding constraint is technical &#8212; tools acquired because competitors acquired them, with no specific business problem they were procured to solve &#8212; the move is to start over with success criteria first. Identify the operational problem, define measurable success, then select technology against those criteria. The tools you have may or may not survive the analysis. Sunk cost is not a strategy.</p><p>If workflow is the binding constraint &#8212; your governance is reasonable, your culture is permissive, your tools are deployed, but nothing has changed in how work actually gets done &#8212; the gap is structural understanding of work itself. Practitioners need to see their own workflows as systems of operations rather than streams of tasks. That capability is not built by tip sheets. It is built by sitting with the work.</p><p>Most executives can answer the diagnostic question honestly if asked plainly. The hard part is not the diagnosis. It is committing to the answer.</p><h2><strong>The line worth remembering</strong></h2><p>Read McKinsey&#8217;s data. Don&#8217;t read McKinsey&#8217;s recommendation. The tide is rising, the playbook you are being sold will leave you stranded, and the boats that take the tide will be the ones whose hulls were built before the water came up.</p><p>There is still time to build. Not much.</p><div><hr></div><p>The work I publish at AndrewLewisWasHere goes deeper on what AI operational readiness actually looks like inside a regulated industry, where you can&#8217;t just deploy and learn. Subscribe and the next piece will land in your inbox.</p>]]></content:encoded></item><item><title><![CDATA[The hidden second clause in the AI productivity story]]></title><description><![CDATA[A new study suggests the trade most of us think we&#8217;re making &#8212; skill for speed &#8212; often delivers neither.]]></description><link>https://andrewlewis.ca/p/the-hidden-second-clause-in-the-ai</link><guid isPermaLink="false">https://andrewlewis.ca/p/the-hidden-second-clause-in-the-ai</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Tue, 28 Apr 2026 12:02:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OzX2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OzX2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OzX2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!OzX2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!OzX2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!OzX2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OzX2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2207219,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/195450206?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OzX2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!OzX2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!OzX2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!OzX2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21dee23b-f302-482e-a71d-3627557b3e85_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Anthropic just published a study that quietly inverts how we talk about AI productivity. They paid 52 professional developers to learn a new Python library in 35 minutes. Half got an AI assistant. Half didn&#8217;t. Then everyone took the same comprehension quiz, with no AI.</p><p>The AI group scored 17% lower. About two grade points. A real effect, not noise &#8212; Cohen&#8217;s d of 0.738 if you care about the statistics.</p><p>That part has been making the rounds. The part that hasn&#8217;t is the productivity finding sitting right next to it.</p><p>On average, the AI group wasn&#8217;t faster.</p><p>People stayed about the same total time on the task. They just shifted from coding to interacting with the assistant. Some of them spent six minutes composing a single query in a 35-minute task. Which means the trade most people assume they&#8217;re making &#8212; accept some skill loss in exchange for speed &#8212; often isn&#8217;t even on offer. They lost the skill. They didn&#8217;t get the speed.</p><h2>The variance is the story</h2><p>What makes this study different from the usual productivity-of-AI-tools paper is that the researchers didn&#8217;t just measure averages. They watched every screen recording. And they found that the average is hiding something more useful than it&#8217;s showing.</p><p>There were six distinct ways people used the AI assistant. The researchers split them by quiz outcome. Three patterns scored between 24 and 39 percent on the comprehension test. Three scored between 65 and 86.</p><p>The low-scoring patterns share a structure. AI Delegation &#8212; handing the whole task over and pasting the result. Progressive Reliance &#8212; starting engaged, drifting into delegation as the clock ticks. Iterative Debugging &#8212; asking the AI to fix things you don&#8217;t understand, over and over.</p><p>The high-scoring patterns share a different structure. Conceptual Inquiry &#8212; asking the AI questions, but writing the code yourself. Hybrid Code-Explanation &#8212; asking for code with the explanation embedded. Generation-Then-Comprehension &#8212; getting the code, then asking the AI to teach you why it worked.</p><p>You can guess what divides the two halves. Cognitive engagement. Not how much AI was used. Not which tool. Not how skilled the developer was going in. Whether they stayed in the work or stepped out of it.</p><p>The participant feedback in the qualitative section is the part that stuck with me. Several developers in the AI group volunteered, unprompted, that they had &#8220;felt lazy,&#8221; or wished they had paid more attention to the explanations the AI gave them, or noticed afterward that there were &#8220;still a lot of gaps in their understanding.&#8221; Cognitive offloading from the inside. They could feel it happening in the moment and didn&#8217;t stop it, because the task pressure was telling them to keep moving.</p><h2>Why debugging is the canary</h2><p>The biggest gap between the groups wasn&#8217;t on the conceptual questions. It was on debugging.</p><p>The no-AI group hit errors. Trio errors specifically &#8212; runtime warnings about coroutines that were never awaited, type errors from passing the wrong kind of object. These errors force you to learn how the library actually works, because you can&#8217;t skip past them without forming a mental model. The AI group skipped past most of them. Their code worked the first time, more often than not. They never built the debugging muscle, because they never had to.</p><p>This is the part that should make organizations nervous.</p><p>The dominant workflow proposal for AI-assisted software development is &#8220;AI writes the code, humans review it.&#8221; Sometimes packaged as &#8220;human in the loop.&#8221; It depends on a workforce that can read code well enough to catch what&#8217;s wrong with it. But if AI assistance during the formation period systematically removes the error-encounter loop that builds review skill, you&#8217;re producing a workforce structurally unable to do the review you&#8217;ve designed your safety story around.</p><p>That isn&#8217;t a future risk. It&#8217;s a workflow currently being deployed.</p><h2>The hidden second clause</h2><p>The dominant claim about AI productivity has been: AI makes you faster. The fine print, based on this evidence, is &#8212; only if you stay cognitively engaged. And the empirical pattern says most people don&#8217;t.</p><p>There&#8217;s a clean way to think about this. AI doesn&#8217;t make you better. It amplifies whether you were doing the thinking in the first place.</p><p>For people who already engage with their work &#8212; who ask why something is structured the way it is, who validate against a model of what good looks like, who treat tools as collaborators rather than dispensers &#8212; AI extends what they can do. For people who don&#8217;t, AI gives them faster ways to look productive while the underlying skill quietly erodes.</p><p>This isn&#8217;t a moral observation. It&#8217;s a structural one. The patterns that preserved learning in the study weren&#8217;t the patterns of people working harder. They were the patterns of people staying in a particular mode of attention.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://andrewlewis.ca/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://andrewlewis.ca/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>What this changes</h2><p>The implications break in three directions.</p><p>For individuals: how you use the tool matters more than which tool you use. Asking for code without asking for explanation is a learning shortcut that costs you compounding interest later. The patterns that work are slower in the moment and pay off in retention.</p><p>For teams: if you&#8217;re scaling junior contributions with AI, you need to deliberately build the engagement loops back in. Errors are not friction to eliminate. They&#8217;re where skill formation actually happens. Removing them feels like productivity. It produces fragility.</p><p>For organizations: the AI workflow you&#8217;re designing assumes a level of human capability you may be actively eroding through the same workflow. That&#8217;s not a contradiction you can prompt your way out of. It&#8217;s an architectural one.</p><p>The original framing &#8212; that AI is a productivity tool &#8212; was always slightly off. AI extends the cognitive work you&#8217;re already doing. Without that work underneath, the extension produces motion without progress.</p><p>AI doesn&#8217;t make you better. It exposes whether you were already doing the work.</p><div><hr></div><p>Read the full Anthropic study: <a href="https://arxiv.org/abs/2601.20245">How AI Impacts Skill Formation</a> by Judy Hanwen Shen and Alex Tamkin.</p><p>If this resonated, subscribe for more notes from inside a regulated industry trying to operationalize AI without breaking what makes it work.</p>]]></content:encoded></item><item><title><![CDATA[Your Brain Is a Judgment Machine Now]]></title><description><![CDATA[AI didn&#8217;t eliminate the hard part of knowledge work. It compressed it into every minute of the day.]]></description><link>https://andrewlewis.ca/p/your-brain-is-a-judgment-machine</link><guid isPermaLink="false">https://andrewlewis.ca/p/your-brain-is-a-judgment-machine</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Mon, 20 Apr 2026 12:02:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uSU1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uSU1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uSU1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!uSU1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!uSU1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!uSU1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uSU1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1774990,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/194358008?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uSU1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!uSU1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!uSU1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!uSU1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38256f1-7890-4ec5-a042-c115b0651bd5_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The dominant narrative about AI and work goes like this: AI handles the production, you handle the thinking, everyone goes home early. It is a clean story. It is also wrong in a way that matters.</p><p>A principal engineer at a large telecom <a href="https://www.reddit.com/r/ClaudeAI/s/0P0PHzx1aZ">posted on Reddit this week</a> about being all-in on agentic coding for two years and thinking about quitting software engineering entirely. Not because the tools don&#8217;t work. Because they work too well. The line that stopped me: <strong>&#8220;The cost of writing code in effort / time was a throttling middleware.&#8221;</strong></p><p>That phrase deserves to sit for a second. Writing code used to be slow enough that your brain could keep up. The effort of typing, compiling, debugging &#8212; all of that created natural pace. You had time to sense when a pattern was wrong. Time to think through the shape of a class or the implications of an architectural choice. The slowness wasn&#8217;t a bug. It was cognitive infrastructure.</p><p>Now that infrastructure is gone. And this engineer &#8212; 13 years of experience, principal level &#8212; reports making ten whiteboard-level architectural decisions before his second cup of coffee. Decisions that used to happen once a sprint, maybe twice, gated by the slow, expensive process of actually building things. The dam broke. The decisions didn&#8217;t disappear. They accelerated.</p><h2>The production layer was never the bottleneck you thought it was</h2><p>When people talk about AI making work faster, they&#8217;re usually describing the production layer &#8212; the part where raw effort turns into output. Drafting, coding, formatting, researching. And yes, AI compresses that layer dramatically. What nobody accounted for is what happens to the <em>other</em> layer &#8212; the judgment layer &#8212; when production speeds up by an order of magnitude.</p><p>Every piece of AI-generated output requires evaluation. Is this right? Is this good enough? Does this fit the architecture? Does this solve the actual problem or just the surface symptom? Those questions existed before AI, but they arrived at a pace your brain could absorb. Production time was thinking time. The gap between &#8220;I need this&#8221; and &#8220;here it is&#8221; gave your mind room to prepare for the decision.</p><p>That gap is gone. And the result is not faster work. It&#8217;s faster <em>judgment demands</em> on a brain that hasn&#8217;t changed speed.</p><p>A <a href="https://hbr.org/2026/03/when-using-ai-leads-to-brain-fry">Boston Consulting Group study published in Harvard Business Review</a> in March 2026 put a name to this: &#8220;AI brain fry.&#8221; They surveyed 1,488 full-time U.S. workers and found that high AI oversight &#8212; the kind that requires reading, interpreting, and evaluating AI output &#8212; was associated with 14% more mental effort, 12% greater mental fatigue, and 19% greater information overload. Workers described a fog or buzzing that forced them to physically step away from their screens. One of the study&#8217;s authors told Fortune the pattern was consistent: people were getting more done but hitting the limits of their cognitive capacity because there were simply too many decisions to make.</p><p>An eight-month study of a 200-person tech firm, led by researchers at UC Berkeley, found the same dynamic from a different angle. AI wasn&#8217;t reducing work. It was intensifying it. Employees processed more information, made more decisions, and experienced more burnout &#8212; not less &#8212; as AI adoption increased.</p><h2>Decision fatigue is not a new concept. The delivery mechanism is.</h2><p>Psychologists have studied decision fatigue for decades. The core finding is straightforward: the quality of your decisions degrades as you make more of them. Roy Baumeister&#8217;s ego depletion research established that decision-making draws from a finite cognitive resource. Make enough decisions and you start defaulting to heuristics, avoiding trade-offs, or simply deferring. The average American adult reportedly makes around 35,000 decisions a day. Most of those are trivial. The ones that matter are the ones that require actual evaluation.</p><p>What AI does is change the ratio. It doesn&#8217;t increase the total number of decisions. It increases the <em>density of consequential ones</em>. When production was slow, your day was a mix of low-stakes mechanical work and occasional high-stakes judgment calls. The mechanical work gave your brain recovery time between the hard decisions. It was boring, but it was load-bearing.</p><p>Remove the mechanical work and what&#8217;s left is a continuous stream of judgment. Architecture choices. Quality assessments. Risk evaluations. Strategic trade-offs. All day. No recovery intervals. The Reddit poster described it precisely: running ten whiteboard-level decisions before morning coffee, decisions that used to be spaced across a sprint. His brain isn&#8217;t slower than it was. The demand on it is faster.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://andrewlewis.ca/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">If you're reading this far, you've probably felt this yourself. I write about this every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>We don&#8217;t understand the cost of running a judgment machine all day</h2><p>This is the part almost nobody is talking about. We&#8217;ve spent two years celebrating AI&#8217;s ability to remove the production burden. We have not spent two minutes thinking about what happens to the human on the other side when the production burden was also a cognitive pacing mechanism.</p><p>The Reddit engineer said something else that stuck: &#8220;I feel like for the devs that have survived layoff rounds, AI has <em>raised</em> the bar of required skills, not lowered it.&#8221; That maps directly to the Jevons Paradox applied to AI &#8212; as AI efficiency increases, the demand for human capability doesn&#8217;t decrease. It increases. The skills that matter shift upward. The judgment, the architectural thinking, the ability to evaluate quality at speed &#8212; those become the job. And the job becomes relentlessly, uninterruptedly hard.</p><p>This isn&#8217;t a coding problem. It&#8217;s a knowledge work problem. Every profession that adopts AI tools effectively will hit this same wall. Lawyers reviewing AI-drafted contracts. Financial analysts evaluating AI-generated models. Marketers assessing AI-produced campaigns. The production layer compresses. The judgment layer concentrates. And the person in the middle has to run their brain at a sustained intensity that the old workflow never required.</p><p>We don&#8217;t have infrastructure for this yet. We don&#8217;t have pacing strategies. We don&#8217;t have cognitive load frameworks adapted for AI-augmented work. We don&#8217;t even have language for the problem &#8212; which is why &#8220;brain fry&#8221; and &#8220;throttling middleware&#8221; resonate so immediately. People recognize the feeling before anyone names it.</p><h2>The work behind the work just got more urgent</h2><p>The conventional response to this problem will be training. Run a workshop on managing AI output. Distribute a tip sheet on decision prioritization. That approach will fail for the same reason it always fails &#8212; it treats the symptom without touching the structure.</p><p>The actual work is harder than a workshop. It&#8217;s developing the ability to see your own workflow clearly enough to know which judgments matter and which don&#8217;t. To calibrate your trust in AI output so you&#8217;re not re-evaluating everything at full intensity. To build the signal discrimination that lets you spot the 5% of output that needs real attention and let the rest move.</p><p>That is not a training problem. That is a capability development problem. And it&#8217;s one that gets more urgent, not less, as the tools get faster.</p><p>Your brain was always a judgment machine. AI just made it the only machine that matters.</p><div><hr></div><p>I'm writing from inside a regulated firm doing AI adoption in real time. Every week I publish what I'm seeing &#8212; the frameworks, the friction, the decisions that actually move work. If that's useful to you, subscribe. It's free, and I'll send you the next one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://andrewlewis.ca/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://andrewlewis.ca/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Your AI Adoption Number Is Lying to You]]></title><description><![CDATA[The metric everyone tracks, the one almost nobody does, and why the gap between them explains everything]]></description><link>https://andrewlewis.ca/p/your-ai-adoption-number-is-lying</link><guid isPermaLink="false">https://andrewlewis.ca/p/your-ai-adoption-number-is-lying</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Thu, 16 Apr 2026 13:11:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dNtp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dNtp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dNtp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dNtp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dNtp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dNtp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dNtp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:346480,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/193750041?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dNtp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dNtp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dNtp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dNtp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7392d3f-3c16-48a2-9345-ab6f30adf143_1024x1024.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Every enterprise AI dashboard in the world has an adoption number on it. Seats activated, prompts per user, tools deployed, percentage of workforce with access. The number goes up every quarter. The board is pleased. The CIO presents it with confidence.</p><p>And almost none of it tells you whether AI is actually working.</p><h2>The metric everyone loves is the metric that matters least</h2><p>Adoption measures activity. Someone logged in. Someone typed a prompt. Someone opened Copilot and asked it to summarize an email they could have read in thirty seconds. All of that registers as adoption.</p><p>What it doesn&#8217;t measure is whether anyone&#8217;s work changed. Whether a contract review that used to take four hours now takes ninety minutes because the lawyer iterated through three rounds of AI-assisted markup. Whether the finance team built an automated reconciliation workflow or just asked ChatGPT to explain what a pivot table does. Whether the senior associate used AI to surface a pattern across two hundred documents that no one would have caught manually &#8212; or whether they used it to rewrite a slightly awkward email.</p><p>Those are fundamentally different behaviors. One is sophisticated use. The other is expensive autocomplete. And the adoption dashboard doesn&#8217;t distinguish between them.</p><h2>The evidence arrived this week, from three directions at once</h2><p>A Wharton School study tracking enterprise AI adoption over multiple years found a widening disconnect between executive enthusiasm and managerial reality. Nearly two-thirds of executives report becoming significantly more optimistic about AI over the past year. The managers implementing those same tools inside actual workflows report something closer to frustration. They see the constraints. They carry the operational burden. They don&#8217;t feel supported.</p><p>A separate survey of 2,400 knowledge workers found that 29% admit to actively undermining their company&#8217;s AI rollout. Among Gen Z workers, that figure reaches 44%. The tactics include feeding proprietary data into unapproved tools, refusing to complete AI training, and in some cases deliberately producing poor-quality output to make the tools look bad.</p><p>And a third study, from HFS Research, found that only 14% of enterprises have a clear AI strategy at all. The other 86% are deploying tools into a vacuum &#8212; no framework for what good use looks like, no feedback loop for what&#8217;s working, no definition of success beyond the adoption dashboard.</p><p>These aren&#8217;t three separate problems. They&#8217;re three symptoms of the same one: organizations measuring the wrong thing, at the wrong level, and mistaking activity for progress.</p><h2>What sophistication actually measures</h2><p>The Conversation Sophistication Score is a framework built on research from KPMG and the University of Texas at Austin. The study analyzed 1.4 million AI interactions across 2,500 employees and identified thirty behavioral characteristics that separate the highest-performing AI users from everyone else.</p><p>The finding that matters: sophistication isn&#8217;t correlated with frequency. The people using AI most often aren&#8217;t the ones using it best. The distinguishing behaviors are things like interaction depth (multi-turn conversations that build on previous outputs), task complexity (applying AI to genuinely difficult problems rather than simple lookups), iterative reasoning (treating AI output as a draft to be refined, not an answer to be accepted), breadth of application (using AI across multiple domains rather than one narrow use case), and fluency signals (adapting communication style and prompt structure to the specific task).</p><p>None of those show up on an adoption dashboard. You can have 95% adoption and 10% sophistication, and your metrics will tell you everything is going great until it very clearly isn&#8217;t.</p><h2>The sequence that actually works</h2><p>The enterprises getting AI right aren&#8217;t doing anything exotic. They&#8217;re solving the structural problems before chasing the visible ones. They define what good AI use looks like before measuring whether it exists. They build governance &#8212; not as a compliance exercise but as a shared understanding of what&#8217;s allowed, what&#8217;s encouraged, and what&#8217;s off-limits. They create conditions where experimentation is safe and failure is data, not career risk.</p><p>And then &#8212; after the structure exists &#8212; they start measuring sophistication alongside adoption. Not instead of it. Alongside it. Because adoption without sophistication is just expensive access. And sophistication without adoption means you have a few brilliant users surrounded by a workforce that&#8217;s opted out.</p><p>The 14% of organizations that have a clear strategy? They&#8217;re the ones building this foundation. The other 86% are wondering why their adoption numbers keep climbing and their results don&#8217;t.</p><p>The dashboard isn&#8217;t broken. The measurement is.</p><div><hr></div><p><em>If you found this useful, consider sharing it with someone leading an AI rollout right now. They probably have an adoption number. They probably don&#8217;t have a sophistication score. That gap is the article.</em></p>]]></content:encoded></item><item><title><![CDATA[Two Conversations About AI, One Building]]></title><description><![CDATA[Executives and managers aren't disagreeing about AI. They're having entirely different discussions.]]></description><link>https://andrewlewis.ca/p/two-conversations-about-ai-one-building</link><guid isPermaLink="false">https://andrewlewis.ca/p/two-conversations-about-ai-one-building</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Thu, 16 Apr 2026 12:00:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Lb3M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Lb3M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Lb3M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Lb3M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Lb3M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Lb3M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Lb3M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated Image" title="Generated Image" srcset="https://substackcdn.com/image/fetch/$s_!Lb3M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Lb3M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Lb3M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Lb3M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c771a72-2553-4b79-8953-ea4edd37e9a4_1408x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Something odd is happening inside organizations right now, and it showed up clearly in a study the Wharton School published this week. The executive suite and the management layer are both talking about AI. They&#8217;re using similar language. They&#8217;re attending the same town halls and reading the same strategy decks. But they are not having the same conversation.</p><p>Nearly two-thirds of executives say they&#8217;ve become significantly more positive about AI over the past year. They see it as a strategic priority. They&#8217;re investing heavily. Some are restructuring their organizations around it. For senior leadership, AI has moved from an interesting capability to an existential commitment &#8212; the question isn&#8217;t whether to go all in, it&#8217;s how fast.</p><p>The managers one or two levels below them? They&#8217;re drowning.</p><p>The gap isn&#8217;t disagreement &#8212; it&#8217;s altitude</p><p>The Wharton/GBK Collective study has been tracking this dynamic for multiple years, and the pattern is consistent: executives experience AI as an opportunity. Managers experience AI as a workload.</p><p>That&#8217;s not because managers are resistant or uninformed. It&#8217;s because they sit at the exact altitude where strategy becomes operations. They&#8217;re the ones who have to reconcile the CEO&#8217;s enthusiasm with the fact that the approved tool doesn&#8217;t integrate with the case management system. They&#8217;re fielding questions from their teams about what&#8217;s allowed and what isn&#8217;t &#8212; often without clear answers, because the governance policy is still being drafted. They&#8217;re absorbing the productivity overhead of learning new tools while still delivering on every existing deadline.</p><p>Executives don&#8217;t see this overhead because it doesn&#8217;t appear in the metrics they review. Adoption dashboards show seats activated and usage trends. They don&#8217;t show the manager who spent three hours this week answering the same AI governance question from four different direct reports because the policy FAQ doesn&#8217;t exist yet.</p><p>The executive sees a line going up. The manager feels the weight behind that line.</p><p>Where the pressure compounds</p><p>This altitude gap produces a specific set of downstream problems that I see repeatedly from inside an organization going through this.</p><p>The first is unfunded mandates. Leadership communicates that AI adoption is a priority. But the time, training, and governance infrastructure required to adopt responsibly aren&#8217;t budgeted separately. They&#8217;re absorbed by the existing management layer on top of everything else. The implicit expectation is that managers will figure it out &#8212; will become AI champions in addition to their actual roles, without reduced workloads or additional resources.</p><p>The second is phantom consensus. Strategy decks present AI adoption as an aligned organizational priority. Everyone nods in the planning meeting. But alignment at the strategy level doesn&#8217;t mean alignment at the implementation level. The manager who nodded in the meeting goes back to a team that has no idea what&#8217;s expected, using tools that half-work, under policies that are still in draft. The strategy deck says &#8220;aligned.&#8221; The floor says &#8220;confused.&#8221;</p><p>The third is consequence asymmetry. The Fortune/Workplace Intelligence survey found that 60% of executives would consider cutting employees who refuse to adopt AI. But only 14% of enterprises have a clear AI strategy. Executives are prepared to enforce adoption of something they haven&#8217;t fully defined. The consequence falls on the people closest to the work, while the strategic ambiguity originates at the top.</p><p>What the sabotage data is actually telling us</p><p>This is where the 44% Gen Z sabotage number from this week&#8217;s headlines becomes less scandalous and more predictable.</p><p>When managers are unsupported, their teams feel it. The confusion rolls downhill. If a manager doesn&#8217;t have clear governance guidance, their team gets inconsistent answers about what&#8217;s allowed. If a manager hasn&#8217;t been given time to understand the tools, they can&#8217;t coach their team on effective use. If a manager is overwhelmed by the implementation burden, their team reads that stress and interprets it &#8212; correctly &#8212; as a signal that AI adoption is creating problems, not solving them.</p><p>The 26% of sabotaging employees who say the strategy is poorly executed aren&#8217;t making an abstract complaint. They&#8217;re reporting what they observe at the manager level every day: pressure without support, mandates without clarity, consequences without strategy.</p><p>The sabotage isn&#8217;t coming from below. It&#8217;s flowing downhill from above.</p><p>The intervention that most organizations skip</p><p>The conventional response is more training, better tools, clearer communication from leadership. Those help. But they miss the structural problem.</p><p>The structural intervention is resourcing the middle. Giving managers dedicated time for AI governance and enablement work. Reducing their operational load during the adoption period rather than adding to it. Creating feedback channels that surface implementation friction to leadership before it calcifies into resistance.</p><p>The organizations I&#8217;ve seen move fastest on AI adoption aren&#8217;t the ones with the biggest budgets or the most ambitious CEOs. They&#8217;re the ones where middle management has actual capacity to do the work that adoption requires &#8212; the unglamorous, invisible, operational work of translating executive vision into something a team of eight people can actually execute on a Tuesday afternoon.</p><p>That work doesn&#8217;t appear on an adoption dashboard. It doesn&#8217;t generate a conference keynote. But it&#8217;s the difference between a strategy that lands and one that produces a 44% sabotage rate.</p><p>The executives and the managers aren&#8217;t enemies. They&#8217;re not even disagreeing. They&#8217;re just standing at different altitudes, describing different views of the same mountain &#8212; and nobody&#8217;s built the trail between them.</p><p>---</p><p><a href="https://youtu.be/zC3WOZ3FNKY">I made a video this week breaking down the full research</a> &#8212; the sabotage data, the Wharton findings, and a framework for measuring sophistication instead of activity. </p>]]></content:encoded></item><item><title><![CDATA[Your AI Metrics Are Measuring the Wrong Thing]]></title><description><![CDATA[A research-backed framework for measuring sophistication, not just activity.]]></description><link>https://andrewlewis.ca/p/your-ai-metrics-are-measuring-the</link><guid isPermaLink="false">https://andrewlewis.ca/p/your-ai-metrics-are-measuring-the</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Wed, 08 Apr 2026 12:20:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!B9my!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B9my!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B9my!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!B9my!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!B9my!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!B9my!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B9my!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1572089,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/193348866?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B9my!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!B9my!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!B9my!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!B9my!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd297b328-56e6-43a7-a2e3-015d9abcaeed_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most organizations measure AI adoption the way they measure gym memberships. How many people signed up. How often they swipe in. Maybe how long they stay. None of which tells you whether anyone is actually getting stronger.</p><p>The AI equivalent: prompt counts, hours logged, tokens consumed, self-assessed skill levels. These numbers are easy to collect, and at most companies they look encouraging. Adoption is up. Usage is growing. The dashboards are green.</p><p>But adoption is not sophistication. And the gap between the two is enormous.</p><h2>What the Research Actually Shows</h2><p><a href="https://kpmg.com/us/en/media/news/utaustin-kpmg-study.html">KPMG and researchers at the University of Texas at Austin spent eight months studying this gap.</a> They analyzed 1.4 million AI prompts from roughly 2,500 professionals &#8212; not surveys, not self-reports, but actual conversation logs at scale. The question was simple: what does sophisticated AI use look like, and how do you tell it apart from routine use?</p><p>The headline finding: about 90% of employees used AI regularly. Only approximately 5% used it in ways that generated differentiated value. That&#8217;s a 17:1 ratio between adoption and sophistication, and most organizations can&#8217;t see it because they&#8217;re measuring the wrong dimension entirely.</p><p>The researchers identified four behavioral patterns that consistently predicted sophisticated use. Not prompt length. Not frequency. Behaviors: treating AI as a reasoning partner rather than accepting first outputs, delegating complex multi-step tasks with clear constraints, applying AI across diverse task types instead of just writing assistance, and sustaining longer working-session-style interactions.</p><p>Here&#8217;s the part that caught me off guard. The most sophisticated users weren&#8217;t the youngest employees &#8212; they were above manager level. The conventional wisdom says junior employees are more natural with these tools. The data says otherwise. There&#8217;s a real difference between being comfortable with AI and being good at getting results from it. Comfort is about familiarity. Sophistication is about judgment.</p><h2>The Problem with Averages</h2><p>When I started building a scoring framework from this research, I ran into an interesting design problem. A weighted average of behavioral dimensions sounds clean, but it lies in predictable ways.</p><p>Consider someone who writes long, detailed initial prompts and sustains multi-turn conversations. Their Interaction Depth score is high &#8212; maybe an 8 or 9. But they never refine outputs. Never push back. Never ask the model to check its reasoning or explore alternatives. Their Iterative Reasoning score is a 2.</p><p>A weighted average might land them at &#8220;Proficient.&#8221; But they&#8217;re not proficient. They&#8217;re just verbose. The length of the prompt isn&#8217;t the signal. What the user does with the output is.</p><p>This is why the framework I built includes gating criteria &#8212; floor rules that prevent misclassification. You can&#8217;t reach the Advanced tier unless both Task Complexity and Iterative Reasoning hit at least 7 out of 10, regardless of what your weighted average says. Those two dimensions are the strongest differentiators in the research, and they carry 55% of the total score.</p><p>The gating mechanism is the single most useful idea in the framework. It forces honest measurement.</p><h2>What This Means for How You Train</h2><p>The dimension-level data tells you something activity metrics never can: where to invest in training.</p><p>If Iterative Reasoning is consistently low across your organization, another &#8220;intro to prompting&#8221; workshop won&#8217;t help. The gap isn&#8217;t in how people write prompts &#8212; it&#8217;s in how they think about the interaction. They need to learn to treat AI as a reasoning partner: assign roles, provide examples, test assumptions, ask the model to verify its own logic.</p><p>If Task Complexity is low, the problem is different. People aren&#8217;t delegating hard enough. They&#8217;re using AI for tasks they could do themselves in roughly the same time, instead of delegating the genuinely complex, multi-step work where AI creates real operating margin.</p><p>The dimension scores give you a specific diagnosis. The diagnosis gives you a specific intervention. That&#8217;s the difference between &#8220;use AI more&#8221; and &#8220;here&#8217;s what to change about how you use it.&#8221;</p><h2>The Uncomfortable Implication</h2><p>If only 5% of users are sophisticated at a firm where 90% are active &#8212; a firm that had invested heavily in AI tools and training &#8212; then sophisticated use doesn&#8217;t happen organically. Making tools available and running training sessions gets you to 90% adoption. It does not get you to sophistication.</p><p>Getting there requires measuring the right things, making specific behaviors visible and expected, and building the feedback loops that help people see the gap between where they are and where they could be. Activity metrics can&#8217;t do that. Behavioral metrics can.</p><p>You can&#8217;t operate what you can&#8217;t measure. And right now, most organizations are measuring the equivalent of gym swipes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://andrewlewis.ca/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://andrewlewis.ca/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>I built a full playbook with the five scoring dimensions, weighted formula, gating rules, score anchors, and a printable worksheet for manual scoring. <a href="https://drive.google.com/file/d/1b-y0wAKx1QkXPS58OPi-VpKpILzVpjTk/view?usp=sharing">I also built a Claude Skill</a> if you want to get sophistication scoring within you conversation, along with tips on how to improve.</p><div class="file-embed-wrapper" data-component-name="FileToDOM"><div class="file-embed-container-reader"><div class="file-embed-container-top"><image class="file-embed-thumbnail-default" src="https://substackcdn.com/image/fetch/$s_!0Cy0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack.com%2Fimg%2Fattachment_icon.svg"></image><div class="file-embed-details"><div class="file-embed-details-h1">Ai Sophistication Scoring Spec</div><div class="file-embed-details-h2">237KB &#8729; PDF file</div></div><a class="file-embed-button wide" href="https://andrewlewiswashere.substack.com/api/v1/file/4ee665d0-fc77-49ce-a0cd-1aa01a75a7bf.pdf"><span class="file-embed-button-text">Download</span></a></div><a class="file-embed-button narrow" href="https://andrewlewiswashere.substack.com/api/v1/file/4ee665d0-fc77-49ce-a0cd-1aa01a75a7bf.pdf"><span class="file-embed-button-text">Download</span></a></div></div><div class="file-embed-wrapper" data-component-name="FileToDOM"><div class="file-embed-container-reader"><div class="file-embed-container-top"><image class="file-embed-thumbnail-default" src="https://substackcdn.com/image/fetch/$s_!0Cy0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack.com%2Fimg%2Fattachment_icon.svg"></image><div class="file-embed-details"><div class="file-embed-details-h1">Ai Sophistication Playbook</div><div class="file-embed-details-h2">37.8KB &#8729; PDF file</div></div><a class="file-embed-button wide" href="https://andrewlewiswashere.substack.com/api/v1/file/cd329de4-280e-4e73-9eff-84681597c62e.pdf"><span class="file-embed-button-text">Download</span></a></div><a class="file-embed-button narrow" href="https://andrewlewiswashere.substack.com/api/v1/file/cd329de4-280e-4e73-9eff-84681597c62e.pdf"><span class="file-embed-button-text">Download</span></a></div></div><p></p><p> </p>]]></content:encoded></item><item><title><![CDATA[The Barriers to AI Adoption Are the Job]]></title><description><![CDATA[Everyone agrees on what&#8217;s slowing AI down. Nobody wants to do the actual work.]]></description><link>https://andrewlewis.ca/p/the-barriers-to-ai-adoption-are-the</link><guid isPermaLink="false">https://andrewlewis.ca/p/the-barriers-to-ai-adoption-are-the</guid><dc:creator><![CDATA[Andrew Lewis]]></dc:creator><pubDate>Wed, 08 Apr 2026 12:02:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7dff!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7dff!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7dff!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!7dff!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!7dff!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!7dff!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7dff!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1596077,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://andrewlewis.ca/i/192676220?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7dff!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!7dff!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!7dff!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!7dff!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b747cc-b72e-43c8-9fb2-fac8f4c4cf12_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Every few weeks, a new report lands cataloguing the barriers to enterprise AI adoption. Poor data management. Incomplete governance. Lack of training. Unclear ROI. Resistance to change.</p><p><a href="https://www.nojitter.com/ai-automation/multiple-roadblocks-impede-generative-ai-adoption">No Jitter published one this week</a>. <a href="https://www.bcg.com/press/27march2026-ai-expectations-rise-in-logistics-scaled-adoption-remains-limited">BCG has one</a>. <a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai">McKinsey&#8217;s latest State of AI survey</a> says 88% of organizations report using AI in at least one business function &#8212; but nearly two-thirds haven&#8217;t begun scaling it across the enterprise.</p><p>An <a href="https://hbr.org/2026/02/why-ai-adoption-stalls-according-to-industry-data">HBR piece from earlier this year</a> put a finer point on it: AI initiatives stall not because the technology fails, but because employees&#8217; anxiety about relevance, identity, and job security drives surface-level adoption without real commitment.</p><p>None of this is new. And that&#8217;s exactly the problem.</p><p>We&#8217;ve been listing the same barriers for two years. The list hasn&#8217;t changed. Which raises an uncomfortable question: if we know what the barriers are, why haven&#8217;t we removed them?</p><div><hr></div><h2>The Barrier List Is Actually a Job Description</h2><p>I lead AI enablement and delivery at a large Canadian law firm. My charter covers generative AI governance, production infrastructure decisions, cross-functional technology guidance, and organizational design for autonomous delivery teams.</p><p>When I read reports listing the top barriers to AI adoption, I don&#8217;t see obstacles. I see my to-do list.</p><p>Data readiness? That&#8217;s the conversation with InfoSec about what categories of data we can and can&#8217;t share with AI tools &#8212; and under what conditions. It&#8217;s the ongoing discussion about zero data retention and data residency requirements that shapes every infrastructure decision we make. And these aren&#8217;t theoretical concerns. The <a href="https://www.theregister.com/2025/07/25/microsoft_admits_it_cannot_guarantee">U.S. CLOUD Act</a> allows American law enforcement to demand data held by U.S.-based providers regardless of where that data is physically stored. A Canadian firm using a U.S.-hosted AI tool with client data on a Montreal server is still exposed. When a <a href="https://winbuzzer.com/2025/07/25/microsoft-admits-it-cannot-guarantee-eu-cloud-data-sovereignty-from-us-government-xcxwbn/">Microsoft executive told the French Senate</a> under oath in June 2025 that he couldn&#8217;t guarantee European data would be safe from U.S. authorities, the same logic applies to every law firm using a U.S.-owned cloud AI service.</p><p>Governance gaps? That&#8217;s the policy work &#8212; figuring out what guardrails are necessary when you&#8217;re building internal infrastructure like MCP servers that connect AI tools to live systems. The answer changes as the tools evolve, which means the governance has to evolve with it.</p><p>The other barriers on the list &#8212; training, middle management alignment, cultural resistance &#8212; those show up too. Every organization doing this work recognizes them. The question isn&#8217;t whether they&#8217;re real. It&#8217;s whether anyone is resourced to actually address them.</p><p>The barrier list isn&#8217;t a research finding. It&#8217;s a scope of work.</p><div><hr></div><h2>The Real Reason the List Doesn&#8217;t Change</h2><p>From the inside, the pattern is pretty clear: the barriers persist because removing them requires a kind of work that most organizations don&#8217;t want to fund, staff, or prioritize.</p><p>Governance work is slow, political, and unglamorous. Nobody gets promoted for writing an AI acceptable use policy. Nobody gets a conference keynote for building a data classification framework. These are infrastructure projects &#8212; essential, invisible, and easy to defer.</p><p>Training is continuous, not one-shot. You can&#8217;t run a lunch-and-learn in Q1 and call it done. AI tools change. Use cases evolve. New risks emerge. Training has to be ongoing, role-specific, and practical. That takes dedicated time from people who are already stretched.</p><p>Middle management is where adoption lives or dies. A <a href="https://www.workera.ai/blog/ai-adoption-will-remain-uneven-in-2026-heres-why-and-how-to-fix-it">Workera analysis from earlier this year</a> noted that most adoption bottlenecks sit in the middle of the organizational chart &#8212; managers who struggle to set expectations for AI-assisted work, don&#8217;t know what &#8220;good&#8221; looks like, and avoid the topic because it raises uncomfortable questions about headcount and value.</p><p>And cultural resistance isn&#8217;t about employees being Luddites. The <a href="https://hbr.org/2026/02/why-ai-adoption-stalls-according-to-industry-data">HBR research</a> showed it&#8217;s really about identity &#8212; people worry that using AI makes them look replaceable, or that their expertise is being devalued. You can&#8217;t train your way out of that. It requires sustained, honest messaging about what AI actually changes and what it doesn&#8217;t.</p><p>Every one of these barriers is addressable. None of them is easy. And most of them require organizational commitment that goes well beyond the technology team.</p><div><hr></div><h2>Where the Work Actually Lives</h2><p>When I look at the barrier list through the lens of what has to happen, the work sorts itself pretty naturally.</p><p>Some of it you can systematize. Governance templates, usage policies, risk assessment frameworks, data classification standards &#8212; these are repeatable artifacts. Build them once, adapt across the organization. AI can even help with its own adoption here. Use it to draft the policies, generate training materials, structure rollout plans.</p><p>Some of it requires people working alongside people. Helping managers evaluate AI-assisted work. Coaching teams on where AI fits their specific workflows. Having the uncomfortable conversations about what changes when a task that used to take four hours now takes forty minutes. You can support this with tools and structured conversations, but you can&#8217;t skip the human part.</p><p>And some of it you just have to do the slow way. Building trust. Shifting culture. Showing through consistent action that efficiency gains won&#8217;t quietly become headcount cuts. That the goal is better work, not cheaper workers. No technology accelerates this.</p><p>Most organizations want to spend their AI budget on the first kind. The actual work is mostly the second and third.</p><div><hr></div><h2>The Patience Problem</h2><p>There&#8217;s an additional dynamic that the barrier reports don&#8217;t capture: the gap between executive expectations and organizational readiness.</p><p>Leadership sees the reports about AI productivity gains. They hear the vendor pitches. They attend the conferences. They come back wanting to know why the organization isn&#8217;t moving faster.</p><p>The honest answer is: because we&#8217;re doing the barrier removal work. And that work doesn&#8217;t produce demo-ready results on a quarterly cadence.</p><p>Governance frameworks aren&#8217;t impressive in a board presentation. Training programs don&#8217;t generate viral LinkedIn posts. The slow, steady work of building organizational readiness for AI is the least visible and most important investment a company can make right now.</p><p>The organizations that are furthest along in AI adoption aren&#8217;t the ones that skipped the barrier work. They&#8217;re the ones that started it earlier and funded it properly. They invested in the boring infrastructure while everyone else was running pilot programs.</p><div><hr></div><h2>What Would Actually Change Things</h2><p>The next time one of these reports drops, resist the urge to nod along and move on.</p><p>Instead, ask: which of these barriers are we actively working to remove? Who owns each one? What resources are behind it? What does progress look like on a 90-day horizon?</p><p>If you can&#8217;t answer those questions, the barrier list isn&#8217;t research. It&#8217;s a mirror.</p><p>These barriers have been known for two years. They&#8217;ll still be known next year if the only response is acknowledging them in another planning deck.</p><p>Somebody has to do the work. In most organizations, that role either doesn&#8217;t exist yet or doesn&#8217;t have the authority to act.</p><p>That&#8217;s the barrier nobody puts on the list.</p><div><hr></div><p><em>I&#8217;d be curious to hear which barrier your organization spends the most time acknowledging and the least time actually working on.</em></p>]]></content:encoded></item></channel></rss>