<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Agentic Brief]]></title><description><![CDATA[An operator's playbook for building services that run on software, not people.]]></description><link>https://www.theagenticbrief.com</link><image><url>https://substackcdn.com/image/fetch/$s_!9bHN!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94dd35a3-2807-4e83-91b6-48bc753a042b_795x790.png</url><title>The Agentic Brief</title><link>https://www.theagenticbrief.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 25 Sep 2026 23:08:16 GMT</lastBuildDate><atom:link href="https://www.theagenticbrief.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Nikhil Gupta]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[nixdotbuilder@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[nixdotbuilder@substack.com]]></itunes:email><itunes:name><![CDATA[Nikhil Gupta]]></itunes:name></itunes:owner><itunes:author><![CDATA[Nikhil Gupta]]></itunes:author><googleplay:owner><![CDATA[nixdotbuilder@substack.com]]></googleplay:owner><googleplay:email><![CDATA[nixdotbuilder@substack.com]]></googleplay:email><googleplay:author><![CDATA[Nikhil Gupta]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Executable Org Chart: Inside the Forward-Deployed Pod]]></title><description><![CDATA[When a pod governs a fleet of agents, roles become runtime policy, handoffs become typed contracts, and human judgment becomes an executable control plane.]]></description><link>https://www.theagenticbrief.com/p/the-executable-org-chart-inside-the</link><guid isPermaLink="false">https://www.theagenticbrief.com/p/the-executable-org-chart-inside-the</guid><dc:creator><![CDATA[Nikhil Gupta]]></dc:creator><pubDate>Tue, 01 Sep 2026 13:45:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!u2dE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!u2dE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!u2dE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png 424w, https://substackcdn.com/image/fetch/$s_!u2dE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png 848w, https://substackcdn.com/image/fetch/$s_!u2dE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png 1272w, https://substackcdn.com/image/fetch/$s_!u2dE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!u2dE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!u2dE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png 424w, https://substackcdn.com/image/fetch/$s_!u2dE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png 848w, https://substackcdn.com/image/fetch/$s_!u2dE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png 1272w, https://substackcdn.com/image/fetch/$s_!u2dE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b1f9f8-c720-4b06-813c-81328a361dd8_1344x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last time I wrote about confidence routing: the system does the work it is sure about, holds what needs review, and escalates what it cannot resolve. That makes the agent safe to put in production. It also creates a new problem.</p><p>The work that gets routed does not resolve itself. Someone has to own the review queue, answer the questions, approve the consequential decisions, and remain accountable for the result. If those responsibilities are spread across a traditional services structure with a project manager, an engineering team and a QA team, you get a narrower version of the same pyramid. The volume moves faster, but judgment still moves through meetings and handoffs.</p><p>A forward-deployed pod is not a smaller consulting team supervising some AI tools. It is the control plane for a fleet of agents. Its job is to define what may run, what must be produced, where work must stop, and who has the authority to let it continue.</p><p>Once you take that seriously, the org chart becomes executable.</p><h2>The executable org chart</h2><p>A normal org chart tells you who reports to whom. An executable org chart tells the system how work is allowed to move.</p><p>Such a pod is best described as a graph. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mekF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mekF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp 424w, https://substackcdn.com/image/fetch/$s_!mekF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp 848w, https://substackcdn.com/image/fetch/$s_!mekF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp 1272w, https://substackcdn.com/image/fetch/$s_!mekF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mekF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp" width="1456" height="841" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:841,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:53400,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.theagenticbrief.com/i/204395809?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mekF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp 424w, https://substackcdn.com/image/fetch/$s_!mekF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp 848w, https://substackcdn.com/image/fetch/$s_!mekF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp 1272w, https://substackcdn.com/image/fetch/$s_!mekF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c3cddf4-bf2e-46e8-84f0-a51d21e0ad74_1456x841.webp 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The nodes are agents. Each node has a role, a charter, a set of skills and tools, a runtime budget, the context it may read, and the other agents it is allowed to contact. </p><p>The edges are the working relationships between them: this agent hands work to that one, this validator checks this output, this orchestrator watches these stages. The edge also carries the conditions for moving forward. It can name the artifact that must be produced, the fields the next agent expects, the validator that must sign off, and whether a person has to approve the transition.</p><p>When a project starts, that description is materialized into isolated agent sessions and workspaces. Each agent receives its own instructions, tools, context, and operating policy. The runtime creates the event log, the handoff chain, the question queues, and the approval state. What looked like an org chart in configuration becomes the actual topology of delivery.</p><p>That distinction matters because a fleet of agents behaves less like a group of employees and more like a distributed system. Agents run concurrently, fail independently, resume from checkpoints, and produce state that other agents consume. Leave their boundaries implicit and you get the usual distributed-systems failures: hidden coupling, missing context, duplicate work, and no clear owner when the process stops.</p><p>The pod is the layer that makes those boundaries explicit.</p><h2>Handoffs are APIs</h2><p>A pod does not eliminate handoffs. It eliminates <em>unstructured</em> handoffs.</p><p>Consider one stage of a data pipeline handing work to the next. "The landing layer is done" is not a useful handoff. Done with which tables? Where are the outputs? What were the row counts? Which mappings were applied? What data-quality checks must the next stage preserve? What grain was inferred, and which assumptions remain open?</p><p>In an executable org chart, those are fields in an I/O contract. That contract does three things at once. It gives the downstream agent enough state to begin without reconstructing the previous agent's reasoning. It gives the validator a precise surface to attack. And it gives the human reviewer a bounded object to approve.</p><p>The pod does not scale because everyone sits close together. It scales because context has a schema.</p><h2>Trust is not authority</h2><p>The trust boundary from the last piece decides whether the evidence is strong enough for an agent's output to proceed. The pod introduces a second boundary: the <strong>authority boundary</strong>.</p><p>This often gets collapsed into a generic &#8220;human-in-the-loop,&#8221; as if any ambiguity should be sent to a person. But an agent can be the right reviewer for many ambiguous cases. A validator can check row-count reconciliation, referential integrity, schema completeness, and whether an answer is grounded in the retrieved evidence. An orchestrator can answer a question that is already settled in an approved project document.</p><p>Humans are needed where the system must assign responsibility for consequence. They approve a business definition, accept a residual data-quality risk, resolve a client-specific tradeoff, or authorize an action against a live system. Their role is not to repeat the agent's analysis. It is to decide whether this exact state transition is allowed.</p><p>The runtime enforces that distinction. When an agent reaches a human-gated edge, it writes its handoff payload and marks the stage ready for review. The validator is notified. The downstream agent cannot start until the gate has an approval record. Once the human approves, the approved payload is delivered to the next agent and the workflow advances automatically.</p><p>For workflows that can change a live system, we go one step further. Approval is bound to a fingerprint of the exact plan the person reviewed. Change the plan and the approval no longer matches. The authorization is single-use and is consumed before execution begins. A general "looks good" cannot silently authorize a different action later.</p><p>That is what human-in-the-loop should mean in production. Not a person hovering over the model, but authority represented as a narrow, inspectable capability in the system.</p><h2>The real limit is review load</h2><p>Once the work is organized this way, the size of the pod becomes an engineering question.</p><p>The naive metric is span of control: how many agents can one forward-deployed engineer supervise? But agent count tells you almost nothing. Ten agents that rarely escalate may require less human attention than one poorly calibrated agent that stops every five minutes.</p><p>The useful quantity is review load. For each workflow:</p><p><code>review load = work rate &#215; escalation rate &#215; average handling time</code></p><p>Add that across the fleet and you have the human demand created by the system. The pod remains stable only while that demand stays below the team's available review capacity. Once it crosses the line, the queue grows, decisions become rushed, and the agents' speed simply produces waiting work faster. You have recreated the base of the services pyramid as an exception queue.</p><p>This is where confidence routing and pod design meet. Better calibration reduces the escalation rate. Better artifacts and handoff contracts reduce the time needed to resolve each escalation. Clear ownership prevents questions from sitting unclaimed. Improving any of those increases the amount of agent work a fixed human team can safely govern.</p><p>It also explains why adding another reviewer is usually the wrong first response. The durable fix is to ask why the system needed that review, whether the evidence could have been gathered automatically, whether the contract was incomplete, or whether the same decision has already been made elsewhere. Headcount would absorb the queue for a while, but only engineering can change its arrival rate.</p><p>Span of control is not how many agents report to you. It is how much unresolved judgment they produce.</p><h2>The control plane and the data plane</h2><p>This gives the pod a cleaner division of labor than "humans do judgment, agents do repetition."</p><p>The agents are the data plane. They profile data, generate mappings, investigate failures, execute approved work, validate outputs, and monitor what was delivered. They do the volume.</p><p>The pod is the control plane. It defines the topology, assigns capabilities, sets thresholds and budgets, owns the contracts, handles the consequential exceptions, and changes the policy when the same exception appears twice. It is accountable for the system without manually touching every unit of work.</p><p>That is precisely why a forward-deployed engineer needs enough technical depth to debug an agent and enough domain judgment to sit across from the customer. But the real skill is connecting the two. Someone who can turn ambiguity into system design: a rule, a required artifact, a validation check, a routing threshold, a permission boundary, or a better handoff contract. The best decision is not merely resolved. It is made cheaper to resolve the next time.</p><h2>Human judgment should compile</h2><p>If every correction stays in the reviewer's head, the pod does not compound. It just performs expert labor through a better interface.</p><p>So the output of human review has to become durable state. Agent sessions write to an append-only event log and can resume from it. Migration artifacts, validation reports, decisions, and handoffs are stored outside the chat transcript so later stages and later engagements can retrieve them. Corrections are captured as feedback against the original output. Where appropriate, a verified outcome can draft a new operating rule, but that rule remains a draft until someone deliberately promotes it.</p><p>This is the learning loop that matters in a services-as-software business:</p><p><code>agent work &#8594; validation &#8594; human decision &#8594; execution &#8594; verified outcome &#8594; policy</code></p><p>The human is not only clearing today's exception. They are producing normative memory for the fleet: what should happen, what must never happen, and what evidence is required before the system may act. Over time, recurring judgment moves out of the review queue and into the platform. The authority boundary stays where the consequence demands, but fewer routine questions need to reach it.</p><p>That is how a small pod remains small. Not because the agents never get stuck, and not because a few exceptional people work infinitely hard. It remains small because each engagement leaves the operating system better than it found it.</p><h2>The pod is part of the product</h2><p>A traditional services firm scales by adding delivery capacity. An AI-native one scales by increasing governed throughput: more work completed per unit of scarce human judgment, without pushing unreviewed risk onto the customer.</p><p>The executable org chart is the organizational form of that idea. Roles become agents and owners. Reporting lines become communication permissions. Handoffs become typed contracts. Reviews become gates. Decisions become durable policy. The structure of the team is enforced by the same runtime that executes the work.</p><p>Put the three pieces together. The data layer makes the agent possible. The confidence layer decides when it can be trusted. The pod supplies the authority, context, and learning loop that let a company stand behind the result.</p><p>Next in the playbook: the first two weeks of an AI-native engagement, where the first deliverable is not a project plan but an evidence-backed account of what is true.</p><p>&#8212; Nikhil</p>]]></content:encoded></item><item><title><![CDATA[Confidence Routing: How to Sell an Outcome Instead of Hours]]></title><description><![CDATA[The old services model bills for effort. To sell the outcome instead, the system has to know when it does not know, and behave differently when it does. A technical look at the layer that makes agents safe to put in production.]]></description><link>https://www.theagenticbrief.com/p/confidence-routing-how-to-sell-an</link><guid isPermaLink="false">https://www.theagenticbrief.com/p/confidence-routing-how-to-sell-an</guid><dc:creator><![CDATA[Nikhil Gupta]]></dc:creator><pubDate>Tue, 21 Jul 2026 13:46:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kjRd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kjRd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kjRd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png 424w, https://substackcdn.com/image/fetch/$s_!kjRd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png 848w, https://substackcdn.com/image/fetch/$s_!kjRd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png 1272w, https://substackcdn.com/image/fetch/$s_!kjRd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kjRd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kjRd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png 424w, https://substackcdn.com/image/fetch/$s_!kjRd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png 848w, https://substackcdn.com/image/fetch/$s_!kjRd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png 1272w, https://substackcdn.com/image/fetch/$s_!kjRd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fae6810-a1e5-4e51-931c-58ceddeedd65_1344x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last time I argued that the hardest part of building AI services is making enterprise data legible, and that the agent is the thin, replaceable layer on top of a much harder foundation. This is about the layer just above that one. It is the difference between an impressive demo and something a business will pay for by the result.</p><p>Here is the shift underneath all of it. A traditional services business bills for effort: hours, headcount, time and materials. You sell the work. An AI-native services business wants to sell the result instead: the migrated system, the resolved ticket, the validated dataset. It charges for that rather than for the time it took. That is a much better business. It is also a much harder promise to keep, because the moment you are paid for outcomes, every wrong answer is your problem, not the client's.</p><p>So the question that decides whether you can sell an outcome is not "how good is the agent on average." It is "what does the agent do when it is wrong, and does it know."</p><h2>You cannot just ask the model how sure it is</h2><p>The naive version of confidence is to ask the model: have it rate its own certainty, or read the token probabilities, and route on that. This does not work, and it fails in the most dangerous way, because language models are confidently wrong. Their self-reported certainty is poorly calibrated, and their token probabilities measure fluency, not correctness. A model will happily assign high probability to a smooth, plausible, completely fabricated answer.</p><p>Route on that signal and you will auto-approve exactly the errors that look most convincing. The calibration has to come from somewhere outside the model's own opinion of itself.</p><h2>What confidence actually is in production</h2><p>We treat confidence as a composite signal, assembled from sources that are allowed to disagree with the model:</p><ul><li><p><strong>A verifier.</strong> A second model whose only job is to check the first one's answer against the evidence, prompted to look for reasons it is wrong rather than reasons it is right. Agreement between a generator and an adversarial verifier is a far stronger signal than either alone.</p></li><li><p><strong>Self-consistency.</strong> Sample the answer several times. If the agent lands in the same place across independent attempts, that stability is evidence; if it scatters, that is a flag, no matter how confident any single run sounded.</p></li><li><p><strong>Grounding.</strong> Is the answer actually supported by the records the agent retrieved, or did it drift off the source? An answer that cites and matches real data beats one that merely reads well.</p></li><li><p><strong>Rule validation.</strong> Cheap, deterministic checks the answer must pass: types, ranges, referential integrity, business constraints. These catch a surprising amount of confident nonsense for almost no cost.</p></li></ul><p>None of these is sufficient alone. Combined, they produce a number that means something, because it is correlated with being right rather than with sounding right.</p><h2>The routing</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hVnK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hVnK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png 424w, https://substackcdn.com/image/fetch/$s_!hVnK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png 848w, https://substackcdn.com/image/fetch/$s_!hVnK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png 1272w, https://substackcdn.com/image/fetch/$s_!hVnK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hVnK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hVnK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png 424w, https://substackcdn.com/image/fetch/$s_!hVnK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png 848w, https://substackcdn.com/image/fetch/$s_!hVnK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png 1272w, https://substackcdn.com/image/fetch/$s_!hVnK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb6283b9-1704-4968-aad3-91226f5fd805_1775x1014.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Once you have a signal you trust, the behavior is simple to state and hard to tune. Three bands:</p><ul><li><p><strong>High confidence: act.</strong> The agent completes the work and moves on. No human touches it.</p></li><li><p><strong>Medium confidence: review.</strong> The work is done but held; a person confirms or corrects it before it takes effect.</p></li><li><p><strong>Low confidence: escalate.</strong> The agent stops and hands the problem to someone, usually with a note on what it was unsure about.</p></li></ul><p>I think of the line between <em>act</em> and <em>review</em> as the <strong>trust boundary</strong>. Everything above it runs as software. Everything below it is still a service delivered by people, with the agent assisting. The whole game is moving that boundary up, safely, over time.</p><h2>Setting the thresholds</h2><p>The thresholds are where the engineering lives, because they encode an economic decision, not just a technical one.</p><p>Set the <em>act</em> threshold too low and the agent auto-completes work it should have flagged. You get speed, but errors leak into production and you are paying for them. Set it too high and almost everything routes to a human. The system is safe, but it is barely software, and your margins start to look like a staffing agency's. The threshold is the dial between those two failure modes.</p><p>You choose the setting with data, not taste. Run the agent against a labeled set where you know the right answers, and at each candidate threshold you can measure two things: how often an auto-approved answer is actually wrong (the error that reaches the client) and how much work is being sent to humans (the cost that eats the margin). That gives you a curve, and you pick the point on it that matches the cost of being wrong in that specific workflow.</p><p>That cost varies wildly, which is why there is no global threshold. An agent suggesting a category for an internal ticket can be wrong sometimes; the cost is a small correction. An agent posting a financial transaction cannot; one wrong auto-action can outweigh a hundred needless escalations. Same architecture, completely different boundary, because the asymmetry between a false <em>act</em> and an unnecessary <em>review</em> is different. You tune to the asymmetry.</p><h2>The human-in-the-loop earns the agent its autonomy</h2><p>The reviews and escalations are not only a safety net. They are the training signal. Every time a person confirms, corrects, or overrides an answer, you learn whether the confidence score was telling the truth, and you feed that back.</p><p>That feedback is what lets the trust boundary move. As the agent accumulates corrected examples on a kind of work, two things improve together: the agent gets better at the task, and the confidence signal gets better calibrated, so you can safely lower the <em>act</em> threshold and let more run as software. The system becomes more autonomous not by decree but by earning it, one verified batch at a time. Week one, most of a new workflow sits in the review band. Month three, the bulk of it is above the trust boundary and a small team handles the long tail.</p><h2>Why this is what lets you sell an outcome</h2><p>Bring it back to the business. A client will not pay for an outcome they cannot trust, and they will not trust a system whose failure mode is silent, confident error. Confidence routing is what makes the failure mode safe: the system does the work it is sure about and visibly flags the rest, so the worst case is a review queue rather than a wrong answer shipped at scale. That is a bounded, inspectable risk, which is the kind a business can actually sign off on.</p><p>So confidence routing is not really a model feature. It is the thing that turns "an agent that is usually right" into "a service you can put your name on and bill by the result." The data layer makes the agent possible. The confidence layer makes it sellable.</p><p>Next in the playbook: the people. When the agent does most of the work and routes the rest, what is the small human team actually for, and how do you organize it: the pod around the fleet.</p><p>&#8212; Nikhil</p>]]></content:encoded></item><item><title><![CDATA[The Real Work in AI Services Is Making Data Legible]]></title><description><![CDATA[Everyone obsesses over the agent. In practice most of the engineering happens underneath it, turning an illegible enterprise data estate into something an agent can actually reason over. A technical look at how, and why it's the whole game.]]></description><link>https://www.theagenticbrief.com/p/the-real-work-in-ai-services-is-making</link><guid isPermaLink="false">https://www.theagenticbrief.com/p/the-real-work-in-ai-services-is-making</guid><dc:creator><![CDATA[Nikhil Gupta]]></dc:creator><pubDate>Tue, 30 Jun 2026 15:18:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!VIxB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VIxB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VIxB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png 424w, https://substackcdn.com/image/fetch/$s_!VIxB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png 848w, https://substackcdn.com/image/fetch/$s_!VIxB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png 1272w, https://substackcdn.com/image/fetch/$s_!VIxB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VIxB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png" width="728" height="269.4453125" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:379,&quot;width&quot;:1024,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:617015,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VIxB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png 424w, https://substackcdn.com/image/fetch/$s_!VIxB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png 848w, https://substackcdn.com/image/fetch/$s_!VIxB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png 1272w, https://substackcdn.com/image/fetch/$s_!VIxB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7870d322-d39f-4d81-a6c8-86a84c58e504_1024x379.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two years ago, growing a services business meant hiring people. The work was done by humans, so more work meant more humans, and the margins were whatever the labor math allowed. We are now building that same kind of business, enterprise system integration, to run mostly on agents instead of people. Engagements that used to take a team of twenty most of a year, we deliver with a small pod and a fleet of agents in weeks.</p><p>When people hear that, they assume the hard part is the agents. It isn't. The agents are the easy part. The hard part, the part that decides whether a project succeeds or quietly falls apart three weeks in, is everything underneath them: making the customer's data legible enough that an agent can reason over it without producing confident nonsense.</p><p>This is the least glamorous engineering I know of, and it is most of the job. Here is what it actually involves.</p><h2>Semantic debt</h2><p>Every enterprise is carrying what I have started calling semantic debt. It is like technical debt, except instead of shortcuts in the code, it is meaning that was never written down.</p><p>A schema tells you a table is called <code>CUST_MASTER</code> and has a column <code>status_2</code> of type <code>VARCHAR</code>. It does not tell you that <code>status_2</code> is the field people actually trust, that <code>status</code> was deprecated in 2017 but never dropped, that the values <code>A</code>/<code>I</code>/<code>H</code> mean active, inactive, and a third state nobody can define anymore, or that a quarter of the rows are test records from a migration that half-finished. None of that is in the database. It lives in the heads of people, some of whom have left.</p><p>Semantic debt shows up in a few consistent shapes:</p><ul><li><p><strong>Schema without semantics.</strong> You have column names and types, but not what they mean, which are authoritative, or how they relate.</p></li><li><p><strong>Entity fragmentation.</strong> The same real-world thing is recorded several ways across systems that were bought rather than built, with no shared key. "Acme Corp," "ACME Corporation," and "Acme Inc." are three rows and one company.</p></li><li><p><strong>Lineage loss.</strong> A number on a report exists, but how it was produced (which tables, which filters, which business rules) is gone. Two reports show different revenue and both are "right," because they define revenue differently and nobody wrote it down.</p></li><li><p><strong>Silent contradiction.</strong> The data disagrees with itself in ways nothing flags: a fact row whose foreign key points at a dimension that does not exist, a date that is sometimes the order date and sometimes the ship date depending on the source.</p></li></ul><p>An agent pointed at this does not see a mess. It sees confident inputs and produces confident outputs, and the outputs are wrong in ways that are expensive to catch later. Garbage that looks like an answer is worse than no answer.</p><h2>Why the obvious approaches fail</h2><p>The first instinct is to hand the model the raw schema: dump the DDL into the context, add a few sample rows, let it figure things out. This works in a demo and fails in production. The model invents joins that are not valid, picks the wrong grain, and chooses whichever of the three revenue definitions sounds most plausible. It has no way to know which <code>status</code> field is real, because that information is not in what you gave it.</p><p>Retrieval over the raw tables has the same problem one layer up. You can embed and search the data all you like; if the meaning was never captured, similarity search just retrieves the same ambiguity faster.</p><p>The thing the agent needs does not exist yet when you arrive. You have to build it.</p><h2>Building the legibility layer</h2><p>The legibility layer is an explicit, verified model of what the customer's data means. It is produced by a pipeline of six stages, with a human-in-the-loop wrapped around all of them.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oB1t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oB1t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png 424w, https://substackcdn.com/image/fetch/$s_!oB1t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png 848w, https://substackcdn.com/image/fetch/$s_!oB1t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png 1272w, https://substackcdn.com/image/fetch/$s_!oB1t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oB1t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oB1t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png 424w, https://substackcdn.com/image/fetch/$s_!oB1t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png 848w, https://substackcdn.com/image/fetch/$s_!oB1t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png 1272w, https://substackcdn.com/image/fetch/$s_!oB1t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa56b4988-4005-486e-a54e-dc46cfaf8de9_2145x1245.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>1. Profile before you interpret.</strong> Before a model is allowed an opinion about a column, we profile it directly against the data. For each column we compute cardinality and null rate, the distribution of values, and a format signature. Do the values look like emails, ISO dates, currency, a short coded enum? We infer candidate primary keys from uniqueness, and candidate foreign keys from value containment: if 98% of <code>orders.cust_ref</code> values appear in <code>cust_master.id</code>, that is a join the declared schema never told you about. Statistics come first, because they ground everything downstream and catch the model the moment it starts to guess.</p><p><strong>2. Infer meaning, then have a second model attack it.</strong> Each table and column goes to an agent along with its profile, a sample of real rows, and the names of its neighbors. The agent proposes what the field means, its semantic type, whether it is authoritative or a deprecated remnant, and how it relates to other tables, attaching a confidence to every claim. A second model runs as an adversarial check: same evidence, but its job is to falsify the first proposal. We route on the result in three bands. Above 0.9 with agreement auto-accepts. Between 0.7 and 0.9 is queued for human review. Below 0.7 escalates. The bands exist because a wrong semantic label here is not a local error; it is paid back, with interest, by everything built on top of it.</p><p><strong>3. Resolve entities.</strong> The three-different-Acmes problem is its own discipline. Comparing every record to every other record is quadratic and hopeless at enterprise scale, so we <em>block</em> first, generating cheap candidate keys (a normalized name plus a postal code, say) so that only plausibly-matching records are ever compared. Each candidate pair is then scored with a mix of deterministic rules and embedding similarity, the scores are thresholded, and the surviving links are clustered with union-find so that "Acme Corp &#8596; ACME Corporation &#8596; Acme Inc." collapse into one entity with one key. Pairs near the threshold, the genuinely ambiguous ones, go to a person. The output is a resolved entity the rest of the system can rely on.</p><p><strong>4. Recover lineage.</strong> Where a metric is produced by transformations (SQL views, stored procedures, pipeline code), we parse them into syntax trees and build a column-level lineage graph: this report field comes from that view, which sums this column filtered by that flag. The business rules buried in the <code>WHERE</code> clause become part of the definition. "Revenue" stops being a word and becomes a specific, traceable computation everyone can point at and argue with.</p><p><strong>5. Assemble the model.</strong> The verified pieces are assembled into a layered model. Raw landing data (bronze) is cleaned and conformed (silver) and then modeled (gold). The gold layer is shaped as a star schema: fact tables declared at an explicit grain, surrounded by conformed dimensions, so that a question like "revenue by region by quarter" has exactly one correct way to be answered. The grain is the contract. Most of the silent contradictions in enterprise reporting are really two numbers computed at two different grains and compared as if they were the same.</p><p><strong>6. Detect what is broken.</strong> The same machinery that builds the model is run as a set of assertions against it: referential integrity (orphaned facts pointing at missing dimensions), grain violations (a "unique" business key that is duplicated), cross-source reconciliation (the same metric from two systems, and the delta between them), and temporal anomalies. A lot of the value we deliver in the first weeks of an engagement is simply telling a customer the truth about their own data, with the specific offending rows attached.</p><p>At the end of this you have something the customer never had: a written, verified account of what their data means. That artifact is the product. The agents that do the actual delivery work (migration, validation, answering questions, monitoring for drift) all stand on top of it.</p><h2>Why we let agents do it, and why we watch them</h2><p>Two things make agents well-suited to this. The work is enormous and repetitive, far too much for a human team to do by hand across thousands of columns and hundreds of tables. And it is exactly the judgment-under-ambiguity that language models are good at, given enough grounding.</p><p>Two things make it dangerous. Models are confident when they are wrong, and this is precisely the place where a wrong answer gets baked into everything downstream. So the agents never run unsupervised. Every inference is grounded in real profiles and samples rather than the model's prior, every output carries a confidence, a second model checks the first, and a human owns the thresholds where the system stops and asks. The goal is not full autonomy. It is a system that does the overwhelming bulk of the work, knows the exact boundary of what it is sure about, and widens that boundary as it earns trust.</p><h2>Why this is the whole game</h2><p>A services business can run on software not because the agents are clever, but because this layer, once built, is leverage. The semantic model is reusable. The entity resolution compounds. The lineage stays recovered. The next engagement at a similar customer, on a similar source system, starts further along than the last one did. The labor you would have spent on twenty people for six months collapses into the work of building and verifying this artifact once, after which everything you deliver runs on it.</p><p>That is why services as software is an engineering story and not a margin story. The margin is a consequence. The cause is whether you can take an illegible enterprise data estate and make it legible faster and more reliably than a room full of consultants could.</p><p>This is the first piece of the playbook. The next ones go up the stack: the delivery agents that run on top of this layer, how confidence routing lets you sell an outcome instead of hours, and how a small pod of people is organized around a fleet of agents. The foundation comes first, because nothing above it works without it.</p><p>&#8212; Nikhil</p>]]></content:encoded></item></channel></rss>