The Last 20%
Essays for the operational 20%
Short, opinionated essays on what stalls between 80% and production. For CIOs, CTOs, and CFOs at $50M-$1B engineering-led companies.
Subscribe on Substack- Essay 01
The Last 20%
Why Every Mid-Market AI Initiative Stalls at the Same Place
An engineering-led IT team ships an AI initiative at 80% within 90 days. The remaining 20% never closes — not because the team is lazy, but because the 20% is structurally different from the 80%, and almost nobody plans for the gap.
#CIO#EnterpriseAI#MidMarketRead on Substack → - Essay 02
BYOK Is the Tell
A 30-Second Vendor Sorting Hat
One 30-second question separates AI vendors who'll be around in 2027 from the ones who won't. Ask it on every bake-off call — and watch the room reorganize itself.
#CIO#AIprocurement#CFO#VendorEvaluation#EnterpriseAIRead on Substack → - Essay 03
Stop Hiring Platform Engineers
Hire One Less
The math on hiring a $200K platform engineer to maintain your AI stack vs. paying a specialist $50K/yr doesn't survive contact with a CFO who can do four-year arithmetic.
#CIO#CFO#ITbudget#PlatformEngineering#FP&ARead on Substack → - Essay 04
Pricing Transparency
$75K + $50K/yr — Here's the Build-vs-Buy Math
Most AI vendors hide pricing behind "talk to sales." Here's exactly what we charge — $75K engagement + $50K/yr — and why publishing the build-vs-buy math is the most defensible move we could make.
#CIO#CFO#AIvendor#PricingTransparency#ITbudgetRead on Substack → - Essay 05
Year 2-3 of the PE Hold
When "We Built It Ourselves" Stops Being a Brag
There's a window in the PE hold cycle when "we built it ourselves" stops being a brag and starts being a liability. Recognize the window before the operating partner does.
#CIO#PE#PrivateEquity#MidMarket#ITdiligenceRead on Substack → - Essay 06
The Power Automate Endgame Problem
Five Cliffs Every Mid-Market Deployment Eventually Hits
Every mid-market Power Automate deployment eventually hits the same five cliffs. The community knows about three of them. The other two only show up at scale.
#PowerAutomate#PowerPlatform#Microsoft365#CIO#ITopsRead on Substack → - Essay 07
What Cyber Insurance Carriers Actually Look at
In an AI Breach Claim
Five artifacts insurance carriers demand when an AI mediates a denied claim. Four of them don't exist in your current pipeline. Here's what they look like and where they go.
#CISO#CIO#CyberInsurance#ITcompliance#RiskManagementRead on Substack → - Essay 08
What I'd Tell a $50M–$1B CIO
If I Wasn't Selling Anything
Twenty minutes of unvarnished advice with no pitch, no demo, no NDA. Most of it is what to stop doing.
#CIO#EnterpriseAI#VendorEvaluation#MidMarket#ITstrategyRead on Substack → - Essay 09
Cognitive Surrender
What Happens When Knowledge Workers Stop Reasoning and Dispatch Problems to AI Instead
A devops engineer named the failure mode of AI without discipline this year. Cognitive surrender is the decision to stop reasoning and let the model decide instead. Seven flavours of vendor-trust collapse, four cautionary cases at four scales, and the structural fix that is not "more AI literacy."
#AIGovernance#CIO#EnterpriseAI#OperationsPartner#AIRead on Substack → - Essay 10
Seven Questions Every AI Vendor Should Be Able to Answer
A Buyer's Diagnostic for AI Vendor Evaluation
Seven questions you can run in a discovery call, RFP, or 30-minute vendor-eval review. One per vendor-trust-collapse flavour, with vendor-honest and vendor-evasion answer shapes laid out so you can tell which one you are hearing.
#AIGovernance#CIO#EnterpriseAI#OperationsPartner#AIRead on Substack → - Essay 11
Two Kinds of AI Exclusion
And the One Your CGL Renewal Won't Tell You About
Headlines about carriers "excluding AI" blur two structurally different exclusions filed with state regulators. One is narrow and applies only to generative AI. The other is broad enough to exclude algorithmic credit scoring. Your renewal preparation needs to know the difference.
#CyberInsurance#AIGovernance#CIO#RiskManagement#InsuranceLawRead on Substack → - Essay 12
What You're Actually Deferring When You Defer Agent 365
Gartner Is Right to Say Wait. The Question Is What Fills the Deferral Window.
Gartner's First Take on Microsoft 365 E7 and Agent 365 says defer, and Gartner is right. The harder question is what you buy instead, and what you commit to operating yourself for the 6 to 18 months of the deferral window. The three honest paths, laid out.
#CIO#Microsoft365#AgentGovernance#ITStrategy#AIRead on Substack → - Essay 13
The Ungoverned Marketing Department
Why Marketing Surrendered to AI First, and the Category That Is Still Empty
Marketing ships in public and the feedback is fast, so cognitive surrender showed up there before anywhere else. The marketing-operator cut of the Cognitive Surrender arc: what the failure mode looks like inside the department, and the governance shape that fixes it.
#Marketing#CMO#AIGovernance#OperationsPartner#B2BMarketingRead on Substack → - Essay 14
Integration Is the Moat
Portfolio Value Creation When Acquisition Velocity Stops Working
The cheap-debt, buy-the-growth roll-up playbook is closing. When you cannot buy the next turn of growth, the only return left is the one you build inside the businesses you already own. Integration is an operating capability, not a transaction.
#PrivateEquity#ValueCreation#PortfolioOps#Integration#AIGovernanceRead on Substack → - Essay 15
The Job Description Is the SOW
The Open Marketing Requisition Is a Statement of Work With a Cleared Budget
Every portfolio company has a marketing requisition that has been open 60 to 120 days. The responsibilities section is a Statement of Work, the salary band is a pre-approved budget envelope, and the hiring manager is the named sponsor. The trial starts Monday.
#PrivateEquity#ValueCreation#PortfolioOps#PortfolioMarketingRead on Substack → - Essay 16
The Vendors Just Named Our Category
Five Billion Dollars in Two Weeks, and the Segment the FDE Model Cannot Reach
In one fortnight of May 2026 the frontier vendors put more than $5B behind the Forward Deployed Engineer delivery shape, and stated publicly which buyers it will not reach. The partner-network half of that admission is the empty square the Operations Partner fills.
#PrivateEquity#ValueCreation#PortfolioOps#AIGovernance#ManagedAIRead on Substack → - Essay 17
The Forward Deployed Engineer Will Not Be Coming to Your Office
Three Structural Responses for the $50M-$1B IT Organisation
The May 2026 FDE capital says who gets the embedded engineer: Fortune-500 regulated verticals. Mid-market IT is routed to partner networks and self-serve APIs. Three structural responses, decidable from engineering bench, time-to-production, and audit-artifact readiness.
#CIO#CISO#EnterpriseAI#AIGovernance#MidMarketITRead on Substack → - Essay 18
Cloud Can't Reach the Work. Local Can't Be Governed.
Why Agents Need a Substrate and a Control Plane — Not a Bigger Model
Every demo of an AI agent runs in a sandbox; real work lives behind a firewall. The market's two default answers are both quietly broken — cloud-only can't reach the work without opening your network, local-only reaches everything but governs nothing. Kubernetes already solved this shape: split the control plane from the data plane and let the workers pull. The tell for any agent platform is the direction of the wire.
#AIGovernance#AIAgents#EnterpriseAI#CISO#ZeroTrustRead on Substack → - Essay 19
Who Operates Is the Moat
Agent Platforms Agree on the Hard Part. They Split on Who Operates.
A wave of agent-orchestration startups — open-source projects and venture-backed seeds alike — have independently converged on the same spine: a governed work-item, a human approval gate, an immutable audit trail. The architecture is basically settled. The question nobody is asking is who actually operates the work — and that operating choice, not autonomy level or model size, is the durable moat.
#AIAgents#AgentOrchestration#AIGovernance#ManagedAI#AgenticAIRead on Substack → - Essay 20
Don't Sandbox the Agent. Govern the Surface.
Why the Local Agent Needs Managing Like an Employee Laptop — Not Sandboxing Like Hostile Code
Caging the local agent is the obvious reflex and the wrong one — a sandbox tight enough to be safe is tight enough to be useless. We already have a discipline for a capable actor with broad reach on a device we don't fully control: an employee with a laptop. Trust the managed device; govern the surface it connects to, where the plane holds the credentials and the gate, and a revoked badge ends the session in seconds.
#AIGovernance#AIAgents#EnterpriseAI#CISO#ZeroTrustRead on Substack → - Essay 21
The Session Is the Tell
The 30-Second Question for Any Vendor Whose AI Agent Runs on Your Machine
When a vendor's agent runs on your machine, something has to authenticate it to your systems — so what credential sits on the box, and where do the real keys live? One 30-second question, three answer shapes: a god-key, an encrypted god-key, or the agent's own short-lived session. The sequel to BYOK Is the Tell — the session is the tell because it reveals blast radius, real per-agent revocability, and where governance actually lives.
#AIGovernance#AIAgents#EnterpriseAI#CISO#ZeroTrustRead on Substack → - Essay 22
The Leads Your Locations Lost Last Month
Your Cost Per Lead Is on a Slide. The Leads That Walked Are in No Dashboard.
Last month your locations took in hundreds of new inquiries — and a real share never got a fast human response, and quietly walked. It's not a demand problem, it's a capture problem: the unowned seam between marketing and operations, a seven-figure leak across a location footprint that shows up in none of your metrics. Here's how to put a per-location number on it.
#MultiLocationMarketing#CMO#MarketingOps#SpeedToLead#FranchiseMarketing#CustomerAcquisitionRead on Substack → - Essay 23
Agency Is Not a Metaphor
When an AI Agent Acts, the Law Looks for a Principal
When an AI "agent" acts, courts don't see a clever tool — they see an agent acting for a principal, and the principal answers for what it did. From Air Canada's chatbot to imposed-liability rulings, the pattern holds: you can't delegate accountability to software. Governance is how you stay a principal you can defend.
#AIGovernance#AILiability#CISO#GeneralCounsel#AIAgentsRead on Substack → - Essay 24
You Can't Incorporate Away the Principal
Argentina Wants AI-Run Companies With No Humans. It Doesn't Work.
Argentina wants to legalise companies run by AI with no humans required — the corporation's last move to shed human liability. It doesn't work: a foreign shell won't hold the bag, and the party who granted the authority and can prove how it was governed is the one the loss comes home to. The sequel to "Agency Is Not a Metaphor."
#AIGovernance#AILiability#GeneralCounsel#Insurance#AIAgentsRead on Substack → - Essay 25
The Evidence Your Regulated Clients Keep Asking For
Run AI on Regulated Data and You've Signed Up to Prove It
Sooner or later a regulated client asks you to prove exactly what your AI did on their data — and that it couldn't have done anything else. A screenshot and "trust us" isn't evidence. Real proof is three things: full provenance, a trail anchored beyond your own reach, and a pack mapped to the client's framework — and the provider who can produce it wins the accounts the others fear.
#MSP#Compliance#HIPAA#SOC2#GovernedAIRead on Substack → - Essay 26
The Audit Log Is the Tell
Can You Believe the Record of What the Agent Did?
Every governed-AI pitch points at an audit log. The question that sorts vendors: if a row were deleted or changed — by an attacker, or by you — could I prove it? Three answer shapes (editable rows / vendor-signed "immutable" / hash-chained and anchored beyond the vendor's reach) tell you whether the record survives its own keeper. Third in the "tell" lineage.
#AIGovernance#AIAgents#EnterpriseAI#CISO#ComplianceRead on Substack → - Essay 27
Who Can Become You?
The Fourth Tell: Who at the Vendor Can Quietly Become You
Every SaaS support team has a "become the customer" button, and almost nobody governs it the way they'd never let you leave the AI ungoverned. The fourth tell, after BYOK, the session, and the audit log: who at your vendor can see and act as you — and would you ever know? The only good answer is the one where you grant it, you see it, and the record can't be rewritten.
#AIGovernance#AIAgents#EnterpriseAI#CISO#TrustRead on Substack → - Essay 28
Who Can Become Your Client?
The Access Question Your Regulated Clients Will Ask Next
Run AI on a regulated client's data and you've inherited the question every client eventually asks: who can become them — see and act as their account? The row nobody reads aloud is the support master key, yours and your AI vendor's. A defensible answer is granted by you, scoped, on a record you can see and can't rewrite — and the provider who has it wins the accounts the others are afraid to take.
#MSP#Compliance#HIPAA#SOC2#GovernedAIRead on Substack → - Essay 29
The Black Box in Your Funnel
Your Call Center Books Half Your Patients. Why Is It Your Least-Measured Channel?
For most multi-location businesses, roughly half of bookings close in a human conversation — the call center — and it is the least-instrumented part of the entire funnel. Last-touch over-credits the digital channels you can see and under-invests in the one actually closing the business; multi-touch attribution quietly breaks when its highest-converting step, the phone, is not instrumented. The call center is not an ops cost center — it is your highest-intent conversion surface, run blind.
#MarketingAttribution#MultiTouchAttribution#CallTracking#MultiLocation#HealthcareMarketing#CMORead on Substack → - Essay 30
The Band Where AI Is Safe Is a Design Problem, Not a Discipline
Everyone now agrees AI shouldn't render the verdict. Agreeing is the easy part — here's how you build the thing that enforces it.
The consensus that AI shouldn't make the final call is the easy part; enforcing it is the hard part. "Be careful" is not a control — it depends on the most distracted person on the worst day. The band where AI is safe has to be a property of the system: scope, gates, and an audit that hold whether or not anyone remembers. Design beats discipline.
#AIGovernance#EnterpriseAI#AILeadership#DecisionMaking#AIStrategyRead on Substack → - Essay 31
"Self-Sufficient When We Leave" Has a Floor
AWS just put a billion dollars behind forward-deployed AI. Its own model names exactly who it cannot serve.
AWS stood up a dedicated Forward Deployed Engineering org with a billion dollars behind it — and its own promise, "self-sufficient when we leave," names exactly who it cannot serve. There is a floor below the Fortune-500-regulated tier where the embedded engineer never comes, and someone still has to operate for everyone beneath it. The category is settled; the floor is the opening.
#AIGovernance#EnterpriseAI#ManagedAI#ForwardDeployedRead on Substack → - Essay 32
What Compounds
A major vendor just open-sourced the agent governance layer. When your differentiator ships as a free download, you find out what your moat actually was.
Databricks open-sourced its entire agent "meta-harness" — Apache-2.0, the whole spine: a model-agnostic harness, a policy engine, a split control/data plane, an approval gate. When the layer you thought was your differentiator ships as a free download, you get an honest answer to what your moat actually was. Read what the open version can't do and two absences jump out, both about time: a record you can prove, and a memory that compounds.
#AIAgents#AIGovernance#AgentOrchestration#ManagedAI#OpenSourceRead on Substack → - Essay 33
A Memory of Judgment
I tore down the five best open-source memory systems. Every one remembers what's true. None remembers whether the agent should be trusted to act — and why nobody built that is the whole business.
I read the five best open-source AI-memory systems line by line — supermemory, mem0, Open Brain, Graphiti, Letta. Every one remembers what's true about the world; not one remembers whether its own past actions were good enough to be trusted again. The gap is structural: you only get that signal if you operate — which makes judgment memory a company, not a library.
#AIAgents#AgentMemory#AIGovernance#ManagedAI#OpenSourceRead on Substack → - Essay 34
The Agents Won't Govern Themselves
Your enterprise won't run one AI agent — it'll run a mixture. I read four line by line. Each guards its own turf well; none governs the fleet. The layer that does is the whole business.
Readers pushed back on the memory teardown: those are libraries — show me the actual agents. So I read four line by line: OpenAI's Codex, two chat-channel assistants (OpenClaw, Hermes), and a self-improvement experiment (Evolver). Per-agent safety is close to solved and the best labs ship it well — but the enterprise runs a mixture of agents, and not one of the four governs the fleet. That governing layer is the whole business.
#AIAgents#AIGovernance#EnterpriseAI#AgentSecurity#ManagedAIRead on Substack → - Essay 35
Nobody Budgets for Governance
A friendly prospect asked three questions I couldn't answer cleanly — and exposed a sequencing error: governance is why customers stay, but productivity is why anyone buys. Accelerate, accumulate, govern. In that order.
Enterprise budgets attach to functions and headcount-shaped work: support ops, the content agency, the SDR tools line. There is no line called "governance" — that line gets created later, by pain. So the sequence is Accelerate → Accumulate → Govern: AI doing real work is why anyone buys; the track record and the risk pile up together; governance is why they stay, expand, and pass audits. What we changed in a week: named offers mapped to budget lines a buyer already has, a fixed pilot price on the website, and governance demoted to one quiet, load-bearing line in every offer.
#AIGovernance#EnterpriseAI#GoToMarket#ManagedAI#FounderLessonsRead on Substack → - Essay 36
Your Agent Needs an Allowance
A dashboard only tells you what already happened — when what happened is "the agent spent the quarter's budget overnight," the dashboard is an obituary. We shipped the alternative: an allowance, with an earned-autonomy ladder for spend.
Everyone's read the runaway-agent bill story, and the resolution is always "we'll watch it more closely." Watching is not a control. Last week we shipped the alternative: a hard 100% wall (new work blocked, in-flight spared — a pause, never a crash), loud pausing with reasons and auto-resume dates, surgical snapshot-then-restore resume, and repeat offenders losing auto-resume entirely. Spend has an earned-autonomy ladder, exactly like actions. Plus a confession about almost putting the gate on the wrong function.
#AIAgents#AIGovernance#FinOps#EnterpriseAI#ManagedAIRead on Substack → - Essay 37
The Law Just Located the Risk
England's judiciary-chaired taskforce just answered who's liable when AI causes harm: whoever granted the autonomy — and can't prove the supervision.
On 7 July the UK Jurisdiction Taskforce published 130 pages on liability for AI harms — a reading of England & Wales private law, chaired by the Master of the Rolls. Read as an operator, it's a map: it says where AI risk legally lives, and the answer is the operations layer. The sentence that decides cases: liability turns on "the level of autonomy and the degree of supervision that is exercised or should have been exercised." Supervision that left no record is legally indistinguishable from supervision that never happened. Foundation-model developers mostly escape; careless deployers don't; professionals face a ratchet; and the only human-review posture that survives is evidenced, genuine review.
#AIGovernance#AILiability#AIAgents#EnterpriseAI#CISORead on Substack → - Essay 38
Gates the Agent Can't Bypass
Our AI agents were told to run the tests before every commit. They did — often. Here's the enforcement layer that replaced the word "often".
My company's code is mostly written by AI agents, and the instruction said: run the type check and the tests before every commit. The agents did — often. Under deadline pressure an agent rationalises exactly like a tired engineer at 6pm, so we stopped asking. Three layers: the commit demands evidence (a marker only a passing gate run can produce), the receipt expires on any post-run edit, and the bypass flag is blocked at a layer the agent can't modify — with a human-only override key. Week one, the enforced gates caught two flaky tests that had been latent for months. An instruction is a request; if your agent governance lives in a prompt, you have a preference with a policy number.
#AIAgents#AIGovernance#SoftwareEngineering#DevTools#PlatformEngineeringRead on Substack → - Essay 39
Nobody Coordinated This
A stranger from a different corner of software independently built the same governance organs I did. When two lineages that never met grow the same structure, you've found a law, not a fashion.
The octopus eye and your eye were invented twice — separate lineages, same machine, because seeing demands a lens. This month I watched the same thing happen in software: an ex-Google engineer building AI for database incident response and I compared notes across unrelated domains, and the notes were the same notes. The audit trail is the product. A log of actions isn't enough — you need a memory of judgment. Certify the harness, not the model. Consistency isn't correctness. Then a third lineage published the same finding, and it wasn't a builder: England's judiciary-chaired AI-liability taskforce concluded liability turns on the supervision you can prove. Convergence doesn't crown a winner — it proves the pressure is real.
#AIAgents#AIGovernance#SoftwareArchitecture#BuildingInPublic#AgenticAIRead on Substack → - Essay 40
The Review Economy
The team that bought a 5x tool and shipped 1.3x. The bottleneck nobody budgets for is the person who has to say "ship it."
Run the arithmetic nobody runs: in a normal org, review is ~10% of total time. Make producers 5x productive and review demand wants to be 50% — an impossibility, so production falls to meet review capacity. You bought 5x; you ship ~1.3x plus unreviewed inventory. AI is the greatest non-bottleneck accelerator ever built. Orgs that don't engineer review pick one of three defaults without noticing: queue collapse, rubber-stamping (an editor's liability with no evidence), or bypass. The fourth option treats review as a system: price it by risk, make each review cheap with evidence-first artifacts, calibrate the judges, record every accept/reject so review compounds, then graduate to supervising the system itself. Plus the three-number diagnostic most organisations can't answer.
#AIAgents#AIGovernance#EngineeringLeadership#EnterpriseAI#AIAdoptionRead on Substack → - Essay 41
The Plausibility Subsidy
AI made looking good free — so looking good stopped meaning anything. The market for lemons is now inside your company, and the fix is evidence that buys freedom.
For the history of knowledge work, polish was expensive, so polish was information — and every manager alive was trained to read that signal. AI just made polish free. What's left is the plausibility subsidy: plausible work costs ~10x less than correct work, and the gap flows to whatever looks good. Adverse selection for slop isn't a character failure; it's the equilibrium. The fix is a design principle: stop trusting assertions about work and read evidence attached to work — produced by the process, not the author. Five instruments, one uncomfortable implication (AI made senior judgment measurable), two traps (Goodhart, and the fatal one: bossware), and the incentive-compatible core: evidence buys freedom.
#AIAgents#AIGovernance#EngineeringLeadership#FutureOfWork#EnterpriseAIRead on Substack → - Essay 42
Context Is Exclusion
Every company has a vector database. Almost none has anyone who decides what belongs in it — and what makes a library a library isn't the pile of books. It's the catalogue, and the discard pile.
A prediction that sounded wrong two years ago and reads like a diagnosis now: bigger context windows made AI worse at most companies. Not technically — operationally. When the window was small, somebody had to choose what went in, and choosing is where the quality lived. The solo operator's AI works for a reason nobody copies: they aren't a retrieval system, they're a relevance system — the asset is knowing what doesn't matter. Context is judgment about what matters, written down, which means context is exclusion. Three failure shapes (the empty RAG, the everything-bagel, the ownership vacuum), the feedstock everyone throws away (every reviewer correction is a context patch), and the library test.
#AIAgents#RAG#KnowledgeManagement#EnterpriseAI#AIEngineeringRead on Substack → - Essay 43
Earned, Not Configured
Every AI rollout configures trust on day one — and configuration can only guess. The organisations that get autonomy right treat trust as a price the agent pays down with evidence: promotion by track record, demotion by incident, ownership by name.
There are two trust systems in every company. The one for people is earned: track record, calibration, graduated responsibility. The one for AI is configured: someone sets permissions on day one and hopes. Configuration sets the ceiling; evidence sets the level. The ladder that fixes it: autonomy earned per task-class by demonstrated calibration, promotion that is slow and boring by design, demotion that is fast and automatic, an owner with a name, and the price that moves only upward — plus portability: the track record survives the model that earned it. Anything else is a guess with uptime.
#AIAgents#AIGovernance#AITrust#EnterpriseAI#AgenticAIRead on Substack → - Essay 44
Everything Implicit Becomes Infrastructure
Why AI works so well for individuals and so poorly for teams: everything that's free when you work alone must be built when you don't. Seven rows, one law.
Why does AI work so well for individuals and so badly for organizations? Everything that makes an individual productive with AI is implicit — invisible, automatic, free. An organization gets none of it for free; every implicit thing must become infrastructure. Seven rows: context (curation is exclusion), trust (autonomy earned, not configured), review (the economy that caps throughput), enforcement (gates, not prompts), attribution (the provable audit trail the law now asks for), incentives (quality made legible), and memory (judgment that compounds). The table is a company — and the diagnostic is to mark each row infrastructure, implicit, or missing. The dangerous mark is implicit: one person is quietly carrying it.
#AIAgents#AIGovernance#EnterpriseAI#EngineeringLeadership#AIAdoptionRead on Substack → - Essay 45
The Grammar of Operations
DSLs make LLMs reliable. The skipped question is where the right DSL comes from. A semi-algorithm exists: compression over real instances — and in operations, the corpus is the audit log.
Unmesh Joshi's piece on martinfowler.com is right: constrain an agent to a small language with a deterministic validator and generate-and-hope becomes generate-check-repair. But it walks past the hard part — where a good DSL comes from. The folklore says "iterate"; the method is compression: start from a corpus of 20-50 real instances (never a concept), factor repeats into constructs and variance into parameters, test round-trip coverage and back-translation fidelity, stop when new instances stop forcing new constructs. What resists the loop is two judgment calls — choosing the corpus and deciding which variance is essential. And operations already holds the corpus: every escalation, approval, and override a governed agent generates is a labeled instance of what the machine may do.
#AIAgents#SoftwareArchitecture#DomainSpecificLanguages#AIEngineering#EnterpriseAIRead on Substack → - Essay 46
Governance You Inherit
The sequel to Everything Implicit Becomes Infrastructure. I built the whole table — and the seven things break identically for every operation, so the floor is reusable, and reusable governance is the actual product.
The anchor essay drew the table: seven things — context, trust, review, enforcement, attribution, incentives, memory — that are free alone and must be built by a team. This is what building the whole table taught me. When you build all seven, for a patent firm and a managed-IT shop and a wealth advisor, you see the thing you can't un-see: the seven break identically. Same failures, same order. Only the one dangerous action at the top differs — a filing deadline, a production push, money moving — and underneath it sits the identical floor. Which means the expensive part, the part every company is told to build for itself, is reusable. You don't commission electricity; you inherit the grid and pay for what you draw. Governance is the same once the floor is built correctly one time: not a project each company repeats, but a utility they inherit. And I can say it's reusable because I'm customer zero — the agents that build the system run under the same seven rows, and row seven is just the trust-until discipline I already used to keep my own memory from lying to me. The last twenty percent, closed once, so nobody I work with has to close it again.
#AIGovernance#ManagedAI#EnterpriseAI#AIAgents#EngineeringLeadershipRead on Substack → - Essay 47
Same Model, Different Floor
I ran the same tasks through two harnesses on an open enterprise benchmark. The frontier model passed everywhere but paid $2.13 a task on raw interfaces against $0.59 on a governed floor. The cheapest model went from 31% to 94% without changing the model. The difference was never intelligence — it was the floor under it.
This month a vendor published a benchmark that is more interesting than its leaderboard. Both lanes run the same model — and when the model is held constant, a leaderboard stops measuring intelligence and starts measuring harnesses. Their claim: moving an agent from curated tools to protocol-realistic APIs costs 18 to 19 points of accuracy and doubles token cost. A vendor claiming its architecture is the special ingredient is marketing. But this claim is checkable, because they published the dataset. So I checked it: same tasks, same judge, same models, through two doors onto the same 32,768-ticket company. Seventy-five trials later, the frontier model passed everything on every interface but paid ~$2 and ~29 turns a task on raw surfaces against $0.59 and 6 turns on the governed tier. The cheapest model went from 5 of 16 on raw APIs — failing by silently dropping records it never knew existed — to 15 of 16 on the best floor per task, including one cell at eight for eight at four cents a run. Fixing the weak model with a bigger model costs about $1.70 a task, forever. Fixing it with a better interface costs a dime. And the decisive layer was not pre-computation: it was two hundred lines of declarative topology — the map of how the sources relate. Halfway through, the experiment audited me back: my own curated tier shipped a misread threshold, and the confession is the best section.
Read on Substack → - Essay 48
The Underwriting Variable
The sequel to The Law Just Located the Risk. The courts put AI liability on supervision you can show; the insurance market just located the same risk on the same axis — and attached a price.
Earlier this month the courts located the risk of autonomous AI: liability turns on the supervision you can show. Last week a second system located the same risk on the same axis, and did what the courts do not, it attached a price. Insurers are filing AI exclusions across general liability, D&O, and E&O. An exclusion does not reduce AI risk, it relocates it from the carrier to the policyholder, and the carve-out for loss arising out of AI turns one question into the whole game at renewal: can you show the loss arose from a licensed human decision, or did it arise out of the model? Same event, opposite sides of the same clause. Two systems that do not coordinate, doctrine and price, reached for the same axis, provable human accountability, so it is not an opinion you can wait out. The market is the sharper signal because it reprices at renewal instead of after a bellwether case, and it converts the audit trail from the line item nobody could justify into the instrument that decides a paid claim from a denied one. Governance stopped being compliance the moment the market made it the underwriting variable.
#AIGovernance#Insurance#EnterpriseAI#RiskManagement#ErrorsAndOmissionsRead on Substack → - Essay 49
The Selection Problem
In July I said AI would explode the variance between workers. The research says the opposite on most work. What it disperses is judgment — and the fear was never replacement.
Three weeks ago I argued that AI amplifies rather than levels, that the mean rises while the variance explodes. Then I read the research and had to take my own argument apart. Put a generative assistant in front of thousands of support agents and the novices gain about 34% while the experts gain almost nothing; a Science paper on writing tasks reports that inequality between workers decreased. These tools compress performance on bounded work, because the model already contains the expert answer and handing it to a novice closes the gap. Where they disperse it is open-ended work, and a field experiment with 640 Kenyan business owners isolates why: identical questions in, identical advice out, opposite outcomes, because of which advice each owner selected and implemented. The variable was judgment. Which reframes the workforce anxiety too. Only 22% worry about losing their job to AI; the top fear, at 51%, is being expected to do more for the same pay. That is not a fear about technology, it is a fear about who banks the gain, and it is the same question at company scale, where metered inputs hand the surplus to the buyer only if the buyer can operate the capability. If judgment is the scarce input, the organization that gets ahead is the one that makes good selection institutional rather than individual, which is what a review record is actually for.
#AI#FutureOfWork#Productivity#AIGovernance#ManagementRead on Substack → - Essay 50
Why AI-Generated Animations Need a Verifier
A model can produce an animation in seconds. Whether it's correct is a coin flip. The gap between those two facts is where all AI-generated work now lives.
Ask a good model to animate a signup confirmation and you get something plausible in seconds — and, roughly half the time, subtly wrong. The research puts unaided LLM animation synthesis around 59% correct; wrap the same generation in a formal verify-and-correct loop and it climbs to about 94%. That gap is not about intelligence — it is about whether anything checks the work. For code the checker is tests; for a financial report it is reconciliation; for animation it was nothing, because animation is a claim about time and the eye is bad at time. Frame-diffing tools cannot evaluate 'the toast slides in before it fades and settles by 1.2 seconds' — a frame has no memory of the frames before it. So I built the missing instrument and put it in the open: Choreo, MIT-licensed, asserts the motion contract and proves the rendered animation satisfies it by sampling the real trace. The verifier taught the usual honest lesson — it is only as good as the contract you give it — and the general point is the whole game now: as agents produce more of everything, the scarce input stops being generation and becomes the checkers. The model is the commodity. The floor under it is what's durable.
#AIGovernance#AIAgents#WebDev#Verification#OpenSourceRead on Substack → - Essay 51
Reviewed, but Not Reconciled
Bookkeepers have a precise name for work that looks checked but was never tied to its source. AI just made that state free to produce — and every profession is inheriting the problem.
There's a thread on r/bookkeeping where a bookkeeper takes over a client's accounts and finds the last bank reconciliation was done in February. It's November — 17,300 uncleared transactions, eight months of books that look finished and prove nothing. The phrase the thread settled on: reviewed, but not reconciled. Reviewed means someone looked — a fact about the reviewer's attention. Reconciled means every number was tied back to its source and every difference explained — a property of the work itself. Almost all knowledge work runs on review, and that mostly worked because producing polished, wrong work used to be expensive, so polish carried information. AI ended that: a confident model answer is a ledger marked reviewed — the look of examined work, tied to nothing — produced at machine speed. The bookkeepers' deeper lesson is where verification bottoms out: the ritual anchors on the bank reconciliation, tying the books to a record nobody inside the company can edit, because the check must come from outside the system being checked. The generator grading its own output is exactly the 'verified' that means a colleague clicked a button. The test to run this week: for each artifact your operation shipped, ask the bookkeeper's question — reviewed, or reconciled? The bookkeepers were early. We're inheriting their problem; we should inherit their standard with it.
#AI#FutureOfWork#AIGovernance#Management#FinanceRead on Substack → - Essay 52
The Verifiability Frontier
The work that automates first is the work someone built instruments to check. Double-entry, software tests, and one real set of books say the automation boundary isn't fixed — it's built.
A $30B software CEO argued last week that the hardest knowledge work automates first because it's the most verifiable — and that we may need new capabilities to test knowledge work. Right axis, wrong verb. Verifiability isn't a property a domain has; it's a property someone installs. Five centuries ago bookkeeping was judgment work, until double-entry and the bank reconciliation turned most of it into arithmetic anyone can check. Code wasn't born verifiable either: fifty years of tests and CI built the harness, and then we forgot it was built. So the useful question about any domain isn't whether it's verifiable — it's what the instrument would look like. The move that works is decomposition: 'are the books right?' is unanswerable by machine, but 'do these ten statement transactions tie to the register?' came back 10/10 matched, difference $0.00 — leaving a human exactly three items that actually deserved judgment. The same books were marked reconciled through a date a year past, with an $8,299 gap: reviewed, never tied out. Where ground truth arrives late, install a scored track record instead — receipts, gates, autonomy earned against outcomes. The automation boundary is the high-water mark of the instruments built so far. It moves when someone moves it.
#AIVerification#KnowledgeWork#AIGovernance#AutomationRead on Substack → - Essay 53
You can't grade generalization on the training set
A system that learned and a system that memorized the answer key get the same benchmark score. The only test that tells them apart is whether it works on what you didn't build it for.
There's one number in AI everyone trusts and no one can audit: the benchmark score, and its strangest property is that it stays true while you fool yourself — a system that learned a skill and one that memorized the answer key produce the identical number. We built a governed layer that made a weak model pass questions it failed raw; the number went up; it felt like a result. Then someone asked whether it worked on the questions we didn't build it for, and it did nothing — twice it did worse. We had mistaken depth (one problem solved to exhaustion) for breadth (many problems handled at all), and the score had happily let us. A benchmaxed result and a general one look identical right up until transfer. For anyone buying AI that's the whole game: a vendor tuned to your pilot data shows a real number, for that data, then collapses on the questions you didn't send. There's only one question that carries information — does it work on what you didn't build it for — so ask which result came from a problem shape they'd never seen before you handed it over, because once is not a demonstration.
#AI#EnterpriseAI#MachineLearning#Benchmarks#AIEvaluationRead on Substack → - Essay 54
Build the space, not the points
A benchmark's questions collapse into a handful of shapes, and the shapes are just relational algebra. Build the algebra once, put a parser in front, and coverage stops being a function of effort.
The narrowness in the last essay wasn't a ceiling — it was a consequence of how we'd built the thing: one hand-written solver per shape of question, coverage growing exactly as fast as we kept writing solvers. Then we wrote the question-shapes down next to each other, vocabulary stripped off, and they stopped looking like a zoo — group, aggregate, join, filter, rank, relational algebra with a few analytic operators bolted on, a short closed list. So we stopped building points and built the space they sit in: the operators once, verified once, with a parser in front that compiles any plain-English question into them. Coverage went from four shapes of nine to nine of nine, most of the last stretch composition rather than code. Generalization stopped being a feeling: the parser wasn't fit to any one question, so every unseen question became a real test of it, with a number that moved as we grounded it. Then it walked into the oldest trap in the box — asked for revenue 'by product area,' it rolled up and discarded the detail the question also wanted — which taught the lesson under the lesson: recognizing the shape is easy, keeping the discipline is hard, so you push the rule down into the engine where a cleverly-worded question can't override it. Breadth is the part you can buy; discipline is the part you build so deep nothing can spend it. The engine is open source.
#AI#OpenSource#LLM#DataEngineering#EnterpriseAIRead on Substack → - Essay 55
The Plane Above the Engine
An individual gets structure, review, and memory for free; a team must build them. Building taught me where they belong: not in a dashboard above the work, but at the substrate where the agent runs.
A month ago I argued that everything a solo operator gets free from AI — context, trust, review, memory — has to be built as infrastructure the moment a team is involved. Then I spent the stretch after building it, and building a thing tells you what arguing about it cannot. Three lessons changed how I see the whole problem. The seven pieces aren't seven features; they hold each other up, so you build one layer, not seven. That layer is a plane, and it sits above the engine on purpose: the machine an agent runs on is right for the work and wrong for managing it — private, temporary, built for one — so the record has to move up to somewhere shared that outlives the run. And a layer above can only review the past, unless you change what the agent may do: so the agent does only the reversible half — gather, draft, propose, record — and hands every irreversible call up to a human. Reversible proceeds; judgment waits. That is why a plane above can govern the live moment: there was never anything to intercept on the machine, because the risky act was never the agent's to take. The engine is turning into a commodity. The layer that manages the labor it performs is the part that is not, and who operates it is the moat.
#AIAgents#AIGovernance#AgenticAI#EnterpriseAI#EngineeringLeadershipRead on Substack → - Essay 56
Everyone Builds the Record. Nobody Builds the Receipt.
Two agent harnesses built the same intent-then-settlement record, for crash recovery rather than compliance. Both stop where the record would have to mean something to someone who wasn't there.
Two agent harnesses landed on my desk this week: pi, a deliberately minimal terminal tool that ships seven capabilities and refuses an eighth, and dsh, DeepSeek's 219-package plugin tree with a kernel-level sandbox. They agree on almost nothing, and both had built the same three-step shape I had built myself — name the intent before you act, settle it after, and keep an identity that survives the gap in between. Neither built it for governance. pi's design document is about crash recovery; dsh's is about replay fidelity. Both were answering the same unglamorous question: the process can die at any moment, so what do I write down, and when? You cannot answer that honestly and end up anywhere else. Write the record after the action and a crash loses the fact that you acted; write it before and you must name what you are about to do, give that intent an identity so the settlement can find it, and mark which effects may be safely repeated and which never can. Work it through and you have built an approval-and-attestation record — not because anyone required one, but because durability taken seriously has the same skeleton as accountability. Then both projects stop at almost the same line, and both are unusually explicit about it. pi declines a permission system outright: run in a container, or build your own confirmation flow. dsh built real enforcement, a Linux sandbox and a fail-closed gate, then refused to define a parent-to-child authority lattice — a parent may own a child whose visible tools are wider than its own. Its approval events are log-only; they never leave the machine that made them. Which is the whole distinction. A log is written for a reader who was present and can assume the context. A receipt is written for someone who was not there: a customer, an auditor, a board asking in eighteen months why the system did something in March. A receipt has to outlive the session, name a person rather than a process, pin what that person could actually see when they decided, and stay legible to someone with no access and no reason to take your word for it. Everyone is building the record. The receipt is a different object.
#AIAgents#AIGovernance#OpenSource#SoftwareArchitecture#AuditabilityRead on Substack → - Essay 57
The Failure That Says Yes
The first run of our verification page reported that our audit chain had proven nothing for a month. What we found next was quieter and worse. On the two failure modes of attestation systems.
The interesting failures in attestation systems are not the ones that refuse to confirm. We learned this by pointing our own instrument at ourselves: the page we built so prospects could verify our audit chain reported, on its first run against production data, that the chain had proven nothing for a month. Verification failed at entry 477 — a rolling deploy had let an event be written by new code and sealed by old code. Written by the future, notarized by the past. Benign, loud, and safe: the chain refused to confirm what it could not confirm. That is the failure that says no. Two days later a signature migration surfaced the other kind. When an account is deleted our cascade removes its audit events, but the chain's bookkeeping survived — so for eleven of a hundred-eighty chains in production, the verifier queried for events, found none, never entered its loop, and returned: valid. A chain whose entire history had been erased reported as pristine. That is the failure that says yes — it fails toward confidence, and it produces exactly the same output as health. Once you have the shape you see it everywhere; we found a third instance in the verification page itself, computing its headline absence claim over a capped query window nobody had proven complete. The fix we refused matters more than the fixes we made: recomputing one hash would have made the chain verify clean, and it is precisely what an attacker would do. Instead the break is signed, disclosed, and permanent — verification continues past it, eight thousand nine hundred and seventy events re-evidenced, and the page shows the discontinuity to prospects. The question that separates governance vendors is not whether they have an audit trail. It is: when did your attestation system last fail, in which direction, and how did you find out? An instrument that can only confirm is not an instrument.
#AIGovernance#AuditTrail#TamperEvidence#AICompliance#TrustButVerifyRead on Substack → - Essay 58
The Inspector Who Wasn't There
A record that satisfies its operator is a practice. A standard begins where someone with no stake can check it - so we published the spec, open-sourced the verifier, and let the census list our debt.
Essay #56 ended at a wall: the record stops mattering exactly where it would have to mean something to someone who wasn't there. Then Louisa Johnson Bullock, founder of Regulayer, gave me the distinction I could not unhear - a control that satisfies the people who run it is a practice; a standard begins when the conditions and entry rules are checkable by someone with no stake in the answer. By that definition my hash-chained, externally anchored audit trail was an excellent practice: every record still needed me in the room, to explain the fields, vouch nothing skipped the logger, run my own verifier on my own export. The operator hides in four places - meaning, verification, completeness, re-execution - so we spent a week on the first three: a schema generated from the same code that writes the records, CI-failed on drift, stating plainly what the records do not prove; an open-source verifier implemented from the published text rather than our code, conformance proven by golden vector; and a census of our own audit coverage that publishes its fifty-two-route debt as a shrink-only ratchet instead of rounding it to zero. Completeness is still testimony - but bounded, enumerated, machine-checked testimony. The sentence an inspector can sign changed countries.
#AIGovernance#AuditTrail#OpenSource#Verification#ComplianceRead on Substack → - Essay 59
The Seven-Block Machine
Every computer has five parts because its arithmetic unit never had to decide anything. A machine whose core unit judges needs two more, and the labs are building neither. Who builds them is the game.
Von Neumann's five blocks - control, arithmetic, memory, input, output - have survived every generation of hardware because they name what any machine that does work needs. Map them onto an agent stack and the obvious version puts the language model in the arithmetic slot. DevRev's public leaderboard shows why that is the wrong reading: the same model through different harnesses lands thirty to forty points apart, while swapping the frontier model for an open-weight one inside a single harness moves the score by under three. The interface effect is roughly ten times the model effect, which is the lesson computing already learned once, in the 1980s, when the compiler and the memory hierarchy overtook the clock. So the model is a fast, stateless judgment unit; the harness is the cache hierarchy; and the thing that lays a company's operands out for it is a compiler. A deterministic machine never needed two more blocks. A machine whose core unit judges does: a gate - a place where certain actions trap to a human, with the trap recorded - and a write-back - a mechanism that decides what an outcome is allowed to make permanent. Four of our own incidents this summer were failures in those two blocks, not in the model. The labs will own the first five blocks. The operator has to own the last two, because they require standing inside the customer's organization over time. Between them sits the compiler, whose input is the one thing the lab does not have. A universal floor is a contradiction in terms. A universal floor compiler is not.
#AIAgents#AIGovernance#EnterpriseAI#AgenticAI#EngineeringLeadershipRead on Substack → - Essay 60
Rings for Judgment
Every agent-safety failure we logged this summer was a permission bug, not a model bug. The fix was forty years old: privilege rings - and a ring the agents cannot rewrite.
We keep a public list of the ways our own safety rules can be beaten, and this summer it gained five entries in ten weeks. Laid side by side, none was a bad answer. A phantom sixteen-hundred-line revert left by a careless reset; two cards carrying a cover-image path for files that did not exist, invisible for six days because the console displayed the path as text; a job that finished but never closed its assignment, so the next session drafted the same post again; a reply built from the wrong lookup table; a subagent's clean commit that rode into a stranger's pull request because its parent switched branches underneath it. Every one was a good worker doing a permitted thing where it should not have been permitted. That is a permission problem, and permission problems have a forty-year-old answer: rings. Ring 0 is the enforcement layer itself - the hook, the guard, the deny rules - which only a human commits, because the thing that is the gate cannot be changed by the thing it gates; an agent prepares the patch script, the founder runs it. Ring 1 is the irreversible or outward act - publish, send, money, protected data - which traps to a person with a recorded decision. Ring 2 is reversible, attributable work behind a pull request. Ring 3 is reading and drafting. Numbering the rings found what prose had missed: a send scope every seat held by default, a policy block that could be approved past without a reason, a commit with no address-space check. Seven of eleven gaps closed in a day, and the patch that closed one of them over-blocked innocent reads until it was fixed the same afternoon. The residuals are named by id. The rule that fell out: every write from a lower ring must be checkable by the ring above it.
#AIGovernance#AgentOps#SecurityRead on Substack → - Essay 61
What Deserves to Be Remembered
An AI model cannot remember what it decided yesterday. The part you build around it decides what becomes permanent, and we found five of our own memories had no rule for what gets in.
On August 21 a recommendation was executed on evidence that had been true the day before and was false by the time anyone acted on it; the execution left a code comment asserting a state of the world that had stopped being true eighteen hours earlier. Nothing broke, but it named the thing. A language model has no write-back path: it cannot carry what it decided into tomorrow. Everything an operator builds around it to do that - registers, trackers, rules files, research corpora, receipts - is a mechanism for deciding what an outcome may make permanent. We audited seventeen of ours. Five had no stated criterion for admission, including the two that agents actually read from when they draft: the operating-rules file, and the research corpus whose only bar was that something had been registered. Our follow-up register verified an item once at sweep time and then served it as true forever, while a sibling ledger gave every lesson a horizon. Receipts were written on every run, quality-blind; lessons were written almost never. The only clean instance was the shrink-only ratchet. The rule we adopted has five clauses in dependency order: a thing becomes durable only if someone specific will act differently because it exists; it enters and leaves with a reference another person can check; it states what would make it false and when it must be re-checked, and past that date it is served as a question; a second occurrence after a fix is an infrastructure failure, never a second item; and durable state is never rewritten, only appended, superseded with a pointer, or lapsed with a date. Activity is not memory. The first clause is the whole selection problem.
#AIAgents#AIGovernance#EnterpriseAI#AgenticAI#EngineeringLeadershipRead on Substack → - Essay 62
The Floor Is a Compiler
I predicted our governance floor would lift a weak model fifteen to twenty-five points. It scored below baseline. What recovered it was not more hand-building. It was a compiler.
In July I wrote that the same model under a different floor scores differently, and that the floor is the product. Then I pre-registered a prediction, in a file with a no-edit rule: our hand-built floor would lift a weak model fifteen to twenty-five points on an open enterprise benchmark. It scored five of fourteen against the baseline's six. The failing traces were not missing views; they were wrong-grain answers delivered confidently in two turns - a nine-hundred-thousand-dollar figure at the parent level where the task wanted seven-hundred-thousand at the part level. My first lesson, structure over computation, was wrong too: advisory text had zero effect at any tier, and advertising a grain parameter sent a task from five of five to zero of five until the parameter was removed. What recovered the floor was computing at the correct grain with no choice left to the agent - twenty-seven to fifty-three percent at the cheap tier over a hundred and forty cells, seventy-one to eighty-six at the frontier tier at sixty percent of the cost - and everything that recovered it turned out to be source-independent: key inference, join planning, grain discipline, deduplication, provenance. So we wrote a compiler. From a derived linkage graph and a six-line spec it reproduces the hand-written view exactly and produces a correct view on a topology it was never built for. We have not re-run the benchmark on compiler output; that is the next registered prediction, and we have submitted nothing publicly. The one-page receipt says thirty-six against forty-three in its third paragraph. Same model, different floor is still true. The product is not the floor. The product is the thing that writes it.
#AIAgents#AIGovernance#EnterpriseAI#Benchmarks#EngineeringLeadershipRead on Substack →
New essays + audio podcast version, weekly
Subscribe to get each essay in your inbox + on the podcast feed.