<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Twistag — Thinking</title>
        <link>https://twistag.com/thinking</link>
        <description>Field notes on AI, agents, data, and modern engineering from the Twistag team.</description>
        <lastBuildDate>Thu, 30 Jul 2026 15:12:58 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <image>
            <title>Twistag — Thinking</title>
            <url>https://twistag.com/images/og-default.jpg</url>
            <link>https://twistag.com/thinking</link>
        </image>
        <copyright>© 2026 Twistag</copyright>
        <item>
            <title><![CDATA[Build, buy, or build-with-you: the third path for AI agents]]></title>
            <link>https://twistag.com/thinking/build-buy-build-with-you</link>
            <guid isPermaLink="false">https://twistag.com/thinking/build-buy-build-with-you</guid>
            <pubDate>Thu, 30 Jul 2026 00:18:09 GMT</pubDate>
            <description><![CDATA[Forrester says 75% of self-build AI agent attempts will fail. Buy locks you to a vendor. The third path: build-with-you, the model that ends with your team owning the system.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: May 2026</em></p><p>Most build-vs-buy frameworks for enterprise AI agents present a binary choice and quietly pick a side. Forrester says 75% of self-build attempts will fail. Buying a packaged platform locks the buyer to a vendor's roadmap and leaves no in-house capability behind. Neither outcome is what the CTO actually wants. This post is about the third option the market is moving to: build-with-you, the embedded-engineering model that ends with the buyer's team owning the system the partner shipped. It's the cluster pillar for our work on capability transfer and how it fits inside <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a>.</p><h3>Key takeaways</h3><ul><li><strong>Build alone fails most of the time.</strong> Forrester predicts 75% of enterprises attempting to build their own agentic systems will fail, citing complexity in retrieval, data architecture, and niche expertise.</li><li><strong>78% of enterprises want AI in-house</strong> (<a href="https://zapier.com/blog/enterprise-ai-statistics/">Zapier, 2026</a>) — 38% don't trust vendor security, 33% fear vendor lock-in. The destination is in-house ownership. The route is the question.</li><li><strong>Buying buys time, not capability.</strong> A packaged agent platform ships in weeks but leaves no proprietary IP, no in-house engineers who can extend it, and a recurring bill tied to a vendor's roadmap.</li><li><strong>Build-Operate-Transfer is older than agentic AI.</strong> BCG, Deloitte, and Indian outsourcers ran the BOT model for years in software delivery. Adapting it to AI delivery is the shape of the third path.</li><li><strong>Big 4 firms now deliver via the <a href="https://openai.com/">OpenAI Frontier Alliance</a></strong> — McKinsey, BCG, Deloitte, Accenture. That changes what a Big 4 engagement is, and creates a specific opening for boutique engineering partners that ship without the Alliance overhead.</li></ul><h3>Why the binary frame is wrong</h3><p>The question "should we build AI agents in-house or buy a platform?" treats two endpoints as the only options. In practice every CTO we work with wants to end up at the same destination: a production AI capability they own, that their team can operate and extend. The disagreement is about the route.</p><p>Pure build assumes the team that will eventually own the agent can also be the team that designs it from scratch. That is rarely true at the start of an agentic programme. Hiring senior agent engineers in 2026 takes six months and the budget for a full team usually doesn't exist before a working system has shipped. Forrester's 75%-fail estimate isn't about engineering talent — it's about the structural impossibility of building production-grade retrieval, evaluation, governance, and integration in parallel while also hiring the team that knows how.</p><p>Pure buy assumes the standardised use cases vendors package will match the enterprise's actual workflow. Sometimes they do. For meeting summaries, support ticket routing, content moderation, document classification — packaged agents work. But the value-creating use cases inside a specific enterprise — the ones that justify a budget — usually don't fit a packaged shape, and configuring around the edges leaves both a vendor dependency and a system the in-house team can't fully control.</p><p>The third path resolves the contradiction: an external team builds, the in-house team learns by working alongside them, ownership transfers throughout the engagement, and the partner leaves when the buyer's team can run the system independently.</p><h3>The build path: what it actually costs</h3><p>The headline numbers for pure-build agentic systems in 2026 are widely reported across industry pricing surveys. A focused custom agent MVP runs $50K-$100K. A full multi-agent system runs $250K-$400K+ in initial build, plus recurring senior-engineer salaries, observability tooling, model inference costs, and infrastructure operations. Build also commits the buyer to an opportunity-cost path: every senior engineer working on the agent is not working on the core product.</p><p>The harder cost is structural. Building a production agent requires expertise in retrieval architecture, prompt and model evaluation, multiagent orchestration, observability, and agent governance — five specialisms that overlap on a CV in maybe one in fifty senior engineers. A team building its first agent will learn the hard parts by hitting them in production. Most stop hitting them at all, because the system gets shelved before it reaches that stage.</p><p>The cases where pure build is the right call are narrow: the use case is genuinely core IP, the team already has a senior agent engineer on staff, the use case is sensitive enough that no partner is acceptable. For most enterprises starting an agentic programme in 2026, none of these conditions hold.</p><h3>The buy path: speed at a cost</h3><p>Packaged agent platforms — Kore.ai, Sierra, Decagon, Aisera, Vertex AI Agent Builder, Amazon Bedrock Agents — solve the time-to-value problem. Configuration takes weeks. Compliance is inherited from the platform. Up-front engineering cost is near zero. For the 90% of use cases that fit a standardised shape, buy is the right answer.</p><p>The trade-offs are familiar but worth naming. <strong>Vendor lock-in</strong> — 33% of enterprise leaders cite this as a buying concern (Zapier, 2026). <strong>No IP</strong> — the prompts, models, and integrations are the vendor's, not the buyer's. <strong>No in-house capability</strong> — the team running the platform is running a vendor's product, not learning the engineering that would let them build the next agent themselves. <strong>Recurring cost tied to vendor roadmap</strong> — pricing rises as usage scales, and the vendor's priorities are not the buyer's priorities.</p><p>For meeting summaries and support routing the trade-offs are acceptable. For the agentic workflows that justify enterprise-scale budgets — the ones that read from and write back to core systems, that touch regulated decisions, that compound competitive differentiation — the trade-offs are not acceptable, and the buy path is the wrong shape.</p><h3>The third option: build-with-you, defined</h3><p>Build-with-you is what the Build-Operate-Transfer model from infrastructure outsourcing looks like when adapted for agentic AI delivery. The shape is the same: an external partner builds the capability, operates it through a transition period, transfers ownership to the buyer's team. The adaptation for AI is in what gets transferred — not just code and infrastructure, but the eval pipelines, the prompt versioning workflows, the agent registry, the observability dashboards, and the in-house team's ability to extend the system without the partner.</p><p>The three properties that distinguish build-with-you from build and from buy:</p><ol><li><strong>Senior engineering depth at the start.</strong> The first six to ten weeks are the partner's senior engineers doing the work, not training the in-house team. The system needs to ship before the team can learn from a working example.</li><li><strong>Continuous capability transfer, not a handover at the end.</strong> Pair working, documented decisions, code reviews of the in-house team's first production changes. The transfer happens throughout the engagement, not in a one-week wrap-up.</li><li><strong>A clear exit signal.</strong> The partner leaves when the in-house team has shipped their own production change without the partner's involvement. Not on a calendar date, not at a budget threshold — on the operational signal that the team can run the system.</li></ol><p>The model is older than agentic AI. We have been running it at <a href="/case-studies/datatalks-customer-data-platform-sports/">Datatalks since 2018</a> — eight years of embedded senior engineering on a customer data platform that the Datatalks team now operates and extends with us still alongside them when scope expands. We ran it more recently at <a href="/case-studies/ai-invoice-automation-aralab-manufacturing/">Aralab</a>, where a Claude Sonnet 4.5 invoice agent was built by our senior engineers and transferred to Aralab's finance and engineering teams across a six-month engagement. We have run it twice with <a href="/case-studies/ai-training-platform-sana-hotels/">SANA Hotels</a> — once for an AI training platform, once for <a href="/case-studies/ai-workforce-optimization-hotel-group-portugal/">staff optimisation</a> — with capability transfer to the SANA operations team built into each phase.</p><h3>How build-with-you compares to the four options on the table</h3><p>Most CTOs have four options to choose from when scoping an agentic engagement. The differences matter because each delivers a different thing at a different cost with a different exit state.</p><table><thead><tr><th>Criterion</th><th>Pure build</th><th>Buy a platform</th><th>Big 4 / Alliance partner</th><th>Build-with-you</th></tr></thead><tbody><tr><td>Time to a production pilot</td><td>9-18 months</td><td>Weeks</td><td>6-12 months</td><td>6-12 weeks</td></tr><tr><td>Up-front investment</td><td>$50K-$400K+ build + team</td><td>Low</td><td>$500K-$5M typical</td><td>$80K-$300K for MVP</td></tr><tr><td>In-house capability at exit</td><td>Built it themselves (if they succeeded)</td><td>None</td><td>Strategy slides + a deployed system</td><td>A working system the team can operate and extend</td></tr><tr><td>Vendor lock-in risk</td><td>None</td><td>High</td><td>Medium (Alliance dependencies)</td><td>Low</td></tr><tr><td>IP ownership</td><td>Client</td><td>Vendor</td><td>Mixed (often client)</td><td>Client owns full stack at handover</td></tr><tr><td>Best for</td><td>Genuinely core IP, existing senior team</td><td>Standardised use cases</td><td>Multi-system enterprise rollouts</td><td>Validated use case, needs production engineering, destination is in-house</td></tr></tbody></table><p>The comparison isn't a value judgment on any path. Each is right for a different question. Build-with-you is the right answer when (a) the use case is validated, (b) the destination is an in-house capability the team can run, and (c) the buyer wants the system fast without the vendor lock-in of buy or the failure rate of build.</p><h3>What changed: the Big 4 are now Alliance delivery partners</h3><p>In 2026 McKinsey, BCG, Deloitte, and Accenture became delivery partners in the <a href="https://openai.com/">OpenAI Frontier Alliance</a> — getting early access to unreleased models, dedicated engineering support, and co-development pathways for enterprise clients. This is a real shift in what a Big 4 AI engagement contains. It also creates a specific opening for boutique engineering partners.</p><p>The Big 4 Alliance model fits enterprises whose primary need is enterprise-scale rollout — multi-business-unit programmes, change management at thousands of seats, governance at parent-company level. The cost matches: Big 4 enterprise AI engagements start at $50K-$500K+ and scale from there.</p><p>The boutique engineering shape — what build-with-you actually is — fits enterprises whose primary need is a working production system shipped quickly by senior engineers without the Alliance overhead. The cost is lower because the engagement is engineering-led rather than advisory-led. The exit state is the same as Big 4 in terms of a deployed system, with a clearer ownership transfer because the engagement was scoped that way from the start.</p><p>Enterprises increasingly choose both: a Big 4 firm for strategy and change management at the parent level, and a boutique build-with-you partner inside specific business units for the production engineering. We work directly with Big 4 partners on this pattern when the structure calls for it. It is not a competition. It is two different shapes of work that pair well.</p><h3>When build-with-you is the wrong call</h3><p>Naming the limits of the model is more useful than overselling it.</p><p><strong>The use case isn't validated yet.</strong> If the buyer is still deciding whether agentic AI is the right answer for a specific workflow, the engagement should start with a strategy advisory or a discovery sprint, not a build-with-you engagement. Build-with-you assumes a validated use case to scope around.</p><p><strong>The destination is permanent outsourcing.</strong> If the buyer's intention is to run the agentic system through an external operator forever, the build-with-you model is over-engineered for the goal. Pure managed services would deliver the same operational outcome at lower cost.</p><p><strong>The buyer doesn't have or want an in-house engineering team.</strong> Build-with-you transfers capability to a team. If there is no team to transfer to, the engagement ends with a system nobody can operate. Better to use a managed-service partner.</p><p><strong>The use case is genuinely commodity.</strong> If the workflow is meeting summaries, document classification, support ticket routing — a packaged platform configured by a vendor partner ships the same outcome in less time at lower cost. Build-with-you exists for the non-commodity case.</p><h3>The handover signal: what tells the partner to leave</h3><p>The cleanest answer to "when does the partner leave?" is also the most operationally honest: when the in-house team ships a production change without the partner's involvement. Not a date on a contract, not a milestone in a slide, not a budget threshold. The operational signal.</p><p>What that looks like in practice. The in-house team writes a new tool definition for an agent — an additional API the agent can call. They write the schema, they write the eval cases, they ship the change through CI, the agent uses it in production correctly. The partner reviewed the PR but didn't write the code. The next change after that, the partner doesn't review. The change after that, the in-house team has shipped three production changes without us.</p><p>That is the signal. It usually arrives between month four and month nine of the engagement, depending on the team's starting expertise and the complexity of the agent. When it arrives, the partner steps out. We have done this at Datatalks (eight years and counting because the scope keeps expanding, but the core CDP capability transferred years ago), at Aralab (transferred at month six, called back for an adjacent module a year later), and across two engagements with SANA Hotels (transferred phase by phase).</p><p>The exit is not the end of the relationship. Build-with-you partnerships often turn into multi-engagement relationships where the partner is called back for the next adjacent system. But the call is the buyer's, not a renewal of a managed-services contract. That difference is the whole point of the model.</p><h3>What this means for the buyer scoping an engagement</h3><p>Three questions clarify whether build-with-you is the right shape:</p><p>The first question is about destination: where does the buyer want the system to live in two years? If the answer is "operated by our team, extended by our engineers, using IP we own," build-with-you fits. If the answer is "operated by a vendor we pay forever" or "operated by an Alliance partner inside a strategic relationship," a different model fits.</p><p>The second question is about validation: is the use case validated, or is it still being scoped? Build-with-you is an engineering engagement, not a strategy engagement. It assumes the use case has been validated through a discovery sprint, a strategy advisory, or the buyer's own internal work.</p><p>The third question is about timeline urgency: how fast does the production pilot need to ship? Build-with-you delivers a production MVP in six to ten weeks. Big 4 engagements take six to twelve months. Pure build takes nine to eighteen months. If the buyer needs the working system in this quarter to justify the budget, the Big 4 path is too slow and pure build is too risky.</p><p>The buyers we work with answer all three questions in a way that points at build-with-you. We have not been the right answer for every CTO we have talked to. We have been the right answer often enough that the model is the operating spine of the <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a> we deliver.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-build-buy-build-with-you.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[EU AI Act deadline moved to December 2027 — what actually changed]]></title>
            <link>https://twistag.com/thinking/eu-ai-act-deadline</link>
            <guid isPermaLink="false">https://twistag.com/thinking/eu-ai-act-deadline</guid>
            <pubDate>Thu, 30 Jul 2026 00:16:33 GMT</pubDate>
            <description><![CDATA[The EU AI Act high-risk deadline just moved from Aug 2026 to Dec 2027. What actually changed, what's still binding, and why the delay is a trap.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Last week the EU Council gave final approval to the AI Omnibus simplification package. The August 2, 2026 deadline for high-risk AI systems moved to December 2, 2027. Embedded AI under Annex I moved to August 2, 2028. The reason is that the harmonised standards required to define what "comply" means are not ready — and the Commission decided to wait for standards rather than enforce against a rulebook that had not been written. This is not "you have more time to do nothing." Some obligations moved up, some are already in force, and the engineering that will land you inside the new dates is a year of work regardless. This is the pillar post for our cluster on EU regulatory readiness. It complements the <a href="/post/production-ai-agent-governance">engineering-grade governance reference we published on the five surfaces</a> with the regulatory-timing view.</p><h3>Key takeaways</h3><ul><li><strong>The Aug 2, 2026 high-risk deadline moved.</strong> Stand-alone Annex III systems now comply by December 2, 2027. Annex I embedded AI (medical devices, machinery, vehicles) by August 2, 2028. <a href="https://www.consilium.europa.eu/en/press/press-releases/2026/05/07/artificial-intelligence-council-and-parliament-agree-to-simplify-and-streamline-rules/">Council final approval June 29, 2026</a>.</li><li><strong>Synthetic content transparency moved <em>up</em>.</strong> The grace period dropped from 6 months to 3. AI-generated content watermarking obligations now apply from <strong>December 2, 2026</strong> — five months from now. Most enterprises that will be affected have not started.</li><li><strong>GPAI obligations have been in force since August 2, 2025.</strong> If your agent stack calls a GPAI model placed on the market after that date, provider obligations apply to that model right now. Commission enforcement powers on GPAI kick in August 2, 2026 — four weeks away.</li><li><strong>The delay is because standards aren't ready.</strong> Harmonised standards under CEN-CENELEC JTC 21 are still being drafted. That work is happening now, and the shape of "comply" is being decided by the standards bodies engaging with the Commission — not later, and not by the enterprises that wait until 2027.</li><li><strong>The engineering hasn't changed.</strong> Whatever "high-risk agent" means when the standards land, it will still require the <a href="/post/production-ai-agent-governance">five engineering surfaces</a> — identity, tool allowlists, evaluation, observability, audit. Starting that build now against a moving target beats starting it in Q3 2027 against a fixed one.</li></ul><h3>What the AI Omnibus actually did</h3><p>On June 29, 2026, the Council of the EU gave final approval to the AI Omnibus simplification package, following the European Parliament's endorsement on June 16 (<a href="https://knowledge.dlapiper.com/dlapiperknowledge/globalemploymentlatestdevelopments/2026/The-Digital-AI-Omnibus-Proposed-deferral-of-high-risk-AI-obligations-under-the-AI-Act">DLA Piper GENIE</a>). The legislative act will be published in the Official Journal shortly and enters into force on the third day after publication. The substantive changes:</p><p><strong>Stand-alone Annex III high-risk systems</strong> — recruitment, credit scoring, education, law enforcement, border control, essential services access — move from August 2, 2026 to <strong>December 2, 2027</strong>. Sixteen months.</p><p><strong>Annex I embedded AI</strong> — AI systems integrated into products already covered by EU sectoral regulation (medical devices under the MDR, machinery under the Machinery Regulation, vehicles under type approval, toys, radio equipment, and more) — move to <strong>August 2, 2028</strong>.</p><p><strong>Transparency for synthetic content</strong> (Article 50) — deepfakes, AI-generated text, AI-generated images and audio — moved the <em>other</em> direction: the grace period was reduced from six months to three, with the new deadline set at <strong>December 2, 2026</strong>. If your agent generates text or images that are consumed as content by end users, the watermarking and disclosure obligations start in five months.</p><p><strong>Regulatory sandboxes</strong> at Member State level — the obligation to establish one by August 2, 2026 was postponed to August 2, 2027.</p><p><strong>What did not move:</strong> GPAI model obligations (in force since August 2, 2025), the AI literacy requirement (in force since February 2, 2025), the prohibited practices (in force since February 2, 2025), and the enforcement powers for the Commission on GPAI models (starting August 2, 2026 as originally planned).</p><h3>What's binding right now</h3><p>The delay makes the headlines. The current obligations do not. Three that are already in force and matter for any enterprise running agents in production:</p><p><strong>GPAI model obligations (August 2, 2025).</strong> If your agent stack calls a general-purpose AI model — a model with significant generality that can perform a wide range of tasks, which covers most frontier models placed on the market after that date — the provider of that model has to publish a public summary of training content, comply with copyright rules, and provide technical documentation to downstream deployers (<a href="https://digital-strategy.ec.europa.eu/en/policies/guidelines-gpai-providers">European Commission GPAI guidelines</a>). Downstream, that means enterprises are entitled to information they were not entitled to before. The <a href="https://digital-strategy.ec.europa.eu/en/policies/guidelines-gpai-providers">GPAI Code of Practice</a> sets out what compliant provider disclosure looks like.</p><p><strong>Commission enforcement on GPAI (August 2, 2026).</strong> Four weeks from now, the Commission gains formal enforcement powers over GPAI providers, including fines. This affects your providers directly and shapes their behaviour toward downstream deployers. If your governance architecture assumes the model provider will be responsive to compliance-driven documentation requests, that assumption becomes safer to make on August 2.</p><p><strong>AI literacy (February 2, 2025).</strong> Every provider and deployer of AI systems must ensure that staff dealing with AI systems have adequate AI literacy. This is not "run a training video." It is a documented ongoing obligation. Enterprises with agents in production should have an audit trail of what training was delivered to which teams.</p><p><strong>Prohibited practices (February 2, 2025).</strong> Social scoring, real-time remote biometric identification in public spaces, emotion recognition in workplaces and schools, and a specific list of other practices are prohibited outright. If your agent's use case sits close to any of these, the classification question should have been answered by now.</p><p><strong>The Colorado AI Act (June 30, 2026).</strong> Not the EU AI Act, but the shape is close enough that most EU-focused governance work satisfies both. Colorado's law requires "reasonable care" for developers and deployers of high-risk AI systems making consequential decisions, and applies to any company doing business with Colorado residents. If your enterprise sells into the US as well as the EU, treat the two frameworks as one engineering programme. The five surfaces satisfy both.</p><h3>Why the delay happened, and why it isn't relief</h3><p>The Omnibus package is not a change of policy. It is a change of enforcement calendar because the technical machinery to enforce is not ready. Harmonised standards under CEN-CENELEC JTC 21 — the standards that define what "comply with Article 9 risk management" means in a testable way — are still being drafted. Notified bodies are still being set up. Conformity assessment procedures for high-risk AI are still being finalised.</p><p>The Commission's calculation was straightforward. Enforcing an obligation against a rulebook that has not been written creates uncertainty for everyone. Better to wait until the standards land, then enforce against them.</p><p>For enterprises, this is a trap. The reason is that the standards are being drafted now. When they land, they will reflect the technical patterns that the standards bodies have been discussing with implementers over the past 18 months — the enterprises that engaged in the working groups, the providers that submitted comments, the delivery partners that shipped early conformance implementations. Enterprises that wait until Q3 2027 to start the engineering will find the standards more prescriptive than they had hoped, and will have less time to adjust the architecture than the sixteen-month calendar suggests. The clock started when the Omnibus passed, not when the deadline arrives.</p><h3>Am I subject to high-risk obligations?</h3><p>Every EU AI Act compliance conversation starts here. The classification is not "we use AI" — it is "our specific AI system is listed in Annex III as stand-alone, or Annex I as embedded, and the specific use case is high-risk under that annex."</p><p>Annex III stand-alone high-risk systems, condensed:</p><ul><li><strong>Employment, workers management, and access to self-employment</strong> — recruitment agents, CV screening, performance evaluation.</li><li><strong>Access to essential private services and public services and benefits</strong> — credit scoring, insurance risk assessment, emergency service dispatch.</li><li><strong>Law enforcement</strong> — evidence evaluation, risk assessment of natural persons.</li><li><strong>Migration, asylum, and border control management</strong> — visa and asylum application processing.</li><li><strong>Administration of justice and democratic processes</strong> — assistance to judicial authorities.</li><li><strong>Biometric identification and categorisation</strong> of natural persons, and emotion recognition.</li><li><strong>Critical infrastructure</strong> — road traffic, water, gas, heating, electricity management.</li><li><strong>Education and vocational training</strong> — determining access, evaluating learning outcomes.</li><li><strong>Product safety components</strong> covered by Annex I regulations.</li></ul><p>If your agentic system reads inputs about a natural person and its output affects a decision in one of these areas, it is likely in scope. If it does not, it is likely not in scope. The safest engineering posture is to run the Annex III classification in Phase 1 of the engagement, not at launch.</p><h3>What to build now regardless of the December 2027 date</h3><p>The engineering that will satisfy the standards when they land is the same engineering that satisfies SOC 2 today, the same engineering that satisfies the <a href="/post/production-ai-agent-governance">Colorado AI Act enforceable June 2026</a>, and the same engineering that made our regulated-industry engagements auditable. The five surfaces from <a href="/post/production-ai-agent-governance">our governance reference</a> cover what needs to exist in code:</p><p><strong>Agent identity.</strong> Every agent has a queryable registry entry. When the standards specify what conformity assessment covers, the assessment will ask for a system inventory. The registry is the answer.</p><p><strong>Tool allowlists and output schemas.</strong> Every tool call is intercepted, validated, and either passed or denied. Article 9 risk management, Article 14 human oversight, and Article 15 accuracy requirements all depend on the allowlist layer.</p><p><strong>Evaluation.</strong> Golden datasets in CI, online LLM-as-a-judge evaluators, human review of edge cases. Article 9 expects ongoing performance monitoring across the lifecycle. The evaluation surface is the ongoing monitoring.</p><p><strong>Observability.</strong> Every planning step, tool call, retrieval, and output emits a structured trace event via OpenTelemetry. Article 12 record-keeping is satisfied by the observability data model.</p><p><strong>Audit.</strong> Append-only logs with cryptographic chain integrity and ten-year retention for high-risk systems. Article 12 record-keeping and Article 15 accuracy requirements both expect this.</p><p>Building these surfaces now takes six to twelve months for a first-time enterprise team, less for a returning team. That fits inside the sixteen-month runway to December 2027 with margin for the specific conformity work when standards land. Starting in Q3 2027 does not.</p><h3>What the regulated engagements have looked like</h3><p>Two Twistag agents shipped into regulated environments where the compliance shape had to be load-bearing on day one.</p><p>The <a href="/case-studies/ai-communications-assistant-uk-water-utility/">AI communications assistant we built for a UK water utility</a> operates in a sector where every customer-facing communication is auditable to the sector regulator. We could not ship the agent without the audit surface being load-bearing from sprint one. A human reviews any outbound communication that fails the confidence threshold; the review is logged with attribution; the agent's input, model version, prompt, and tool calls are all queryable if asked. When the EU AI Act obligations land, the utility's compliance officer will not have new engineering work — the surfaces are already there. That is what "engineering compliance in" means.</p><p>The <a href="/case-studies/regtech-platform-european-ingredient-brands/">RegTech platform we built for European ingredient brands</a> is the clearest case where the audit surface was not next to the agent — it was the product. Every decision mapped to a specific regulatory clause and produced evidence on demand. The customers of the platform are enterprise brands whose own compliance officers query the platform daily. Building the audit surface as a product is what made the platform credible to those customers.</p><h3>The timeline moved. The work didn't.</h3><p>The Omnibus is regulatory realism. Standards are not ready, enforcement calendars have to accommodate that, and the Commission chose to wait rather than enforce against uncertainty. For enterprises, the substantive change is small: the deadline moved from a date most teams were not going to hit to a date most teams still won't hit if they start in Q3 2027.</p><p>The engineering that lands you inside the deadline is the same engineering that makes an agent auditable, evaluable, governed, and operable — the <a href="/post/production-ai-agent-governance">five surfaces we build into every regulated engagement</a>. The engagements that ship on time will be the ones that started building in 2026 against a moving target, not the ones that started in 2027 against a fixed one.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-eu-ai-act-deadline.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[What an agent operation center actually contains: an engineering tour]]></title>
            <link>https://twistag.com/thinking/agent-operation-center</link>
            <guid isPermaLink="false">https://twistag.com/thinking/agent-operation-center</guid>
            <pubDate>Thu, 30 Jul 2026 00:16:06 GMT</pubDate>
            <description><![CDATA[McKinsey named it. Most vendors have never shown one. Here's what actually goes inside an agent operation center, section by section, from production agents we run.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>McKinsey named the surface in its late-2025 tech services paper. Most vendor demos have not shown one. An agent operation center is where an enterprise's human and agentic workforces are managed on the same screen — where the SVP sees the joint output of both, where the on-call engineer replays a failed agent decision, where the compliance officer answers a regulator's question about what happened last Tuesday. This post is the engineering tour: five sections that go inside the surface, the data models behind each, and the case study patterns we build them from. It's a cluster child of our pillar on <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a> and it operationalises the <a href="/post/production-ai-agent-governance">five engineering surfaces from our governance reference</a>.</p><h3>Key takeaways</h3><ul><li><strong>The agent operation center is one screen with five sections.</strong> Live agent activity, human review queue, approval chains, agent registry, and governance dashboard. Every enterprise agent deployment eventually converges on this shape.</li><li><strong>65% of high-performing AI adopters have defined human-in-the-loop validation, versus 23% of others</strong> (<a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/seizing-the-agentic-ai-advantage">McKinsey State of AI 2025</a>). The operation center is where that validation actually happens — a queue, not a policy document.</li><li><strong>The joint dashboard is the point.</strong> Human and agentic workforce results appear together. The business manager, VP, and SVP evaluate the combined performance of both. Splitting them into two screens loses the value.</li><li><strong>Low-risk autonomous, high-impact approved.</strong> The approval-chain configuration is the operating model, not the ML model. It is where the enterprise's risk tolerance is encoded.</li><li><strong>The center is not an off-the-shelf product.</strong> It is a set of surfaces built on top of your agent stack's telemetry, wired into your existing identity, ticketing, and BI systems. Shipping it is a Product Engineering deliverable, not a vendor licence.</li></ul><h3>What an agent operation center is (and isn't)</h3><p>The agent operation center is the enterprise operating surface for a fleet of AI agents. It shows what agents are doing, what they need from humans, how they are performing, and what the compliance and cost picture looks like — all on the same screen and all queryable across the same identity model as the rest of the enterprise's operational tooling.</p><p>It is not an observability platform. Observability platforms like <a href="https://langfuse.com/">Langfuse</a> or Arize show engineers the traces behind agent decisions. The operation center consumes those traces but presents them to the operator — the business manager, the compliance officer, the on-call reviewer — in a shape that fits their workflow. Observability is engineering-facing. The operation center is enterprise-facing.</p><p>It is not an AI chat sidebar. The operation center is where the humans running the workflow are, not where the humans consuming the agent's output are. A customer-service agent's supervisor sits in the operation center; the customer sits in the chat.</p><h3>Section 1: Live agent activity</h3><p>The top of the operation center is a real-time view of what the agent fleet is doing. Each agent, each in-flight task, each recent completion, each failure. Filterable by owner team, by agent type, by user (or upstream agent) on behalf of whom the action ran, by cost, by latency.</p><p>Under the hood, this section reads from the observability trace store. Each trace event carries the <a href="/post/production-ai-agent-governance">agent identity</a> that names which agent produced it and which team owns it. Aggregations happen on the read side — one query gives you "everything the invoice automation agents did in the last hour, grouped by outcome." Drilling into a specific event opens the full trace: planning steps, tool calls, retrievals, outputs, cost in tokens and dollars.</p><p>This section becomes load-bearing the moment the fleet includes more than one agent. Multi-agent handoffs are opaque without a joint activity view — an operator watching agent A cannot tell whether agent B is upstream, downstream, or blocked. Making the joint pipeline legible is the section's real job.</p><h3>Section 2: Human review queue</h3><p>The queue is where agent outputs that fail confidence thresholds land for human decision. Each item in the queue carries the agent's proposed action, the confidence signal that triggered the review, the input that produced it, the model version and prompt that ran, and the reviewer's action buttons: approve, reject, edit-and-approve, escalate.</p><p>The threshold is configurable per agent and per output type. A tool call that would refund a customer gets a lower autonomy threshold than one that formats a response. A well-tuned queue typically sees single-digit percentages of agent outputs — the long tail of low-confidence cases, cases where output schema validation failed, or cases where an extracted entity did not match a known value. The reviewing team clears the queue as part of their daily workflow, and the reviewed cases feed back into the golden dataset that trains the next iteration of the eval pipeline.</p><p>Two design notes matter. First, the queue is time-bounded. Items expire if not reviewed within a configured window — the agent then either falls back to a default action or blocks the workflow, depending on the policy. Second, the queue emits the same trace events as any other agent decision, so the audit surface captures the human review as a first-class event, not as a Slack-thread annotation.</p><h3>Section 3: Approval chains and escalations</h3><p>Some agent actions never run autonomously, regardless of confidence. Outbound customer communications above a size threshold, database writes to systems flagged as regulated, spend decisions above a dollar amount — these require a named human to approve before the agent proceeds. The approval-chain section is where those requests queue.</p><p>The design pattern is a rules engine attached to the tool allowlist. Each tool declaration includes an autonomy scope: "auto" for actions the agent can take without approval, "approve" for actions that need a named reviewer, "escalate" for actions that need an out-of-band decision. When the agent proposes an "approve" action, the request appears in the section with the reviewer(s) named. When the request is approved, the agent proceeds; when denied, the agent takes the fallback action defined in the tool declaration.</p><p>This is where the enterprise's risk tolerance actually lives. Not in a policy PDF — in the autonomy-scope configuration on the tools the agent can call. Changing "auto" to "approve" on a tool is the enterprise deciding it needs more oversight on that action. Making the change in the operation center means the change is auditable, versioned, and rollback-safe.</p><h3>Section 4: Agent registry</h3><p>The registry is a queryable inventory of every agent running in the environment. For each agent: a unique ID, the team that owns it, the tools it is allowed to call, the data sources it can read, the model versions it has run on, the prompt versions it has used, the current status.</p><p>The registry is where the compliance officer answers questions. "Which agents are touching customer PII?" is a filter. "Which agents were running on model version X when incident Y happened?" is a filter and a time range. "Which agents does team Z own, and who is on call for them?" is two filters and a link to the on-call rotation.</p><p>Every entry in the registry is a resource in the same identity model as the rest of the enterprise's infrastructure. Agents inherit access grants from their owning team, and the audit surface attributes every action to the specific agent and to the identity on behalf of whom it ran. The registry is the connective tissue between the agent fleet and the enterprise's existing governance tooling.</p><h3>Section 5: Governance dashboard</h3><p>The last section is where the enterprise sees whether the agent fleet is meeting its quality, cost, and compliance targets. Quality: golden-dataset pass rate, online LLM-as-a-judge score trend, human-review override rate. Cost: token spend per agent, per team, per task; trend against budget. Compliance: EU AI Act obligation coverage, SOC 2 control status, incidents open and closed.</p><p>The governance dashboard is not built for the engineering team. It is built for the SVP asking "are we on track?" and the compliance officer preparing for a review. Numbers matter more than traces. Trends matter more than snapshots. A quality regression that lasted an hour and self-corrected is a trace event; a persistent trend down over three days is a dashboard event, and this is where it shows up first.</p><p>This section is also where the joint view — human and agentic workforce performance on the same screen — matters most. The dashboard rolls up the outputs of both. The invoice team's throughput includes the human-reviewed invoices and the auto-approved ones. The customer-support team's SLA includes the tickets the agent closed and the ones the humans handled. Splitting the view into "human team" and "agent team" hides the interaction. Combining it makes the interaction visible.</p><h3>What the operation center is not: an observability tool</h3><p>Observability platforms — Langfuse, Arize, <a href="https://aws.amazon.com/bedrock/">AWS Bedrock AgentCore</a>, Datadog LLM observability — are the layer under the operation center. They store the traces, run the eval pipelines, and provide the raw query surface engineers use to debug. The operation center consumes those traces and presents them to non-engineering users in a shape their workflow fits.</p><p>The distinction matters at build time because it changes what the buyer is asking for. "We want to see what our agents are doing" is a request for the operation center. "We want to debug why our agent is failing" is a request for observability. Most enterprise engagements need both, and the operation center's job is to sit above the observability tool as the presentation layer for the business and compliance stakeholders.</p><h3>Three shapes we build the operation center in</h3><p>The operation center is not one architecture. It is a set of five sections that can be assembled in three deployment shapes depending on what the enterprise already runs.</p><p><strong>Isolated.</strong> A standalone operation center built as its own application, with its own data model, identity, and UI. The five sections are first-class pages. This shape fits enterprises that do not yet have an observability platform in place and want the operation center to be the trace store as well as the presentation layer. It is the heaviest lift, and the shape with the most flexibility over what data the operation center sees.</p><p><strong>On the side, alongside existing telemetry.</strong> The operation center runs as its own application but reads from an existing observability platform — Langfuse, Arize, <a href="https://aws.amazon.com/bedrock/">AWS Bedrock AgentCore</a>, Datadog LLM observability — as the trace source. Engineering keeps its debugging tool. Business, compliance, and operations get their own surface. The two applications share identity and a data pipeline, and the operation center's job is to translate engineering-facing traces into operator-facing sections.</p><p><strong>On top of Langfuse (or another telemetry tool).</strong> The operation center runs inside the observability platform's extension surface, as custom dashboards, custom pages, and custom review queues built on the platform's own primitives. This is the lightest shape when the enterprise has already standardised on a specific observability tool and wants the operation center to be part of the same product experience. Langfuse's dashboarding and evaluation surfaces are flexible enough to host the five sections directly.</p><p>Which shape fits is a Phase 1 conversation, not a launch decision. The isolated shape suits enterprises whose observability tooling is not yet mature. The on-the-side shape suits enterprises with an established platform that engineers want to keep as-is. The on-top shape suits enterprises that have already invested in the observability platform and want a single pane of glass. All three ship the same five sections; the difference is where the sections live.</p><h3>The operation center is where trust becomes measurable</h3><p>Every enterprise agent deployment eventually needs its operation center. Teams that plan for it in Phase 1 ship it as part of the MVP, and their operating teams get to see the agent working from day one. Teams that plan for it later ship it as an ops emergency six months in, when the compliance officer or the SVP asks a question the engineering team cannot answer from a Slack thread.</p><p>McKinsey's naming of the surface has helped enterprises put a word on the thing they were already trying to build. What the word points at is a set of five sections that live above the agent stack's telemetry and connect to the enterprise's identity and BI systems. Building it takes the same <a href="https://twistag.com/product-engineering">Product Engineering craft</a> that builds any operational tool. It is not exotic. It is not an off-the-shelf licence. It is the operating surface that makes the agent fleet legible to the humans responsible for it — and that legibility is where trust becomes measurable.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-agent-operation-center.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[AI agent token economics: what actually moves the bill]]></title>
            <link>https://twistag.com/thinking/ai-agent-token-economics</link>
            <guid isPermaLink="false">https://twistag.com/thinking/ai-agent-token-economics</guid>
            <pubDate>Thu, 30 Jul 2026 00:15:50 GMT</pubDate>
            <description><![CDATA[Agents use 3-10x more tokens than chatbots by design. Five engineering levers cut the bill by 40-70% in 30 days. What actually moves the numbers, from production.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>An AI agent uses 3-10x more tokens than a chatbot handling the same query. The reason is structural: a chatbot returns a response, but an agent plans, calls tools, verifies outputs, and often retries. A single agent task can burn $5-8 in API fees against a frontier model with no engineering effort applied. The good news is that most of that cost is optional — teams that run the six-stage cost sequence (audit, baseline, compress, route, cache, monitor) commonly cut token spend by 40-70% within a month while holding quality steady (<a href="https://fast.io/resources/ai-agent-token-cost-optimization/">Fastio, 2026</a>). This post walks through the five engineering levers that actually move the bill. It's a cluster child of our pillar on <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a>.</p><h3>Key takeaways</h3><ul><li><strong>Agents cost 3-10x more per task than chatbots</strong> because they make more LLM calls per user request — planning, tool selection, execution, verification, response. Some agentic workloads run 50x more tokens than the equivalent chat completion.</li><li><strong>Prompt caching cuts 45-80% of the input token cost</strong> where it applies. OpenAI's cached token rate is 50% off list; Anthropic and Bedrock offer similar discounts. Every long system prompt should be caching by default.</li><li><strong>Model routing is the highest-impact lever.</strong> A frontier reasoning model can cost 190x more than a fast small model for the same task class. Routing simple queries to the cheap model and hard queries to the expensive one is often the single biggest change teams make.</li><li><strong>Observability at the trace level is not optional.</strong> Without per-agent, per-task, per-tool cost attribution, cost optimisation is guesswork. Every trace event includes the tokens consumed and the dollars spent.</li><li><strong>Cost is not model choice. Cost is architecture.</strong> The cheapest model wired into a badly-shaped agent will cost more than an expensive model wired into a well-shaped one.</li></ul><h3>Why agents cost more than chatbots (and how much more)</h3><p>A chatbot receives a message and returns a response. One LLM call, one input, one output. The token bill is proportional to the length of the exchange.</p><p>An agent receives a task and produces an outcome. The path from task to outcome runs through planning ("what should I do?"), tool selection ("which function do I call?"), execution ("what did the tool return?"), verification ("does this look right?"), and response generation ("how do I explain this to the user?"). Each step is an LLM call. Some steps loop — the agent retries when a tool call fails, or replans when the initial approach did not work. Multi-agent architectures amplify the pattern: agent A hands off to agent B, and each handoff carries context to be re-processed.</p><p>Industry data puts the multiplier at 3-10x for typical agentic workflows, and up to 50x for complex ones (<a href="https://leanopstech.com/blog/agentic-ai-cost-runaway-token-budget-2026/">LeanOps, 2026</a>). Unconstrained, an agent solving a moderately complex task against a frontier model runs $5-8 in API fees per task. At production volumes — say, ten thousand invoice-processing tasks per day — that is a $50-80K daily bill for the naïve architecture.</p><p>The naïve architecture is not the ceiling. It is the starting point. The five levers below are what teams apply to move the bill.</p><h3>Lever 1: Prompt caching (45-80% input cost reduction)</h3><p>The single easiest cost lever in 2026 is prompt caching. When a long prompt — a system prompt, a large context document, an agent's tool definitions — is re-used across many requests, the LLM provider caches the processed representation and charges a discounted rate for subsequent reads. OpenAI charges 50% of list price for cached tokens. Anthropic and AWS Bedrock offer similar discounts, with cache write costs typically 25% above list and cache read costs at 10% of list (<a href="https://www.anthropic.com/pricing">Anthropic pricing</a>).</p><p>For agents, the pattern is universal. The system prompt is the same across every call. The tool definitions are the same across every call. The output schema is the same across every call. Only the user-specific context differs. Caching the shared prefix cuts 45-80% of the input token cost, and improves time-to-first-token by 13-31% as a side benefit.</p><p>The engineering effort is small — configure caching on the provider SDK, ensure the system prompt is stable enough that the cache stays warm — but the payoff is enormous. Every long-prompt agent should be caching by default. This is the first cost lever we implement in a build and the one that pays back fastest.</p><h3>Lever 2: Model routing (up to 80% total spend reduction)</h3><p>Not every step in an agent's workflow needs the frontier model. Planning a multi-step task benefits from a reasoning model. Formatting a response, extracting a field, classifying an intent — these are small models' natural territory. Routing the right task to the right model is where the biggest cost reductions land.</p><p>The gap is stark. A frontier reasoning model can cost 190x more than a fast small model for a task both can handle (<a href="https://www.requesty.ai/blog/ai-agent-cost-optimization-how-to-cut-llm-spend-by-80-percent-with-routing">Requesty, 2026</a>). OpenAI's GPT-5 architecture routes internally between a fast model and a reasoning model based on query complexity; teams building on top of frontier models can implement the same pattern explicitly, using cheaper models for the majority of calls and reserving the expensive model for the steps that actually need it.</p><p>The engineering pattern is a router in front of the LLM call. The router classifies the task — often via a small classifier LLM or a rules engine — and selects the model. Cheap models handle field extraction, entity resolution, format conversion, simple classification. Expensive models handle planning, evaluation, and cases where the cheap model returned low confidence. Well-routed agentic workloads report 60-80% total spend reduction versus a naïve "everything on the frontier model" architecture.</p><h3>Lever 3: Batch and async processing</h3><p>Not every agent task is user-facing in real time. Invoice processing, report generation, data enrichment, overnight reconciliation — these tasks tolerate latency, and the LLM providers have priced tolerance in. Anthropic's Batch API charges 50% of list. OpenAI's Batch API charges 50% of list with a 24-hour SLA. Bedrock offers similar pricing for asynchronous workloads.</p><p>For agentic workflows with a mix of synchronous and asynchronous tasks, routing the asynchronous work to the batch API halves the cost of that half of the workload. The engineering effort is a queue and a scheduler — the same primitives most production systems already have.</p><p>The pattern that recurs in production is a two-path architecture. The real-time path — user-facing agent responses — runs on the standard API with prompt caching. The batch path — reconciliation, reporting, overnight enrichment, drift checks against historical data — runs on the batch API at half the cost. Both paths share the same agent code, the same tool definitions, and the same evaluation pipeline; only the API endpoint and the SLA differ.</p><h3>Lever 4: Structured outputs and prompt compression</h3><p>Two smaller levers with real impact.</p><p><strong>Structured outputs.</strong> When an agent's response must fit a JSON schema, forcing structured output at the LLM level (rather than freeform text plus post-hoc parsing) cuts output tokens by 30-50% and reduces retry rates on schema-validation failures. Every frontier model now supports structured output at the API level — Anthropic's tool use, OpenAI's structured outputs mode, Bedrock's guarded generation. The engineering effort is defining the schema; the payoff is fewer tokens and fewer retries.</p><p><strong>Prompt compression.</strong> Long context is expensive. Techniques like LLMLingua (compresses prompts by 3-5x while preserving semantic meaning) and semantic caching (returns cached responses for semantically similar queries) trade a small quality risk for a large cost reduction. These are per-workload judgment calls rather than universal wins, but they land another 10-30% cost reduction on prompts that were oversized to begin with.</p><h3>Lever 5: Cost observability at the trace level</h3><p>None of the levers above work without visibility. Every trace event in the agent's observability platform — Langfuse, Arize, AWS Bedrock AgentCore, whatever the stack uses — must carry the tokens consumed and the dollars spent, attributed to the specific agent, the specific task, the specific tool call, the specific model version.</p><p>Cost observability turns cost optimisation from guesswork into engineering. The team sees which agents burn the most, which tasks are the most expensive, which models are being over-used. The routing rules get tuned against real usage patterns. The caching hit rate gets monitored. Regression is caught within a day, not at the end of the billing cycle.</p><p>At the operation center level, the <a href="/post/agent-operation-center">governance dashboard section</a> rolls the trace-level cost data up into per-agent, per-team, per-task views. The SVP sees the fleet's total spend. The engineering manager sees which agents deviate from their budget. The compliance officer sees which model versions are being called for regulated workflows. Every stakeholder gets the view they need from the same underlying trace store.</p><h3>What a cost-optimised architecture looks like in practice</h3><p>The five levers compound when applied together. A representative cost-optimised architecture, drawn from patterns we see in production, looks like this:</p><p>The system prompt is 3,000-5,000 tokens (tool definitions plus a stable rule set) and cached — so incremental task processing costs 10% of the input-token list price rather than 100%. Structured output is enforced at the model level, so extracted fields either match the schema or the agent retries with an explicit correction prompt; there is no free-form parsing that would burn output tokens and then require re-work. Tasks where the confidence signal on an extracted value falls below threshold route to the <a href="/post/agent-operation-center">human review queue</a>; those cost near-zero in additional tokens because the model does not attempt to over-reason a low-confidence case. Real-time tasks run on the standard API; batch tasks — reconciliation, reporting, drift checks — run through the batch API at 50% of list.</p><p>The result is a per-task cost that lands in cents rather than dollars at production volumes measured in thousands per day. The economics work because the architecture was designed around the five cost levers from Phase 1, not because a special model or trick made the cost disappear.</p><h3>The lever that beats all others</h3><p>The five levers above compound. Prompt caching times model routing times batch processing times structured output times cost observability gets the naïve architecture down 60-80% in the first 30 days. That is real money at production volumes. It is also inside the reach of any team willing to instrument the traces and iterate.</p><p>The lever that beats all of them is the one that decides not to use an agent for the task. A packaged model, a rules engine, a straight LLM call — many workflows are handled better and cheaper by a simpler architecture. The teams that ship agents into production successfully are the same teams that resist using them for problems that do not need them. The token bill is the receipt of that discipline.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-ai-agent-token-economics.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Capability transfer in AI delivery: what actually transfers, and how]]></title>
            <link>https://twistag.com/thinking/capability-transfer-ai-delivery</link>
            <guid isPermaLink="false">https://twistag.com/thinking/capability-transfer-ai-delivery</guid>
            <pubDate>Thu, 30 Jul 2026 00:15:34 GMT</pubDate>
            <description><![CDATA[The build-with-you model ends with your team owning the system. Here's what actually transfers, how pair working works, and what the exit signal looks like in code.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Capability transfer is the operational spine of the build-with-you delivery model. Our <a href="/post/build-buy-build-with-you">pillar on build-with-you</a> named the market thesis and positioned it against pure build, pure buy, and Big 4 Alliance delivery. This post is the operational answer to the next question buyers ask: what actually transfers, mechanically, from the partner's engineers to the in-house team, and how does the partner know when to leave. The Build-Operate-Transfer (BOT) model has been running in software delivery for two decades. Adapting it to agentic AI delivery in 2026 means transferring artefacts the offshore-GCC playbook did not have to worry about — eval pipelines, prompt versioning, agent registries, model provenance. This post walks through the six categories of artefact that transfer, the pair-working pattern that makes the transfer stick, the exit signal that says the transfer is done, and the failure modes when a partner cuts corners.</p><h3>Key takeaways</h3><ul><li><strong>Capability transfer is six categories of artefact, not a wiki page.</strong> Code, prompts, eval datasets, infrastructure-as-code, runbooks, and pair-working knowledge. All six transfer, or none of them meaningfully transfer.</li><li><strong>The BOT model is not new — the artefact list is.</strong> Classic BOT transfers people, processes, and a running operation. AI delivery transfers eval pipelines, prompt versioning discipline, agent registries, model provenance. The engineering artefacts are what agentic delivery adds.</li><li><strong>Pair working is the transfer, not a nice-to-have.</strong> Documentation transfers what you already know how to look up. Pair working transfers what you did not know to ask.</li><li><strong>The exit signal is operational, not calendar-based.</strong> The transfer is complete when the in-house team ships a production change without the partner's involvement. Not on a date, not at a budget threshold, not at a milestone review.</li><li><strong>The failure mode is a "handover event" at the end.</strong> Handovers concentrated in a single week never work. Transfer distributed across the whole engagement does.</li></ul><h3>The six categories of artefact that transfer</h3><p>The transfer manifest is not one repository. It is a set of six discrete categories, each with its own transfer mechanism and its own way to verify the transfer landed.</p><p><strong>Code.</strong> The obvious one. Source code, tests, CI configuration, deployment scripts. Transferred through git repositories the in-house team owns from day one — not a partner-owned repo that gets pushed at the end. The transfer verification is the in-house team's first pull request that merges without the partner's review.</p><p><strong>Prompts and prompt versioning.</strong> Agentic systems live and die by prompt discipline. Every prompt is versioned, every prompt change goes through the eval pipeline, every deployed prompt maps to a specific commit. The transfer artefact is the prompt registry — a versioned store, sometimes an internal service, sometimes a directory in the repo with a specific structure — and the CI checks that gate prompt changes on eval regression. Verification is the in-house team modifying a prompt, running the eval, and shipping the change.</p><p><strong>Eval datasets.</strong> Golden datasets in CI. Adversarial cases for context contamination. The rubric definitions the LLM-as-a-judge evaluator uses. The human review cases that fed the last golden update. All of this is data the in-house team has to be able to extend. Transferred as datasets in the repo (or in a designated data store) with clear ownership and clear rules for when new cases get added. Verification is the in-house team adding a new eval case in response to a production failure and closing the loop.</p><p><strong>Infrastructure as code and secrets.</strong> The infrastructure the agent runs on — Kubernetes manifests, Terraform, cloud IAM, database migrations, secret management — all as code, all in a repo the in-house team owns, all deployable without the partner. Secrets rotate through the in-house team's secret management (Vault, cloud KMS, whatever they use) from Phase 2. Verification is the in-house team deploying a new environment from scratch.</p><p><strong>Runbooks and on-call.</strong> The playbooks for what to do when something goes wrong. Incident response procedures. The list of dashboards to check. The escalation tree. The recurring failure modes and their fixes. Transferred as documentation the in-house team maintains (not the partner's wiki), and refined by pair working during Phase 3. Verification is the in-house team handling a real production incident without paging the partner.</p><p><strong>Pair-working knowledge.</strong> The part that does not fit in a document. Why we picked this retrieval architecture over the other one. What happens when the confidence threshold drops. Which prompt patterns work well with Claude versus GPT. The trade-offs we already tried and rejected. This category transfers through the pair-working mechanism below — it does not transfer any other way.</p><h3>Pair working is the transfer mechanism</h3><p>Documentation transfers the parts you already know how to look up. Pair working transfers the parts you did not know to ask. Every serious capability-transfer engagement runs on pair working through Phase 2 and Phase 3.</p><p>The pattern. Partner engineer and in-house engineer are paired on a specific piece of work — a new agent, an integration extension, a prompt refactor, an eval-set expansion. They work at the same screen (or same VS Code Live Share session, or same paired repo branch), and the partner narrates decisions as they make them. The in-house engineer questions choices they would not have thought to make. Over the course of the engagement, the balance shifts: the partner does more of the driving early, the in-house engineer does more of the driving later, and by month four to six the partner is a reviewer rather than a driver.</p><p>Two things make pair working work. <strong>First</strong>, the in-house engineer has real capacity for it. Capability transfer without dedicated in-house time is theatre. <strong>Second</strong>, the partner engineer is genuinely senior. A junior partner engineer transfers a junior's understanding, which is not what the buyer is paying for.</p><p>This is where the partner-selection question we covered in the <a href="/post/build-buy-build-with-you">build-with-you pillar</a> actually shows up in practice. Senior engineering depth is not a nice sales-page claim. It is what determines whether pair working is a transfer mechanism or an expensive shared-screen exercise.</p><h3>The exit signal is operational</h3><p>Every capability-transfer engagement needs an exit criterion that is not "the calendar says we are done." The signal we use is operational: the in-house team has shipped a production change without the partner's involvement.</p><p>What that looks like specifically. The in-house team writes a new tool definition for an agent. They write the schema for its arguments and its return type. They write the eval cases that cover the new tool's behaviour. They ship the change through CI, the eval passes, the agent uses the new tool in production correctly. The partner reviewed the pull request but did not write the code. That is one signal.</p><p>The next change after that, the partner does not review. The change after that, the in-house team has shipped three production changes without any partner involvement. That is the signal that the transfer is done.</p><p>The signal usually arrives between month four and month nine, depending on the team's starting expertise and the complexity of the system. When it arrives, the partner steps out. Not on a date, not at a budget threshold. On the operational reality that the team can run the system.</p><h3>The failure modes when transfer is done badly</h3><p>Three patterns account for most of the transfer failures we have seen or heard about across the industry.</p><p><strong>The "handover event" at the end.</strong> The partner runs the whole engagement, then schedules a week of "knowledge transfer" at the finish. The wiki gets written. The team gets briefed. The partner leaves. Two months later the in-house team encounters a failure mode the partner would have recognised in a heartbeat and cannot resolve without paging them back. This pattern fails because transfer is a distributed practice across many decisions, not a concentrated event.</p><p><strong>The one-way dependency on partner infrastructure.</strong> The partner's monitoring dashboard, the partner's eval runner, the partner's staging environment. All of it convenient during the engagement. All of it a hostage situation when the partner leaves. The fix is to make the in-house team's tooling primary from Phase 2, even when the partner's tools would be faster in the short term.</p><p><strong>Undocumented tacit knowledge.</strong> The partner senior engineer holds the model of why this specific retrieval architecture was picked, why the confidence threshold sits at that value, why this prompt pattern works and that one does not. If the tacit knowledge never gets pair-worked out, the in-house team inherits a system whose choices they do not understand. Six months later they revert one of the choices because they do not know why it was made, and the system regresses.</p><h3>What has to be true at engagement start for transfer to work</h3><p>Two conditions determine whether capability transfer succeeds, and both have to be true in Phase 1.</p><p>The in-house team exists and has bandwidth. Transfer requires attention. If the receiving team is also running production for the existing systems, and there is no dedicated time booked for pair working and code reviews, the transfer will not happen. Honest bandwidth planning in Phase 1 is what makes Phase 3 possible.</p><p>The partner is engineering-led, not delivery-management-led. Transfer happens through pair working with senior engineers. If the partner's model is to send a delivery manager and staff junior engineers under them, the tacit knowledge does not transfer. The seniority of the people doing the actual work is the strongest predictor of whether the exit signal ever arrives.</p><p>Both conditions surface in Phase 1 discovery. Neither can be manufactured mid-engagement. Buyers that get these two right end up with the working system and the team that can run it. Buyers that get either one wrong end up with a working system that the partner has to keep running.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-capability-transfer-ai-delivery.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[AI agent observability: the OpenTelemetry + Langfuse pattern]]></title>
            <link>https://twistag.com/thinking/ai-agent-observability</link>
            <guid isPermaLink="false">https://twistag.com/thinking/ai-agent-observability</guid>
            <pubDate>Thu, 30 Jul 2026 00:15:11 GMT</pubDate>
            <description><![CDATA[OpenTelemetry's GenAI semantic conventions are the 2026 standard. Here's what a compliant agent trace looks like, what it captures, and the shipping pattern.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Agent observability used to be a vendor debate. In 2026 it is a specification: the <a href="https://opentelemetry.io/blog/2026/genai-observability/">OpenTelemetry GenAI Semantic Conventions</a>, maintained by a Special Interest Group under the CNCF, now defines what a trace event captures across six layers — LLM client calls, agent orchestration, MCP tool calls, workflow composition, content capture, and quality evaluation. Langfuse, Arize, AWS Bedrock AgentCore, and Datadog LLM observability all now consume traces on the standard OTLP endpoint. The vendor question has flipped from "which format" to "which backend fits your other tooling." This post is the engineer's tour of what a compliant agent trace actually contains, the pattern for shipping observability from day one, and where observability stops — because eval is the layer above it. It's a cluster child of our pillar on <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a> and it operationalises the observability surface from our <a href="/post/production-ai-agent-governance">governance reference</a>.</p><h3>Key takeaways</h3><ul><li><strong>OpenTelemetry GenAI is the 2026 standard.</strong> The <a href="https://opentelemetry.io/blog/2026/genai-observability/">GenAI SIG</a> formed in April 2024 and the spec is now at v1.41. Every serious agent framework instruments against it; every serious observability platform consumes it. If your stack does not emit OTel-compliant traces, you are on the wrong side of the standard.</li><li><strong>The span tree is invoke_agent at the top, chat and execute_tool as children.</strong> The nesting mirrors the agent's control flow. An operator reading the trace sees planning, tool calls, retrievals, and outputs as a single navigable timeline.</li><li><strong>Six standard attributes carry the load.</strong> <code>gen_ai.request.model</code>, <code>gen_ai.usage.input_tokens</code>, <code>gen_ai.usage.output_tokens</code>, <code>gen_ai.response.finish_reasons</code>, and the corresponding cost and latency fields. Everything downstream — cost dashboards, quality alerts, compliance audits — reads from these.</li><li><strong>Content capture is off by default.</strong> Prompts, tool arguments, and outputs contain sensitive data. The spec keeps them out of traces unless you opt in per environment. Opting in is a security decision, not a default.</li><li><strong>Langfuse is now an OTel backend, not an alternative to OTel.</strong> Its <code>/api/public/otel</code> endpoint consumes OTLP traces from any instrumented agent. Same for Arize, Datadog, and the AWS-managed observability service. The lock-in that used to exist at instrumentation time is gone.</li></ul><h3>What a compliant agent trace looks like</h3><p>Every action an agent takes emits a span. Spans nest to form a tree. Reading the tree tells you what the agent did, in what order, with what result, at what cost.</p><p>The top-level span is <code>invoke_agent</code>. It represents the whole task — a user request, an incoming event, a scheduled trigger. Its attributes carry the agent identity, the user or upstream agent on whose behalf the action ran, the total token count, the total cost, the total latency, and the final outcome.</p><p>Inside <code>invoke_agent</code>, the agent's planning and execution steps emit child spans. <code>chat</code> spans wrap LLM calls — one per call, with the model version, the input and output token counts, the finish reason, and (if content capture is enabled) the prompt and response. <code>execute_tool</code> spans wrap tool invocations — the tool name, the arguments, the return value, and the latency. <code>retrieval</code> spans (for RAG steps) carry the query, the retrieved chunks, and (where the platform supports it) the relevance score.</p><p>Multiagent architectures nest the pattern. When agent A calls agent B, agent B's <code>invoke_agent</code> span is a child of one of agent A's spans. The context of who called whom propagates through the trace, so an operator debugging a failure can follow the call chain from the top-level task down to the specific tool call that returned unexpected data.</p><h3>The six layers OTel GenAI now covers</h3><p>The original spec covered LLM client calls only. Since 2024 it has expanded to cover the full agent surface:</p><p><strong>Layer 1 — LLM client call.</strong> The chat span. Model, tokens, finish reason, latency, cost. This is where the spec started and where every implementation still begins.</p><p><strong>Layer 2 — Agent orchestration.</strong> The invoke_agent span and its children. Planning steps, tool selection, control flow. This is where the multi-step agent's behaviour becomes legible.</p><p><strong>Layer 3 — MCP tool calling.</strong> <a href="https://modelcontextprotocol.io/">Model Context Protocol</a> tool invocations get their own execute_tool spans with MCP-specific attributes. As MCP has become the standard for exposing enterprise data to agents in 2026, this layer is where enterprise-context traces show up.</p><p><strong>Layer 4 — Workflow composition.</strong> For multi-step or multi-agent workflows, the spec now captures workflow-level events — approval waits, human-review handoffs, batch job completions.</p><p><strong>Layer 5 — Content capture.</strong> Opt-in. When enabled, spans include the full prompt, the full response, the tool schemas, the tool arguments, the tool results. This is what makes traces useful for debugging quality issues. It is also what makes them a regulated data source, so the opt-in default is deliberate.</p><p><strong>Layer 6 — Quality evaluation.</strong> Newly added in v1.41. Spans can attach evaluation scores — the LLM-as-a-judge score, the golden-dataset match, the human-review verdict. Observability and eval used to be separate systems. In the current spec they are the same event graph.</p><h3>The attributes every trace carries</h3><p>The spec defines a small set of standard attributes that every instrumented agent emits. The list is short by design:</p><ul><li><code>gen_ai.request.model</code> — the model called, e.g. <code>claude-sonnet-4-5</code> or <code>gpt-5-mini</code></li><li><code>gen_ai.request.max_tokens</code>, <code>temperature</code>, <code>top_p</code> — the request parameters</li><li><code>gen_ai.response.finish_reasons</code> — why the model stopped: <code>stop</code>, <code>tool_calls</code>, <code>length</code>, <code>content_filter</code></li><li><code>gen_ai.usage.input_tokens</code> and <code>gen_ai.usage.output_tokens</code> — the token counts</li><li><code>gen_ai.usage.cost</code> — the cost in dollars (added by the instrumentation layer, not the model)</li><li><code>gen_ai.agent.name</code> and <code>gen_ai.agent.id</code> — the agent identity</li><li><code>gen_ai.tool.name</code>, <code>gen_ai.tool.arguments</code>, <code>gen_ai.tool.result</code> — tool call details (arguments and result gated by content capture)</li></ul><p>Everything downstream — the token economics dashboards from our <a href="/post/ai-agent-token-economics">cost post</a>, the human-review queue from our <a href="/post/agent-operation-center">operation center post</a>, the audit trail from our <a href="/post/production-ai-agent-governance">governance reference</a> — reads from these attributes. Standardising them at the trace layer is what makes those downstream surfaces portable across observability backends.</p><h3>The privacy default: content capture is off</h3><p>The spec keeps prompts, tool arguments, and outputs out of traces by default. Turning content capture on is a per-environment decision, made against the specific regulatory frame the workload runs in.</p><p>Two patterns work in production. <strong>Environment-scoped capture:</strong> content capture is on in development and staging, off in production. Engineers debug against staging traces with full content; the production trace store never sees sensitive input. This works for high-throughput consumer workloads where the value of debugging any individual production event is low.</p><p><strong>Redaction-in-flight:</strong> content capture is on in production, but a redaction step in the trace pipeline replaces PII, credentials, and regulated field values with hashes or category markers before the trace lands in the store. This works for enterprise workloads where debugging real production events matters more than absolute content minimisation. The redaction pipeline itself becomes an audit surface.</p><p>Which pattern fits is a Phase 1 conversation between engineering, security, and the compliance officer. The right answer varies by industry, by jurisdiction, and by the specific data classes the agent handles.</p><h3>Langfuse, Arize, Bedrock AgentCore, Datadog — the backend layer</h3><p>The observability backend is where traces live, dashboards render, and alerts fire. Four options are common in 2026:</p><ul><li><strong><a href="https://langfuse.com/integrations/native/opentelemetry">Langfuse</a>.</strong> Open-source, OTel-compliant, strong on eval integration and prompt versioning. Fits teams that want visibility into prompt-level changes as first-class events.</li><li><strong>Arize AI.</strong> Enterprise-focused, strong on drift detection and continuous evaluation. Fits teams with ML operations maturity already in place.</li><li><strong>AWS Bedrock AgentCore.</strong> Managed service, tightly integrated with the Bedrock agent stack. Fits enterprises whose agents run on Bedrock and whose observability tooling is otherwise AWS-native.</li><li><strong>Datadog LLM observability.</strong> Native to the Datadog stack. Fits enterprises whose infrastructure observability is already Datadog and want the LLM layer to unify with everything else.</li></ul><p>All four consume OTel-compliant traces on the OTLP endpoint. Switching between them costs an environment variable, not a re-instrumentation. That is the practical benefit of the standard: the backend choice is deferrable, and reversible when your enterprise's tooling shifts.</p><h3>The pattern for shipping observability from day one</h3><p>Observability is a launch requirement, not a launch follow-up. The pattern we use in delivery:</p><p><strong>Instrument at the framework layer, not the application.</strong> The agent framework (LangChain, CrewAI, Google ADK, the Microsoft Agent Framework) is the natural instrumentation point. Native OTel support in the framework means the application code stays clean and the traces stay consistent across every agent in the fleet.</p><p><strong>Emit to OTLP from Phase 2.</strong> The OTLP collector runs in the same infrastructure as the agent. Traces stream to it, then out to whichever backend the enterprise picked. If the backend choice is not finalised in Phase 2, the collector buffers to disk while the decision lands — the collection layer is not blocked on the decision.</p><p><strong>Enable content capture in staging only.</strong> Production stays privacy-conscious by default. Content capture in production is a Phase 3 decision, and it requires the redaction pipeline to be shipped first.</p><p><strong>Ship the trace-to-alert path early.</strong> A latency spike, a cost regression, a finish-reason distribution shift — the alerting rules for these get defined in Phase 2 so they fire from the first day of production traffic. Fixing the alerts later means missing the first two weeks of production behaviour, which is when the most surprises hit.</p><h3>Where observability stops (and evaluation begins)</h3><p>Observability tells you what the agent did. Evaluation tells you whether what it did was right. The two systems overlap — v1.41 of the OTel spec explicitly includes evaluation scores as span attributes — but they are not the same layer.</p><p>Observability's job is to record. Every span captures what happened. The observer decides what to look at.</p><p>Evaluation's job is to judge. Golden datasets in CI catch regressions before deploy. Online LLM-as-a-judge evaluators score live traffic continuously. Human review of edge cases feeds the next golden update. The evaluation results attach back to the trace as attributes, so the two systems form one queryable graph.</p><p>The teams that ship agents into production successfully treat observability as necessary and evaluation as sufficient. The trace shows you the agent's decision. The eval score tells you whether the decision was correct. Both belong in the same event store; both belong in the same operator's workflow. That is what OTel GenAI v1.41 makes possible, and what the pattern above lands.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-ai-agent-observability.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Conversational commerce architecture: what replaces e-commerce search]]></title>
            <link>https://twistag.com/thinking/conversational-commerce-architecture</link>
            <guid isPermaLink="false">https://twistag.com/thinking/conversational-commerce-architecture</guid>
            <pubDate>Thu, 30 Jul 2026 00:14:55 GMT</pubDate>
            <description><![CDATA[Keyword search is 20 years old and losing to shoppers who want to describe what they need. Here's the architecture that actually ships — layer by layer.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Classical e-commerce search is a keyword box that returns a ranked list. It has been that way for twenty years and it is losing the shopper who types "shoes I can run in that are good for flat feet under 150 euros" and gets nothing useful. Conversational commerce is what replaces that shopper's experience: a system that understands the description, plans the searches that answer it, reasons over the results, and returns a recommendation with the shopper's constraints explained. This post is the architecture behind that experience — the three layers that ship, the retrieval patterns underneath, the cost pattern that keeps it affordable, and the boundary between conversational commerce and a chatbot bolted onto a product catalogue. It's the pillar for our cluster on AI-native product engineering.</p><h3>Key takeaways</h3><ul><li><strong>Conversational commerce is not a chatbot on a catalogue.</strong> A chatbot fragments the shopping journey (browse → open chat → browse again). Conversational commerce integrates natural-language input into the core discovery surface. The shopper's description IS the query.</li><li><strong>Three layers ship: intent understanding, retrieval, response generation.</strong> Intent understanding classifies and normalises the query. Retrieval combines vector search, keyword, and structured filters. Response generation produces a curated recommendation grounded in the catalogue, not a ranked list.</li><li><strong>Classify before you invoke the LLM.</strong> Simple queries ("Nike Pegasus 41") route to keyword search — no LLM needed. Complex intent queries ("shoes for flat feet, wide toe box, under €150") route to the reasoning layer. This alone cuts LLM cost by 60-80% (<a href="https://alhena.ai/blog/generative-discovery-beyond-vector-search-ecommerce/">Alhena AI, 2026</a>).</li><li><strong>Vector search is necessary but not sufficient.</strong> Vector search solved relevance. Generative discovery solves decisions. The LLM reasons over the retrieved candidates and explains the recommendation to the shopper. Both layers ship together.</li><li><strong>Grounding in the catalogue is non-negotiable.</strong> Every recommendation resolves to a specific SKU, a specific stock level, and a specific price. Hallucinated products are a category-killing failure mode. Structured output enforced against the catalogue schema is the guard.</li></ul><h3>Why keyword search is losing</h3><p>The classical e-commerce search pipeline is a tokeniser, an inverted index, a ranking function, and a facets sidebar. It works when the shopper knows the product's proper name. It fails when the shopper knows what they want the product to do.</p><p>The gap is widening in 2026 because shoppers have been trained by LLM chat interfaces to describe rather than name. A shopper who spent thirty minutes chatting with an AI assistant last Tuesday is not going to shrink their vocabulary when they arrive on a retailer's product page. They will type in the same descriptive shape. If the retailer's search returns "no results" or an obviously irrelevant list, the shopper leaves.</p><p>The industry response has been three-fold. Semantic search, which uses vector embeddings to match on meaning rather than tokens. Chatbots, which sit alongside the catalogue as a separate surface. And a growing set of retailers building conversational commerce as the primary product discovery experience — the shopper's description IS the query, and the response is a curated recommendation grounded in the retailer's catalogue.</p><p>The first two are half-measures. Semantic search still returns a ranked list; the reasoning does not happen. Chatbots fragment the journey — the shopper opens a chat to describe the problem, then has to close it to browse the recommended products. Conversational commerce is the integration.</p><h3>What conversational commerce actually is</h3><p>Conversational commerce is a product discovery surface where the shopper describes what they need in natural language, and the system returns a curated recommendation — one to five products, each grounded in the catalogue with stock, price, and reasoning about why it fits. The interaction can be single-turn ("I need a rain jacket for city cycling") or multi-turn ("also under €200"), and the recommendations refine as the conversation progresses.</p><p>Three properties distinguish it from a chatbot bolted onto a catalogue:</p><p><strong>The input surface is primary, not secondary.</strong> The shopper does not have to open a chat window. The search box accepts descriptive queries and routes them through the same understanding layer. Existing keyword queries still work — the shopper does not have to change behaviour.</p><p><strong>The output is a recommendation, not a list.</strong> Classical search returns twenty ranked results and leaves the ranking rationale opaque. Conversational commerce returns one to five products with a reason attached to each: "this pair has a wide toe box, is under €150, and reviewers with flat feet rated it highly." The reasoning is what turns discovery into a decision.</p><p><strong>The catalogue grounding is enforced.</strong> Every recommendation resolves to a real SKU. The LLM does not invent products, does not misstate stock, does not misquote price. Structured output against the catalogue schema is the guardrail. Hallucinated products in a retail context are worse than no result — they destroy trust in the recommendation surface.</p><h3>The three-layer architecture</h3><p>Every production conversational commerce system we have looked at converges on the same three-layer shape.</p><h4>Layer 1: Intent understanding</h4><p>The intent layer classifies and normalises the query. Its job is to decide what the query needs before invoking any expensive downstream layer.</p><p><strong>Classification.</strong> Is this a keyword query ("Nike Pegasus 41")? A descriptive query ("running shoes for flat feet")? A multi-turn continuation ("also in a smaller size")? A comparison request ("this one vs the other one")? A support question ("when will this ship")? Each class routes to a different pipeline.</p><p><strong>Normalisation.</strong> For descriptive queries that will go to the reasoning layer, the intent layer extracts structured constraints — price range, size, colour, category, brand — and converts them into filters. The remaining descriptive content stays as free text for the retrieval and reasoning layers to work on.</p><p><strong>The cost pattern lives here.</strong> Classification is a cheap model call (small open-weight model or a rules-based classifier). Only complex intent queries route to the expensive reasoning layer. Keyword queries skip the LLM entirely. This is the <a href="/post/ai-agent-token-economics">token economics discipline from our cost post</a> applied to conversational commerce.</p><h4>Layer 2: Retrieval</h4><p>The retrieval layer's job is to produce a set of candidate products the reasoning layer can decide over. Three retrieval mechanisms combine, always in parallel:</p><p><strong>Vector search</strong> for semantic relevance. Product descriptions, review summaries, and structured attributes get embedded into a vector store (Pinecone, Weaviate, Qdrant, pgvector, or the retailer's existing choice). The query embedding retrieves the nearest neighbours in embedding space. This layer handles the "wide toe box" part of "wide toe box running shoes" — the store has products described in terms of "roomy forefoot" or "generous toe area," and vector search matches on meaning.</p><p><strong>Keyword search</strong> for exact-name and identifier matching. Product names, SKUs, brand names — a shopper who typed "Nike Pegasus 41" should get the Pegasus 41, not something semantically similar. Elasticsearch, OpenSearch, or the vector store's hybrid mode. This layer is fast and cheap and handles a large share of real traffic.</p><p><strong>Structured filters</strong> for the constraints the intent layer extracted. Price under €150, size 42, in stock. Filters apply as hard cuts on top of the vector and keyword results. This is where the catalogue schema pays back — every product has known-good structured fields, and filtering is deterministic.</p><p>The combined result is a candidate set of 20-100 products, ranked by a hybrid score that combines vector similarity, keyword relevance, and structured-filter compliance. This candidate set is what the reasoning layer decides over.</p><h4>Layer 3: Reasoning and response generation</h4><p>The reasoning layer takes the candidate set and produces the recommendation. This is where the LLM lives, and where the decision replaces the ranked list.</p><p><strong>The reasoning step.</strong> The LLM is prompted with the shopper's original query, the extracted constraints, the top candidate products (with their structured attributes and short descriptions), and any conversation history. It selects one to five products from the candidate set and explains the selection.</p><p><strong>The grounding constraint.</strong> The LLM's response fits a structured output schema. Each recommended product references a specific SKU from the candidate set — the LLM cannot invent products, cannot misstate attributes. Post-generation validation checks that every referenced SKU exists in the catalogue and that stock and price match the current state.</p><p><strong>The response shape.</strong> One to five products with a per-product reason ("wide toe box, under €150, high rating from flat-foot reviewers") plus a short overall summary of the shopper's constraint set. The output can render as a curated grid, a comparison table, or a chat-like sequence, depending on where the surface lives in the retailer's UI.</p><h3>The cost pattern that makes it viable</h3><p>Conversational commerce at retail scale sees millions of queries per day. Running every query through the reasoning layer would cost more than the incremental revenue justifies. The industry has converged on the classify-first pattern to solve this:</p><ul><li><strong>Keyword queries</strong> ("Nike Pegasus 41") route to keyword search only. No LLM call. Sub-cent cost per query.</li><li><strong>Structured queries</strong> ("running shoes size 42 under €150") route to keyword + filter. Still no LLM. Cheap.</li><li><strong>Simple descriptive queries</strong> ("cheap running shoes") route to vector search only, with a small classifier LLM re-ranking. Cents per query.</li><li><strong>Complex descriptive queries</strong> ("running shoes for flat feet with wide toe box under €150 that reviewers rate for long distances") route to the full three-layer pipeline. This is where the reasoning LLM runs. Highest cost — but the query volume is a small percentage of total traffic.</li></ul><p>The classification step itself is a cheap small-model call — sub-cent per query. The result: LLM cost per query lands in single-digit cents on average across all traffic, even though the complex-query subset costs closer to ten cents each. This is the pattern that makes conversational commerce a positive-margin feature rather than a cost sinkhole.</p><h3>Where conversational commerce stops (and chatbots begin)</h3><p>There is a legitimate boundary between conversational commerce and post-purchase support. Conversational commerce is a discovery surface — its job is to help the shopper choose. Post-purchase questions (shipping status, returns, warranty) belong in a support chat, not in the discovery surface. Conflating the two produces a discovery experience that also does support badly and a support experience that also does discovery badly.</p><p>The pattern that works is a shared identity layer with two surfaces. The shopper's session, their catalogue history, and their preferences persist across both. When they land in the discovery surface, the recommendations know their history. When they land in support, the support agent knows what they bought. But the surfaces themselves are optimised for different jobs — recommendation and reasoning versus troubleshooting and resolution.</p><h3>The implementation pattern</h3><p>Shipping a conversational commerce system into production tends to follow the same sequence:</p><p><strong>Phase 1 — Catalogue readiness audit.</strong> Do the products have structured attributes? Are the descriptions rich enough for vector embedding? Are stock and price data reliable? A catalogue that is not agent-ready cannot support conversational discovery, regardless of the front-end architecture. This is the <a href="/post/agentic-ai-workflow-services-defined">agent-readable data governance</a> work extended to retail.</p><p><strong>Phase 2 — Retrieval layer first.</strong> Ship vector search, keyword search, and structured filtering as the underlying retrieval before touching the reasoning layer. The retrieval layer alone is already better than the classical search for a large share of queries.</p><p><strong>Phase 3 — Intent classification and cost gating.</strong> Ship the classification layer that routes queries to the right pipeline. This is what makes the reasoning layer economically viable.</p><p><strong>Phase 4 — Reasoning layer and grounded output.</strong> Add the LLM reasoning step with structured output enforcement. Start with a narrow query class (descriptive complex queries) and expand once the eval pipeline is stable.</p><p><strong>Phase 5 — Continuous evaluation.</strong> Golden datasets in CI catch regressions. LLM-as-a-judge evaluators score live traffic. Human review of edge cases feeds the next iteration. The <a href="/post/ai-agent-observability">observability and evaluation pattern from our OTel post</a> applies unchanged — a conversational commerce system is an agent, and it needs the same instrumentation.</p><h3>What this shifts for the retailer</h3><p>Conversational commerce is not a bolted-on feature. It is a change to what the search surface does. The retailer that ships it well will see three things move:</p><p><strong>Discovery conversion goes up on descriptive queries</strong> — the shoppers who used to get "no results" and leave now get a curated recommendation and buy. This is the primary upside.</p><p><strong>Support tickets on "how do I find X" go down</strong> — the discovery surface answers the question the shopper would have asked support. This is a downstream saving.</p><p><strong>The catalogue investment pays back differently.</strong> Rich product descriptions, structured attributes, and quality review summaries all matter more than they did before. The retailer that has been investing in catalogue quality gets a compounding return. The retailer that has been treating the catalogue as a data dump discovers that conversational discovery does not work on a dump.</p><p>Every retailer running an e-commerce store in 2026 is on the wrong side of at least one of these shifts. Naming which one you are on wrong side of is the first Phase 1 conversation.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-conversational-commerce-architecture.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Model Context Protocol (MCP) for enterprise: when it pays off]]></title>
            <link>https://twistag.com/thinking/mcp-for-enterprise-when-it-pays-off</link>
            <guid isPermaLink="false">https://twistag.com/thinking/mcp-for-enterprise-when-it-pays-off</guid>
            <pubDate>Thu, 30 Jul 2026 00:14:42 GMT</pubDate>
            <description><![CDATA[Model Context Protocol went from experimental to enterprise-default in 18 months. Here's when MCP replaces custom REST wiring — and when it doesn't.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Model Context Protocol went from experimental spec (November 2024) to enterprise-default in about eighteen months. As of mid-2026, <a href="https://andrew.ooo/answers/mcp-model-context-protocol-enterprise-adoption-july-2026/">78% of enterprise AI teams have MCP-backed agents in production</a>, the ecosystem holds more than fourteen thousand servers, and governance has moved to the Linux Foundation's <a href="https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation">Agentic AI Foundation</a>, co-founded by Anthropic, Block, and OpenAI with support from Google, Microsoft, AWS, Cloudflare, and Bloomberg. That is the fastest a technical protocol has crossed the enterprise threshold in recent memory. The question for a CTO in 2026 is no longer "should we look at MCP" — it is "for which of our agent surfaces does MCP pay off, and where does custom REST wiring still win." This post is the decision framework we apply. It is a cluster child of our pillar on <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a>.</p><h3>Key takeaways</h3><ul><li><strong>MCP pays off when the agent needs to call many tools with governance in the middle.</strong> One agent, many servers, an approval and audit layer between them. Custom REST wiring gets brittle past three or four tools; MCP normalises the shape.</li><li><strong>MCP does not pay off on the latency-critical path.</strong> MCP adds <a href="https://andrew.ooo/answers/mcp-model-context-protocol-enterprise-adoption-july-2026/">600ms to 3 seconds of baseline latency</a>. For anything real-time (checkout, fraud scoring, live conversation) keep the direct call.</li><li><strong>The trust boundary is the server, not the tool.</strong> A poisoned tool description can hijack an agent's behaviour across every other server it has access to. Enterprise MCP deployments require a registry, a review gate, and per-server allowlists.</li><li><strong>OAuth 2.1 is the auth default.</strong> Dynamic Client Registration, Protected Resource Metadata, Resource Indicators — the 2026 spec raises the bar beyond conventional web login. If your identity stack does not do OAuth 2.1 yet, MCP forces you to fix that first.</li><li><strong>Registry and observability are inseparable from the protocol.</strong> The Linux Foundation registry is not a nice-to-have; it is what makes an enterprise MCP deployment auditable. Every tool call needs to be logged with parameters, identity, and result hashes.</li></ul><h3>What MCP actually is (in one paragraph)</h3><p>MCP is a protocol that standardises how an agent calls tools that live outside its own process. The agent (the "client") talks to one or more servers, each of which exposes a set of tools with typed inputs, typed outputs, and metadata. The wire format is JSON-RPC; the transport is either standard input/output (for local servers) or HTTP with Server-Sent Events (for remote servers). What it replaces is the pattern where every agent stack invented its own tool-calling convention against a hand-wired set of internal REST endpoints. What it does not replace is the underlying APIs — MCP servers wrap existing APIs; they do not become them.</p><h3>The three deployment shapes</h3><p>Enterprise MCP deployments settle into three shapes. Each has different economics.</p><p><strong>Intelligence layer.</strong> MCP servers sit adjacent to critical paths. The agent uses them to read state, summarise, plan — but the actual write or transactional call goes through the classical, low-latency path. This is where MCP pays off most reliably. The latency cost of MCP does not sit on the customer-facing action, and the observability and governance benefit sit on every read.</p><p><strong>Sidecar.</strong> MCP wraps a specific system (a data warehouse, a ticketing system, a CI pipeline) and exposes it to any agent that needs it. One team owns the server, many teams consume it. This is the shape that generalises fastest inside an enterprise — the <a href="https://hidekazu-konishi.com/entry/mcp_server_ecosystem_reference_2026.html">Cloudflare directory</a> of 13 remote servers, Anthropic's directory of 200-plus Custom Connectors, and the GitHub/Stripe/Notion/Linear vendor-operated servers are all sidecar shapes at the vendor level.</p><p><strong>Batch.</strong> MCP for offline, high-volume, non-latency-sensitive work. Overnight reconciliation, weekly reporting, bulk enrichment. The latency cost does not matter; the write governance and audit trail do.</p><p>The pattern that fails, consistently, is putting MCP on the synchronous, customer-facing hot path. The 600ms-3s baseline shows up as a support ticket. Every enterprise pattern that shipped successfully in 2026 keeps MCP one hop off that path.</p><h3>When MCP pays off vs custom REST wiring</h3><p>The economic question. Custom REST wiring is faster in a sprint. MCP pays back over a fleet.</p><p>MCP pays off when:</p><ul><li><strong>The agent needs to call five or more distinct tools.</strong> Below three or four, custom wiring is quick and clean. Above that, the ceremony of maintaining tool schemas, authentication, retries, and cost tracking per tool starts to dominate the delivery cost. MCP standardises the shape, which lets one governance layer sit on all of them.</li><li><strong>Multiple agents will call the same tool.</strong> Sidecar pattern. Write the server once, consume it from every agent that needs the capability. Without MCP, every agent invents its own client for that endpoint.</li><li><strong>You need auditable governance on the tool layer.</strong> Regulated industries. Any workload where the audit says "prove which agent called which tool with which arguments and what the result was." MCP servers make this a first-class thing to instrument; custom wiring makes it a per-endpoint retrofit.</li><li><strong>The vendor already ships an MCP server.</strong> <a href="https://hidekazu-konishi.com/entry/mcp_server_ecosystem_reference_2026.html">GitHub, Stripe, Cloudflare, Linear, Notion, Atlassian</a> — the vendor-shipped servers are typically better maintained than a hand-rolled client for the same API. Consuming the official server is usually the winning move.</li></ul><p>MCP does not pay off when:</p><ul><li><strong>The agent calls one internal endpoint.</strong> Direct call. The MCP layer adds latency and complexity for no fleet benefit.</li><li><strong>The path is latency-critical.</strong> Anything under a second of round-trip budget. Keep the direct call and use MCP for observability only.</li><li><strong>The workload is a single-agent, single-team artefact.</strong> No fleet, no governance surface. Wrapping it in MCP is protocol overhead that never pays back.</li></ul><h3>The trust boundary — where every enterprise deployment gets it wrong</h3><p>The single most-misunderstood property of MCP: the trust boundary is the server, not the tool. A malicious or compromised MCP server can serve tool descriptions that hijack the agent's behaviour on subsequent tool calls — including calls to other, unrelated servers the same agent has access to. Invariant Labs demonstrated a scenario where a poisoned tool description silently exfiltrated a user's message history through an apparently benign tool invocation.</p><p>The implication for enterprise deployment. Every MCP server the agent can talk to has to be a server the enterprise trusts. Not the vendor's marketing page — the actual code, the actual hosting, the actual update path. That means three surfaces:</p><p><strong>A registry.</strong> The Linux Foundation registry gives public servers signed metadata and version tracking, but enterprises need a private registry for their own servers plus an approval gate for public ones they consume. No agent connects to a server that is not on the enterprise registry.</p><p><strong>A review process for third-party servers.</strong> Before a public MCP server hits the agent's allowed-servers list, someone has read the code, checked the update mechanism, and verified the identity of who ships it. This is the same review that already exists for third-party libraries in most enterprises. The bar for MCP servers has to be at least that high.</p><p><strong>Per-server allowlists at the agent level.</strong> A given agent has access to a specific list of servers, not "any server the platform offers." Cross-server prompt injection is a real failure mode; the mitigation is scope reduction. This surface lives in the <a href="/post/production-ai-agent-governance">production agent governance reference</a> and gets exercised on every MCP-backed workload.</p><h3>The auth and identity story</h3><p>The 2026 MCP spec settles authentication on OAuth 2.1. What that means practically for the enterprise:</p><p><strong>Dynamic Client Registration.</strong> MCP clients register with the authorization server at runtime. The enterprise's identity provider has to support DCR; if it does not, this becomes the blocking work.</p><p><strong>Protected Resource Metadata.</strong> The client discovers the authorization server from the resource server's own metadata. This makes the wiring lighter but also means the metadata itself becomes a trust surface — the discovery endpoint has to be authenticated.</p><p><strong>Resource Indicators.</strong> Tokens are bound to a single resource server and cannot be replayed elsewhere. This is what stops a token issued for the ticketing MCP server from being usable against the finance MCP server.</p><p><strong>Session-scoped authorization.</strong> For write actions, the emerging pattern is a session-scoped token that expires when the human-defined session ends. The agent cannot renew the session on its own; a human explicitly approves a new session. This is the mechanism that lets an enterprise say "yes, agents can write to production" without giving them permanent write credentials.</p><p>For enterprises whose identity stack still lives in older SAML-and-basic-auth territory, MCP forces a modernisation conversation. That conversation is a Phase 1 architecture item, not a Phase 3 surprise.</p><h3>Observability is not optional</h3><p>Every enterprise MCP deployment we have seen ship successfully in 2026 has full trace-level observability on every tool call. The parameters, the identity of the calling agent, the identity of the human on whose behalf it acts, the result, and (where feasible) a cryptographic hash of the response. This is the <a href="/post/ai-agent-observability">OpenTelemetry GenAI pattern from our observability post</a> applied to MCP specifically — <code>execute_tool</code> spans with MCP-specific attributes, the <a href="https://opentelemetry.io/blog/2026/genai-observability/">protocol got its own trace layer</a> in the 2026 spec.</p><p>The reason this matters is not just debugging. Between January and February 2026 alone, security researchers filed <a href="https://promptention.ai/blog/mcp-security-guide-2026/">more than thirty CVEs</a> targeting MCP servers, clients, and infrastructure components. The highest-severity finding scored 9.6 on CVSS. Trace-level observability is what lets an enterprise say, after an incident, exactly which agents ran which tools and what they returned. Without it, incident response on an MCP-related issue is a forensic archaeology exercise.</p><h3>The pattern for shipping MCP into an enterprise</h3><p>The sequence we run when an enterprise adopts MCP for the first time:</p><p><strong>Phase 1 — Identity modernisation.</strong> OAuth 2.1 support in the enterprise IdP. Dynamic Client Registration. Resource Indicators enforced. If this is already there, this phase is a review; if not, it is the biggest single piece of preparatory work.</p><p><strong>Phase 2 — Registry and review gate.</strong> Stand up the private registry. Define the review criteria for both internal and third-party servers. Assign the review owner. No server connects to an agent until this is running.</p><p><strong>Phase 3 — First sidecar server.</strong> Pick one internal system with high agent demand and expose it as an MCP server. Typically the ticketing system, the code repository, or the data warehouse. Ship the server, ship the observability, ship the review.</p><p><strong>Phase 4 — Fleet the pattern.</strong> Add servers to the registry. Add agents that consume them. The economics of MCP show up here — a new agent that needs the ticketing system reuses the sidecar server, not a new client. The governance layer already exists; the new agent inherits it.</p><p><strong>Phase 5 — Public server consumption.</strong> With the review gate mature, begin consuming public servers from the Linux Foundation registry. GitHub, Stripe, Cloudflare — the ones where the vendor operates the server and the enterprise's job is trust review, not implementation.</p><p>Phases 1 and 2 are what stop MCP from becoming a security incident later. The temptation is to skip them and go straight to Phase 3 because the sidecar shows the fastest value. That temptation is what created the 2026 CVE surge.</p><h3>The decision, in one sentence</h3><p>MCP pays off when you need one governance layer over many tool calls and can afford the latency to get it. When either half of that sentence is false, keep the direct call. Everything else in the enterprise MCP conversation — the registry, the OAuth 2.1 stack, the observability, the review gate — flows from getting that decision right.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-mcp-for-enterprise-when-it-pays-off.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Modernising legacy systems for AI agent access]]></title>
            <link>https://twistag.com/thinking/modernising-legacy-systems-for-ai-agent-access</link>
            <guid isPermaLink="false">https://twistag.com/thinking/modernising-legacy-systems-for-ai-agent-access</guid>
            <pubDate>Thu, 30 Jul 2026 00:14:24 GMT</pubDate>
            <description><![CDATA[Legacy systems don't need a rewrite to be agent-accessible. Here's the shim-layer pattern, the anti-corruption layer, and when each pays off.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Most enterprises trying to deploy AI agents in 2026 hit the same wall. The agents work. The models are ready. The vendor pitches are compelling. And then the target system — the ERP that runs on a database from 2007, the ticketing platform with no API, the underwriting engine written in a language the current team does not speak — cannot be called by anything modern. Legacy system integration is now <a href="https://viston.tech/ai-agent-integration-with-legacy-systems-a-practical-guide-for-businesses-in-2026/">named as the first of three core obstacles to agentic AI</a> because older systems lack the APIs, real-time execution, and modularity that agents need. The engineering response is not a rewrite. Rewrites take years, the enterprise cannot wait, and the operational risk is unacceptable. The response is a shim layer — a modern, agent-accessible surface that sits between the agent and the legacy system, translating in both directions, protecting each side from the other. This post is the pattern for building that shim layer. It is a cluster child of our pillar on <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a>.</p><h3>Key takeaways</h3><ul><li><strong>The shim layer is the default modernisation pattern for agent access.</strong> Not a rewrite. Not a lift-and-shift. A translation layer that makes the legacy system agent-accessible without changing it.</li><li><strong>The anti-corruption layer (ACL) is the shim's job.</strong> <a href="https://netalith.com/blogs/microservices-architecture/anti-corruption-layer-pattern-legacy-integration-2026">It converts legacy formats and contracts into modern standards</a> and stops legacy technical debt from bleeding into the agent's world. The ACL is the structural guarantee.</li><li><strong>The interface choice — REST, GraphQL, MCP — depends on the fleet.</strong> One agent, one tool: REST. Many agents, one tool: MCP. Many agents, many tools, need for query composition: GraphQL.</li><li><strong>Agent-readable is more than API-readable.</strong> The shim also has to expose schema, entity resolution, and provenance the agent can reason over. An endpoint that returns raw legacy field names is a technically valid API and an operationally useless one.</li><li><strong>The strangler-fig pattern retires the shim over time.</strong> The shim is designed to eventually replace the legacy system, not stay forever. Every service the shim owns is a candidate for a proper rewrite once the agent workload proves the shape.</li></ul><h3>The five failure modes when you skip the shim</h3><p>Every enterprise that tries to point an agent directly at a legacy system encounters some subset of the same failures:</p><p><strong>The legacy schema leaks into the agent's world.</strong> The agent's prompts contain field names like <code>CUST_ADR_LN1_TXT</code> because the legacy database schema flows straight through. The prompt becomes a cryptography exercise. The <a href="https://netalith.com/blogs/microservices-architecture/anti-corruption-layer-pattern-legacy-integration-2026">ACL exists specifically to prevent this bleed</a>.</p><p><strong>Rate limiting on the legacy system breaks the agent.</strong> The legacy system was sized for a handful of concurrent human users. The agent runs a hundred queries in a burst and takes production down.</p><p><strong>Legacy authentication is incompatible with modern identity.</strong> The legacy system uses a static service account and cannot represent per-user identity. The agent has no way to preserve the identity of the human on whose behalf it acts. Audit becomes impossible.</p><p><strong>Response shapes are unstable.</strong> A field that returns a number in most cases returns a string sometimes, or is missing under some legacy code path. The agent's downstream reasoning breaks on the edge case.</p><p><strong>No sandbox environment.</strong> The legacy system has one environment: production. The agent development team is testing against live customer records.</p><p>The shim layer's job is to solve all five. It is not just an API on top of an API. It is a governance, identity, and reliability surface that the legacy system did not have.</p><h3>The shim layer's five responsibilities</h3><p>Every shim layer that ships successfully in 2026 does the same five things:</p><p><strong>Translation.</strong> The shim converts between the legacy system's data shape and a modern, agent-readable shape. Legacy field names get normalised. Legacy status codes become semantic strings. Legacy timestamps get standardised to UTC ISO-8601. This is the classical ACL work.</p><p><strong>Identity propagation.</strong> The shim accepts modern identity — OAuth 2.1 bearer tokens, session-scoped credentials — and translates that into whatever the legacy system understands (service-account impersonation, per-user credential mapping, whatever the legacy stack supports). The audit trail on the shim side records who initiated each call.</p><p><strong>Rate limiting and back-pressure.</strong> The shim protects the legacy system from agent traffic. A hundred-query burst becomes a queued, paced sequence. Circuit breakers stop cascading failures when the legacy system starts responding slowly.</p><p><strong>Response shape stabilisation.</strong> The shim normalises response variability. Missing fields get defaulted or explicitly nulled. Type coercion happens on the shim side. Response schemas are versioned and validated. Agents reason over a stable contract, not the legacy system's edge cases.</p><p><strong>Sandbox and replay.</strong> The shim exposes a sandbox surface that returns realistic responses without touching the legacy system. This lets the agent's development, evaluation, and CI pipelines run without needing production access to the legacy stack.</p><p>Skip any of the five and the shim becomes another leaky abstraction. All five make the shim a real integration layer.</p><h3>The interface choice — REST, GraphQL, MCP</h3><p>Every shim layer exposes an interface. In 2026 there are three defensible choices, and the choice depends on the shape of the fleet.</p><p><strong>REST.</strong> Sensible when the fleet is small and the queries are predictable. One agent, one tool, one endpoint. The shim exposes a small number of hand-crafted endpoints tuned for the specific queries the agent runs. Fast to ship, easy to debug, familiar to every backend engineer. This is still the winning choice for a large share of workloads.</p><p><strong>GraphQL.</strong> Sensible when the agent needs to compose queries across multiple legacy entities. Order + customer + line items + shipment status, in one round-trip, without inventing a new REST endpoint for each combination. GraphQL's schema also gives the agent an introspectable surface — the agent can discover what queries are possible. This is the shape that pays back when the shim is fronting a system with many related entities.</p><p><strong>MCP.</strong> Sensible when many agents will call the shim, and governance sits on the tool-call layer. The <a href="/post/mcp-for-enterprise-when-it-pays-off">MCP decision framework from our protocol post</a> applies unchanged — MCP pays off when the fleet benefits from one governance layer over many tool calls. For a shim in front of a widely-consumed legacy system, this is often the right choice.</p><p>Nothing prevents combining these — a shim can expose a GraphQL surface for exploratory work and an MCP surface for governed tool calls over the same underlying data. But the choice sets the expectations for how the fleet consumes the shim, and that choice deserves the Phase 1 conversation.</p><h3>Agent-readable is more than API-readable</h3><p>An API that returns valid JSON is not automatically usable by an agent. Agent-readable requires three properties on top of a working endpoint:</p><p><strong>Schema the agent can reason over.</strong> Every field has a name that means something outside the legacy context, a type, and a docstring that describes what it represents in business terms. <code>customer_lifetime_value_eur</code> beats <code>CLV_AMT</code>. An agent reading a tool description with a well-annotated schema will use it correctly the first time.</p><p><strong>Entity resolution.</strong> The shim gives every entity a stable identifier the agent can use across tool calls. A customer returned by the "get customer" call carries an identifier that the "get orders" call accepts. Without entity resolution, the agent has to reason about which representation of a customer it is holding — a task humans get wrong and models get more wrong.</p><p><strong>Provenance.</strong> The shim's responses include where the data came from, when it was last updated, and how confident the shim is in the value. This lets the agent reason about staleness and reliability — critical when the legacy system is a batch-updated data warehouse rather than a live transactional system.</p><p>These three properties turn a technically valid API into a surface an agent can operate on without hand-holding. They are the concrete meaning of the <a href="/post/agentic-ai-workflow-services-defined">agent-readable data governance definition from our pillar on data readiness</a> — the pattern applied to a legacy system.</p><h3>The strangler-fig endgame</h3><p>The shim is not a permanent architecture. It is a bridge, and the bridge has a planned decommissioning.</p><p>The <a href="https://viston.tech/ai-agent-integration-with-legacy-systems-a-practical-guide-for-businesses-in-2026/">strangler-fig pattern</a> describes what happens next. As the shim proves the agent-facing shape of each legacy capability, the corresponding legacy system component becomes a candidate for a proper modernisation. Not a lift-and-shift — a re-implementation of the capability, with the shim's agent-readable surface as the specification. The new implementation lands behind the shim; the shim continues to expose the same contract; the legacy system component gets retired. From the agent's perspective, nothing changes.</p><p>This is what makes the shim investment pay back over multiple years. In year one, the shim makes agents possible against the legacy system. In years two and three, the shim's contract becomes the target architecture for the modernisation. In year five, the legacy system is gone and the shim is the system.</p><p>The alternative — pointing agents directly at the legacy system and betting on a future rewrite — leaves the enterprise stuck. The rewrite gets postponed indefinitely, the agent workload accumulates on brittle direct integrations, and the eventual modernisation project has to re-instrument every agent. The shim pattern avoids that trap.</p><h3>The five-phase build</h3><p>The sequence we run when we ship a shim layer into an enterprise:</p><p><strong>Phase 1 — Contract design.</strong> Sit with the agent team and the legacy system owners. Define the agent-readable contract: what entities exist, what queries are needed, what response shapes look like. This is the highest-impact phase and the one enterprises are most tempted to skip.</p><p><strong>Phase 2 — Interface choice and skeleton.</strong> Pick REST vs GraphQL vs MCP. Stand up the shim skeleton with a small number of full paths — enough to prove the pattern, not enough to lock in the design.</p><p><strong>Phase 3 — ACL and legacy integration.</strong> Wire the shim to the legacy system. Translation, identity, rate limiting, response stabilisation. Ship one path from the agent all the way to the legacy system and back.</p><p><strong>Phase 4 — Sandbox and CI.</strong> Ship the sandbox surface. Get the agent's CI pipeline running against the sandbox. This is what unblocks the agent team from waiting on production legacy access.</p><p><strong>Phase 5 — Fleet the pattern.</strong> Add more paths. Add more agents. Retire direct legacy access. This is the phase where the shim starts paying back the investment.</p><p>Phase 1 is where projects succeed or fail. A shim built without a rigorous contract design ends up mirroring the legacy system's shape, which is exactly the failure mode the shim was supposed to prevent. Every hour spent on the contract pays back tenfold when the agents actually start calling it.</p><h3>The one question that decides the whole architecture</h3><p>Every legacy modernisation-for-agents project turns on one question: is the shim the future, or a bridge?</p><p>If the shim is a bridge, it is designed to be strangled out. The contract is what matters; the implementation is temporary. Investment goes into the contract quality and the sandbox. The legacy system's replacement is planned from day one.</p><p>If the shim is the future, it is designed to become the system. The contract is what matters; the implementation gets industrial-grade over time. The legacy system may never fully retire — parts of it might live behind the shim for years — but the shim is where the enterprise's forward architecture lives.</p><p>Both answers are legitimate. The wrong answer is not deciding. A shim built without a stance on this question ends up with neither the disciplined bridge economics nor the industrial-grade future architecture. Naming the answer in Phase 1 is what makes the rest of the work coherent.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-modernising-legacy-systems-for-ai-agent-access.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[What a 70%-AI-generated codebase actually looks like]]></title>
            <link>https://twistag.com/thinking/70-percent-ai-generated-codebase</link>
            <guid isPermaLink="false">https://twistag.com/thinking/70-percent-ai-generated-codebase</guid>
            <pubDate>Thu, 30 Jul 2026 00:14:04 GMT</pubDate>
            <description><![CDATA[Industry average is 25-40%. We're at 70%. Here's what changes in review layers, PR shape, and test coverage when AI writes most of the code.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>The industry data on AI-generated code is inconsistent because everyone measures it differently. Sonar's <a href="https://www.sonarsource.com/state-of-code-developer-survey-report.pdf">2026 developer survey</a> puts the share of AI-generated code across all committed code at close to 50% early this year. <a href="https://larridin.com/developer-productivity-hub/developer-productivity-benchmarks-2026">Larridin's engineering benchmarks</a> put the "sweet spot for mature teams" at 25-40%. GitHub Copilot's own telemetry puts the accepted-suggestion rate at about 30%. Twistag runs closer to 70% AI-generated on greenfield engagements — a number worth being specific about, because at that percentage the review layers, the PR shape, the test discipline, and the definition of a senior engineer's job all change. This post is what changes. It is the pillar of our cluster on AI-native product engineering (F).</p><h3>Key takeaways</h3><ul><li><strong>70% AI-generated is a delivery outcome, not a target.</strong> The number falls out of a specific review discipline. Missing the discipline and pushing the AI harder produces slop faster; missing the discipline and pushing it less loses the productivity.</li><li><strong>The review layer count doubles.</strong> Classical review is one layer (senior engineer reads the PR). AI-native review is four layers (AI author agent → AI reviewer agents → senior engineer → integration tests as executable spec). Each layer catches a different failure mode.</li><li><strong>The PR shape shrinks and the PR count grows.</strong> A 70%-AI-generated feature ships in five or six small PRs, not one large one. Each PR is small enough for a reviewer agent and a human to fully understand. This is the biggest visible change in the codebase.</li><li><strong>Tests are written first, more often than not.</strong> Not classical TDD. The pattern is: senior engineer writes the spec and the test cases; AI writes the implementation; the tests are the acceptance signal. Test coverage on AI-authored code is systematically higher than on human-authored code from the same team.</li><li><strong>The senior engineer's job becomes editorial.</strong> Less typing, more judgement about what to accept, what to reject, what to redesign. This is a bigger shift in the role than most engineering leaders have priced in.</li></ul><h3>The percentage question — what "70%" actually means</h3><p>Any measurement of AI-generated code is contested because there are at least four things you could count. The pattern in 2026 that produces the least ambiguous number:</p><ul><li>The unit of measurement is <strong>lines of code in merged PRs</strong> (not suggestions, not accepted suggestions, not "AI-assisted" self-reports).</li><li>The classification is <strong>whether the line was originally authored by an AI tool</strong> or a human — including lines the human later edited. If the human touched a line the AI wrote, it still counts as AI-originated.</li><li>The scope is <strong>a specific service or repository over a specific time window</strong> (a sprint, a month, a release). Aggregating across the whole engineering org washes out the interesting variation.</li></ul><p>By that definition, our greenfield engagements in 2026 sit around 70%. Brownfield engagements sit closer to 40-50%, because integration and migration work has more one-off human logic. The <a href="https://larridin.com/developer-productivity-hub/developer-productivity-benchmarks-2026">Sonar and Larridin numbers</a> are believable industry averages — 25-40% for mature teams. The gap between 40% and 70% is not tool selection; the tools are converging. The gap is the review and integration discipline.</p><h3>The four review layers</h3><p>At 70% AI-generated, the classical single-reviewer PR model breaks. The bugs it lets through are different from human-authored bugs — subtle logic errors on paths that look plausible, off-by-one issues on boundary conditions, incorrect assumptions about the shape of upstream data. The review response is four layers.</p><p><strong>Layer 1 — the author agent's self-check.</strong> The AI that wrote the code is prompted to review its own output before it opens the PR. It runs the tests, reads its own diff, catches trivial issues (missing imports, obvious type mismatches, unreachable branches). This layer removes noise from what reaches the reviewer.</p><p><strong>Layer 2 — reviewer agents.</strong> Anthropic's <a href="https://www.infoq.com/news/2026/04/claude-code-review/">Claude Code Review multi-agent system</a>, launched in March 2026, is the reference implementation. When a PR opens, a fleet of specialised agents examines the diff — one for logic errors, one for boundary conditions, one for API misuse, one for security issues, one for project-convention compliance. A verification step tries to disprove each finding before it gets posted. The surviving findings become inline PR comments.</p><p><strong>Layer 3 — the senior engineer.</strong> A human reads the diff, reads the reviewer agents' comments, decides what to accept and what to override. This is the editorial layer, and it is where most of the seniority now sits. The senior engineer is not typing the code; they are judging the code.</p><p><strong>Layer 4 — integration tests as executable spec.</strong> The tests that were written first (see below) are the final gate. If they pass, the change ships. If they fail, the AI is prompted to fix — often successfully in one iteration, sometimes escalating back to layer 3 for a design change.</p><p>Skipping any of the four layers is where the failures happen. Skip layer 1 and the reviewer agents drown in noise. Skip layer 2 and the human reviewer misses the subtle logic errors. Skip layer 3 and the codebase drifts from the team's convention. Skip layer 4 and edge cases ship broken.</p><h3>The PR shape change</h3><p>The single biggest visible difference in an AI-native codebase is the size and count of pull requests. Classical shape: one PR per feature, 400-1500 lines, one reviewer. AI-native shape: five or six PRs per feature, 80-200 lines each, layered review on every one.</p><p>The reason is the review economics. A human reviewer at layer 3 can hold about 200 lines of a diff in their head with real understanding. Above that, they start reviewing the shape and skimming the content — which is exactly where subtle AI-authored bugs slip through. A reviewer agent at layer 2 is equivalent — its per-PR accuracy drops on large diffs. Small PRs keep both reviewers at peak performance.</p><p>The count change matters for engineering leadership because velocity metrics that count PRs shipped or merged look inflated at first glance. They are not inflated; the unit got smaller. Merged-lines-per-week or merged-features-per-week are the metrics that stay comparable across the AI-native shift.</p><h3>Tests first — but not classical TDD</h3><p>Classical TDD is: write the test, watch it fail, write the implementation, watch it pass. It is a discipline for humans working alone. AI-native testing looks different:</p><p><strong>The senior engineer writes the acceptance criteria first.</strong> In plain English, or in a spec document, or as the top of the AI's context window. This is the design step — deciding what "done" means.</p><p><strong>The AI writes the tests before the implementation.</strong> Given the acceptance criteria, the AI generates unit tests, integration tests, and edge cases. The human reviews the tests. This step is where design errors get caught — if the tests are wrong, the implementation will be wrong.</p><p><strong>The AI writes the implementation.</strong> Iterates until the tests pass. If it cannot make them pass, it escalates to the human — which usually means the tests captured a constraint the implementation cannot satisfy without a design change.</p><p><strong>The tests become the merge signal.</strong> Layer 4 of the review. Green tests means the change ships. Red tests block.</p><p>The result is test coverage that is systematically higher than a human-authored codebase from the same team. Not because humans are lazy — because AIs are relentless. The AI does not skip a test case because it is boring; it writes them all.</p><h3>The seniority shift</h3><p>The role change is the biggest thing to plan for. In a classical codebase, the senior engineer's day is a mix of typing (implementation), thinking (design), and reviewing (peers). The typing is the fungible part; the thinking and reviewing are where the seniority lives.</p><p>At 70% AI-generated, the typing collapses. The senior engineer no longer implements most of what they design; the AI does. What is left is thinking and reviewing — the parts that were always the highest-value anyway, but were mixed into a day that also included a lot of typing.</p><p>The specific skill shifts:</p><p><strong>Judgement about what to accept.</strong> Every AI-generated diff is a proposal. The senior engineer's job is to accept it, edit it, reject it, or redesign what was asked for. This is editorial work — closer to a book editor's day than a typing engineer's day.</p><p><strong>Framing the problem for the AI.</strong> The AI is only as good as the specification it receives. Writing acceptance criteria that are unambiguous, complete, and testable is the new craft skill. It looks like technical writing; it feels like design.</p><p><strong>Recognising the failure modes.</strong> The subtle AI failure modes — plausible-looking wrong code, off-by-one on boundaries, wrong assumptions about upstream shape — are trainable. Senior engineers who have spent six months in an AI-native codebase spot them fast. Junior engineers still miss them. This is why the <a href="https://larridin.com/developer-productivity-hub/developer-productivity-benchmarks-2026">senior review burden went up 20-35%</a> in 2026 as junior engineers leaned harder on AI.</p><p>The engineering managers who priced this shift correctly staffed differently. They kept the senior count, cut the mid-level count, and stopped hiring juniors into typing roles — because there are no typing roles left.</p><h3>What does not change</h3><p>Three things stay classical even at 70% AI-generated:</p><p><strong>Architecture.</strong> The AI is bad at greenfield architecture. It follows the shape it is given and iterates locally. The senior engineer still designs the service boundaries, the data model, the interface between the new work and the rest of the system. This is where the design decisions get made.</p><p><strong>Security-critical code paths.</strong> Authentication, authorisation, cryptographic operations, PII handling. AI assistance is used here — it is faster to generate the code — but the review discipline tightens further. Two humans review every security-critical PR, and the reviewer agent set expands to include a dedicated security agent.</p><p><strong>Domain knowledge.</strong> The AI does not know the enterprise's business rules. Encoding those rules — into acceptance criteria, into tests, into the code — is human work. The AI helps type it up, but the senior engineer holds the model of what the business actually needs.</p><p>At 70% AI-generated, most of the codebase looks and reads familiar. The change is not in the code's style. It is in who wrote it, how it got reviewed, and what the senior engineer was doing while it happened.</p><h3>What this shifts for the enterprise</h3><p>The enterprise that hires an engineering partner in 2026 is buying delivery velocity that depends on this discipline being real. Ask three questions in the qualification:</p><ul><li><strong>What percentage of the partner's delivered code is AI-authored, by their own definition of the measurement?</strong> The answer should be specific. "A lot" is not an answer. Vagueness is a signal that the discipline is not in place.</li><li><strong>What are their review layers?</strong> Four layers is the current best-practice shape. Two is what most agencies still do. The number correlates with defect rates in production.</li><li><strong>How do they staff?</strong> Senior-heavy is the AI-native shape. Junior-heavy is the traditional shape. Both work at different price points, but they produce different codebases.</li></ul><p>The 70% figure is a delivery outcome of a review discipline. Any partner claiming the number without being able to explain the discipline is claiming a productivity gain they cannot systematically produce.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-70-percent-ai-generated-codebase.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Design engineering in an AI-native pipeline]]></title>
            <link>https://twistag.com/thinking/design-engineering-ai-native-pipeline</link>
            <guid isPermaLink="false">https://twistag.com/thinking/design-engineering-ai-native-pipeline</guid>
            <pubDate>Thu, 30 Jul 2026 00:12:02 GMT</pubDate>
            <description><![CDATA[The design-to-code loop used to be a translation problem. In an AI-native pipeline it collapses to a review problem. Here's what changes.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>The classical design-to-code loop is a translation problem. A designer produces a Figma file. An engineer reads the Figma file, invents implementation details Figma cannot express, ships the code. The design gets translated twice — once when Figma had to compress the designer's intent into layers and constraints, again when the engineer had to invent what Figma left out. Every translation loses fidelity. In an AI-native pipeline, the translation problem collapses. Figma and the coding agents share a rich enough contract that most of the loss disappears; what is left is a review problem — the designer and the engineer both look at the shipped component and judge whether it matches the intent. This post is what changes in the design pipeline, and what the design engineer's role becomes in the middle of that change. It is a cluster child of our <a href="/post/70-percent-ai-generated-codebase">70%-AI-generated codebase pillar</a>.</p><h3>Key takeaways</h3><ul><li><strong>The design-code loop compressed from days to hours.</strong> Not because designers work faster. Because the AI coding tools can read Figma files with enough fidelity that the implementation is a review, not a translation.</li><li><strong>Design tokens became agent-readable in 2026.</strong> Figma's variables and code connections are the specification. The AI reads them as structured data, not as pixel comps. This is the single biggest shift.</li><li><strong>The design engineer role is now the high-value seat.</strong> Not a designer, not a frontend engineer — a person who owns both the design token discipline and the AI-coding pipeline that consumes it. Rare, valuable, expensive to hire.</li><li><strong>Design systems get stricter, not looser.</strong> With AI generating the components, ambiguity in the design system produces ambiguous components. The design system's job shifts from "guide humans" to "specify unambiguously for AI."</li><li><strong>The prototype-to-production gap collapses.</strong> A Figma prototype that would previously have taken a sprint to become production-worthy code now becomes production-worthy code in an afternoon — as long as the design token discipline is in place.</li></ul><h3>Why the classical loop had translation loss</h3><p>The classical Figma-to-code loop:</p><ul><li>The designer produces layouts, components, and states in Figma.</li><li>The engineer reads the Figma file and interprets it — what is a component vs a variant, what should be responsive, what state transitions are implied.</li><li>The engineer writes the code, invents the details Figma did not specify (accessibility semantics, hover behaviour beyond the hover frame, keyboard focus order, dark mode variants that were not designed).</li><li>The engineer ships the code. The designer opens it. The designer notices twenty small discrepancies with the design. Some get fixed, some do not.</li></ul><p>Each step lost information. The designer had to compress their intent to fit Figma. Figma compressed further to fit its layer model. The engineer had to expand back out, and their expansion was necessarily a guess about the designer's intent for the things Figma could not carry. The finished component had drift from the design, and the drift was invisible to anyone who was not looking at both the Figma file and the running code side by side.</p><p>The AI-native pipeline does not eliminate the translation. It moves it to a place where both parties can see it and judge it.</p><h3>What changed in 2026</h3><p>Three things shifted in the design tooling and the coding agents that together closed most of the loss:</p><p><strong>Figma's structured tokens became a first-class contract.</strong> Design tokens (colours, spacing, typography, motion) live as Figma variables. They map to CSS variables, Tailwind config, or design-system code in the target framework via Figma's Code Connect and similar surfaces. The AI reads this mapping as structured data. When the coding agent produces a button, it uses the design token names, not hex codes.</p><p><strong>Coding agents got fluent at reading Figma.</strong> <a href="https://help.figma.com/hc/en-us/articles/32132100833559-Guide-to-the-Figma-MCP-server">Figma's own MCP server</a>, plus a growing set of third-party integrations, lets an AI coding tool read a Figma frame with high fidelity — components, variants, tokens, states. The AI can extract the layer tree, the component hierarchy, and the tokens in a single call. This is what makes the design-to-code loop hours-not-days.</p><p><strong>Design system code generation caught up.</strong> The AI can generate framework-specific components (React, Vue, Svelte, native platform) that reference the design system's primitives rather than hard-coding the styling. The generated component slots into the design system's shape, not next to it. This matters because a generated component that does not use the design system's primitives creates the same drift problem the classical loop had.</p><p>Together, the loop looks different. The designer produces the design and the tokens. The engineer (or the design engineer — see next section) points the coding agent at the Figma frame with the design tokens configured. The agent produces a first cut of the component in the design system's shape. The designer opens the running component. The engineer and the designer judge together what needs to change. Adjustments cycle back through the coding agent. The finished component looks like the design because it was generated from the design's structured data, not translated from a pixel comp.</p><h3>The design engineer role</h3><p>The role that emerged in this pipeline is the design engineer. Not a designer, not a frontend engineer — a person who owns the boundary between them.</p><p>The specific responsibilities:</p><p><strong>Design token discipline.</strong> Every colour, every spacing value, every typography ramp lives as a token with a name and a mapped implementation. The design engineer owns the naming, the mapping, and the review process that stops one-off values from accumulating. A design system without token discipline produces AI-generated components that are visually correct and structurally chaotic.</p><p><strong>AI coding pipeline configuration.</strong> The Figma-to-code AI works well only when the coding agent has the right context — which components exist, which tokens map to what, what the framework conventions are. The design engineer configures this context and maintains it as the design system evolves.</p><p><strong>Boundary judgement calls.</strong> The parts of a component that are design decisions vs the parts that are engineering decisions were always contested. In an AI-native pipeline, the design engineer holds the call — they know what the designer intended, and they know what the code needs, and they judge the trade-off in the moment.</p><p><strong>Design system testing.</strong> Component tests that catch design regressions — snapshot tests, visual regression, accessibility tests. The design engineer owns the test suite that stops the design from drifting silently.</p><p>Hiring for this role is hard because it is genuinely a hybrid — a candidate who was a good frontend engineer but never designed will miss the design half; a candidate who was a good designer but never wrote production code will miss the engineering half. Teams that get this hire right compress their design-to-production cycle by a large factor. Teams that leave the seat empty and split the responsibilities across a designer and an engineer keep some of the translation loss.</p><h3>The design system's new job</h3><p>Design systems used to be documentation for humans — a set of principles, patterns, and examples that a designer or engineer read to make consistent choices. In an AI-native pipeline, the design system is a specification for the AI.</p><p>The concrete shifts:</p><p><strong>Unambiguous naming.</strong> Two tokens named <code>spacing-md</code> and <code>space-3</code> are the same value serving different purposes and will confuse the AI. The design system consolidates to one canonical name per concept.</p><p><strong>Fewer optional parameters.</strong> A component with 30 optional props is a component the AI will produce inconsistently. The design system reduces the surface — required props for the essentials, one or two optional variants, no more.</p><p><strong>Explicit anti-patterns.</strong> The design system documents not just what to do but what not to do. "Do not put an icon-only button next to a labelled button in the same row." The AI reads these as constraints and does not violate them.</p><p><strong>Machine-readable examples.</strong> Every component in the design system has a canonical usage example in code. The AI uses these examples as reference when generating new instances. Without them, the AI invents its own reference, and consistency drifts.</p><p>Design systems that were maintained casually in 2024 need to be maintained with engineering discipline in 2026. The upside is that the AI does the tedious work of applying the system correctly; the downside is that the system's ambiguities become the AI's errors.</p><h3>What the loop looks like from start to finish</h3><p>The full-cycle shape when the pipeline is in place:</p><ul><li><strong>Designer</strong> produces a Figma frame using design tokens. Adds a component to the file, wires the tokens, sets the states.</li><li><strong>Design engineer</strong> points the coding agent at the frame. Provides the framework context (React, Tailwind, project conventions).</li><li><strong>Coding agent</strong> generates a first-cut component that uses the design system's primitives and references the tokens by name.</li><li><strong>Design engineer</strong> reviews the component in a running preview. Adjusts what is off. Adds the accessibility semantics and keyboard behaviour if the AI missed them (it usually does not, but sometimes).</li><li><strong>Designer</strong> opens the running component. Compares to the intent. Requests changes. The design engineer cycles the changes through the AI.</li><li><strong>PR opens.</strong> Review agents from the <a href="/post/70-percent-ai-generated-codebase">multi-agent review pattern</a> examine the code. The design engineer signs off. The PR merges.</li><li>Total elapsed time: hours, not days.</li></ul><p>The parts that stay slow:</p><ul><li>Genuinely new design patterns that are not in the design system. The AI cannot generate what has not been specified. The design engineer has to add the pattern to the system first, then generate against it.</li><li>Complex interactions with a lot of state (drag-and-drop, animation-heavy flows, real-time collaboration). The AI generates the shape; the engineer still writes the state logic.</li><li>Accessibility for genuinely custom components. The AI does the standard patterns; anything outside them still needs the engineer to think through the semantics.</li></ul><p>The parts that get faster are most of the work — routine components, variants, state-based styling, dark-mode variants, responsive breakpoints. Which is why the loop compresses.</p><h3>What this shifts for the buyer</h3><p>The enterprise buying design and engineering services from a partner in 2026 should look for three things:</p><p><strong>A design engineer on the team.</strong> Not two people, one who designs and one who engineers, split across the boundary. One person who owns both. If the partner does not have this role, the loop has translation loss.</p><p><strong>A published design token strategy.</strong> The partner should be able to describe how they set up tokens in Figma, how the tokens map to the framework, and how the coding agents consume them. A partner that cannot answer this is not running an AI-native pipeline; they are running the classical loop with AI as a productivity boost.</p><p><strong>Cycle time in hours, not sprints.</strong> For components that fit the design system, the design-to-shipping cycle should be measured in hours. If the partner needs a sprint to ship a button, they are still in the classical loop.</p><p>The design-code loop is the part of product engineering where AI-native discipline shows up fastest. A partner that has not converged on the pipeline in 2026 is delivering at the classical loop's speed, which is now roughly 3-5x slower than the AI-native shape.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-design-engineering-ai-native-pipeline.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Prototype-in-a-day: the AI-native product discovery loop]]></title>
            <link>https://twistag.com/thinking/prototype-in-a-day</link>
            <guid isPermaLink="false">https://twistag.com/thinking/prototype-in-a-day</guid>
            <pubDate>Thu, 30 Jul 2026 00:11:45 GMT</pubDate>
            <description><![CDATA[One day from idea to running prototype. Here's what fits, what doesn't, and where the boundary with production sits.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>The product discovery loop used to be measured in weeks. A designer would sketch, a product manager would spec, an engineer would build a scaffolded prototype, and by the end of a two-week sprint the team had something a user could click through. The AI-native version compresses that loop to a day — and by day we mean a real working day, not a stretched sprint. Idea in the morning, running prototype by afternoon, first user feedback by end of day. This post is what fits in the day, what does not, where prototype-in-a-day ends and production begins, and the common failure modes when the discipline slips. It is a cluster child of our <a href="/post/70-percent-ai-generated-codebase">70%-AI-generated codebase pillar</a>.</p><h3>Key takeaways</h3><ul><li><strong>Prototype-in-a-day is discovery, not delivery.</strong> The prototype exists to test an idea, not to be shipped. Confusing the two produces prototypes that get rushed into production and production systems that get scoped like prototypes.</li><li><strong>The loop has three phases in one day: shape, build, test.</strong> Morning shape (spec, sketches, design tokens). Middle build (AI coding agents produce the running prototype). Afternoon test (real user in front of it, learning captured).</li><li><strong>What fits in the day is a single hypothesis, one flow, one persona.</strong> Multiple hypotheses do not fit. Prototypes that try to test three flows in one day test none of them well.</li><li><strong>The prototype-to-production gap is real and it is the point.</strong> The prototype cheats on the things that make production expensive (auth, edge cases, error states, observability). Naming what it cheats on is what stops the "just ship the prototype" trap.</li><li><strong>The AI stack that makes it possible is small and specific.</strong> Design tokens in Figma. A coding agent with access to a starter framework. A test surface the user can hit without deploying. Nothing exotic.</li></ul><h3>What "a day" actually is</h3><p>The day is not eight hours of continuous work by one person. It is a specific choreography:</p><p><strong>Morning (2 hours) — shape.</strong> Product lead frames the hypothesis in one paragraph: "we think X user needs Y capability to accomplish Z." Designer produces enough Figma to specify the surface — one flow, three or four screens, using the existing design tokens. No pixel-perfect polish; enough to be readable by the coding agent and the human reviewer. Product lead lists five specific questions the prototype needs to answer.</p><p><strong>Middle (3-4 hours) — build.</strong> Design engineer points the coding agent at the Figma frames. The agent produces a running prototype in the team's starter framework (Next.js on Vercel, or the equivalent). The prototype uses fake data, skips real auth, does not persist state across sessions. What matters is that the flow works and the surface looks close enough to the intended experience that a user can react to it. The <a href="/post/design-engineering-ai-native-pipeline">design engineering pipeline from our sibling post</a> is what makes this possible.</p><p><strong>Afternoon (2 hours) — test.</strong> Real user (or realistic proxy) sits in front of the prototype for 30-45 minutes. Product lead facilitates, designer observes. The five questions from the morning get answered. The learning is captured in writing before the session ends — because the session is where the loop earns its keep.</p><p><strong>End of day</strong> — decision. Continue with this direction, pivot to a new hypothesis, or drop the initiative. Whichever it is, the next day starts with clarity that would previously have taken two weeks.</p><p>The whole day involves three or four people, but only one of them (the design engineer) is heads-down the whole time. Product and design pair through morning and afternoon, and hand off to the design engineer during the build.</p><h3>What fits (and what does not)</h3><p>The single most important discipline is scoping the hypothesis to fit the day.</p><p><strong>Fits:</strong></p><ul><li>A new user surface for an existing capability. "What if the checkout flow started with a single question instead of a form?"</li><li>A new capability in a familiar shape. "What if we added a recommendation panel to the product listing?"</li><li>A choice between two design directions on the same problem. Build both, test both, pick one.</li></ul><p><strong>Does not fit:</strong></p><ul><li>A whole product. Prototype-in-a-day tests parts. Testing a whole product takes multiple days on multiple parts.</li><li>A capability with real backend logic. If the hypothesis is "does the search algorithm feel right," the prototype needs the search algorithm — which is not a day's work. Scope the hypothesis to something a fake result set can test.</li><li>Anything requiring real user data. The prototype cannot legitimately test on real customer data in a day, and testing on synthetic data invalidates half the learning.</li></ul><p>The pattern for scoping down: if the answer to "what will we learn in the afternoon test" is more than two sentences, the hypothesis is too big for a day. Cut it until the answer is two sentences.</p><h3>The AI stack that makes it work</h3><p>The tools that show up in a prototype-in-a-day run:</p><ul><li><strong>Figma</strong> with the team's design system tokens configured. This is the specification the coding agent reads.</li><li><strong>A coding agent</strong> — Claude Code or Cursor Composer per the <a href="/post/claude-code-cursor-ai-coding-stack">AI-coding stack post</a>. Cursor for smaller prototypes, Claude Code when the surface spans many files.</li><li><strong>A starter framework</strong> the team maintains. Whatever the team ships production with (Next.js, SvelteKit, Remix), but stripped down — no auth, no database wiring, no analytics. The starter exists specifically for prototypes.</li><li><strong>A no-config hosting surface.</strong> Vercel preview URLs, Netlify, or the equivalent. The prototype has to be reachable by the user's browser without SSH and without a build pipeline.</li><li><strong>A user-testing surface.</strong> Session recording, live share, or in-person. Whichever fits the user; the key is that the observer sees what the user sees.</li></ul><p>The stack is deliberately not exotic. Prototype-in-a-day works because the team has practised the choreography, not because they have a novel technology. Teams that try to introduce a new tool during a prototype day burn the day on the tool.</p><h3>Where the prototype-to-production gap sits</h3><p>The prototype is designed to cheat on the things that make production expensive. Naming what it cheats on is what stops the "just ship it" trap.</p><ul><li><strong>Authentication.</strong> The prototype uses a hardcoded logged-in user or bypasses auth entirely. Production needs real login, session management, and OAuth flows.</li><li><strong>Persistence.</strong> The prototype fakes writes. Real writes need database schema, migrations, referential integrity, and backup.</li><li><strong>Error states.</strong> The prototype assumes the happy path. Production has to handle network failures, upstream timeouts, validation errors, and user mistakes.</li><li><strong>Observability.</strong> The prototype has no metrics, no logs, no error tracking. Production needs all three.</li><li><strong>Edge cases.</strong> The prototype handles the median user. Production handles the tail.</li><li><strong>Accessibility.</strong> The prototype uses reasonable defaults. Production has to meet the accessibility bar the team commits to.</li><li><strong>Security.</strong> The prototype accepts inputs at face value. Production sanitises, authorises, and audits.</li></ul><p>Every one of those is a real cost. The prototype's job is to answer whether the idea is worth paying those costs. If the answer is yes, the delivery team picks up the work with the prototype as the specification. If the answer is no, the team saved the delivery cost.</p><p>The failure mode is skipping the discipline and letting the prototype become the production system. This is a version of technical debt that a competent product lead can spot in Phase 1 of any engagement. The tell: someone says "the prototype is already working, why don't we just clean it up." The clean-up is where the cost lives, and the clean-up is more expensive than starting over.</p><h3>The failure modes when the day slips</h3><p>Three failure modes account for most of the "our prototype day did not work" stories we hear.</p><p><strong>Scope creep during the morning.</strong> The hypothesis grew from one to three. The morning went from 2 hours to 5 hours. The build phase started at 2pm, not 11am. The user test got moved to tomorrow. The day broke. The fix is a strict two-sentence rule on the hypothesis; if it grows, cut it.</p><p><strong>The coding agent got stuck.</strong> The prototype needed a component the design system did not have. The AI improvised, made three attempts, none looked right. The design engineer took over and hand-coded. The build took six hours. The fix is to include "if the design system does not have it, we skip that part of the flow" in the day's scope.</p><p><strong>The user did not show.</strong> No user, no test, no learning. The day produced a prototype with nothing to react to. The fix is to book the user in advance and treat the slot as immovable — the day is shaped around the afternoon, not the other way around.</p><p>The teams that get prototype-in-a-day to work reliably do all three things right and treat the choreography as the product. The teams that treat it as an aspiration ship prototypes on some days and lose others to the failure modes above.</p><h3>What this shifts for the enterprise</h3><p>Enterprises that adopt prototype-in-a-day as their discovery pattern find three things move:</p><ul><li><strong>Discovery cost drops sharply.</strong> A hypothesis that would previously have cost two weeks now costs a day. Portfolios of hypotheses become affordable to test.</li><li><strong>Failure gets cheaper.</strong> A discarded hypothesis costs a day, not a sprint. Teams take more risks because the downside is bounded.</li><li><strong>The gap between product and engineering shrinks.</strong> The design engineer and the product lead are in the same room all day. The classical hand-off between product spec and engineering scoping compresses to nothing.</li></ul><p>The enterprise that has not adopted the pattern in 2026 is running product discovery at 5-10x the cost of a partner that has. Not because the partner is smarter — because the partner has practised the choreography enough times that it happens without ceremony.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-prototype-in-a-day.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[How we run an AI-native agency]]></title>
            <link>https://twistag.com/thinking/how-we-run-an-ai-native-agency</link>
            <guid isPermaLink="false">https://twistag.com/thinking/how-we-run-an-ai-native-agency</guid>
            <pubDate>Wed, 29 Jul 2026 14:33:01 GMT</pubDate>
            <description><![CDATA[Not the marketing pitch. The specific operating model — internal ops, staffing shape, estimation, pricing — that produces 70% AI-generated delivery.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>"AI-native" has become the marketing claim every services firm makes in 2026. This post is the specific operating model behind the claim, at Twistag, as of July 2026. It is honest about what we automated, what we did not, what our team shape looks like now, how we estimate, how we price, and where we are still figuring things out. It is a template other agencies and internal delivery teams can adapt — not a copy-paste playbook, but a set of shapes that we have found stick. It is the pillar of our cluster on AI-native delivery process (G).</p><h3>Key takeaways</h3><ul><li><strong>AI-native is an operating model, not a tool selection.</strong> The tools are the visible layer. The invisible layer — how work gets scoped, staffed, estimated, priced, and reviewed — is where 70% code generation actually comes from.</li><li><strong>We automated the internal ops that produce the biggest time savings: proposal generation, meeting synthesis, spec drafting, code review, delivery reporting.</strong> We did not automate the parts that require judgement or relationship — sales, delivery leadership, senior engineering, client conversation.</li><li><strong>Our team shape changed. Senior count stayed. Mid-level count compressed. Junior hiring changed shape entirely.</strong> Nobody gets hired into a typing role anymore. Every role has an editorial or judgement component that AI does not do.</li><li><strong>Estimation changed twice. First it got faster and looser, then it got tighter again.</strong> The right shape is: we estimate less, we estimate more accurately, we revisit weekly. Fixed-scope engagements got easier; hourly engagements got harder to justify.</li><li><strong>Pricing is moving from time-and-materials toward outcome-based.</strong> Not all the way, not overnight, but the direction is clear. When velocity is 5-10x, T&amp;M no longer represents what the client is buying.</li></ul><h3>What we automated (and what we did not)</h3><p>The exercise most agencies run wrong is trying to automate the client-facing surface first. That is where relationships live and where AI has the least payback. The internal ops surface — the stuff nobody outside the agency ever sees — is where the automation actually pays off.</p><p><strong>Automated internal ops (net time saved is substantial):</strong></p><ul><li><strong>Proposal generation.</strong> The <a href="/product-engineering">Twistag proposal skill</a> takes a discovery-call transcript, a scope note, and produces a first-draft proposal in the correct brand shape. A senior consultant edits it in 30 minutes rather than writing it in 3 hours.</li><li><strong>Meeting synthesis.</strong> Every internal and external meeting gets transcribed and summarised automatically. Action items surface without a manual capture step. The synthesis is close enough to right that a two-minute review catches what matters.</li><li><strong>Spec drafting.</strong> Product managers use AI to draft feature specs from acceptance criteria. Same pattern as the coding tools — the AI produces the first draft, the human edits.</li><li><strong>Code review layer.</strong> The four-layer review from our <a href="/post/70-percent-ai-generated-codebase">70%-AI-generated codebase pillar</a> — author self-check, reviewer agents, senior review, integration tests. The reviewer-agent layer is the biggest single time-saver.</li><li><strong>Delivery reporting.</strong> Client status reports draft themselves from the sprint's telemetry and merged PRs. The delivery lead edits, adds narrative, sends.</li></ul><p><strong>Not automated (deliberately):</strong></p><ul><li><strong>Sales conversations.</strong> The prospect wants a person. AI helps prep, but the conversation is human.</li><li><strong>Delivery leadership.</strong> The delivery lead's judgement about pace, quality, and client mood is the point of the role. AI cannot do it.</li><li><strong>Senior engineering.</strong> Architecture, hard trade-offs, incident response. The AI helps, but the senior engineer is accountable.</li><li><strong>Client relationship management.</strong> The relationship is a person-to-person contract. AI cannot own it.</li><li><strong>Hiring interviews.</strong> The interview is a judgement about whether a person will thrive on the team. AI could screen; we do not use it that way.</li></ul><p>The pattern: automate everything that is repetitive knowledge work with an editable output. Do not automate anything that trades on trust or judgement.</p><h3>The team shape change</h3><p>Our staffing shape at the end of 2025 vs mid-2026 shows the biggest structural change of the year:</p><p><strong>Senior engineers: count stayed roughly flat.</strong> The role changed (see the <a href="/post/70-percent-ai-generated-codebase">70%-codebase pillar</a>) but the headcount is the same. Each senior is more productive; we absorbed the productivity into taking on more or larger engagements rather than shedding seniors.</p><p><strong>Mid-level engineers: count compressed by about a third.</strong> The typing work that used to fill a mid-level's day is now AI-generated. Mid-levels who transitioned to editorial / review roles stayed; mid-levels who were doing well but leaned on typing over judgement did not.</p><p><strong>Junior engineers: hiring stopped for pure typing roles.</strong> We still hire junior engineers — but into apprenticeship-style pairings with a senior, where the junior is learning judgement, not producing typing volume. It is a smaller pipeline and a slower ramp; the juniors we do bring in ramp to editorial productivity in about six months.</p><p><strong>Design engineers: new role, small count.</strong> The role we describe in the <a href="/post/design-engineering-ai-native-pipeline">design engineering pipeline post</a>. Two or three across the agency. Very high value.</p><p><strong>Product managers, designers: shape held.</strong> These roles were already judgement-heavy. AI helped them ship faster; the count did not change.</p><p><strong>Delivery leads: count grew slightly.</strong> Because we run more engagements per senior engineer, we need slightly more delivery-lead capacity. The role is more relational than before — less spec-writing, more client-facing.</p><p><strong>Internal ops: count dropped by half.</strong> Proposal ops, meeting ops, delivery ops — the roles that were coordinating the internal surface are largely automated. The people who did them either moved into delivery leadership or moved on.</p><p>The visible outcome: a leaner org that ships more per person. The invisible outcome: a much higher requirement on judgement across every remaining role.</p><h3>Estimation, the second time round</h3><p>Our estimation practice went through a shape change in 2026 that is worth naming.</p><p><strong>First shift (early 2026).</strong> Estimation got faster and looser. AI could produce a plausible estimate from a spec in minutes. We ran with the AI estimate. Some engagements went well, some went 20% over. The variance was higher than the classical estimation.</p><p><strong>Second shift (mid-2026).</strong> We tightened back. The AI estimate is now the first-cut, not the final. A senior engineer reviews it against the specific engagement risks — integrations, data quality, team familiarity. The estimate ships accurate to within 10% on most engagements.</p><p>The pattern that produces the tight estimate:</p><ul><li><strong>AI produces the first cut</strong> from the spec, using our historical delivery data as reference.</li><li><strong>Senior engineer overrides the AI's assumptions</strong> on the specific risks that are not in the historical data.</li><li><strong>We revisit weekly.</strong> If the estimate is drifting, the client knows within a week, not at the end of the engagement.</li></ul><p>Fixed-scope engagements became easier to sell because we can honestly commit to them. The client knows what they are getting; we know what we are delivering. This is the shape <a href="/post/70-percent-ai-generated-codebase">outcome-based pricing sits on top of</a>.</p><h3>Pricing, still moving</h3><p>Our pricing model was mostly time-and-materials through 2024, mostly fixed-scope through 2025, and moving toward outcome-based in 2026. The direction is set; the exact shape is still being worked out.</p><p><strong>Time-and-materials.</strong> Broke because our velocity was hard to describe. When one senior engineer with an AI stack ships in a day what previously took a week, the client is paying for the same outcome at a fraction of the hours. Clients notice, and either negotiate a discount or feel underpaid-for. The model does not survive this asymmetry.</p><p><strong>Fixed-scope.</strong> Our main model in 2026. We commit to a scope, we commit to a price, we deliver. The AI-native operating model is what makes this work — we can estimate accurately and we can absorb small overruns because our unit cost is lower. The client is paying for the outcome, not the hours.</p><p><strong>Outcome-based.</strong> Emerging on a subset of engagements. The client pays a base fee for the engagement and a bonus tied to a specific business metric. The pattern only works where the business metric is clean and attributable, which is a smaller set of engagements than the marketing suggests.</p><p>The shape we ended up at: fixed-scope for most work, outcome-based where the metric is clean, T&amp;M only for open-ended discovery work where the scope genuinely cannot be nailed down.</p><h3>Where we are still figuring things out</h3><p>An honest section, because the AI-native operating model is not fully solved and we do not want to pretend it is.</p><ul><li><strong>Junior engineer development.</strong> We stopped hiring for typing roles, but we have not fully worked out the apprenticeship model that gets a junior to editorial competence in six months. The pattern works when it works; it does not scale as cleanly as the classical junior-to-mid pipeline did.</li><li><strong>Estimation on genuinely novel work.</strong> Our historical data covers what we have done before. When the engagement is architecturally novel, the AI estimate is worse than the classical estimate. Senior judgement overrides it, but that means the estimation cost is still real.</li><li><strong>Client understanding of the model.</strong> Some clients still buy time. When we quote a fixed scope that reflects our AI-native velocity, the client sometimes reads it as a low quote and worries about corners cut. The conversation about why the price is what it is happens on every engagement.</li><li><strong>Internal ops that we have not automated yet.</strong> Contract review, some finance operations, some parts of the delivery pipeline still have manual work that has an AI-native shape we have not built yet. Prioritising them against client work is a recurring choice.</li></ul><h3>What this means if you are running an agency</h3><p>Three things transfer from our operating model to any agency (or internal delivery team) trying to become AI-native.</p><p><strong>Start with internal ops, not client-facing surfaces.</strong> The payback lives inside. The client-facing polish comes later — and comes for free — once the internal ops are running.</p><p><strong>Shift the team shape deliberately, not by attrition.</strong> Deciding what your team should look like at the end of the shift is a Phase 1 conversation. Waiting for the shape to emerge through hiring and departures takes years and produces regret.</p><p><strong>Fix the estimation and pricing together.</strong> They move together. A T&amp;M business model with AI-native delivery pace is a losing shape. A fixed-scope model with pre-AI estimation is a different losing shape. The internal discipline changes both at once.</p><p>The operating model shift is the real work. The tools are the easy part. The agencies that will still exist in 2028 have started the operating-model work in 2026. The ones that are still marketing "AI-native" without having done the internal shift will find that clients notice the gap by 2027.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-how-we-run-an-ai-native-agency.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Estimating AI-augmented engineering]]></title>
            <link>https://twistag.com/thinking/estimating-ai-augmented-engineering</link>
            <guid isPermaLink="false">https://twistag.com/thinking/estimating-ai-augmented-engineering</guid>
            <pubDate>Wed, 29 Jul 2026 14:32:33 GMT</pubDate>
            <description><![CDATA[Story points break at 10x velocity. Here's what replaces them, what stays hard, and how to sanity-check a partner's AI-augmented estimate.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Story points, t-shirt sizing, hours estimates — the classical engineering estimation vocabulary calibrated to a specific idea of how fast a senior engineer could ship. When the AI-augmented senior engineer ships five to ten times faster on the parts where AI does the typing, the calibration breaks. A three-point story is now half a day. A "large" t-shirt is a morning. The number stayed the same; the underlying unit shifted. Teams that did not notice keep planning as if a sprint holds twenty story points, when it actually holds sixty. This post is what breaks, what replaces it, what stays hard even with AI, and how to sanity-check a partner's AI-augmented estimate. It is a cluster child of our <a href="/post/how-we-run-an-ai-native-agency">AI-native agency operating model pillar</a>.</p><h3>Key takeaways</h3><ul><li><strong>Story points broke because they measured typing effort, which AI now does.</strong> The unit had to shift from typing hours to editorial hours. Every team's point calibration is now wrong until they recalibrate.</li><li><strong>The parts that stayed hard did not get faster.</strong> Design, architecture, integration, incident response, business-rules encoding — all the parts that always were the hard part of engineering. AI helps at the margins; they still take real time.</li><li><strong>The right unit is now a feature or a slice, not a point.</strong> Estimation in features + delivery pace = confidence interval. Points are two calibration steps removed from anything a stakeholder cares about.</li><li><strong>Weekly re-estimation replaces sprint-boundary estimation.</strong> The velocity is high enough that a two-week sprint estimate is stale by end of week one. The estimate has to live in the same working cycle as the code.</li><li><strong>Sanity-checking a partner's estimate has three specific tests.</strong> Their breakdown of what is AI-typing vs judgement time, their weekly re-estimation cadence, and their variance track record.</li></ul><h3>What broke about story points</h3><p>Story points were a shorthand. They compressed complexity, effort, and risk into one number that a team could estimate in aggregate. The compression worked when the underlying components were stable — a team that consistently shipped 20 points a sprint had a real signal.</p><p>Two things changed in 2026:</p><p><strong>The typing effort collapsed.</strong> A story that would previously have taken a senior engineer three days of typing now takes a day of editorial time. The point was calibrated to the three days; the shipped work took a day. Points inflated silently.</p><p><strong>The editorial effort became the actual constraint.</strong> The senior engineer's day is now bounded by how many AI-generated diffs they can meaningfully review, not by how many lines they can type. A point that reflects typing effort does not reflect the actual constraint.</p><p>The result: story points that used to have a fuzzy but useful relationship to elapsed time now have no consistent relationship. A three-point story might be half a day (all typing, AI does it) or three days (design decisions plus integration). Teams that keep points as their unit are not lying; they are using a broken instrument.</p><h3>What replaces points</h3><p>The unit we settled on: <strong>features shipped per week</strong> as the top-line measure, with <strong>hours of editorial time per feature</strong> as the diagnostic.</p><p><strong>Features shipped per week.</strong> A team ships some number of user-visible or system-user-visible features each week. This is what the stakeholder cares about, and it is what the team can commit to. In an AI-native team the number is typically 3-8 features per week for a five-person delivery unit. In a classical team the same unit shipped 1-2 per week.</p><p><strong>Hours of editorial time per feature.</strong> Under the features, the diagnostic that catches drift. If a feature normally takes 4-6 hours of senior editorial time and this one is taking 12, something is off — usually design ambiguity or an unclear acceptance criterion. The diagnostic is a signal, not a plan.</p><p>The two measures compose into an honest estimate for a stakeholder. "This scope is 15 features. Our team ships 5 per week. Three weeks." When the actual work involves 4 hours of editorial per feature on the routine ones and 20 hours on two tricky ones, the plan holds. When the tricky ones balloon to 40 hours, the plan updates weekly and the stakeholder knows in week one, not week three.</p><p>Story points are missing from this vocabulary deliberately. If a team wants to use points internally as a shorthand for editorial hours, that is fine; the stakeholder-facing number is features and elapsed time.</p><h3>What stayed hard</h3><p>The parts of engineering that always were the hard parts did not get faster in proportion to the typing collapse.</p><p><strong>Design.</strong> Deciding what to build. What the entities are, what the flows are, what the user's mental model needs to be. AI helps produce artefacts (specs, diagrams, prototypes) but the decisions are human. Design time did not compress in a 10x way; it compressed maybe 1.5-2x.</p><p><strong>Architecture.</strong> Service boundaries, data model, integration points. Same shape as design — the AI helps produce, the human decides. Architecture time compressed maybe 2x.</p><p><strong>Integration.</strong> Getting the new work to talk to the existing systems. Legacy APIs, upstream data quality, undocumented conventions. AI helps with the typing but the discovery is human. Integration time compressed maybe 2-3x.</p><p><strong>Incident response.</strong> When something breaks in production, judgement about severity and remediation is human work. AI helps with hypothesis generation and diagnostics; it does not shorten the debug loop as much as the marketing suggests.</p><p><strong>Business rules encoding.</strong> Turning the enterprise's actual rules into code. The rules live in someone's head or in a fragmented spec. Getting them out and coded correctly is human interviewing plus editorial work.</p><p>Add up the hard parts of a real engagement and they still take most of the elapsed time. AI-augmented engineering is not "everything is 10x faster." It is "the typing is 10x faster, and the hard parts got a little faster." A partner that estimates as if everything is 10x is going to overrun on the parts that stayed hard.</p><h3>The weekly re-estimation cadence</h3><p>Classical estimation was a sprint-boundary event. The team estimated at planning, ran the sprint, retrospected at the end. Two-week feedback loops.</p><p>At AI-native velocity, a two-week loop is stale. In two weeks the team ships 10-16 features rather than 2-3, so the estimate at planning is far from the reality at the retrospective. The cadence has to move to weekly:</p><p><strong>Monday.</strong> Look at what remains in scope. Look at what shipped last week. Update the estimate.</p><p><strong>Wednesday.</strong> Mid-week check on the current features. Anything that is running long? Reforecast.</p><p><strong>Friday.</strong> What actually shipped. What did not. Why. Feed the answer into Monday's estimate.</p><p>This is a lightweight cadence — the meetings are 20 minutes each — but the discipline is real. A team that runs weekly re-estimation catches drift in week one and can renegotiate scope in week two. A team that runs sprint-boundary estimation catches drift at the end and has to scramble to save the sprint.</p><h3>Sanity-checking a partner's estimate</h3><p>The enterprise buying delivery services from a partner in 2026 will get AI-adjusted estimates from every serious partner. Three tests separate the honest partners from the marketing.</p><p><strong>Test 1: The AI-typing vs judgement breakdown.</strong> Ask the partner to break the estimate into "what AI writes" and "what humans decide." An honest partner has the breakdown ready. The AI-writing portion is typically 30-50% of the elapsed time on greenfield work, less on brownfield. A partner whose estimate says "AI writes 80% of it" is either lying or does not understand their own delivery.</p><p><strong>Test 2: The weekly re-estimation cadence.</strong> Ask how they will keep the estimate current. An honest partner has a weekly cadence and a specific meeting. A partner who says "we will let you know if it slips" has a sprint-boundary cadence and will notice slippage too late.</p><p><strong>Test 3: The variance track record.</strong> Ask what percentage of their last ten engagements delivered within 10% of the estimate. An honest partner has the number and it is 70-90%. A partner who says "all of them" is lying. A partner who says "we do not track" is not disciplined enough to trust with a fixed scope.</p><p>The three tests are cheap to run in a discovery call and expensive to fake in real delivery. Any partner that fails one of the three has an estimation problem that will show up in the engagement.</p><h3>What this shifts for the buyer</h3><p>The enterprise's own estimation practice needs an update too. Two moves:</p><p><strong>Recalibrate the internal team's velocity numbers.</strong> If the internal engineering team is using AI coding tools, their story-point calibration is now off. A quarterly reset is the minimum. Weekly is better.</p><p><strong>Change what the roadmap conversation measures.</strong> The classical roadmap conversation was "how many story points fit in a quarter." The new conversation is "how many features can we ship in a quarter, at what confidence." The unit shift is small; the conversation quality it produces is large.</p><p>The estimation shift is one of the biggest visible signs of an AI-native operating model. A team still counting points at sprint boundaries is not AI-native, no matter what the marketing says. A team that measures features per week, re-estimates weekly, and tracks variance is doing the work.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-estimating-ai-augmented-engineering.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Capability transfer as the default engagement shape]]></title>
            <link>https://twistag.com/thinking/capability-transfer-default-engagement</link>
            <guid isPermaLink="false">https://twistag.com/thinking/capability-transfer-default-engagement</guid>
            <pubDate>Wed, 29 Jul 2026 14:32:18 GMT</pubDate>
            <description><![CDATA[Project delivery ends when the code ships. Capability transfer ends when the team can run it. Here's why we default to the second.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Classical project delivery ends when the code ships. The partner walks away with a running system and a set of runbooks; the client team either owns the code or does not, but the partner's obligation is done. Capability transfer ends when the client's team can run the system without the partner. Every engagement we shape at Twistag defaults to the second, and the shift has changed three specific things: the contract shape, the staffing shape, and the tooling shape. This post is what changes when capability transfer is the default engagement shape, not the exception. It is a cluster child of our <a href="/post/how-we-run-an-ai-native-agency">AI-native agency operating model pillar</a> and it operationalises the exit-signal framing from our earlier <a href="/post/capability-transfer-ai-delivery">capability-transfer child post</a>.</p><h3>Key takeaways</h3><ul><li><strong>The default matters because most engagements do not name a transfer plan on day one.</strong> Making it the default forces the Phase 1 conversation about who runs the system after we leave.</li><li><strong>Contract shape shifts from deliverables to outcomes plus transfer.</strong> The classical statement of work listed features. The transfer-default SOW lists features and names the transfer criteria — the specific things the client team must be able to do independently for the engagement to close.</li><li><strong>Staffing shifts from a partner-heavy team to a joint team.</strong> In year one the partner has more seats. By month four to six the ratio inverts. If the ratio does not invert, the transfer is not happening.</li><li><strong>Tooling shifts from partner-preferred to client-primary.</strong> Every tool the client team will use in production has to be a tool the client owns. Partner-hosted convenience becomes partner-locked-in convenience.</li><li><strong>The default sends a signal even before the engagement starts.</strong> Sales conversations that start with "who owns this in a year" filter for clients who want a partner, not a permanent vendor. Which is the client relationship we want.</li></ul><h3>Why we default to transfer</h3><p>The alternative — permanent partnership — is legitimate for some engagements. But defaulting to it hides the harder conversation. If the client and partner never explicitly agree on whether the client will run the system, both sides make assumptions. The client assumes the partner will keep running it because "we paid you to build it." The partner assumes the client will pick it up because "we finished." Six months in, an incident happens and both sides discover the assumption gap in the worst possible way.</p><p>Defaulting to transfer forces the Phase 1 conversation. The client says explicitly whether they want to run it or whether they want a managed relationship. The partner scopes accordingly. Both sides know the shape of the engagement's end state before the engagement starts.</p><p>For a services firm the default has a second consequence: it filters clients. Clients who want a permanent vendor and are not up for a Phase 1 conversation about transfer often self-select out. This is not a loss. The clients who stay through that conversation are the ones we can serve well.</p><h3>What changes in the contract</h3><p>The classical statement of work has a features list, an acceptance criteria list, a timeline, and a payment schedule. The transfer-default SOW has all of that plus a specific transfer criteria section.</p><p>Transfer criteria are behavioural, not documentary. Not "the partner will deliver runbooks" — every engagement delivers runbooks. Instead:</p><ul><li><strong>The client team shipped three production changes without partner involvement.</strong> Named signal from our <a href="/post/capability-transfer-ai-delivery">capability-transfer-in-AI-delivery post</a>.</li><li><strong>The client team handled one real production incident without paging the partner.</strong> The incident is real because it happens on its own schedule; the criterion is that when it happens, the client team owns it.</li><li><strong>The client team extended the eval pipeline for a new failure mode.</strong> For AI-native engagements specifically. The eval is where the discipline lives; owning eval extension is the sign the team can operate the system.</li></ul><p>The transfer criteria are worded as observable events. The engagement closes when they are met, not when the calendar says so. This has an interesting side effect: it aligns the partner's incentives. The partner wants to close the engagement — the sooner the criteria are met, the sooner the partner is free to take on new work — so the partner has skin in making the transfer happen.</p><p>Payment schedules follow the same shape. A portion of the fee is tied to the transfer criteria. Not enough to distort the engagement, but enough that the partner cannot ignore the transfer surface.</p><h3>What changes in the staffing</h3><p>In a classical partner-heavy engagement, the delivery team is mostly partner staff with a couple of client engineers observing. In a transfer-default engagement, the ratio shifts over time.</p><p><strong>Month 1-2.</strong> Partner team leads. Client team has one or two people paired closely with senior partner engineers, doing real work under supervision. The partner takes on most of the load because they can move fastest; the client presence is real, not observational.</p><p><strong>Month 3-4.</strong> The ratio starts to shift. Client engineers own specific components. Partner engineers review their work but do not write it. New capability additions increasingly get paired between a client engineer and a partner engineer, with the client engineer doing most of the driving.</p><p><strong>Month 5-6.</strong> Client team leads. Partner engineers are on-call for specific hard problems (architecture decisions, hard debugging) but not on the day-to-day. The exit signal from the <a href="/post/capability-transfer-ai-delivery">capability-transfer post</a> — three production changes shipped without partner involvement — usually arrives here.</p><p><strong>Month 7+.</strong> Partner engagement is asynchronous. Advisory, escalation, occasional deep dives. Not day-to-day delivery.</p><p>The ratio shift has to happen for the transfer to be real. If the partner is still driving day-to-day at month six, either the transfer criteria were unrealistic (in which case renegotiate) or the client team is not being given room to lead (in which case fix that). Missing the shift is the single most common way a transfer engagement quietly turns into a permanent partnership.</p><h3>What changes in the tooling</h3><p>The client's tools are the primary tools from Phase 2, not Phase 5.</p><p><strong>Repositories are on the client's Git host.</strong> Not a partner org that gets transferred later. The client's engineers have the same access from day one that they will have in perpetuity.</p><p><strong>Environments are in the client's cloud.</strong> Development, staging, production. The partner deploys into the client's cloud from the beginning, so the operational surface is familiar to the client team as soon as they start operating it.</p><p><strong>Observability is in the client's stack.</strong> If the client already has Datadog, we do not introduce Grafana Cloud "just for the engagement." Everything the client team will run against post-transfer is what they run against during the engagement.</p><p><strong>AI stack is on the client's plane.</strong> For AI-native engagements this is subtle. The partner's Claude Code subscriptions, prompt caches, and evaluation datasets need to transfer or be re-created in the client's plane. If they do not, the client team has to rebuild the AI-native discipline from scratch after the partner leaves.</p><p>The convenience cost is real. Partner-preferred tooling is faster for the partner. Client-primary tooling is slower for the partner in months one and two. But it is what makes the transfer real, and skipping it is where partner lock-in creeps in.</p><h3>The signal to the market</h3><p>Defaulting to capability transfer is a positioning choice as much as a delivery discipline. It signals to a prospective client:</p><ul><li>We do not want to be your permanent vendor for this. We want to build the capability and hand it to you.</li><li>We are willing to write the transfer criteria into the contract. This is a commitment, not a promise.</li><li>We are staffed to make the transfer real. The seniors on the engagement are seniors, not a delivery manager and a bench of juniors.</li></ul><p>The signal filters. Prospects who read it and want to continue the conversation are the prospects who value engineering depth and want a partnership shape that respects their internal team. Prospects who read it and go quiet are prospects who wanted a permanent vendor — a legitimate need, but one that is served better by a different kind of firm.</p><p>For a services firm in 2026, filtering aggressively for the clients we can serve well is a survival strategy. The market is more crowded than it has been. The differentiation lives in the shape of the engagement, not the marketing.</p><h3>What this shifts for the enterprise</h3><p>The enterprise evaluating a partner in 2026 can use the capability-transfer default as a qualifier. Two questions surface it fast:</p><ul><li><strong>Will you write transfer criteria into the SOW, worded as observable events?</strong> A partner that will not is not committed to transfer. The words in the marketing are cheaper than the words in the contract.</li><li><strong>How does your staffing shape change from month one to month six?</strong> A partner whose team stays the same shape is not transferring anything. A partner whose team ratio inverts is doing the work.</li></ul><p>The enterprises that get the capability-transfer default right end up with working systems and teams that can run them. The enterprises that skip the conversation end up with working systems and a vendor relationship that is harder to leave than it should be.</p><p>Which is the outcome the whole default was designed to avoid.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-capability-transfer-default-engagement.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Value pricing for AI-native delivery: what replaces time-and-materials]]></title>
            <link>https://twistag.com/thinking/value-pricing-ai-native-delivery</link>
            <guid isPermaLink="false">https://twistag.com/thinking/value-pricing-ai-native-delivery</guid>
            <pubDate>Wed, 29 Jul 2026 14:32:03 GMT</pubDate>
            <description><![CDATA[Time-and-materials breaks when velocity is 10x. Here's why fixed-scope and outcome-based pricing are replacing it — and how to structure each.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>Time-and-materials pricing survived four decades in professional services because the customer was buying labour. When the labour cost per unit of output dropped 5-10x for parts of an engagement in 2026, T&amp;M stopped representing what the customer was buying and started representing what the provider was spending. Two very different things. The pricing conversation moved. This post is what the new conversation looks like — the mechanics of fixed-scope pricing when scope is credible, the mechanics of outcome-based pricing where the metric is clean, and the shrinking role T&amp;M still plays for genuinely open-ended discovery work. It is a cluster child of our <a href="/post/how-we-run-an-ai-native-agency">AI-native agency operating model pillar</a>.</p><h3>Key takeaways</h3><ul><li><strong>T&amp;M broke because it exposed the wrong number.</strong> When one senior engineer with an AI stack ships in a day what previously took a week, T&amp;M invoices the day at the wrong hourly rate for the wrong deliverable.</li><li><strong>Fixed-scope became the default because the </strong><a href="/post/estimating-ai-augmented-engineering"><strong>estimation discipline caught up</strong></a><strong>.</strong> The partner can commit to a scope and price honestly. The client gets an outcome, not an invoice.</li><li><strong>Outcome-based pricing works where the metric is clean and attributable.</strong> Where it is, it is the fastest-growing pricing model in 2026. Where it is not, forcing it produces theatre.</li><li><strong>T&amp;M still fits open-ended discovery.</strong> Genuinely exploratory work where the scope legitimately cannot be nailed down at the start. Small share of engagements. Legitimate share.</li><li><strong>The pricing conversation is a signal to the buyer about the partner.</strong> A partner that will only work T&amp;M in 2026 is either hiding a variance problem or has not done the AI-native operating-model work. Neither is a good signal.</li></ul><h3>Why T&amp;M broke</h3><p>The classical services deal was: the customer buys a certain number of engineer-hours at a certain rate. The provider bills the hours. Both sides trust that the hours correspond to the output.</p><p>The correspondence held because the labour was the constraint. A senior engineer wrote code at a specific pace. Hours in, code out. The rate reflected the market for the engineer's skill. The unit of output — a working feature, a shipped integration — had a stable relationship to hours because the engineer's typing was most of the work.</p><p>The AI-augmentation shift broke the correspondence. The senior engineer is now producing 3-5x the output per hour on the parts where AI does the typing, and roughly the same on the parts where judgement dominates. The hour is no longer a stable unit relative to the output. Two things happen:</p><p><strong>The customer notices the hours dropping.</strong> Not immediately, but eventually. When a T&amp;M engagement that used to burn 40 hours a week starts burning 25 hours for more output, the customer asks why the invoice is smaller. Then they ask why the hourly rate was set for the pre-AI world. The rate conversation ends badly for the provider.</p><p><strong>The provider notices the risk asymmetry.</strong> The provider is now delivering the same outcome at a fraction of the hours. Under T&amp;M, the provider's revenue falls. But the provider still holds the risk of overrun, integration surprises, and the parts of the work AI does not compress. The provider is worse off than they were pre-AI. Providers respond by raising the hourly rate. Customers see the higher rate and negotiate down. Neither side is happy.</p><p>The T&amp;M model does not survive both sides being unhappy. It survives when both sides trust the number. When the AI-augmentation shift breaks that trust, T&amp;M breaks.</p><h3>Fixed-scope, done right</h3><p>Our main pricing model in 2026 is fixed-scope. The mechanics that make it work:</p><p><strong>Honest estimation.</strong> From the <a href="/post/estimating-ai-augmented-engineering">estimation post</a>: AI produces the first cut, senior engineer overrides the assumptions, the team commits to a scope and a number. Variance track record 70-90% within 10% of estimate.</p><p><strong>Scoped exclusions in writing.</strong> Everything that is not in scope, named. Integrations to specific systems, feature variants, data cleanup work. If something is outside the fixed-scope wall, it is a change order.</p><p><strong>Change order economics.</strong> Change orders are priced the same way as the initial scope — AI-first estimate, senior review, commit. Not billed at a T&amp;M multiplier. This is what stops the "fixed-scope plus lots of change orders" pattern that used to make fixed-scope a fiction.</p><p><strong>Milestone payments.</strong> Payment schedule follows delivery milestones, not calendar dates. The client pays for output, not elapsed time. A milestone that lands early gets paid early; a milestone that lands late gets paid late.</p><p><strong>Capability transfer criteria.</strong> From the <a href="/post/capability-transfer-default-engagement">capability-transfer-default post</a>. A portion of the final payment is tied to the transfer criteria being met. The engagement is not closed until the client team can run the system.</p><p>The result is a contract that both sides can trust. The client knows the maximum they will spend and the outcome they will get. The provider knows the revenue they will earn and the exit criteria they need to meet. Neither is surprised at the end.</p><h3>Outcome-based pricing, done carefully</h3><p>Outcome-based pricing ties a portion of the fee to a specific business metric. When the metric is clean, this is the fastest-growing pricing model in 2026 — clients like it because they pay for value, providers like it because their upside is uncapped when they deliver well.</p><p>The metric criteria that make it work:</p><p><strong>Clean.</strong> The metric is a specific business number — conversion rate, revenue per user, cost per transaction. Not "customer satisfaction," not "productivity." Numbers that already exist in the client's dashboard, not new numbers invented for the contract.</p><p><strong>Attributable.</strong> The change the partner ships is causally responsible for moving the metric. If the metric moves for reasons the partner did not cause (a marketing campaign, a competitor's price change, seasonal effect), the pricing model is not tracking the value of the work.</p><p><strong>Baselined.</strong> The pre-engagement baseline is measured for enough time to be honest. Metrics have natural variance; a baseline of two weeks does not represent the underlying rate.</p><p><strong>Timeboxed.</strong> The measurement window is finite. Six months, twelve months. Not "in perpetuity" — that is a licence deal, not an outcome-based delivery deal.</p><p>When those four conditions hold, outcome-based works. The mechanics: a base fee that covers the delivery cost with a modest margin, plus a bonus tied to metric movement. The bonus is a percentage of the metric-driven value, structured so the client pays more when they benefit more.</p><p>When any of the four conditions fail, outcome-based produces theatre. The pricing model exists, but neither side actually believes the number, so the base fee grows to cover the risk and the bonus becomes ceremonial. Better to use fixed-scope than to fake outcome-based.</p><h3>Where T&amp;M still fits</h3><p>Genuinely open-ended discovery work. The scope cannot be nailed down at the start because the point of the engagement is to figure out what the scope should be. AI advisory, technology-strategy work, deep-dive research engagements.</p><p>T&amp;M for this work is honest. The client pays for the discovery hours, gets a recommendation, decides what to do next. Neither side is pretending to know a scope that they do not. The hourly rate reflects the seniority of the person doing the discovery — usually a principal-level engineer or advisor, priced accordingly.</p><p>Small share of our engagements. Legitimate share. The pattern that fails is using T&amp;M for delivery work where a fixed scope is possible but the partner does not have the estimation discipline to commit.</p><h3>The pricing conversation as a signal</h3><p>An enterprise evaluating a partner in 2026 can read the pricing conversation as a signal about the partner's operational maturity.</p><p><strong>A partner that offers only T&amp;M</strong> for delivery work has either not adjusted to the AI-augmentation shift or is hiding a variance problem. Neither is a good signal.</p><p><strong>A partner that offers fixed-scope with a clean change-order process</strong> has the estimation discipline. Their variance track record is worth asking about.</p><p><strong>A partner that offers outcome-based on specific engagements</strong> where the metric is clean is genuinely aligning their incentive with the client's. Their willingness to put revenue at risk is a signal of confidence.</p><p><strong>A partner that offers all three</strong> — T&amp;M for discovery, fixed-scope for delivery, outcome-based where the metric is clean — has the operational maturity to match the shape of the pricing to the shape of the work. This is what mature AI-native services firms look like in 2026.</p><p>The enterprise that reads the pricing conversation this way ends up with partners who match the actual shape of the engagement. The enterprise that treats pricing as a discount negotiation gets partners who are optimising for hours billed, not for value delivered. The former is where the AI-native shift is going. The latter is where the pre-AI professional services market was, and where it will not stay.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-value-pricing-ai-native-delivery.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[How generative AI connects enterprise data systems]]></title>
            <link>https://twistag.com/thinking/generative-ai-connects-the-dots</link>
            <guid isPermaLink="false">https://twistag.com/thinking/generative-ai-connects-the-dots</guid>
            <pubDate>Wed, 08 Jul 2026 13:50:46 GMT</pubDate>
            <description><![CDATA[Generative AI connects enterprise data through RAG, tool calling, and MCP. With only 29% of enterprise apps integrated, the connection layer decides success.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p>
<p>Generative AI connects enterprise data systems through three mechanisms: retrieval that assembles context at query time, tool calls that read and write live systems, and protocols such as MCP that standardise both. The connection layer, not the model, is where most programmes stall: on average, only <a href="https://www.salesforce.com/news/stories/connectivity-report-announcement-2025/">29% of enterprise applications are integrated</a> (MuleSoft, 2025). This post explains how each mechanism works and which one to fix first.</p>
<p><strong>Key takeaways</strong></p>
<ul>
<li>MIT's NANDA initiative found that <a href="https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/">95% of enterprise generative AI pilots deliver no measurable P&amp;L impact</a>, and the study blames integration that fails to adapt to workflows, not model quality.</li>
<li>The average enterprise runs 897 applications, and only 29% of them are connected (MuleSoft Connectivity Benchmark, 2025). An AI system can only reason over the slice it can reach.</li>
<li><a href="https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025">Gartner predicted at least 30% of generative AI projects would be abandoned</a> after proof of concept by the end of 2025. Poor data quality is the first cause it lists.</li>
<li>Three mechanisms do the connecting: retrieval-augmented generation for knowledge, tool calling for actions, and the Model Context Protocol for standardised access, replacing one custom connector per system pair with one open standard.</li>
<li>Across the retrieval systems we shipped in 2025 and 2026, data preparation consumed roughly two thirds of the engineering effort. The model was never the bottleneck.</li>
</ul>
<h2>Why do AI initiatives fail without connected data?</h2>
<p>They fail because a model can only reason over what it can reach. MIT's NANDA initiative analysed 300 public deployments and found <a href="https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/">95% of enterprise GenAI pilots produce no measurable P&amp;L return</a>; the authors point at flawed integration, tools that neither learn from nor adapt to the workflows around them, rather than model capability. <a href="https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025">Gartner reached a similar verdict from the cost side</a>, predicting that at least 30% of GenAI projects would be abandoned after proof of concept, with poor data quality first on its list of causes. A pilot that answers well from a clean demo corpus degrades on contact with the real estate: silos, stale copies, undocumented permissions.</p>
<h2>What are the three mechanisms that connect AI to data?</h2>
<p>Every production system we have built or audited uses some combination of three: context assembly, retrieval, and tool calling. They solve different problems, and they fail in different ways.</p>
<h3>Context assembly</h3>
<p>A language model is stateless. Everything it appears to know about your business in a given exchange was placed into its context window: the conversation so far, retrieved documents, schemas, instructions. Context windows are large but finite. They hold hundreds of pages, not the petabytes an enterprise stores, so the engineering problem is selection: deciding, per request, which fragments of which systems earn a place. Poor selection is invisible in demos and expensive in production, because the model answers confidently from whatever it was given.</p>
<h3>Retrieval-augmented generation</h3>
<p>RAG retrieves relevant records from your stores at query time, through semantic search over embeddings, keyword indexes, or both, and injects them into the context before the model answers. It is the standard pattern for grounding output in current, private data without retraining a model. Its ceiling is set upstream: if the corpus contains three conflicting versions of a policy document, retrieval will faithfully surface the conflict. Deduplication, ownership, and access control are retrieval-quality problems before they are governance problems.</p>
<h3>Tool calling</h3>
<p>Tool calling lets the model invoke functions against live systems: query a database, create a ticket, post a journal entry. This is the step from answering to acting, and it is the foundation of the <a href="/ai-and-agents">AI and agent systems we build</a>. It is also where connection quality becomes a safety property. A wrong retrieval produces a bad answer; a wrong write produces a bad ledger. Schema validation, retries, and a human-review path for unexpected results are part of the integration, not extras.</p>
<h2>When does MCP replace custom integrations?</h2>
<p>When the number of system-to-AI connections grows past a handful. Before protocols, every data source needed its own connector for every AI application, an N-by-M problem that Anthropic's <a href="https://www.anthropic.com/news/model-context-protocol">Model Context Protocol</a>, open-sourced in November 2024, collapses into N servers speaking one standard. Anthropic's framing of the underlying problem is blunt: even frontier models are "constrained by their isolation from data — trapped behind information silos and legacy systems." MCP does not fix bad data; it standardises access to whatever data you have. The economics of adopting it, and where a plain internal API is still the better call, are covered in our post on <a href="/post/mcp-for-enterprise-when-it-pays-off">when MCP pays off for enterprise integration</a>.</p>
<h2>What makes data agent-readable?</h2>
<p>An AI system can use a data source when four conditions hold: the schema is documented, identifiers are consistent across systems, permissions are machine-checkable, and access runs through an API rather than a screen. Most estates fail at least two of the four, which is why <a href="https://www.salesforce.com/news/stories/connectivity-report-announcement-2025/">90% of IT leaders report that data silos create business problems</a> (MuleSoft, 2025). For older platforms, the API condition is usually the hard one. The architecture for exposing them without a rewrite is its own topic, covered in <a href="/post/modernising-legacy-systems-for-ai-agent-access">modernising legacy systems for AI agent access</a>.</p>
<h3>What this looks like in practice</h3>
<p>A document-processing agent we run for a European manufacturer spent its entire first engineering month on identifier reconciliation between the ERP and the document store. None of that month went to prompts or model selection. Once supplier IDs resolved consistently, retrieval accuracy stopped being a debate and the agent went to production the following quarter. That sequencing, data first, is the default in our <a href="/data-cloud">data and cloud engineering</a> work because reversing it means rebuilding.</p>
<h2>Is a data warehouse the same as connected data?</h2>
<p>No. Warehouses and lakes copy data on a schedule for analytics; the connection layer AI needs is live, operational, and permission-aware. A nightly extract answers "what were yesterday's orders" but cannot support an agent that checks stock before confirming today's order, and it usually strips the row-level permissions that decide what a given user's assistant is allowed to see. The two investments are complementary rather than interchangeable: the warehouse serves reporting, while retrieval and tool calling need governed paths into the systems of record themselves. Enterprises that spent the last decade centralising copies still have integration work ahead, which is exactly what the 29% application-connectivity figure measures.</p>
<h2>How much of the work is data work?</h2>
<p>Most of it. MuleSoft's benchmark puts <a href="https://www.salesforce.com/news/stories/connectivity-report-announcement-2025/">39% of IT team time into designing, building, and testing custom integrations</a> before any AI enters the picture. Our own ratio across retrieval projects shipped in 2025 and 2026 is close to two thirds data work: deduplicating corpora, mapping permissions, reconciling identifiers, setting freshness rules. Teams budget for model work because it is novel; the estimate that needs the padding is the plumbing.</p>
<h2>Where should an enterprise start?</h2>
<p>Start narrow and connect only what the first workflow needs.</p>
<ol>
<li>Pick one workflow and inventory its systems. A quoting assistant might touch the CRM, the price book, and a contracts folder. Three connections is a project; 897 applications is a decade.</li>
<li>Make the retrieval corpus trustworthy. Deduplicate, assign ownership, and map who may see what before embedding anything, because retrieval inherits every upstream defect.</li>
<li>Add one governed tool call. Give the system a single write action with validation and a review path, and measure its failure rate for a month before adding the second.</li>
<li>Standardise access once patterns repeat. When the third team requests the same connection, that is the signal to move it behind a shared protocol server rather than copy the integration.</li>
</ol>
<h2>What separates the teams that ship?</h2>
<p>The pattern across our 2025 and 2026 deliveries is consistent: teams that spent their first quarter on data plumbing shipped agents that survived production, and teams that spent it on prompts rebuilt in the second quarter. The next protocol cycle, with agent-to-agent communication already following the MCP playbook, will reward the same preparation, because every new layer assumes the one below it is connected.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-generative-ai-connects-the-dots.jpg" length="0" type="image/jpg"/>
        </item>
        <item>
            <title><![CDATA[Pulse: building autonomous teams with less process and more impact]]></title>
            <link>https://twistag.com/thinking/building-autonomous-teams-with-less-process-and-more-impact</link>
            <guid isPermaLink="false">https://twistag.com/thinking/building-autonomous-teams-with-less-process-and-more-impact</guid>
            <pubDate>Wed, 08 Jul 2026 13:50:46 GMT</pubDate>
            <description><![CDATA[Pulse is our AI delivery methodology: five agents, an eight-check review layer, and one dashboard behind the 90% of code we ship AI-generated by design.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p>
<p>Ninety percent of the code we ship at Twistag is spec-driven and AI-generated by design. The other ten percent is where our engineers spend their judgment. Pulse is the system that makes that ratio safe: five delivery agents that move work from brief to production, an eight-check agentic review layer that gates every pull request, and one dashboard that replaced five disconnected tools. Smaller teams shipping more, with less process and more impact.</p>
<p><strong>Key takeaways</strong></p>
<ul>
<li>Ninety percent of the code Twistag ships is spec-driven and AI-generated; senior engineers spend their judgment on the ten percent that needs it.</li>
<li>The 2024 DORA report found a 25% increase in AI adoption came with an estimated 7.2% drop in delivery stability. Review has to scale with generation.</li>
<li>Every pull request that touches a client environment passes eight sequential checks; a failure halts the pipeline, and every override lands in an audit trail.</li>
<li>One dashboard with seven categories (quality, security, cost, activity, people, adoption, ops) replaced Snyk, Sonar, Codacy, LinearB, and the tabs between them.</li>
<li>A five-person squad running Pulse delivers what traditionally requires eight to ten engineers.</li>
</ul>
<h2>Why did code review break at AI volume?</h2>
<p>Shipping is fast now. Controlled experiments put the speed-up in numbers: developers working with an AI pair programmer completed tasks <a href="https://arxiv.org/abs/2302.06590">55.8% faster than a control group</a> (Peng et al., 2023). Things that used to take a sprint take a couple of days. The part nobody wants to talk about is what that volume does to the people reviewing it.</p>
<p>When most of what an engineer is reading was not written by another human, the reading itself changes. Eyes skim. Attention drifts toward formatting because the formatting always looks tidy. AI does not write bad code most of the time. It writes <em>plausible</em> code. Code that passes the tests, looks fine, and quietly reintroduces the auth pattern the team killed six months ago.</p>
<blockquote>
<p><strong>"If your AI usage went up 10x and your review process didn't, you don't have a faster team. You have a slower incident waiting to happen."</strong>
— Fred Sarmento, founder, Twistag</p>
</blockquote>
<p>The industry data backs the instinct. The <a href="https://dora.dev/research/2024/dora-report/">2024 DORA report</a> found that a 25% increase in AI adoption was associated with an estimated 7.2% drop in delivery stability and a 1.5% dip in throughput. Individual developers got faster and happier; delivery got shakier. Generation scaled. The systems around it did not.</p>
<p>We saw that incident coming early and decided not to wait for it to cost a client. For years our stack was Snyk, Sonar, Codacy and LinearB. Each one solved a slice. None of them were built for a world where most of the code is not written by a person. And on top of that, none of them talked to each other.</p>
<p>This is where João came in.</p>
<blockquote>
<p><strong>"Five tools. None of them talked to each other. One morning we said enough."</strong>
— João Belo, VP of Engineering, Twistag</p>
</blockquote>
<p>What we built has three pillars. The first delivers the work. The second reviews the work. The third tells us whether either of the first two is doing its job. Together they are Pulse.</p>
<h2>What is Pulse, and what isn't it?</h2>
<p>Pulse is a proprietary AI delivery methodology built around senior engineers. It is not an autopilot, a code generator, or a wrapper around someone else's tooling. It is the operating layer Twistag uses to deliver client work, designed for the reality that the volume of generated code now outpaces the human capacity to read it line by line. It is also the delivery layer of <a href="/post/how-we-run-an-ai-native-agency">how we run an AI-native agency</a> more broadly.</p>
<p>Three pillars carry the methodology:</p>
<table>
<thead>
<tr>
<th>Pillar</th>
<th>What it does</th>
<th>Who owns the final call</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Five delivery agents</strong></td>
<td>Convert intent into shipped software across spec, architecture, code, quality and deployment</td>
<td>Senior engineer</td>
</tr>
<tr>
<td><strong>Eight-check agentic review layer</strong></td>
<td>Gate every pull request against the patterns tired reviewers miss</td>
<td>Senior engineer (after the layer clears)</td>
</tr>
<tr>
<td><strong>Pulse Dashboard</strong></td>
<td>Surface quality, security, cost, activity, people, adoption and ops in one screen</td>
<td>Engineering leadership</td>
</tr>
</tbody>
</table>
<p>The agents handle volume. The review layer enforces consistency at that volume. The dashboard tells leadership where judgment is most needed. The engineer still owns every call — approves, overrides, pushes back. What Pulse changes is what review looks like at this scale.</p>
<h2>What do the five delivery agents do?</h2>
<p>The five agents map to the phases of the delivery lifecycle. Each one removes a specific kind of friction that slows engineers down.</p>
<table>
<thead>
<tr>
<th>Agent</th>
<th>Phase</th>
<th>What it removes</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Spec</strong></td>
<td>Discovery → Specification</td>
<td>Ambiguous requirements, missing acceptance criteria, undocumented edge cases</td>
</tr>
<tr>
<td><strong>Arch</strong></td>
<td>Architecture</td>
<td>Late-binding architecture debates, undocumented trade-offs, scale assumptions made in PR comments</td>
</tr>
<tr>
<td><strong>Code</strong></td>
<td>Implementation</td>
<td>Boilerplate, scaffolding, repetitive convention enforcement</td>
</tr>
<tr>
<td><strong>Guard</strong></td>
<td>Quality assurance</td>
<td>Drift between what was specified and what was built, regressions, performance surprises</td>
</tr>
<tr>
<td><strong>Ship</strong></td>
<td>Deployment</td>
<td>Manual rollback planning, environment drift, post-deploy monitoring set up after the incident</td>
</tr>
</tbody>
</table>
<h3>One agent per phase, brief to production</h3>
<p><strong>Spec</strong> turns briefs, calls and stakeholder conversations into structured specifications with acceptance criteria and edge cases written down before a single line of code is written. Engineers start sprints with clarity, not questions.</p>
<p><strong>Arch</strong> evaluates that specification against the project's existing architecture, technical constraints, and scale requirements. It proposes implementation paths, flags trade-offs, and documents decisions. The architecture conversation happens before the pull request, not during code review.</p>
<p><strong>Code</strong> pairs with engineers during development. It generates implementation scaffolding, writes tests alongside features, and enforces project conventions across the codebase so senior engineers move at higher velocity on the parts that actually need them.</p>
<p><strong>Guard</strong> runs continuous quality analysis across code, tests, and deployment configuration. It catches regressions, security issues, and performance degradation before they reach staging. Every merge request meets the same standard regardless of who wrote it.</p>
<p><strong>Ship</strong> orchestrates deployment pipelines, environment configuration, and release sequencing. It handles rollback planning, feature flags, and post-deployment monitoring setup. What used to take a dedicated DevOps cycle now runs as part of every sprint.</p>
<h3>What the agents leave to humans</h3>
<p>The agents do not replace engineers. They remove the work that slows engineers down so the senior people Twistag hires spend their time on architecture decisions, complex problem-solving, and client collaboration, not on boilerplate, documentation, or deployment checklists. The same agents are what let us take a concept to <a href="/post/prototype-in-a-day">a working prototype in a day</a> when a client needs to see the idea before funding the build.</p>
<p>That is the delivery side. Now the part that broke at AI volume: review.</p>
<h2>How does the eight-check agentic review layer work?</h2>
<p>The agentic review layer sits between Code and a human approver. It is not a linter. It is not a security scanner with a pretty UI. It is eight specialised checks, each one tuned to the failure modes that show up when most of the code in a pull request was not typed by a person. Every PR that touches a client environment passes through this layer before a human signs off.</p>
<p>The eight checks are sequential. A failure halts the pipeline; the engineer either fixes the issue or files an explicit override that lands in the audit trail. No silent passes.</p>
<table>
<thead>
<tr>
<th>#</th>
<th>Check</th>
<th>What it catches</th>
<th>Why it matters at AI volume</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td><strong>Spec conformance</strong></td>
<td>Drift between the diff and the agreed spec or acceptance criteria</td>
<td>AI tends to over-deliver: it adds a flag here, a helper there, none of it asked for. Conformance keeps scope honest.</td>
</tr>
<tr>
<td>2</td>
<td><strong>Security regression</strong></td>
<td>Known vulnerable patterns, secrets in code, deprecated auth or authorization patterns reintroduced</td>
<td>Plausible code can pass tests and quietly bring back a pattern the team killed months ago. This is the check that keeps that from shipping.</td>
</tr>
<tr>
<td>3</td>
<td><strong>Convention drift</strong></td>
<td>House style violations: naming, error handling, logging, file structure, module boundaries</td>
<td>AI writes code that looks tidy in isolation but fights the project's conventions in aggregate. Drift compounds across files.</td>
</tr>
<tr>
<td>4</td>
<td><strong>Anti-pattern repetition</strong></td>
<td>The same poor pattern echoed across multiple files in the same PR</td>
<td>One bad pattern is a code smell. The same bad pattern in eight files is a future refactor on the engineering budget.</td>
</tr>
<tr>
<td>5</td>
<td><strong>Performance pitfalls</strong></td>
<td>N+1 queries, hot-loop allocations, missing indexes, unbounded recursion, accidental quadratic work</td>
<td>Tests pass at test scale. Performance fails at production scale. The check models the difference.</td>
</tr>
<tr>
<td>6</td>
<td><strong>Test integrity</strong></td>
<td>Tautological assertions, tests that verify the implementation rather than the contract, missing edge cases</td>
<td>AI is excellent at writing tests that confirm whatever it just wrote. The check forces tests back onto the spec.</td>
</tr>
<tr>
<td>7</td>
<td><strong>Dependency &amp; supply-chain hygiene</strong></td>
<td>New packages, version pinning, license risk, transitive CVEs, unused additions</td>
<td>A new dependency is a permanent decision made in a five-second autocomplete. The check makes that decision explicit.</td>
</tr>
<tr>
<td>8</td>
<td><strong>Data &amp; API safety</strong></td>
<td>Breaking API contract changes, unsafe migrations, PII handling, schema compatibility</td>
<td>The change that breaks a client integration is rarely the one a human reviewer notices at 5pm on Friday. The check does.</td>
</tr>
</tbody>
</table>
<h3>What a verdict looks like</h3>
<p>Each check produces a structured verdict — pass, fail with required fix, or pass-with-warning — and writes the result back to the PR with the offending lines highlighted. A senior engineer then reads the verdicts, not the diff cold.</p>
<blockquote>
<p><strong>"AI doesn't write bad code, mostly. It writes plausible code. Passes the tests, looks fine, and quietly reintroduces an auth pattern the team killed six months ago."</strong>
— Fred Sarmento, founder, Twistag</p>
</blockquote>
<h3>Who guards the layer itself</h3>
<p>Two principles guard the layer itself. First, the human still owns every call. The review layer surfaces what to look at; it does not approve or merge. Second, every override is logged, attributed, and visible on the dashboard. If a team is overriding the same check repeatedly, that is signal — either the check is wrong, or the team is taking on debt with its eyes open. Either way, leadership sees it.</p>
<h2>What does the Pulse Dashboard show?</h2>
<p>The eight checks fixed half of the review problem. The other half was visibility. Five tools, five tabs, five different opinions about what a healthy project looks like, and code review that dragged into Friday because the picture was always one tab away. That coordination tax is not unique to us: HBR's research across more than 300 organisations found time spent on collaborative activities has <a href="https://hbr.org/2016/01/collaborative-overload">ballooned by 50% or more</a> over two decades, much of it coordination rather than judgment.</p>
<p>So we built one dashboard. Seven categories. Every angle, one screen.</p>
<blockquote>
<p><strong>"Judgment goes where it matters. For the first time, we can see where that is."</strong>
— João Belo, VP of Engineering, Twistag</p>
</blockquote>
<table>
<thead>
<tr>
<th>Category</th>
<th>What it surfaces</th>
<th>The decision it informs</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Quality</strong></td>
<td>Escaped defects per release, P0/P1 bug count, regression rate, agentic-review pass rate per check, code-review escape rate</td>
<td>Where reviewer time pays back, which of the eight checks needs tightening</td>
</tr>
<tr>
<td><strong>Security</strong></td>
<td>Open CVEs by severity, time-to-patch, secrets-exposure incidents, license risk, dependency drift, authz regressions caught at the agentic layer</td>
<td>When to escalate, what to ship, what to hold</td>
</tr>
<tr>
<td><strong>Cost &amp; ROI</strong></td>
<td>Cost per shipped feature, AI inference spend per squad, infra cost trend, build and test minutes consumed, hours saved by Pulse against baseline</td>
<td>Where Pulse pays back, where it doesn't, where the budget is being spent on motion rather than progress</td>
</tr>
<tr>
<td><strong>Activity</strong></td>
<td>PRs opened and merged, deploy frequency, lead time for changes, review turnaround, work-in-progress, cycle time</td>
<td>Bottlenecks, flow health, whether the team is actually shipping or just busy</td>
</tr>
<tr>
<td><strong>People</strong></td>
<td>Workload distribution, on-call rotation load, focus-time vs. meeting-time, after-hours commit activity, knowledge concentration</td>
<td>Where burnout is brewing, where to redistribute, where one person knows too much that nobody else does</td>
</tr>
<tr>
<td><strong>Adoption</strong></td>
<td>Feature usage by released capability, activation rate of new releases, customer engagement on Pulse-shipped work, retention by cohort</td>
<td>Whether the things we ship are the things that get used</td>
</tr>
<tr>
<td><strong>Ops</strong></td>
<td>Uptime, error rate, P95 and P99 latency, alert volume, MTTR, MTTD, on-call incident count</td>
<td>Production health, where to invest reliability work next</td>
</tr>
</tbody>
</table>
<p>The categories are not independent. A spike in <em>Activity</em> (PRs flying through) without a corresponding spike in <em>Adoption</em> (customers actually using what shipped) is a warning. A green <em>Quality</em> board with a red <em>People</em> board means the team is paying for that quality with sleep. The dashboard is wired so leadership can read those relationships in seconds, not in three Slack threads.</p>
<p>The dashboard is also the audit trail. Every override on the agentic review layer, every escaped defect, every infra cost spike is attributed and timestamped. If a client asks why we shipped a particular release on a particular day, the answer is on the dashboard.</p>
<h2>What changed for engineers and for clients?</h2>
<p>Two shifts came out of all of this, and they reinforce each other.</p>
<p>For engineers, judgment moved up the stack. They are no longer reading every diff cold at 5pm because the agentic layer has already pre-read it for them. They are reading the verdicts, the overrides, the high-signal sections the layer flagged. The work is still demanding, arguably more so, because every interaction is non-trivial. But the volume of low-signal review is gone. The engineers Twistag hires shipped enough systems before joining to know when "looks fine" isn't fine. AI just made that instinct more valuable, not less.</p>
<p>For clients of our <a href="/product-engineering">product engineering</a> practice, two things change. Smaller teams deliver more: a five-person Twistag squad with Pulse produces what traditionally requires eight to ten engineers. And consistency stops depending on individual heroics — the agentic review layer enforces the same eight checks on sprint one and sprint twenty, regardless of which engineer is on duty.</p>
<p>The numbers we watch closely are the unglamorous ones. Override rate per check, by team, over time. Escape rate of issues that cleared the agentic layer and reached production. Time from PR opened to merge, separated by whether the PR touched a client environment. None of these are headline metrics. They are the metrics that tell us whether Pulse is actually working or whether we are fooling ourselves with motion.</p>
<h2>Which principles transfer beyond Pulse?</h2>
<p>Three principles hold the system together. They are also the parts a buyer in another industry can take and apply tomorrow.</p>
<p><strong>One: assume the code was not written by a human, and review accordingly.</strong> When you stop expecting human authorship, you stop relying on the heuristics that depend on it (intent, comments, naming choices) and start building checks that work on the artefact regardless of who or what produced it. Most review processes still assume human authorship. Most code now isn't.</p>
<p><strong>Two: every override is data, not friction.</strong> A check that fails and gets overridden is the most useful event in the system, because it tells you whether the check is calibrated, whether the team is taking on debt knowingly, and whether a particular pattern is becoming load-bearing. Overrides are routed to the dashboard, not buried in a CI log.</p>
<p><strong>Three: one screen for engineering health beats five tabs.</strong> The cost of context-switching between tools is paid by the people whose judgment you most want focused. Pulling the most decision-relevant signal from each tool into one view is not about prettiness — it is about putting judgment in the place where the trade-offs are visible.</p>
<p>None of this works without experienced product engineers in the loop. The tooling is sharp, but the final call still belongs to someone who has shipped enough systems to know when plausible code is hiding a real problem. Twistag's hiring bar exists because of this, not in spite of it. AI made that bar more important, not less.</p>
<h2>What should buyers ask an engineering partner?</h2>
<p>For organisations evaluating engineering partners, the question is no longer whether a team uses AI. Everyone uses AI. The question is whether AI is embedded in the delivery system — with review, visibility, and accountability around it — or bolted on top of a process that was designed for ten percent of today's code volume.</p>
<p>Pulse is the system. The five delivery agents move work forward. The eight-check review layer keeps it production-grade. The dashboard tells us where judgment is being spent and whether it is paying back. Twistag built it because the alternative was waiting for the slow incident to land. We would rather not. The next iteration is already underway: feeding the override data back into the checks themselves, so the layer recalibrates from every call an engineer makes.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-building-autonomous-teams-with-less-process-and-more-impact.webp" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[How agentic AI changes the unit of work in enterprise teams]]></title>
            <link>https://twistag.com/thinking/agentic-ai-the-new-unit-of-work-in-the-modern-workplace</link>
            <guid isPermaLink="false">https://twistag.com/thinking/agentic-ai-the-new-unit-of-work-in-the-modern-workplace</guid>
            <pubDate>Wed, 08 Jul 2026 13:50:46 GMT</pubDate>
            <description><![CDATA[Agentic AI shifts the unit of work from tasks people do to outcomes agents own. 46% of leaders already automate whole workstreams; teams must redesign roles.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p>
<p>The unit of work is changing. For decades, teams organised around tasks a person executes; agentic AI reorganises them around outcomes an agent owns and a person reviews. Already, <a href="https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born">46% of business leaders</a> say their organisation uses agents to fully automate workstreams. This post covers what that shift means for roles, workflow design, and the operating model behind the <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a> we define in our pillar guide.</p>
<h2>Key takeaways</h2>
<ul>
<li>46% of leaders say their organisation already uses AI agents to fully automate entire workstreams, and 82% expect digital labour to expand workforce capacity within 12 to 18 months (Microsoft Work Trend Index, 2025).</li>
<li>Usage data leans collaborative: Anthropic's Economic Index measured 57% of AI interactions as augmentation and 43% as automation. The unit of work is shared, not surrendered.</li>
<li>Gartner predicts 15% of day-to-day work decisions will be made autonomously by agentic AI in 2028, up from 0% in 2024. Review capacity, not build capacity, becomes the team's bottleneck.</li>
<li>Gartner also expects over 40% of agentic AI projects to be cancelled by the end of 2027. The survivors redesign the workflow around the agent instead of bolting an agent onto the old one.</li>
<li>From our own delivery work: a Semantic Kernel agent that converts plain-language questions into SQL against Snowflake replaced a reporting queue measured in days with answers in seconds, and moved the analyst's job from writing queries to validating them.</li>
</ul>
<h2>What is the unit of work, and why is it changing?</h2>
<p>The unit of work is the smallest chunk of value a team plans, assigns, and tracks. For fifty years that chunk was a task performed by a person: write the query, reconcile the ledger, draft the summary. An agent changes the shape of the chunk. Software that plans, calls tools, observes results, and retries can own an outcome, with a person setting the goal and judging the output.</p>
<p>The shift is collaborative rather than wholesale. <a href="https://www.anthropic.com/news/the-anthropic-economic-index">Anthropic's Economic Index</a>, built from millions of anonymised Claude conversations, measured 57% of usage as augmentation, where the model checks, iterates, and teaches alongside a person, against 43% as direct task automation. Teams are not handing work over. They are splitting each unit of work into a part the agent executes and a part a human owns.</p>
<h2>What changes for team roles when agents own tasks?</h2>
<p>Three responsibilities appear on every team that delegates to agents: setting scope, reviewing output, and owning exceptions. None of them sat on an org chart in 2023, and all three are becoming ordinary. In Microsoft's <a href="https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born">2025 Work Trend Index</a>, a survey of 31,000 workers across 31 countries, 82% of leaders said they expect to use digital labour to expand workforce capacity within 12 to 18 months, and leaders expect their teams to be training agents (41%) and managing them (36%) within five years. Microsoft's name for the emerging role is the "agent boss."</p>
<p>The practical consequence: job descriptions shift from throughput to judgment. A senior analyst who once produced twenty reports a week now defines the templates, sets the guardrails, and reviews the drafts an agent produced overnight. The skill being paid for is no longer execution speed. It is knowing what correct looks like, and catching the 1-in-20 output that isn't.</p>
<h2>How does workflow design change?</h2>
<p>Delegating to an agent is not assigning a ticket to a junior engineer. Three design changes separate workflows that hold up in production from demos that stall.</p>
<h3>Every delegation needs a contract</h3>
<p>An agent needs the goal, the boundaries, and the escalation rule written down: what it may read, what it may change, and when it must stop and ask. Across the agents we shipped in 2025-26, the workflows that failed internal review were rarely failing on model quality. They failed because nobody had defined who is accountable when the agent is wrong. Write the contract before the prompt.</p>
<h3>Exceptions become the design centre</h3>
<p>Human workflows hide their exceptions inside experienced heads; the veteran in accounts knows which supplier's invoices always arrive malformed. An agent forces those cases into the open. The real design work is deciding which inputs route to a human queue and making that queue somebody's job. In the production agents we run, a low-single-digit percentage of items routing to human review is a healthy pattern. Zero usually means the thresholds are too loose to catch anything.</p>
<h3>Review becomes a first-class activity</h3>
<p>Gartner predicts <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">15% of day-to-day work decisions will be made autonomously through agentic AI in 2028</a>, up from 0% in 2024. Read the other side of that number: 85% of decisions still involve a person, and a growing share of them are reviews of agent output. Teams that do not budget review time convert it into unplanned rework. Plan review capacity the way you plan compute.</p>
<h2>What does the new unit of work look like in production?</h2>
<p>One example from our own delivery work. For a European enterprise customer, we used Microsoft's Semantic Kernel to build an agent that takes a plain-language question, interprets the intent, writes the corresponding SQL, executes it against the company's Snowflake warehouse, and returns the answer with the underlying data attached. It qualifies as an agent because it plans and selects its own tools, even though it mostly interacts with itself.</p>
<p>The engineering is deliberately unremarkable: the language model handles language, SQL handles retrieval, and the warehouse's existing permissions handle access. The operating-model change is the point. A reporting request that previously queued for days behind a data team now returns in seconds, and the data team's unit of work moved up a level, from writing queries to curating the semantic layer and reviewing the queries the agent writes. We reused the same framework for a second customer to generate documents and working HTML prototypes on demand. The pattern transfers because the delegation contract, not the use case, is the reusable part.</p>
<h2>What stays with humans?</h2>
<p>Accountability, ambiguous goals, and anything that crosses an organisational boundary. An agent can reconcile the ledger; it cannot decide that the reconciliation policy is wrong. It can draft the customer response; it should not decide to waive the fee. The dividing line we apply in practice: agents own outcomes whose success criteria can be written down and tested, people own outcomes where the criteria themselves are in dispute. That line moves every year as evaluation tooling improves, but it moves by deliberate decision, not by drift. Letting it drift is how organisations end up in the incident reviews.</p>
<h2>How do you adopt the new unit of work without joining the 40%?</h2>
<p>Start with one workflow, redesign it around the agent, and staff the review loop before scaling to a second. Gartner expects <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">over 40% of agentic AI projects to be cancelled by the end of 2027</a> on escalating costs, unclear business value, or inadequate risk controls. McKinsey's State of AI research points at the same root cause from the other direction: workflow redesign correlates with EBIT impact more than any other organisational factor, yet only <a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai">21% of organisations using gen AI have redesigned any workflows at all</a>. The pattern that fails is adding an agent to an unchanged process and expecting the process to improve.</p>
<p>This operating-model work is what our <a href="/ai-and-agents">AI and agents practice</a> is built around, and it pairs with the team-shape question we examine in <a href="/post/building-autonomous-teams-with-less-process-and-more-impact">building autonomous teams with less process and more impact</a>: smaller teams, clearer ownership, fewer handoffs for agents to inherit.</p>
<p>The same Gartner forecast says 33% of enterprise software applications will include agentic AI by 2028, up from less than 1% in 2024. When that lands, the unit of work stops being a metaphor and becomes a line in the resource plan: teams will estimate, price, and staff around outcomes delegated to agents. The teams practising that arithmetic on one workflow today are the ones who will find 2028 unremarkable.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-agentic-ai-the-new-unit-of-work-in-the-modern-workplace.jpg" length="0" type="image/jpg"/>
        </item>
        <item>
            <title><![CDATA[The business case for modernizing legacy systems in the AI era]]></title>
            <link>https://twistag.com/thinking/modernizing-legacy-systems-the-key-to-unlocking-business-potential</link>
            <guid isPermaLink="false">https://twistag.com/thinking/modernizing-legacy-systems-the-key-to-unlocking-business-potential</guid>
            <pubDate>Wed, 08 Jul 2026 13:50:46 GMT</pubDate>
            <description><![CDATA[The case for legacy modernization now includes AI eligibility. Tech debt eats 20-40% of technology estate value, and agents cannot act on closed systems.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p>
<p>Legacy modernization used to be justified by two numbers: run cost and risk. AI adds a third, larger term to the equation, because systems that agents cannot read or act on are excluded from the productivity gains every board is now budgeting for. The baseline is already expensive: CIOs estimate tech debt at <a href="https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/tech-debt-reclaiming-tech-equity">20 to 40 percent of the value of their entire technology estate</a> (McKinsey). Here is how to weigh standing still, wrapping, and rebuilding.</p>
<p><strong>Key takeaways</strong></p>
<ul>
<li>CIOs put tech debt at 20 to 40 percent of technology estate value, and report that 10 to 20 percent of the budget meant for new products gets diverted to servicing it (McKinsey).</li>
<li>The US federal government spends about <a href="https://www.gao.gov/products/gao-25-107795">80% of its $100+ billion annual IT budget operating and maintaining existing systems</a>, some of them 60 years old (GAO, 2025).</li>
<li>Developers lose <a href="https://stripe.com/files/reports/the-developer-coefficient.pdf">42% of their working week, 17.3 of 41.1 hours, to maintenance and bad code</a>, roughly $85 billion a year in global opportunity cost (Stripe).</li>
<li>Gartner predicts <a href="https://www.gartner.com/en/newsroom/press-releases/2026-06-18-gartner-predicts-more-than-70-percent-of-mainframe-exit-projects-will-fail-due-to-overestimation-of-generative-ais-capabilities">more than 70% of mainframe exit projects initiated in 2026 will fail</a> because teams overestimate what GenAI migration tooling can deliver.</li>
<li>MIT's NANDA research found 95% of enterprise GenAI pilots produce no P&amp;L impact, with poor integration into existing systems as the leading structural cause.</li>
</ul>
<h2>What does standing still actually cost?</h2>
<p>More than the maintenance line item shows. The clearest public benchmark is the US federal government, which spends about <a href="https://www.gao.gov/products/gao-25-107795">80% of its $100+ billion annual IT budget on operating and maintaining existing systems</a>; GAO's 2025 review of the 11 most critical legacy systems found eight running on outdated languages and seven carrying known, unpatchable security vulnerabilities. Private estates follow the same shape. <a href="https://stripe.com/files/reports/the-developer-coefficient.pdf">Stripe's developer survey</a> measured 17.3 hours of a 41.1-hour engineering week going to maintenance, debugging, and bad code, a 42% tax on the most constrained resource most companies have. And <a href="https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/tech-debt-reclaiming-tech-equity">McKinsey found</a> that 10 to 20 percent of budget nominally allocated to new products quietly drains into tech-debt remediation. Standing still is not a neutral option; it is a recurring charge with a rising rate.</p>
<h2>Why does AI change the modernization math?</h2>
<p>Because AI initiatives inherit the estate they land on. MIT's NANDA research found <a href="https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/">95% of enterprise GenAI pilots deliver no measurable P&amp;L return</a>, and its diagnosis is integration: tools that cannot connect deeply to existing systems and workflows stall. An agent can only automate a process whose systems expose state through APIs with checkable permissions. A platform that only speaks through a green-screen session or a nightly batch export is invisible to it. That converts modernization from a cost-reduction project into an eligibility question: every roadmap item that assumes AI-assisted operations silently depends on <a href="/data-cloud">data and systems the AI can actually reach</a>. The cost of standing still now includes the compounding value of everything the estate disqualifies you from building.</p>
<h2>Should you modernize or wrap?</h2>
<p>Wrap when the core is stable and the need is access; modernize when the system itself has to keep changing. This post covers the investment decision. The engineering side, <a href="/post/modernising-legacy-systems-for-ai-agent-access">how to open legacy systems to AI agents without a full rewrite</a>, is a sibling post worth reading alongside it.</p>
<h3>When wrapping wins</h3>
<p>Wrapping puts an API layer, and increasingly an MCP server, in front of the existing system so newer software and agents can read from it and write to it under governance. It wins when the core logic is correct and rarely changes: think a stable billing engine or a warehouse system that does its one job well. Wrapping costs months rather than years, retires no risk inside the core, and buys you AI eligibility now. It is the right first move for most estates because it generates evidence about which systems actually matter before you commit rebuild money.</p>
<h3>When rebuilding wins</h3>
<p>Rebuilding wins when change frequency is high, when the platform is decaying, or when the people who understand it are leaving. GAO's markers are a useful checklist: of the 11 most critical US federal legacy systems, <a href="https://www.gao.gov/products/gao-25-107795">eight run on outdated languages such as COBOL and assembly, four sit on unsupported hardware or software, and seven operate with known security vulnerabilities that cannot be remediated without modernization</a>. Any two of those conditions on a revenue-critical system make wrapping a delay tactic, not a strategy.</p>
<h3>Why full-exit projects keep failing</h3>
<p>The tempting third option, a wholesale AI-powered migration off the old platform, has the worst track record. <a href="https://www.gartner.com/en/newsroom/press-releases/2026-06-18-gartner-predicts-more-than-70-percent-of-mainframe-exit-projects-will-fail-due-to-overestimation-of-generative-ais-capabilities">Gartner predicts more than 70% of mainframe exit projects initiated in 2026 will fail to produce the intended benefits</a> because teams overestimate what GenAI code-conversion tooling can do against complex legacy code. Gartner's own advice for many environments is to use GenAI for modernization in place rather than migration off the platform. Our experience matches: AI assistance compresses the rewrite of a well-understood module, and does nothing for the module nobody understands. Incremental replacement, one bounded slice at a time behind the wrapper, keeps each failure small.</p>
<h2>What does modernization unlock for AI adoption?</h2>
<p>Three returns, in the order they arrive.</p>
<ol>
<li>Engineering capacity comes back first. McKinsey found companies that actively manage down tech debt free their engineers to spend <a href="https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/tech-debt-reclaiming-tech-equity">up to 50% more time on work that supports business goals</a>. That capacity is what staffs the AI roadmap.</li>
<li>Agent eligibility follows. Once a system exposes clean APIs and permissions, it becomes a legitimate target for the <a href="/post/modernising-legacy-systems-for-ai-agent-access">agent-access architecture</a> and the automations that follow.</li>
<li>Compounding delivery speed is worth the most. A production-planning system we rebuilt for a European industrial client in 2025 moved from an overnight batch to a queryable service; the AI-assisted scheduling features that followed a quarter later were a small increment on top, not a second project.</li>
</ol>
<p>This ordering is why our <a href="/product-engineering">product engineering</a> teams scope modernization and AI features as one programme, not two: the first pays for the second, and the second justifies the first.</p>
<h2>Who should own the modernization decision?</h2>
<p>The P&amp;L that the system serves, not the technology function alone. McKinsey's first principle for managing tech debt is to treat it as a business issue: trace ownership of each system's debt to the profit and loss it supports, and reflect the "interest" cost in that unit's numbers. In our engagements this changes behaviour faster than any architecture review, because a product owner who sees the drag in their own margin stops treating modernization as an IT request competing with features. It also keeps scope honest. A business owner funds the slice that unblocks their roadmap; only a central IT budget funds a three-year replatforming with no committed consumer.</p>
<h2>How do you build the case?</h2>
<p>Price all three terms, not just the first one.</p>
<ol>
<li>Price the recurring charge. Sum the maintenance line, the Stripe-style engineering tax, and the 10 to 20 percent of new-product budget being diverted. This is what standing still costs per year.</li>
<li>Price the risk retirement. Unsupported components and unpatchable vulnerabilities carry a probable-loss figure your security team can estimate; GAO treats these as first-order modernization triggers, and so should the business case.</li>
<li>Price AI eligibility. List the roadmap items blocked because a system cannot expose state to an agent, and attach the value already claimed for them in board decks. This term is usually the largest and the least often written down.</li>
<li>Then choose the smallest slice that moves all three numbers, wrap the rest, and re-run the case each budget cycle with production evidence instead of estimates.</li>
</ol>
<h2>What happens to this market next?</h2>
<p>The market will force this discipline soon enough. Gartner expects <a href="https://www.gartner.com/en/newsroom/press-releases/2026-06-18-gartner-predicts-more-than-70-percent-of-mainframe-exit-projects-will-fail-due-to-overestimation-of-generative-ais-capabilities">75% of mainframe exit vendors to pivot or cease operations by 2030</a> as one-size-fits-all migration promises collapse. The durable position is owning an incremental modernization capability inside your own delivery organisation, and the window to build one is the next two budget cycles, while the estate still has people who remember how it works.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-modernizing-legacy-systems-the-key-to-unlocking-business-potential.jpg" length="0" type="image/jpg"/>
        </item>
        <item>
            <title><![CDATA[Cut process complexity before you add AI agents]]></title>
            <link>https://twistag.com/thinking/how-cutting-complexity-can-drive-real-innovation</link>
            <guid isPermaLink="false">https://twistag.com/thinking/how-cutting-complexity-can-drive-real-innovation</guid>
            <pubDate>Wed, 08 Jul 2026 13:50:46 GMT</pubDate>
            <description><![CDATA[Automating a broken process hardens the breakage. With 95% of GenAI pilots showing no P&L impact, simplify the workflow before you deploy AI agents on it.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p>
<p>Automating a broken process does not fix it; it hardens the breakage and executes it faster. Before deploying AI agents on a workflow, delete the steps that should not exist. The case for that sequencing is stark: <a href="https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/">MIT research found 95% of enterprise GenAI pilots</a> deliver no measurable P&amp;L impact, and the stalls trace to workflow fit, not model quality.</p>
<h2>Key takeaways</h2>
<ul>
<li>MIT's NANDA initiative found 95% of enterprise GenAI pilots deliver no measurable P&amp;L impact; the tools stall because they do not learn from or adapt to the workflows they sit in (Fortune, 2025).</li>
<li>Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls.</li>
<li>Only 21% of organisations using gen AI have redesigned any workflows, yet McKinsey finds fundamental workflow redesign is the organisational change most correlated with EBIT impact.</li>
<li>The advice is 36 years old: Michael Hammer's 1990 rule, "don't automate, obliterate," applies with more force to agents than it did to the enterprise IT it was written about.</li>
<li>Across the agent projects we shipped in 2025-26, deleting process steps before building removed roughly a third of the planned integration scope. The cheapest engineering hours are the ones you never buy.</li>
</ul>
<h2>Why does automating a broken process make it worse?</h2>
<p>Because automation is an amplifier. Michael Hammer made the argument in 1990 in <a href="https://hbr.org/1990/07/reengineering-work-dont-automate-obliterate">Reengineering Work: Don't Automate, Obliterate</a>: companies were embedding outdated processes in silicon and software instead of eliminating them, and the heavy IT spending of that era disappointed accordingly. Agents raise the stakes on the same mistake. A script executes a bad process rigidly, so its failures at least look identical. An agent executes a bad process at volume and with variation, working around obstacles a human would have questioned, and producing confident output at every broken step. This is why simplification sits at the front of the production patterns in our pillar guide to <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a>: the process is the specification, and agents inherit its bugs.</p>
<h2>Why do agent projects inherit process debt?</h2>
<p>An agent has to encode every rule of the workflow it runs, including the rules nobody wrote down. Every undocumented approval, dormant exception branch, and duplicate handoff becomes an integration to build, an evaluation case to test, and a failure mode to monitor. That cost structure explains the cancellation data. Gartner attributes its <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">40% cancellation forecast</a> to escalating costs and unclear business value, which is what happens when a team discovers mid-build that the process has twice as many rules as anyone believed. The MIT finding rhymes: pilots stall when tools fail to fit how work actually flows. Complexity is not just a drag on the process. It is the line item that kills the project.</p>
<h2>What should you delete before you build?</h2>
<p>Map the full workflow with the people who run it, then work through four categories of candidate deletions before writing any agent code. The order below reflects how often each category shows up in the workflows we map, most frequent first.</p>
<h3>Approvals that never say no</h3>
<p>If an approval step approves more than 99% of what reaches it, it is not a control. It is a log with latency. In the workflow-mapping sessions we run before agent builds, these are the first candidates found and the easiest to remove: replace them with a notification and an audit trail, and reserve genuine approval for the cases that historically got rejected.</p>
<h3>Handoffs that only move information</h3>
<p>A handoff earns its place when the receiving party transforms the work or adds a judgment. A handoff that only relocates information adds queue time and an error surface, and once an agent arrives it calcifies into an API boundary someone has to build and maintain. Collapse those steps first; each one deleted is an integration you never pay for.</p>
<h3>Exception branches nobody can explain</h3>
<p>Ask the process owner when each branch fires. If no one can answer, an agent cannot either, and the branch will surface later as an unexplained eval failure. Either document the trigger precisely or retire the branch and let the exception route to a human queue. A named queue beats a mystery rule.</p>
<h3>Outputs nobody reads</h3>
<p>Reports, status emails, and reconciliation files with no downstream consumer are pure waste today and worse tomorrow, because an agent will generate them tirelessly and at cost. Deletion beats automation on every metric that matters here. If the output vanishes for a month and nobody asks, it was already gone.</p>
<h2>How do you know a process is simple enough for an agent?</h2>
<p>Three tests, applied in order. First, the process fits on one page, exceptions included; if it cannot be described, it cannot be delegated. Second, it has a named owner who can state what a correct outcome is, because that statement becomes the evaluation set. Third, the exception rate is known, because that number sets the human-review threshold and the staffing behind it. McKinsey's State of AI research shows why the bar is worth clearing: <a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai">only 21% of organisations have redesigned workflows</a> when adopting gen AI, yet redesign is the change most correlated with EBIT impact. The one-page version of the process is the redesign.</p>
<h2>Isn't this just business process reengineering again?</h2>
<p>The diagnosis is the same; the economics are not. Reengineering in the 1990s meant multi-year consulting programmes because mapping and rebuilding processes was manual and slow, and many efforts collapsed under their own weight. The agent era inverts the cost curve. The build itself forces the map to exist, because every rule must be written down to be delegated, and a simplification pass beforehand is a two-week exercise, not a two-year one. Simplification stopped being a transformation programme and became a delivery phase, which is exactly how our <a href="/product-engineering">product engineering practice</a> treats it: deletion is scoped, scheduled, and signed off like any other milestone.</p>
<h2>What does this look like in practice?</h2>
<p>A composite from our production work. A document-heavy invoice workflow arrived for automation with twelve steps, four of them approvals. Mapping showed two approvals that had never rejected anything and two handoffs that only moved files between teams. We deleted those four steps before building, cutting the agent's integration scope roughly in half, and shipped against the five steps that remained. The agent now routes 2-3% of documents to human review, a ratio that is only manageable because the redundant approval layers went first; run through the original twelve steps, the same volume would have buried the review queue.</p>
<p>The lesson generalises beyond agents, and it is the process-side twin of the team-design argument in <a href="/post/building-autonomous-teams-with-less-process-and-more-impact">building autonomous teams with less process and more impact</a>: fewer steps means fewer handoffs, clearer ownership, and less for any system, human or agent, to inherit.</p>
<p>Gartner's cancellation window runs to the end of 2027. Between now and then, every enterprise agent budget will quietly fund one of two things: encoding a process as found, or encoding a process worth keeping. The deletion audit that decides which takes about two weeks. Schedule it before the agent budget is committed, not after the eval suite starts failing.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-how-cutting-complexity-can-drive-real-innovation.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[What enterprise product teams can learn from agency delivery]]></title>
            <link>https://twistag.com/thinking/what-enterprise-product-managers-can-steal-from-software-agencies-to-build-game-changing-products</link>
            <guid isPermaLink="false">https://twistag.com/thinking/what-enterprise-product-managers-can-steal-from-software-agencies-to-build-game-changing-products</guid>
            <pubDate>Wed, 08 Jul 2026 13:50:46 GMT</pubDate>
            <description><![CDATA[Enterprise teams close the delivery gap with agency habits: weekly shipping, one-day prototypes, teams under 10. Pendo found 80% of features go unused.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p>
<p>The delivery gap between enterprise product teams and agencies is operating discipline, not talent: shipping cadence, prototype-first validation, small-team ownership, and scope-cutting. The cost of skipping that discipline is measurable — <a href="https://www.pendo.io/resources/the-2019-feature-adoption-report/">Pendo's analysis</a> of anonymised product usage data found 80 percent of features in the average software product are rarely or never used. This post breaks the four habits down, with the evidence behind each.</p>
<p><strong>Key takeaways</strong></p>
<ul>
<li>Pendo found 80 percent of features in the average software product are rarely or never used; for a $50 million-revenue software company that is roughly $8.4 million a year spent on unwanted features.</li>
<li>McKinsey's Developer Velocity Index, from 440 large organizations, links top-quartile delivery capability to revenue growth four to five times faster than bottom-quartile peers.</li>
<li>Google's DORA research finds elite delivery teams are twice as likely to meet or exceed their organizational performance goals.</li>
<li>Amazon caps teams at roughly 10 people with single-threaded ownership; Gallup data cited by AWS shows engagement of 42 percent or higher in sub-10-person groups versus under 30 percent in larger organizations.</li>
<li>A prototype built in one day typically costs about 1 percent of the production build and kills bad ideas before they consume a roadmap quarter — first-party numbers from our own prototype sprints.</li>
</ul>
<h2>Why do agencies ship faster than enterprise product teams?</h2>
<p>Because the constraint is different, not the people. An agency lives on fixed budgets and dated contracts: if nothing demonstrable exists by Friday, the client asks why. That constraint produces habits — weekly releases, prototypes before commitments, hard scope decisions — that most enterprise teams never form because their funding model does not force it. The gap this creates is documented. <a href="https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/developer-velocity-how-software-excellence-fuels-business-performance">McKinsey's Developer Velocity Index</a>, built from surveys of technology executives at 440 large organizations across 12 industries, found top-quartile companies grew revenue four to five times faster than bottom-quartile peers between 2014 and 2018, with 60 percent higher total shareholder returns and 55 percent higher scores on innovation. Delivery discipline is not an engineering detail. It shows up in the P&amp;L.</p>
<h2>What does a weekly shipping cadence change?</h2>
<p>Everything downstream of it. Shipping weekly forces small batches, small batches force clear priorities, and clear priorities expose scope that should not exist. <a href="https://cloud.google.com/blog/products/devops-sre/using-the-four-keys-to-measure-your-devops-performance">Google's DORA research</a> finds elite delivery teams are twice as likely to meet or exceed their organizational performance goals. The <a href="https://dora.dev/research/2024/dora-report/">2024 DORA report</a> adds a warning that lands squarely on cadence: AI-assisted coding raises individual productivity but degrades delivery stability and throughput wherever fundamentals like small batch sizes slip. Cadence is the fundamental.</p>
<p>Enterprise teams tend to assume cadence needs new tooling or headcount. It mostly needs permission — we wrote about the operating model behind that in <a href="/post/building-autonomous-teams-with-less-process-and-more-impact">building autonomous teams with less process</a>. The mechanical version:</p>
<ul>
<li>Demo something running every week, to a stakeholder empowered to say no</li>
<li>Split any work item that cannot produce a visible increment within two weeks until it can</li>
<li>Treat the release as the status report: if it shipped, it is done; if it did not, no slide deck changes that</li>
</ul>
<h2>Why do agencies prototype before committing?</h2>
<p>Because most feature ideas fail, and a prototype is the cheapest place to find out. Pendo's analysis of usage data across 615 product deployments found that <a href="https://www.pendo.io/resources/the-2019-feature-adoption-report/">80 percent of features are rarely or never used</a>, and estimated $29.5 billion in development spend at public cloud software companies sits in those features. An agency cannot absorb that hit rate — unused output is unbilled learning — so validation happens before the build, not after launch.</p>
<p>The working method is a prototype with a one-day budget: real interface, faked backend, put in front of a user or a sponsor before any architecture discussion. We documented the format in <a href="/post/prototype-in-a-day">prototype in a day</a>. Our own numbers from those sprints: the one-day version costs about 1 percent of the production build and answers the two questions that kill features — does anyone want this, and is it technically sane. Enterprise teams usually invert the sequence, validating with documents and committees, which is why the 80 percent figure exists.</p>
<h2>How does small-team ownership change delivery?</h2>
<p>Amazon's two-pizza rule — no team bigger than two pizzas can feed, in practice <a href="https://aws.amazon.com/executive-insights/content/amazon-two-pizza-team/">fewer than 10 people</a> — is the most copied version of the principle. The part enterprises skip is what Amazon pairs it with: single-threaded ownership, one team owning one service across the full lifecycle, from idea through operations. The evidence is older than the slogan. AWS points to the Ringelmann effect, where individual output drops as group size grows, and to Gallup workplace data showing engagement of 42 percent or higher in organizations under 10 people against under 30 percent in larger ones.</p>
<p>The enterprise translation is not renaming departments into squads. It is deleting handoffs: the team that decides also builds, and the team that builds also operates. Every handoff removed returns days per cycle. Every approval gate retained should have to justify itself against that cost. A useful audit takes an afternoon: trace one recent feature from decision to production and count the teams it touched. Anything above three is coordination tax, and the tax compounds with every release.</p>
<h2>Why is cutting scope a discipline rather than a failure?</h2>
<p>Agencies hold the deadline fixed and treat scope as the variable, because on a fixed-price contract an overrun comes out of margin. Enterprise teams usually invert it — scope fixed by the annual plan, timeline elastic — which is exactly how unused features accumulate: the plan was written twelve months before anyone could test its assumptions. The discipline in practice:</p>
<ul>
<li>Rank the backlog by evidence, not stakeholder seniority; features validated by a prototype or usage data ship first</li>
<li>When the timeline slips, cut from the bottom of the list; never move the demo</li>
<li>Re-decide scope at every release, because each release generates the usage data the annual plan never had</li>
</ul>
<p>Scope-cutting only reads as failure when output is the metric. Measured on outcomes, it is the fastest cost reduction available to a product organization.</p>
<h2>What do agencies get wrong about enterprise delivery?</h2>
<p>The comparison cuts both ways, and pretending otherwise would be dishonest. Agencies routinely underestimate the enterprise integration surface: a feature that touches four internal systems and a compliance review is not slow because the team is slow. Agencies also under-price operational maturity. A platform carrying ten years of customer data has failure costs a greenfield build never faces, and a cadence tuned for greenfield work breaks against quarterly release windows and regulated change control. The habits transfer; the calendar does not always. In our enterprise engagements the workable compromise is a weekly internal demo cadence feeding a slower external release train, so the team keeps the forcing function even when production deploys are gated.</p>
<h2>Which habit should enterprise teams adopt first?</h2>
<p>Cadence, because it forces the other three: a weekly demo makes big batches impossible, exposes teams too large to coordinate, and turns scope debates into evidence debates. It also requires no reorganization — one team, one quarter, one visible increment a week is a decision a director can take on Monday. This is the operating model we bring into enterprise work through our <a href="/product-engineering">product engineering practice</a>: agency cadence applied inside enterprise constraints. The next two years raise the stakes. As AI-assisted development pushes individual output up, DORA's 2024 data shows the gains accrue to teams whose delivery fundamentals can absorb the extra throughput. The distance between shipping weekly and shipping quarterly is about to get more expensive, not less.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-what-enterprise-product-managers-can-steal-from-software-agencies-to-build-game-changing-products.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Where AI agents actually pay off in logistics operations]]></title>
            <link>https://twistag.com/thinking/ai-in-logistics-revolutionizing-efficiency-and-redefining-innovation</link>
            <guid isPermaLink="false">https://twistag.com/thinking/ai-in-logistics-revolutionizing-efficiency-and-redefining-innovation</guid>
            <pubDate>Wed, 08 Jul 2026 13:50:46 GMT</pubDate>
            <description><![CDATA[AI agents pay off in logistics document processing, exception handling, and tracking queries. McKinsey measures documentation lead-time cuts up to 60%.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p>
<p>AI agents in logistics are software systems that read freight documents, resolve shipment exceptions, answer track-and-trace queries, and support planning decisions with limited human input — the same <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a> applied to freight operations. The economics are measurable: <a href="https://www.mckinsey.com/capabilities/operations/our-insights/beyond-automation-how-gen-ai-is-reshaping-supply-chains">McKinsey finds</a> generative AI cuts documentation lead time by up to 60 percent. This post maps where agents pay off first and what deployment looks like in practice.</p>
<p><strong>Key takeaways</strong></p>
<ul>
<li>Documents are the first win: McKinsey measures up to 60 percent shorter lead times for producing shipping documentation and a 10 to 20 percent lighter workload for logistics coordinators.</li>
<li>Exception handling scales: one last-mile operator with a fleet of more than 10,000 vehicles saved $30 million to $35 million with virtual dispatcher agents, on roughly $2 million invested (McKinsey).</li>
<li>Altana AI's chief science officer reports efficiency gains of 30 to 50 percent on existing global-trade processes, with some workflows running ten times faster (Maersk).</li>
<li>Gartner predicts over 40 percent of agentic AI projects will be canceled by the end of 2027; narrow scope and a measured baseline separate the survivors.</li>
<li>Gartner also expects agentic AI inside 33 percent of enterprise software applications by 2028, up from under 1 percent in 2024.</li>
</ul>
<h2>What counts as an AI agent in logistics?</h2>
<p>An agent differs from the forecasting models logistics teams have run for a decade in one way: it acts. A demand-forecast model outputs a number. An agent reads the bill of lading, extracts the twenty fields your TMS needs, drafts the customs entry, and escalates the two fields it could not verify. The distinction matters because the market is crowded with relabeled products: <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">Gartner estimates</a> only about 130 of the thousands of vendors selling "agentic AI" have real agent capability, and calls the rebranding of chatbots and RPA "agent washing." When a vendor pitches agents for your operation, apply one test: does the system take an action in your systems of record, and can you audit that action afterwards?</p>
<h2>Where do agents pay off first in logistics operations?</h2>
<p>Four workflows dominate the published evidence: document processing, exception handling, track-and-trace queries, and planning support. They share a profile — high volume, structured outcomes, and a clear escalation path to a human when the agent is unsure.</p>
<h3>Document processing</h3>
<p>Freight runs on paperwork: commercial invoices, bills of lading, customs declarations, dangerous-goods certificates. McKinsey's operations practice reports that generative AI can cut the lead time for producing shipping documentation <a href="https://www.mckinsey.com/capabilities/operations/our-insights/beyond-automation-how-gen-ai-is-reshaping-supply-chains">by up to 60 percent</a>, while reducing logistics coordinators' workload by 10 to 20 percent. This matches what we see outside the sector. The document-processing agents Twistag ships in manufacturing translate directly: a production invoice agent extracts and validates fields automatically and routes roughly 2 to 3 percent of documents to a human-review queue, with nothing failing silently downstream. The pattern is identical for a customs entry or a proof-of-delivery. Only the schema changes.</p>
<h3>Exception handling</h3>
<p>Exceptions are where dispatcher hours go: missed pickups, vehicle breakdowns, refused deliveries, address failures. McKinsey documents a last-mile operator with a fleet of more than 10,000 vehicles that saved $30 million to $35 million by deploying virtual dispatcher agents for driver troubleshooting and roadside assistance, on an investment of about $2 million. The same research describes a carrier with just over 150 vehicles that saved $3.5 million using an AI-mediated three-way messaging layer connecting drivers, dispatchers, and customers. The economics work because exceptions are frequent, individually small, and mostly resolvable from data the operator already holds. The agent handles the routine 90-plus percent; dispatchers keep the genuinely novel cases.</p>
<h3>Track-and-trace queries</h3>
<p>"Where is my shipment" is the highest-volume question every logistics operator answers. An agent grounded in TMS and telematics data answers it directly, with the escalation path reserved for disputes and claims. The gains compound at the operations level: Peter Swartz, chief science officer at Altana AI, <a href="https://www.maersk.com/insights/digitalisation/2024/07/02/ai-in-logistics-and-supply-chains">reports efficiency gains of 30 to 50 percent</a> on existing global-trade processes, with some running ten times faster. In the same Maersk discussion, the company's chief data officer describes port-berthing plans that once took days of manual permutation now taking hours against a live digital twin of the terminal.</p>
<h3>Planning support</h3>
<p>Planning is where AI in logistics started, and agents extend it rather than replace it. McKinsey's supply-chain research found that early adopters of AI-enabled supply-chain management <a href="https://www.mckinsey.com/industries/metals-and-mining/our-insights/succeeding-in-the-ai-supply-chain-revolution">improved logistics costs by 15 percent, inventory levels by 35 percent, and service levels by 65 percent</a> relative to slower-moving competitors. DHL's <a href="https://www.dhl.com/us-en/home/innovation-in-logistics/logistics-trend-radar/gen-ai.html">Logistics Trend Radar</a> tracks generative AI use cases from freight-route optimization to warehouse-layout generation and automated report drafting. The agent layer adds the final step: instead of a planner reading a dashboard and opening a spreadsheet, the agent drafts the stock transfer or the reroute and asks for approval.</p>
<h2>What does deploying a logistics agent actually look like?</h2>
<p>Less impressive than the demo, more valuable than the pilot. Gartner predicts <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">more than 40 percent of agentic AI projects will be canceled</a> by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. A January 2025 Gartner poll of 3,412 practitioners found only 19 percent of organizations had made significant agentic AI investments; 42 percent were investing conservatively. Most of the market is still hedging.</p>
<h3>Why logistics agent projects get canceled</h3>
<p>The failure modes are consistent with what Gartner describes and what we see in delivery. Teams pilot an agent against a demo dataset, then discover the real workflow spans a TMS, a WMS, an ERP, and the EDI streams between them — and the integration budget dwarfs the model budget. Or the project starts from the technology ("we need agents") rather than a workflow with a measurable baseline, so nobody can say whether it worked. Gartner's advice matches our experience: rethinking the workflow around the agent usually beats bolting an agent onto the workflow as it stands.</p>
<h3>What the surviving projects share</h3>
<p>The projects that reach production share three properties:</p>
<ul>
<li>A narrow workflow with a baseline metric captured before launch — documents processed per coordinator-hour, exceptions closed without escalation, query resolution time</li>
<li>Structured access to the systems of record rather than screen-scraping, so every agent action is written back and auditable</li>
<li>A human-review path sized honestly: 2 to 3 percent of volume routing to a review queue is normal for document workflows, and pretending it will be zero is how trust dies</li>
</ul>
<h2>How should a logistics operator start?</h2>
<p>Pick the workflow where volume is high and judgment is low — document intake usually wins — and instrument the baseline before the agent ships. A first production agent on a scoped workflow is a matter of weeks, not quarters; we broke down <a href="/post/how-long-to-ship-an-ai-agent">how long it takes to ship a production AI agent</a> in detail, and the timeline holds for logistics. The build is the smaller half of the work. Evaluation sets, fallback design, and operator training are what make the agent stick, and they are the core of our <a href="/ai-and-agents">AI and agents practice</a>: the same delivery pattern, pointed at freight.</p>
<h2>What changes in logistics by 2028?</h2>
<p>The forecasts are aggressive but directionally consistent. Gartner expects 15 percent of day-to-day work decisions to be made autonomously through agentic AI by 2028, up from zero in 2024, and agentic capability inside 33 percent of enterprise software applications by the same date. In logistics that lands first in the back office — the document, exception, and query workflows above, where the published evidence already exists. The operators positioned to capture it are spending 2026 on unglamorous groundwork: clean EDI streams, structured document archives, and audit trails an agent can act on. When agent layers arrive inside the major TMS platforms, those operators switch them on. Everyone else starts a data project.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-ai-in-logistics-revolutionizing-efficiency-and-redefining-innovation.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[The ROI of design: what the evidence shows and how to measure it]]></title>
            <link>https://twistag.com/thinking/the-roi-of-design-aligning-creativity-with-business-impact</link>
            <guid isPermaLink="false">https://twistag.com/thinking/the-roi-of-design-aligning-creativity-with-business-impact</guid>
            <pubDate>Wed, 08 Jul 2026 13:50:46 GMT</pubDate>
            <description><![CDATA[Design-led companies grew revenue 32 percentage points faster over five years (McKinsey). What design ROI is, the evidence behind it, and how to measure it.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p>
<p>The ROI of design is the measurable business value that design decisions produce: revenue growth, retention, conversion, and lower cost to operate. The evidence is unusually strong. McKinsey tracked 300 companies over five years and found that top-quartile design performers grew revenue 32 percentage points faster than their industry peers (<a href="https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-business-value-of-design">The Business Value of Design</a>, 2018). This post covers what design ROI means, the research behind it, and how to measure it in a shipping product.</p>
<p><strong>Key takeaways</strong></p>
<ul>
<li>Top-quartile companies in the McKinsey Design Index grew revenue 32 percentage points faster and total shareholder returns 56 points faster than industry peers over five years.</li>
<li>The DMI Design Value Index found that design-led companies outperformed the S&amp;P 500 by 211% over ten years.</li>
<li>Forrester measured a 301% ROI on IBM's design thinking practice, with design and alignment time cut by 75% and design defects halved.</li>
<li>Four metrics carry most design measurement: conversion rate, retention, task completion time, and engagement.</li>
<li>Returns concentrate at the top: McKinsey found the market barely distinguishes average design from below-average, but pays the top quartile disproportionately.</li>
</ul>
<h2>What is the ROI of design?</h2>
<p>The return on investment of design is the measurable value that design contributes to a business. Every design decision, from a reworked onboarding flow to a simplified data table, either moves a business number (conversion, retention, operating cost, revenue per user) or it does not. Design ROI is the discipline of knowing which, before and after the work ships.</p>
<p>In practice, design ROI lives at the intersection of two forces that often pull apart. Users ask for one thing; the business is focused on growth or cost-efficiency. Good design finds the point where both are served, and the returns show up as lower churn, higher conversion, and loyalty that compounds. Design that serves only one side of that equation, however polished, does not produce a return.</p>
<p>The definition matters because design is still widely bought on taste. In McKinsey's research, more than 40 percent of companies said they do not talk to their end users during development, and over half admitted they have no objective way to assess what their design teams produce. A function that is not measured gets funded last and cut first. The research below is the case for measuring it.</p>
<h2>How strong is the evidence for design ROI?</h2>
<p>Three independent research programmes reach the same conclusion from different angles: companies that treat design as a measured business function outperform companies that treat it as a styling pass.</p>
<table>
<thead>
<tr>
<th>Study</th>
<th>Finding</th>
<th>Scope</th>
</tr>
</thead>
<tbody>
<tr>
<td><a href="https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-business-value-of-design">McKinsey, The Business Value of Design</a> (2018)</td>
<td>Top-quartile design performers grew revenue 32 percentage points faster and shareholder returns 56 points faster than industry peers over five years</td>
<td>300 listed companies, two million+ financial data points, three industries</td>
</tr>
<tr>
<td><a href="https://www.dmi.org/page/DesignValue">DMI Design Value Index</a> (2015)</td>
<td>Design-led companies outperformed the S&amp;P 500 by 211% over ten years</td>
<td>16 US public companies meeting six design-management criteria</td>
</tr>
<tr>
<td><a href="https://www.ibm.com/downloads/cas/Z4WBDR8Q">Forrester, Total Economic Impact of IBM's design thinking practice</a> (2018)</td>
<td>301% ROI over three years; design and alignment time cut by 75%</td>
<td>Four enterprise clients interviewed plus 60 executive survey responses</td>
</tr>
</tbody>
</table>
<h3>What the numbers actually say</h3>
<p>The McKinsey result carries the most weight because of its method: two million pieces of financial data and more than 100,000 recorded design actions across medical technology, consumer goods, and retail banking. The correlation held in all three industries, which suggests the effect is not a software-sector artefact. It applies whether the product is a device, a service, or an app.</p>
<p>The sharpest detail hides in the quartiles. Revenue and shareholder-return differences between the fourth, third, and second quartiles were marginal. The market pays for design excellence, not design adequacy. Being slightly better than average returns almost nothing; reaching the top quartile returns the full 32 points.</p>
<h3>What the numbers do not say</h3>
<p>All three studies measure correlation, not causation. Companies that run design rigorously tend to run everything rigorously. That does not weaken the practical conclusion, because the mechanisms McKinsey identified (measuring design like revenue, cross-functional teams, continuous user testing) are operational habits, and operational habits can be adopted.</p>
<h2>How do you measure the ROI of design?</h2>
<p>Measure design the way you measure any investment: pick the business metric the design change is supposed to move, record the baseline, ship the change, and compare. Four metrics cover most product work:</p>
<ol>
<li><strong>Conversion rate:</strong> the share of users who complete a target action, such as sign-up, purchase, or activation.</li>
<li><strong>Retention:</strong> whether the change keeps users coming back, read as churn or cohort retention.</li>
<li><strong>Task completion time:</strong> how fast users get through the workflows the product exists to serve.</li>
<li><strong>Engagement:</strong> depth of interaction with the features the change touched.</li>
</ol>
<p>The instrument layer is A/B tests, funnel analytics, and session data. The discipline that separates measurement from theatre is the baseline: record the metric before the redesign, define the attribution window before you ship, and resist the urge to credit design for lifts that coincided with a pricing change or a traffic spike. A number without its baseline is a press release, not a measurement.</p>
<p>McKinsey documented an online gaming company where a small usability improvement to the home page was followed by a 25 percent increase in sales, and where polish beyond that point added almost nothing to users' value perception, so the team stopped. Both halves are the lesson. Measurement tells you when design pays and when to stop spending.</p>
<h2>How does design ROI change as a company grows?</h2>
<p>The metric that matters shifts with the stage of the business. What earns a return at launch wastes money at scale, and the reverse.</p>
<h3>Early stage: speed to signal</h3>
<p>Before product-market fit, design ROI is measured in learning per week. The job is to ship testable versions fast, watch real users, and kill weak directions early. Polish is negative ROI at this stage; every hour spent perfecting a screen that user testing will invalidate is an hour lost.</p>
<h3>Growth stage: friction removal at scale</h3>
<p>Once the model works, small percentages become large numbers. A one-point conversion improvement on a checkout that processes millions of orders is real revenue. Design work shifts to removing friction from proven journeys, building design systems so quality scales without headcount, and keeping the experience consistent across platforms.</p>
<h3>Maturity: differentiation</h3>
<p>In crowded markets, design becomes the moat. This is where the McKinsey quartile finding bites: adequate design returns nothing at maturity because every competitor has it. The return comes from experiences competitors have not matched, which requires the continuous user research and iteration habits the top quartile share.</p>
<h2>Why does design ROI depend on engineering?</h2>
<p>A design decision produces zero return until it ships. That makes engineering throughput a multiplier on design ROI, and it is the part most design-ROI conversations skip. In the Forrester study, better design understanding upstream cut development and testing time by 33 percent and halved design defects. The return showed up in the engineering budget, not the design budget.</p>
<p>This is why at Twistag we run design and <a href="/product-engineering">product engineering</a> as one pipeline rather than a handoff. Handoffs are where design intent decays: the flow that tested well gets simplified under deadline, the empty states never get built, and the version you measure is not the version you designed. We have written about <a href="/post/design-engineering-ai-native-pipeline">how design and engineering merge in an AI-native pipeline</a>, where the implementation cost of a design decision drops far enough that iteration speed stops being the constraint.</p>
<p>That drop is what changes the economics next. The 32-point spread McKinsey measured comes from a period when shipping a design change took a sprint. When AI-assisted implementation ships an interface change in a day, a team can test design decisions weekly instead of quarterly — and the compounding advantage will belong to the teams already measuring which changes pay.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-the-roi-of-design-aligning-creativity-with-business-impact.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Claude Code, Cursor, and the AI-coding stack we reach for]]></title>
            <link>https://twistag.com/thinking/claude-code-cursor-ai-coding-stack</link>
            <guid isPermaLink="false">https://twistag.com/thinking/claude-code-cursor-ai-coding-stack</guid>
            <pubDate>Tue, 07 Jul 2026 15:40:15 GMT</pubDate>
            <description><![CDATA[70% of senior engineers now use 2-4 AI coding tools daily. Here's the stack our team reaches for — and where each one wins.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: July 2026</em></p><p>The AI-coding tool question that consumed 2024 and 2025 has effectively resolved. <a href="https://larridin.com/developer-productivity-hub/developer-productivity-benchmarks-2026">70% of senior engineers in 2026 use two to four AI coding tools simultaneously</a>, not one. The interesting question is no longer "which tool wins" — it is "which tool for which task." This post is the stack our senior engineers actually reach for through a normal working day, why each one wins its slot, and where the boundaries between them sit. It is a cluster child of our <a href="/post/70-percent-ai-generated-codebase">70%-AI-generated codebase pillar</a>.</p><h3>Key takeaways</h3><ul><li><strong>The stack is layered by task type, not chosen by preference.</strong> IDE-inline for quick edits, CLI agents for heavy refactors and greenfield features, PR review agents for the merge gate, autonomous agents for scoped background work. Each tool wins at what it is shaped for.</li><li><strong>Cursor for daily editing, Claude Code for heavy sessions.</strong> This is the pattern that has settled among senior teams. Cursor stays in the editor and moves fast on the small stuff; Claude Code takes over when the change is architectural or spans many files.</li><li><strong>Subagents are what makes CLI tools scale.</strong> A single AI agent hits its context limit fast on a real codebase. The subagent pattern — dispatch specialised sub-agents for specific classes of work, aggregate their outputs — is how CLI tools handle work that IDE agents cannot.</li><li><strong>Cost per PR is now a metric worth tracking.</strong> Prompt caching, subagent choice, and model routing between them change per-PR cost by 5-10x. Teams that do not track it are leaving substantial spend on the table.</li><li><strong>The IDE is not going away.</strong> Autonomous coding agents got most of the 2026 press, but the daily driver for senior engineers is still an IDE with an AI assistant. The reason is judgement — the human wants to see what happens as it happens, not read a report at the end.</li></ul><h3>The four slots in the stack</h3><p>Every senior engineer we talk to describes a similar four-slot mental model.</p><p><strong>Slot 1 — IDE-inline assistant.</strong> The always-on, per-keystroke completion and inline suggestion tool. Cursor, GitHub Copilot, JetBrains AI. This is the slot the engineer spends the most hours in, and the tool has to be low-latency, non-intrusive, and easy to override.</p><p><strong>Slot 2 — CLI coding agent.</strong> The heavier tool the engineer reaches for when the work is beyond what fits into inline suggestion. Claude Code, Codex CLI, Gemini CLI. Multi-file changes, greenfield features, architectural refactors. Runs in a terminal alongside the IDE, with access to the whole repo.</p><p><strong>Slot 3 — PR review agent.</strong> The multi-agent reviewer that runs when a PR opens. Anthropic's <a href="https://www.infoq.com/news/2026/04/claude-code-review/">Claude Code Review</a>, GitHub's Copilot PR review, third-party review tools. Runs asynchronously, comments inline, gets merged into the review discipline described in the <a href="/post/70-percent-ai-generated-codebase">70%-AI-generated codebase pillar</a>.</p><p><strong>Slot 4 — Autonomous background agent.</strong> The tool that takes a scoped task and runs to completion without engineer intervention. Devin, various Anthropic and OpenAI offerings, in-house agents built on top of Claude Code. Used for bounded, well-specified work — dependency updates, test coverage expansion, boilerplate migrations.</p><p>The stack is layered because the tasks are different. Trying to make one tool win all four slots is what produced most of the 2024 disappointment with AI coding. Accepting that the tools specialise is what makes the stack productive.</p><h3>Slot 1: IDE-inline — Cursor is the default, GitHub Copilot the fallback</h3><p>The pattern that has settled in 2026: Cursor for daily editing, Copilot only when the customer's environment does not allow Cursor. The differentiators.</p><p>Cursor's Composer surface — multi-file inline editing without leaving the IDE — is where senior engineers spend most of their AI-coding time. It fits the review discipline: the engineer sees each suggestion, accepts or rejects it, moves on. The latency is low enough that it does not break flow. The scope is bounded enough that the engineer does not lose track of what changed.</p><p>GitHub Copilot is the fallback because it works everywhere GitHub works — including in enterprises that have not approved Cursor. The daily-driver experience is a step behind Cursor's, but Copilot has better organisation-wide policy controls, which matters for enterprise procurement.</p><p>The one place inline assistants lose is when the task is bigger than a few files. Then the engineer moves to slot 2.</p><h3>Slot 2: CLI — Claude Code owns the heavy slot</h3><p>The heavy slot is the one that changed the most in 2026. Claude Code emerged as the default for architectural work and multi-file refactors, with Cursor's Composer covering the middle ground. The differentiators.</p><p>Claude Code runs in a terminal, has access to the whole repository, and — most importantly — supports the subagent pattern. When a task requires context that will not fit in one agent's window, Claude Code dispatches specialised sub-agents for specific classes of work. A refactor across 30 files becomes ten paired-up sub-agents rather than one overwhelmed one.</p><p>The other Claude Code feature that matters: <code>/loop</code>, the automated test-run-and-refactor cycle. The engineer gives it a failing test and Claude Code iterates until the test passes. This is what makes tests-first at 70% AI-generation practical — the AI writes the implementation, runs the tests, fixes what breaks, and reports back when it is green or when it cannot make it green.</p><p>Codex CLI and Gemini CLI compete in the same slot. The choice between them is usually about the model the team already has enterprise access to, not about a substantive capability gap.</p><h3>Slot 3: PR review — the multi-agent shape has won</h3><p>The pattern that settled in 2026: multi-agent review at PR open time. <a href="https://www.infoq.com/news/2026/04/claude-code-review/">Anthropic's Claude Code Review</a> is the reference implementation, launched March 2026. When a PR opens, a fleet of specialised agents examines the diff — one per class of issue (logic errors, boundary conditions, API misuse, security, project conventions). A verification step tries to disprove each finding before it posts. The surviving findings become inline PR comments the human reviewer sees at review time.</p><p>The design decision worth naming: the review agents are separate from the coding agents. Not because the model cannot review its own work — it can — but because the review discipline benefits from the review agents having a different system prompt, a different set of allowed tools, and a different failure mode. The reviewer that catches a bug is not the writer that made it.</p><p>GitHub Copilot ships a similar surface, tuned for GitHub-native workflows. Third-party tools compete on specific niches (security-focused review, compliance-focused review). The multi-agent shape is the winning pattern across all of them.</p><h3>Slot 4: Autonomous background — bounded scope only</h3><p>The slot that got the most 2026 press and delivers the least reliably. Autonomous coding agents work well when the task is bounded, well-specified, and does not depend on judgement calls that only the senior engineer can make. They work badly when either half of that is missing.</p><p>The tasks that succeed:</p><ul><li><strong>Dependency updates.</strong> Well-scoped, testable, mostly boilerplate. Autonomous agents nail it.</li><li><strong>Test coverage expansion.</strong> Given a code file and a coverage target, an autonomous agent generates the tests. The human reviews the diff.</li><li><strong>Boilerplate migrations.</strong> Framework upgrades, API version changes, style-guide sweeps. Bounded, high-volume, low-judgement — exactly the shape.</li></ul><p>The tasks that fail:</p><ul><li><strong>Anything with a design decision embedded.</strong> The autonomous agent picks a design that looks plausible, gets it wrong in a way that only shows up two weeks later.</li><li><strong>Anything that spans a merge conflict.</strong> Autonomous agents cannot yet reliably navigate a codebase where their assumptions have been invalidated by another engineer's work.</li><li><strong>Anything security-critical.</strong> The autonomous agent cannot own the judgement about acceptable risk. That is a Layer 3 human call — see the <a href="/post/70-percent-ai-generated-codebase">70%-codebase pillar</a>.</li></ul><p>The teams that get autonomous agents to work in 2026 use them for the bounded tasks only, with a strict review discipline on the diff.</p><h3>The cost economics</h3><p>Cost per PR became a real metric in 2026. The moves that change it:</p><p><strong>Prompt caching.</strong> Claude Code and Cursor both support prompt caching. On a repo where the same set of files gets re-read across many AI calls, caching cuts cost by 60-80% — the same numbers from our <a href="/post/ai-agent-token-economics">token economics post</a> applied to the coding case.</p><p><strong>Subagent choice.</strong> A subagent that runs on a smaller, cheaper model for straightforward classification tasks and reserves the top-tier model for hard reasoning saves substantial cost. This is model routing at the coding-tool layer.</p><p><strong>Autonomous vs interactive.</strong> Autonomous agents that run to completion tend to burn more tokens than the equivalent interactive session, because the engineer would have stopped an unproductive path earlier. Not a reason to avoid autonomous — a reason to bound the scope tightly.</p><p>A team running at 70% AI-generation in 2026 typically sees cost-per-merged-PR in the low tens of dollars for CLI-heavy work, low single dollars for IDE-inline-heavy work. Teams that do not track it discover, sometimes at the quarter-end invoice, that the number is much higher than they thought.</p><h3>The tool churn is the real risk</h3><p>The tools we describe in this post will not be the tools we describe in this post twelve months from now. That is the pace of the market. The mitigation is not to pick the tools that will last — it is to pick the discipline that will last.</p><p>The four-slot mental model. The review discipline from the <a href="/post/70-percent-ai-generated-codebase">70%-codebase pillar</a>. The subagent pattern. The cost-tracking. All of these outlive the specific vendors. An engineering leader who invests in the discipline can swap the tools in and out as the market shifts and lose very little. An engineering leader who invests in a specific tool and skips the discipline has to re-learn the pattern every time the market shifts.</p><p>Which is roughly every six months.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
        </item>
        <item>
            <title><![CDATA[Agentic AI workflow services: what's in scope, what's not]]></title>
            <link>https://twistag.com/thinking/agentic-ai-workflow-services-defined</link>
            <guid isPermaLink="false">https://twistag.com/thinking/agentic-ai-workflow-services-defined</guid>
            <pubDate>Tue, 07 Jul 2026 14:19:06 GMT</pubDate>
            <description><![CDATA[Agentic AI workflow services are the engineering capabilities that take agentic AI from pilot to production. Here's what's in scope, what's not, and how delivery actually runs.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: May 2026</em></p><p>Agentic AI workflow services are the engineering capabilities enterprises need to take agentic AI from pilot to production. The category covers multiagent architecture, agent orchestration, agent governance and security, integration with the systems agents need to read from and write back to, and the operating model that lets humans and agents share the same workflow. This post defines the category, draws the line between what's in scope and what isn't, and walks through what a delivery engagement actually contains.</p><h3>Key takeaways</h3><ul><li><strong>Agentic AI workflow services are an engineering category, not an advisory one.</strong> The deliverable is a running system, not a slide deck or a roadmap.</li><li><strong>Three layers define the scope:</strong> integration with existing systems, multiagent architecture and orchestration, and agent governance and security. Skip one and the agent stalls.</li><li><strong>Most agent pilots fail at integration, not at the model.</strong> Only about 12% of enterprise agent initiatives reach production at scale (<a href="https://www.composio.dev/">Composio AI Agent Report 2025</a>), and the failure pattern is consistent: brittle connectors, weak governance, no human-review surface.</li><li><strong>Capability transfer is part of the brief.</strong> A vendor that builds <em>for</em> you leaves a system and a dependency. A team that builds <em>with</em> you leaves a system and a team that can run it.</li><li><strong>The EU AI Act's August 2026 high-risk deadline turns much of this into a procurement question, not a technology question.</strong> Documentation, agent registries, and human oversight have to be in the codebase, not the appendix.</li></ul><h3>What are agentic AI workflow services?</h3><p>Agentic AI workflow services are the engineering capabilities required to design, ship, and operate AI agents inside an enterprise. The term names a category of work that sits between strategy advisory (which produces recommendations) and packaged AI products (which produce a tool, but not the integration into your systems).</p><p>The category exists because the gap between an agent demo and a running production agent is wider than most teams expect. A demo answers a prompt. A production agent reads from real systems, calls real tools, writes back to real databases, fails gracefully under tool errors, gets evaluated continuously, gets governed against compliance obligations, and hands clean state to a human when something escalates. The engineering between those two states is what agentic AI workflow services covers.</p><p>We treat it as three layers:</p><ol><li><strong>Integration with existing systems</strong> — the APIs, connectors, queues, and data plumbing the agent calls into.</li><li><strong>Multiagent architecture and orchestration</strong> — the design that decides which agent does what, with shared state and clear handoffs.</li><li><strong>Agent governance and security</strong> — the guardrails, evaluation, observability, and audit that let an enterprise operate the system day after day.</li></ol><p>A delivery that ships any one of those layers without the other two will not survive contact with production. The three layers are how we scope, estimate, and deliver every engagement on the <a href="https://twistag.com/ai-and-agents">/ai-and-agents</a> service line.</p><h3>What's in scope?</h3><p>Six capability areas sit inside agentic AI workflow services. Each is a real engineering deliverable, not a workshop or a recommendation.</p><p><strong>Multiagent architecture and orchestration.</strong> A multiagent system is a set of specialised agents that hand off subtasks to each other under an orchestration layer. In production it has four parts: the orchestrator that decides which agent does what, the specialised agents each with a narrow tool allowlist, the shared state layer for memory, retrieval, and audit logs, and the human-review surface where exceptions escalate. We design and ship all four. The PepTalk agentic operations system we <a href="https://twistag.com/case-studies/agentic-ai-operations-peptalk-speaker-agency/">built for the speaker agency</a> runs this pattern on AWS Bedrock with pgvector retrieval and Argilla for evaluation.</p><p><strong>Custom agent development.</strong> Domain-specific agents tailored to the enterprise context, not packaged agents or third-party deployments. The work covers LLM integration and fine-tuning, RAG and retrieval pipelines, tool use and function calling, and production deployment with monitoring. The <a href="https://twistag.com/case-studies/ai-invoice-automation-aralab-manufacturing/">Aralab invoice automation agent</a> is a Claude Sonnet 4.5 agent that reads manufacturing invoices, calls into GCP and Firebase, and traces every decision through LangFuse. All of it is tailored to the way Aralab's finance team actually operates.</p><p><strong>Joint human-agent operating models.</strong> Agents alongside humans, not in place of them. The work designs the operating model: agent handoffs, exception escalation, approval chains, audit and review surfaces. The team that runs the workflow has to trust what the agent did. We worked through this pattern across two engagements with <a href="https://twistag.com/case-studies/ai-training-platform-sana-hotels/">SANA Hotels</a> — once for an AI training platform, once for <a href="https://twistag.com/case-studies/ai-workforce-optimization-hotel-group-portugal/">staff optimisation</a> — where the human-review surface and the operations team's daily workflow had to be designed together, not bolted on.</p><p><strong>Capability transfer to internal teams.</strong> The end state is an in-house capability, not a permanent dependency. We embed senior engineers, ship the system, document it, and transfer ownership to the team that will run it. Source code, prompts, evaluation datasets, infrastructure-as-code: all of it transfers. The model is sometimes called build-operate-transfer in adjacent industries; in agentic AI it's the difference between an engagement that compounds value for the client and one that compounds dependency.</p><p><strong>Agent governance, security, and observability.</strong> Production agents read from systems and write back to them. That's a different threat surface than a chatbot. The work covers tool allowlists, output schemas, content filters, golden-dataset evaluation, regression tests in CI, full trace capture, latency and cost dashboards, drift alerts, and audit logs designed to satisfy SOC 2 and the EU AI Act. We treat agents like services, not prompts.</p><p><strong>AI-ready data foundations.</strong> Agents only work as well as the data they read. AI-ready data means clean entity resolution (the agent has to know that "Acme Corp" and "Acme Corporation" are the same record), retrieval-friendly storage (vector indexes for unstructured data, well-modeled tables for structured queries), and governance that an agent can introspect (clear access boundaries, audit logs, schemas the agent can read). This is the work on the <a href="https://twistag.com/data-cloud">/data-cloud</a> service line, and it sits underneath every agentic delivery.</p><h3>What's out of scope (and why this matters)</h3><p>Naming what isn't in scope is more useful than naming what is, because most failed agent programmes died inside the gap between what the buyer thought they were getting and what the engagement actually shipped.</p><p>These are <em>not</em> agentic AI workflow services:</p><ul><li><strong>AI strategy advisory or maturity assessments.</strong> Useful work, but the deliverable is recommendations. No code, no integration, no agent. If a buyer needs a roadmap, the right partner is a strategy firm, not an engineering team.</li><li><strong>Reselling packaged agent products.</strong> Configuring a vendor's pre-built agent inside an enterprise is implementation work, not engineering. It can be valuable, but it's a different scope, with different risks (vendor lock-in, opaque governance).</li><li><strong>One-off LLM integrations without governance.</strong> A chat feature that calls OpenAI and returns text is not an agent. Without tool calls, evaluation, observability, and audit it's a chatbot with extra steps.</li><li><strong>Fine-tuning a model on enterprise data, then leaving.</strong> Fine-tuning is a technique inside an agentic build, not a deliverable on its own. A fine-tuned model with no orchestration, no governance, and no integration is a model card, not a service.</li><li><strong>Pure prompt engineering.</strong> Prompts are the cheapest part of an agent. Treating an engagement as "prompt engineering" undersells the work and produces systems that fail at every other layer.</li><li><strong>Anything involving "agentic RPA."</strong> Wrapping an LLM around a robotic-process-automation script does not produce an agent. It produces a fragile script with non-deterministic failure modes. The category exists to be more rigorous than this, not less.</li></ul><p>The category exists because the buyer is moving on from these adjacent options and is asking for an engineering team that ships the running system. Pretending the category covers any of the above dilutes the answer to the question the buyer actually has.</p><h3>Why the category exists</h3><p>The 2026 conversation is about why agentic AI mostly hasn't shipped yet. McKinsey's December 2025 survey of 200 C-suite executives (<a href="https://www.mckinsey.com/">referenced in the agentic services framing</a>) identified three blockers keeping enterprise agentic AI in pilot: integration complexity, internal expertise gaps, and agent security and governance concerns. The survey reflects what every operator already sees in their own programme.</p><p><strong>Integration complexity.</strong> Most agent pilots stall at the integration layer. The agent can plan, but it can't act because the tools it needs to call live behind APIs nobody documented. Tool calling fails between 3 and 15 percent of the time in production (<a href="https://arize.com/blog/common-ai-agent-failures/">Arize, 2026</a>), and every integration point is a place where that failure rate compounds. The work isn't building a smarter agent. It's building the agent and the integration layer together. The Aralab engagement was an integration project as much as an agent project. The agent ships only because GCP, Firebase, and the manufacturing finance system can be reached cleanly.</p><p><strong>Internal expertise gaps.</strong> Multiagent systems, agent governance, and production-grade evaluation are skills enterprises don't yet have on staff. Hiring is slow. The market for senior agent engineers is tight, and the budget for a full in-house team usually doesn't exist before a few systems have shipped. The pragmatic answer is to embed senior engineers, ship the system, document it, and transfer the capability. That's what capability transfer is designed to do, and why it sits inside the scope rather than next to it.</p><p><strong>Agent security and governance.</strong> Production agents read from systems and write back to them. That's a different threat surface than a chatbot. Microsoft's open-source <a href="https://opensource.microsoft.com/blog/2026/04/02/introducing-the-agent-governance-toolkit-open-source-runtime-security-for-ai-agents/">Agent Governance Toolkit</a> is one of the first vendor-neutral attempts to standardise the runtime controls that production agents need: identity, allowlisting, output validation, audit. The EU AI Act's high-risk obligations from August 2026 turn many of these controls into legal requirements rather than engineering nice-to-haves.</p><p>The category is the engineering response to all three blockers in the same engagement.</p><h3>How agentic AI workflow services compare to adjacent services</h3><p>A buyer choosing between options usually has four on the table. The differences matter, because each option produces a different deliverable, on a different timeline, with different ownership.</p><table><thead><tr><th>Criterion</th><th>Agentic AI workflow services</th><th>AI strategy consulting</th><th>System integrator (SI)</th><th>Packaged agent platform</th></tr></thead><tbody><tr><td><strong>Primary deliverable</strong></td><td>A running production agent + transfer</td><td>Strategy, roadmap, business case</td><td>Configured systems and integrations</td><td>A vendor product, configured</td></tr><tr><td><strong>Time to a production pilot</strong></td><td>6–12 weeks for a focused MVP</td><td>N/A — output is a plan</td><td>6–18 months typical</td><td>Days to weeks (configuration only)</td></tr><tr><td><strong>IP and code ownership</strong></td><td>Client owns full stack at handover</td><td>Client owns the deliverable (the slides)</td><td>Mixed; SI usually retains some IP</td><td>Vendor owns the platform</td></tr><tr><td><strong>Capability transfer to in-house team</strong></td><td>Built into the engagement</td><td>Not part of the deliverable</td><td>Sometimes, often as a separate SOW</td><td>Limited — vendor-managed</td></tr><tr><td><strong>Vendor lock-in risk</strong></td><td>Low — model and framework agnostic</td><td>None</td><td>Medium to high</td><td>High</td></tr><tr><td><strong>Best fit for</strong></td><td>Validated use case, needs production engineering</td><td>Pre-validation, exec alignment</td><td>Large multi-system rollouts</td><td>Standardised use cases</td></tr></tbody></table><p>The comparison is not a value judgment. Each option is right for a different question. Agentic AI workflow services is the right answer when the use case is validated, the buyer wants a running system fast, and the destination is an in-house capability the team can operate without the partner.</p><h3>What a delivery engagement actually looks like</h3><p>Most agentic AI workflow services engagements share the same shape, even when the use case is different. We run them in four phases.</p><p><strong>Phase 1 — Discovery and architecture (1–2 weeks).</strong> A senior agent engineer and a senior data engineer scope the use case, the integrations, the data sources, the governance obligations, and the in-house team that will eventually run the system. The output is a one-page architecture, a named risk list, and a delivery plan with concrete milestones. The buyer always talks to the engineer who will lead the work, not a sales person.</p><p><strong>Phase 2 — Production MVP (6–10 weeks).</strong> The team ships the agent into production for a focused scope: narrow use case, one or two integrations, one user role. Multiagent architecture if the scope calls for it, single-agent if it doesn't. Evaluation, observability, and the human-review surface ship in this phase, not after — they're how we know the agent works.</p><p><strong>Phase 3 — Scale and capability transfer (3–6 months).</strong> Multi-team rollout, additional roles, additional integrations, and continuous capability transfer to the in-house team. Pair working, documentation, and review of the in-house team's first production changes. Governance hardens against the obligations the use case is subject to (SOC 2, EU AI Act, sector-specific).</p><p><strong>Phase 4 — Ongoing partnership or full handover.</strong> The end state is the buyer's choice. Some clients keep us embedded for years. The <a href="https://twistag.com/case-studies/datatalks-customer-data-platform-sports/">Datatalks engagement</a> is in its sixth year and still shipping. Others transfer fully and call us back when the next system needs building. Both are good outcomes when the call is the buyer's.</p><p>The phases are not sacred. The principle behind them is: ship something real every phase, transfer capability throughout, and let the buyer decide where the engagement ends.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-agentic-ai-workflow-services-defined.jpg" length="0" type="image/jpg"/>
        </item>
        <item>
            <title><![CDATA[How long does it take to ship an AI agent to production?]]></title>
            <link>https://twistag.com/thinking/how-long-to-ship-an-ai-agent</link>
            <guid isPermaLink="false">https://twistag.com/thinking/how-long-to-ship-an-ai-agent</guid>
            <pubDate>Sat, 04 Jul 2026 17:31:11 GMT</pubDate>
            <description><![CDATA[95% of GenAI pilots never reach production. The honest answer depends on five variables. Real phase-by-phase timelines from production agents we shipped.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: June 2026</em></p><p>The honest answer is that it depends on five variables, and the five variables are knowable in the first week. A focused production MVP ships in six to ten weeks. A multi-team enterprise rollout runs three to six months on top of that. Most pilots that drag past twelve weeks never graduate to production at all. This post breaks down the timeline by phase, names the five scope variables that move it, and walks through the actual delivery shape of three production agents we shipped. It's a cluster child of our pillar on <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a>, and it picks up where the <a href="/post/why-ai-agents-fail-in-production">failure-mode analysis in our production-failures post</a> explained what stops the pilot from reaching day two.</p><h3>Key takeaways</h3><ul><li><strong>A production MVP ships in 6-10 weeks</strong> for a focused scope: one workflow, one or two integrations, one user role. Multi-team enterprise rollout adds three to six months for evaluation, observability, access control, and capability transfer.</li><li><strong>95% of GenAI pilots never reach production</strong> (<a href="https://web.media.mit.edu/~bcohen/projects/nanda/">widely cited MIT NANDA estimate, 2025</a>). The reason is rarely the model. It is integration debt, governance gaps, and a 12-week pilot window that closed without a path to production.</li><li><strong>The 12-week graduation rule.</strong> Pilots that pass week 12 without a concrete production handover plan tend not to graduate. The number is empirical across enterprise deployments — beyond that window, the cost-of-attention math turns against the project.</li><li><strong>Five variables move the timeline.</strong> Integration count, data readiness, governance burden, evaluation depth, and team familiarity. Everything else is downstream of these five.</li><li><strong>The build phase is rarely what slows you down.</strong> It is Phase 2 — data and integration readiness — that runs over budget on most projects.</li></ul><h3>The five variables that move the timeline</h3><p>Every CTO asking "how long" is actually asking five sub-questions. Naming them up front turns a vague estimate into a scope conversation.</p><p><strong>Integration count.</strong> Each external system the agent reads from or writes back to adds days, sometimes weeks. A single-integration agent (read a knowledge base, return an answer) ships fastest. An agent that calls into a CRM, a finance system, an inventory database, and a customer email service costs four integration cycles. Most agentic workflows we ship at production sit at two to four integrations.</p><p><strong>Data readiness.</strong> Is the data the agent needs already in a form it can read? Clean entity resolution, retrieval-friendly storage, agent-readable schemas — these are the data-platform work covered in the data foundations of agentic AI workflow services. Where the data is not ready, Phase 2 (below) grows by weeks. Most timeline overruns we see come from this variable, not from the build itself.</p><p><strong>Governance burden.</strong> A consumer agent and a regulated-industry agent have wildly different governance budgets. Tool allowlists, output schemas, audit logs, agent registry, EU AI Act mapping — the <a href="/post/agentic-ai-workflow-services-defined">five engineering surfaces in our governance reference</a> all take engineering time. A high-risk agent under EU AI Act obligations adds two to four weeks of governance work that a low-risk internal agent skips.</p><p><strong>Evaluation depth.</strong> A workflow where wrong answers cost the business money needs golden datasets, online LLM-as-a-judge evaluators, and a human-review queue, all built before launch. A lower-stakes workflow ships with lighter eval and matures it in production. The choice is a deliberate trade-off, not an afterthought.</p><p><strong>Team familiarity.</strong> Is the in-house team that will eventually own the agent already familiar with the patterns, or is this their first agentic system? First-time teams need more pair-working time during the build, which slows the partner's pace but accelerates the eventual transfer. Returning teams ship faster.</p><h3>Phase 1: Discovery and architecture (1-2 weeks)</h3><p>A senior agent engineer and a senior data engineer scope the use case, the integrations, the data sources, the governance obligations, and the in-house team that will eventually run the system. The output is a one-page architecture diagram, a named risk list, and a delivery plan with concrete weekly milestones. The buyer always talks to the engineer who will lead the work, not a sales person.</p><p>The scope conversation answers the five variables. If any answer is unresolvable in week one, that itself is a finding — the use case is not validated enough for an MVP engagement and the right next step is a discovery sprint rather than a build.</p><h3>Phase 2: Production MVP (6-10 weeks)</h3><p>The team ships the agent into production for a focused scope. Narrow use case, one or two integrations, one user role. Multiagent architecture if the scope calls for it, single-agent if it doesn't. Evaluation, observability, and the human-review surface ship in this phase, not after — they are how we know the agent works.</p><p>A six-week MVP is realistic when integrations are one or two, data is reasonably ready, and governance burden is moderate. A ten-week MVP is the upper bound when integrations are three or four, or when governance is heavy (regulated industry, EU AI Act high-risk). Pushing past ten weeks usually means a variable was mis-scoped in Phase 1, and the right response is a checkpoint conversation rather than a quiet slip.</p><p>The MVP is in production — real users, real traffic, real consequences for wrong answers — at the end of this phase. Not a sandbox demo. Not a staging environment. Production.</p><h3>Phase 3: Scale and capability transfer (3-6 months)</h3><p>Multi-team rollout, additional user roles, additional integrations, and continuous capability transfer to the in-house team. Pair working, documented decisions, code reviews of the in-house team's first production changes. Governance hardens against the obligations the use case is subject to (SOC 2, EU AI Act high-risk obligations from August 2026, sector-specific).</p><p>This phase is also when the eval pipelines accumulate enough production data to be genuinely informative. The golden dataset grows. The LLM-as-a-judge calibration tightens. The human-review queue surfaces patterns that feed the next sprint of improvements.</p><p>Three to six months is the typical band. Heavier rollouts — multi-region, multi-business-unit, multi-language — extend toward six months. Tighter rollouts complete in three.</p><h3>Phase 4: Hardening and ongoing</h3><p>The exit signal is the in-house team shipping a production change without the partner's involvement. From that point the system is operated and extended by the buyer's team, with the partner available for the next adjacent system if and when the buyer calls. Hardening — performance optimisation, cost optimisation, governance evolution as regulations change — continues indefinitely, owned by the in-house team.</p><h3>Three real timelines</h3><p>The phases above are not theoretical. Three production agents we shipped:</p><p>The <a href="/case-studies/ai-invoice-automation-aralab-manufacturing/">Aralab invoice automation agent</a> is a Claude Sonnet 4.5 agent reading manufacturing invoices and writing back to GCP, Firebase, and Aralab's manufacturing finance system. Three integrations, moderate governance, returning team. Phase 1 took ten days. Phase 2 (MVP into production with Aralab's finance team) ran nine weeks. Phase 3 (scale + capability transfer to Aralab's engineering team) ran four months. Total from kickoff to handover: roughly seven months. The agent has been live for over a year and Aralab's team has shipped multiple production extensions without us.</p><p>The <a href="/case-studies/agentic-ai-operations-peptalk-speaker-agency/">PepTalk agentic operations system</a> is a multiagent architecture on AWS Bedrock with pgvector retrieval and Argilla for evaluation. The scope was wider — multi-agent orchestration across the speaker-engagement workflow — and that mattered. Phase 1 took two weeks. Phase 2 ran ten weeks (the upper bound, because the multi-agent eval pipeline was non-trivial). Phase 3 ran five months. The system is now operated by PepTalk with us alongside on adjacent work.</p><p>The <a href="/case-studies/ai-training-platform-sana-hotels/">SANA Hotels AI training platform</a> used AI avatars and Anthropic Claude to personalise training delivery. Two integrations, moderate governance, first-time team. Phase 1 took ten days. Phase 2 ran eight weeks. Phase 3 ran three months. A second engagement for <a href="/case-studies/ai-workforce-optimization-hotel-group-portugal/">staff optimisation</a> followed and shipped faster because the team was now familiar.</p><h3>What makes a timeline slip</h3><p>Three patterns account for most slips.</p><p><strong>Data is less ready than the buyer thinks.</strong> This is the dominant cause. Phase 1 should surface it and either replan Phase 2 around data-foundation work or sequence the engagement so the data work runs in parallel. Slips happen when Phase 1 was too short to catch the gap.</p><p><strong>Governance scope expanded mid-build.</strong> A change in the use case (now the agent touches customer PII, now it makes recommendations that affect credit, now it is in scope for the EU AI Act) adds engineering weeks that were not budgeted. The fix is a checkpoint conversation with explicit scope amendment, not a quiet timeline drift.</p><p><strong>The in-house team has less bandwidth than promised.</strong> Capability transfer requires the in-house team's attention. When the team is also running production for the existing systems, the transfer slows. The fix is honest bandwidth planning in Phase 1.</p><h3>The budget is the lock, not the timeline</h3><p>The CTO's question "how long does it take?" usually has a second question behind it: how much will it cost? In agentic AI delivery, those two are tightly coupled — engineering weeks are the largest cost line — and the 12-week graduation rule applies to budget too. Pilots that drag past twelve weeks without a production path tend to consume the entire pilot budget and produce nothing operationally useful.</p><p>The shape of the conversation that works is the inverse. Lock the budget. Pick a scope that fits inside the six-to-ten-week MVP band. Ship into production. Then decide whether to extend. The phases above are designed for this conversation, and the agents we have shipped this way are the proof that the shape works.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-how-long-to-ship-an-ai-agent.jpeg" length="0" type="image/jpeg"/>
        </item>
        <item>
            <title><![CDATA[Production AI agent governance: a SOC 2 / EU AI Act-aware reference]]></title>
            <link>https://twistag.com/thinking/production-ai-agent-governance</link>
            <guid isPermaLink="false">https://twistag.com/thinking/production-ai-agent-governance</guid>
            <pubDate>Sat, 04 Jul 2026 17:31:11 GMT</pubDate>
            <description><![CDATA[Audits don't fail because the agent was bad. They fail because nobody can answer questions about it. An engineer's reference for production AI agent governance.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: June 2026</em></p><p>Audits don't fail because the agent was bad. They fail because nobody can answer questions about it. Which model version produced this decision. What data did it see. Who approved this prompt change. Was there human oversight when the agent took this irreversible action. Most agent governance content reads as a compliance lawyer's view of these questions. This post is the architect's view: the five engineering surfaces that make production governance possible, how they map to the <a href="https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/">OWASP Top 10 for Agentic Applications 2026</a>, and how they satisfy the EU AI Act high-risk obligations that take effect August 2026 and the SOC 2 trust services criteria that show up in every enterprise procurement. It's a child of our pillar on <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a> and it picks up where the <a href="/post/why-ai-agents-fail-in-production">governance-gaps failure mode in our production-failures post</a> left off.</p><h3>Key takeaways</h3><ul><li><strong>Governance is five engineering surfaces, not a compliance document.</strong> Identity, tool allowlists and output schemas, evaluation, observability, and audit. Build all five and an audit becomes a query, not a fire drill.</li><li><strong>The OWASP Top 10 for Agentic Applications 2026</strong> (peer-reviewed by 100+ industry experts) names the ten threats: planning, tool use, identity, supply chain, code execution, memory, inter-agent communication, cascading failures, human-agent trust, and rogue agents. Treat it as your threat checklist.</li><li><strong>Microsoft's <a href="https://opensource.microsoft.com/blog/2026/04/02/introducing-the-agent-governance-toolkit-open-source-runtime-security-for-ai-agents/">Agent Governance Toolkit</a></strong> (open-source, April 2026) is the first stack to cover all ten OWASP threats with sub-millisecond runtime policy enforcement. It is also the first toolkit with explicit EU AI Act and SOC 2 compliance grading built in.</li><li><strong>The EU AI Act high-risk obligations take effect August 2026.</strong> Agentic systems often touch high-risk categories — credit, hiring, regulatory reporting, infrastructure operation. Documentation, agent registry, and stop-and-revoke controls have to be in the codebase, not the appendix.</li><li><strong>Governance is not a launch follow-up.</strong> Teams that bolt it on after a successful pilot are six months late to the actual scaling problem. Build it in from sprint one and the pilot is the production system.</li></ul><h3>What governance actually means in code</h3><p>The word "governance" in enterprise AI conversations usually means a slide deck. In production it means specific, observable engineering decisions made in five places. We treat these as five surfaces because that is what they look like when you architect them: each one is a layer with concrete artefacts, deployment patterns, and audit queries.</p><p>The five surfaces are: <strong>agent identity</strong>, <strong>tool allowlists and output schemas</strong>, <strong>evaluation</strong>, <strong>observability</strong>, and <strong>audit</strong>. Skipping any one of them is the structural reason most production agents fail compliance reviews. We saw it in the governance-gaps failure mode across multiple engagements, and we now build all five from the architecture phase forward.</p><h3>Surface 1: Agent identity — who is this thing, what does it know, who owns it</h3><p>Every production agent has an identity that exists in three places at once: in an agent registry the team can query, in an identity system that resolves to a service principal or workload identity, and in the audit trail next to every action the agent takes.</p><p>The agent registry is a list — usually a CMDB entry or a Backstage component — of every agent running in the environment. For each agent: a unique ID, the team that owns it, the tools it is allowed to call, the data sources it is allowed to read, the model versions it has run on, the prompt versions it has used. The registry is queryable. When an auditor asks "which agents are touching customer PII," the answer is a single query, not a Slack thread.</p><p>The identity in the workload sense matters because agents act on behalf of users, on behalf of other agents, or on behalf of the system itself. Tracking which is the OWASP Agentic Top 10's third risk (ASI03 — identity). Without an agent identity, every action looks like a privileged service call from the same source, and there is no way to revoke one agent without affecting the rest.</p><h3>Surface 2: Tool allowlists and output schemas — what can the agent do, what can it return</h3><p>Production agents do not have unbounded access to tools. They have an allowlist declared at configuration time and enforced at runtime. The allowlist names exactly which functions the agent can call, with what argument shapes, with what authorisation scope. Anything not on the list is rejected at the orchestration layer before the agent ever sees the option.</p><p>Microsoft's Agent Governance Toolkit ships this layer as a stateless policy engine with sub-millisecond enforcement latency (p99 below 0.1ms). It hooks into LangChain callbacks, CrewAI task decorators, Google ADK plugins, and the Microsoft Agent Framework middleware pipeline. The pattern is the same regardless of toolkit: every tool call is intercepted, validated against the allowlist, and either passed through or denied with a structured event.</p><p>Output schemas matter for the inverse reason. The agent's response to the LLM has to fit a schema before any downstream action takes it as input. JSON Schema with required fields, allowed enums, length and format constraints. If the LLM returns a malformed response, the schema validator rejects it, the agent retries, and the human-review queue catches the long tail. The <a href="/case-studies/ai-invoice-automation-aralab-manufacturing/">Aralab invoice automation agent</a> routes around 2-3% of invoices into a structured human-review queue because their output failed the schema — and that is what keeps the other 97% trustworthy enough to write back to the manufacturing finance system without human review.</p><h3>Surface 3: Evaluation — golden datasets, online judges, regression in CI</h3><p>Evaluation is the surface most teams under-invest in. Three layers, all running continuously:</p><ol><li><strong>Golden datasets in CI.</strong> A fixed test set of input-output pairs that every prompt or model change has to pass before deploy. When a developer updates a prompt, the golden set runs in CI; if the pass rate drops, the deploy is blocked. The set evolves over time as new failure modes get added.</li><li><strong>Online LLM-as-a-judge evaluators.</strong> A separate LLM scores live agent responses against quality criteria, sampled across production traffic. The score is logged with the trace, and a drop triggers an alert. This is how you catch the regression that the golden set missed.</li><li><strong>Human review of edge cases.</strong> Cases the judge flagged as low confidence, or that landed in a fall-back queue, get reviewed by a human. The reviewed cases feed back into the next version of the golden set.</li></ol><p>The combination gives you a deterministic baseline (golden), a continuous quality signal (LLM judge), and a human ground truth (review). All three are required to satisfy the EU AI Act Article 9 risk management obligations, which expect ongoing performance monitoring rather than a one-time pre-deployment evaluation.</p><h3>Surface 4: Observability — every decision emits a trace</h3><p>Standard infrastructure monitoring tells you the agent service is running. It does not tell you what the agent did, why, or whether the result was right. For that you need agent observability: every planning step, every tool call, every retrieval, every model invocation, every output emits a structured trace event.</p><p>The industry has converged on <a href="https://opentelemetry.io/">OpenTelemetry</a> as the wire protocol and on platforms like Langfuse for the trace store and the UI. Each trace event includes the agent identity, the user (or upstream agent) on behalf of whom the action ran, the model version, the prompt version, the tool definition, the inputs, the outputs, and the cost in tokens and dollars. Traces are correlated across the multi-step plan, so an analyst can replay a failed task from prompt to outcome.</p><p>Observability is also what makes the OWASP Top 10's eighth risk (ASI08 — cascading failures) detectable. When agent A's misclassification leads agent B down the wrong workflow path, the trace shows you the exact handoff and the confidence signal that should have triggered a human review. Without traces, the failure is silent.</p><h3>Surface 5: Audit — append-only logs designed for regulator questions</h3><p>Audit is what observability becomes when the consumer is a regulator instead of an engineer. The same trace events are written to an append-only log with the properties an audit demands: cryptographic chain integrity (each entry includes a hash of the previous), time-of-action timestamps that cannot be backdated, and immutable retention for the period the applicable framework requires (SOC 2 commonly seven years for financial-impact controls, EU AI Act ten years for high-risk systems).</p><p>The audit log answers regulator questions directly: which agent decision affected this customer, which model version produced it, what data did the agent see, what tool calls did it make, was a human in the loop. The <a href="/case-studies/ai-communications-assistant-uk-water-utility/">AI communications assistant we built for a UK water utility</a> operates in a sector where every customer-facing communication is auditable, and we had to ship a working agent whose every decision was traceable to an input, a model version, a prompt, and a tool call. Audit was not a launch follow-up — it was the launch criterion.</p><h3>How the five surfaces map to the OWASP Agentic Top 10</h3><p>The <a href="https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/">OWASP Top 10 for Agentic Applications 2026</a> is peer-reviewed by over 100 industry experts and names ten distinct threat categories. The five engineering surfaces above cover all ten.</p><table><thead><tr><th>OWASP threat (2026)</th><th>Primary surface</th><th>Notes</th></tr></thead><tbody><tr><td>ASI01 — Planning</td><td>Evaluation + observability</td><td>Golden-set tests on multi-step plans; trace replay</td></tr><tr><td>ASI02 — Tool use</td><td>Tool allowlists</td><td>Allowlists, schema validation, scope enforcement</td></tr><tr><td>ASI03 — Identity</td><td>Agent identity</td><td>Registry + workload identity + audit attribution</td></tr><tr><td>ASI04 — Supply chain</td><td>Identity (registry) + audit</td><td>Model and dependency provenance tracked in registry, audited on change</td></tr><tr><td>ASI05 — Code execution</td><td>Tool allowlists + runtime sandbox</td><td>Allowlist denies non-approved code execution; runtime ring isolation</td></tr><tr><td>ASI06 — Memory</td><td>Tool allowlists + audit</td><td>Memory write scopes are tool definitions; memory reads are traced</td></tr><tr><td>ASI07 — Inter-agent communication</td><td>Identity + audit</td><td>Each agent-to-agent handoff is an identified action in the trace</td></tr><tr><td>ASI08 — Cascading failures</td><td>Evaluation + observability</td><td>Whole-pipeline golden tests; trace correlation across handoffs</td></tr><tr><td>ASI09 — Human-agent trust</td><td>Evaluation + observability</td><td>Human-review surface; confidence thresholds; replay</td></tr><tr><td>ASI10 — Rogue agents</td><td>Identity + audit</td><td>Registry deviation triggers; kill switch from the registry</td></tr></tbody></table><p>Microsoft's <a href="https://github.com/microsoft/agent-governance-toolkit">Agent Governance Toolkit</a> is the first open-source stack to address all ten with deterministic runtime enforcement. We use it on engagements where the toolkit ecosystem (LangChain, CrewAI, Microsoft Agent Framework) matches the rest of the buyer's stack, and we use the same patterns with custom orchestration where the buyer's stack is different. The patterns matter more than the toolkit.</p><h3>How the five surfaces map to EU AI Act and SOC 2</h3><p>The EU AI Act's high-risk AI obligations take effect August 2026. The Colorado AI Act becomes enforceable June 2026. SOC 2 shows up in every enterprise procurement regardless of geography. The five surfaces satisfy the technical obligations across all three.</p><p><strong>EU AI Act Article 9 — risk management system.</strong> Continuous evaluation across the lifecycle. <em>Satisfied by:</em> evaluation surface (golden + online + human review).</p><p><strong>EU AI Act Article 12 — record-keeping and logs.</strong> Automated logging across the lifetime of the high-risk system. <em>Satisfied by:</em> audit surface (append-only, cryptographic chain, ten-year retention).</p><p><strong>EU AI Act Article 13 — transparency and information to users.</strong> Operators understand and use the system properly. <em>Satisfied by:</em> agent registry surface (which agents are running, what they can do, who owns them).</p><p><strong>EU AI Act Article 14 — human oversight.</strong> Effective oversight by natural persons. <em>Satisfied by:</em> tool allowlists (output schemas force human review for low-confidence outputs) + audit (human-in-the-loop decisions are logged with attribution).</p><p><strong>SOC 2 Common Criteria CC6 (logical access)</strong> and <strong>CC7 (system operations).</strong> Access controls, monitoring, change management. <em>Satisfied by:</em> identity + tool allowlists + audit.</p><p><strong>SOC 2 CC9 (risk mitigation).</strong> Risk identification and response. <em>Satisfied by:</em> evaluation + observability.</p><p>Microsoft's <a href="https://opensource.microsoft.com/blog/2026/04/02/introducing-the-agent-governance-toolkit-open-source-runtime-security-for-ai-agents/">Agent Compliance package</a> ships pre-built compliance grading against these frameworks, mapped to the toolkit's runtime telemetry. We do not need a third party to grade us; we run the grading in CI alongside the unit tests.</p><h3>What this looks like in regulated case studies</h3><p>Three Twistag engagements where the five surfaces were the engagement.</p><p>The <a href="/case-studies/regtech-platform-european-ingredient-brands/">RegTech platform we built for European ingredient brands</a> is the clearest case. Every agent decision had to map to a specific regulatory clause and produce evidence on demand. The audit surface was not a logging layer next to the agent — it was the product. The agent registry, the tool allowlists, and the evaluation pipelines were what made the regulatory product credible to its enterprise customers.</p><p>The <a href="/case-studies/ai-invoice-automation-aralab-manufacturing/">Aralab invoice automation agent</a> runs Claude Sonnet 4.5 with LangFuse for traces and a structured human-review queue for outputs that fail schema validation. The audit surface answers Aralab's finance team's questions ("which invoice did the agent reject and why?") and would answer a tax authority's questions if asked. The same trace events feed both queries.</p><p>The <a href="/case-studies/ai-communications-assistant-uk-water-utility/">AI communications assistant for a UK water utility</a> operates in a regulated sector where every customer-facing communication is auditable. We could not ship the agent without the audit surface being load-bearing from sprint one. A human reviews any outbound communication that fails the confidence threshold; the review is logged with attribution; the agent's input, model version, prompt, and tool calls are all queryable by the regulator if asked. The agent has been in production since 2024 because the governance shipped with it.</p><h3>The second-day problem</h3><p>Governance is the engineering that decides whether day two of the production agent's life is possible. Most agent pilots that fail did not fail on day one — they failed when day two arrived in the form of an audit, a regression in quality, a regulatory deadline, or a model upgrade that broke the prompt. The teams that ship into regulated environments have learned to treat governance as a launch criterion, not a launch follow-up.</p><p>The five surfaces are the operating shape of that decision. Identity makes the system queryable. Tool allowlists and output schemas make the actions constrained. Evaluation makes quality measurable. Observability makes failure debuggable. Audit makes the regulator's question answerable. Build all five and the agentic system is something the enterprise can operate. Skip any one of them and day two is when the programme stops shipping.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-production-ai-agent-governance.jpeg" length="0" type="image/jpeg"/>
        </item>
        <item>
            <title><![CDATA[Five engineering reasons AI agents fail in production]]></title>
            <link>https://twistag.com/thinking/why-ai-agents-fail-in-production</link>
            <guid isPermaLink="false">https://twistag.com/thinking/why-ai-agents-fail-in-production</guid>
            <pubDate>Sat, 04 Jul 2026 17:31:11 GMT</pubDate>
            <description><![CDATA[Only 12% of enterprise AI agents reach production. The five engineering failure modes we see most often, and the patterns that fix them, with real examples.]]></description>
            <content:encoded><![CDATA[<p><em>Last updated: May 2026</em></p><p>About 97% of executives report deploying AI agents over the past year, but only 12% of those initiatives reach production at scale (<a href="https://www.composio.dev/">Composio AI Agent Report 2025</a>). The gap is an engineering problem, and the failure modes are predictable. This post walks through the five we see most often, the case studies where we ran into them, and the patterns that fix them. It's the postmortem companion to our pillar on <a href="/post/agentic-ai-workflow-services-defined">agentic AI workflow services</a>.</p><h3>Key takeaways</h3><ul><li><strong>The failure rate is structural, not technical.</strong> McKinsey finds that 60% of AI projects fail because of workflow misalignment and governance gaps, not model performance.</li><li><strong>Three patterns dominate the failure data:</strong> poor retrieval (dumb RAG), broken tool integrations (brittle connectors), and errors that compound across multi-step plans (<a href="https://www.shakudo.io/blog/enterprise-ai-agent-production-failures">Shakudo, 2026</a>). Governance gaps and context engineering misses show up later.</li><li><strong>Tool calling fails between 3 and 15 percent of the time in production</strong> (<a href="https://arize.com/blog/common-ai-agent-failures/">Arize, 2026</a>). Plan for retries, schema validation, and a human-review surface for the long tail.</li><li><strong>The engineering response is consistent across failure modes:</strong> observable, evaluated, governed, recoverable. Treat agents like services, not prompts.</li></ul><h3>1. Dumb RAG: retrieval that demos well and breaks under real queries</h3><p><strong>The failure mode.</strong> Retrieval-augmented generation gets built once, against a curated dataset, with no evaluation pipeline. The system pulls whatever the embedding model returns, feeds it to the LLM, and trusts the answer. It demos beautifully. Then real users arrive with real queries and the agent surfaces wrong products, irrelevant documents, or stale context. The failure is silent — the agent confidently produces an answer that happens to be grounded in the wrong source.</p><p><strong>Where we saw it.</strong> Any retrieval-heavy build is at risk. On the <a href="/case-studies/conversational-ecommerce-ai-product-discovery/">conversational commerce engagement we ran for a global sportswear brand</a>, the question wasn't "does retrieval work" but "does it return the <em>right</em> product when a shopper asks in their own language." Pinecone returned semantically similar matches; the question was whether that matched buyer intent.</p><p><strong>The engineering response.</strong> Hybrid retrieval (vector plus keyword plus filters) almost always beats pure vector search for commercial queries. Chunking gets evaluated, not assumed. Every retrieval gets logged with the query, the chunks returned, and a relevance score from an LLM-as-a-judge evaluator. The eval pipeline becomes the gate: a retrieval change ships when the golden-dataset score holds and the LLM-judge score on live traffic does not regress. The fix is not a smarter embedding model — it's treating retrieval as a system you measure, not a library you call.</p><h3>2. Brittle connectors: the integration-layer tax</h3><p><strong>The failure mode.</strong> The agent can plan but cannot act. The tools it needs live behind APIs nobody documented, schemas drift without notice, and a single failed call propagates into a wrong answer because nothing validates what came back. Tool calling fails between 3 and 15 percent of the time in production (<a href="https://arize.com/blog/common-ai-agent-failures/">Arize, 2026</a>). At ten tool calls per task, those failure rates compound into double-digit task failure if nothing handles them.</p><p><strong>Where we saw it.</strong> The <a href="/case-studies/ai-invoice-automation-aralab-manufacturing/">Aralab invoice automation agent</a> is a Claude Sonnet 4.5 agent that reads manufacturing invoices and writes back to GCP, Firebase, and the manufacturing finance system. The agent was the easier half; the harder half was making the integration layer survive real-world messiness: malformed invoices, intermittent third-party outages, schema drift. Without that discipline, the agent shipped a clean demo and quietly mishandled the long-tail invoices that don't fit the happy path.</p><p><strong>The engineering response.</strong> Wrap every tool with an explicit interface, not raw API calls. Validate inputs and outputs against schemas. Retry idempotent calls with backoff and a circuit breaker. Route anything that doesn't validate into a structured human-review queue rather than letting the agent invent a plausible answer. Log every tool call and every fall-back as a first-class trace event. The unsexy framing matters: agents that work in production look more like reliable distributed systems than like prompt engineering.</p><h3>3. Compounding errors in multi-step plans</h3><p><strong>The failure mode.</strong> The agent makes a small mistake on step one. Step two builds on step one. Step three builds on step two. By step five, the agent is confidently executing the wrong plan because nothing along the way questioned whether step one was right. The pattern is particularly vicious in multiagent systems, where each agent treats the prior agent's output as ground truth.</p><p><strong>Where we saw it.</strong> The <a href="/case-studies/agentic-ai-operations-peptalk-speaker-agency/">PepTalk agentic operations system</a> runs as a multiagent architecture on AWS Bedrock with pgvector for retrieval and Argilla for evaluation. Multiple specialised agents handle different parts of the speaker-engagement workflow. The early version produced impressive whole-pipeline demos and occasional cliff-edge failures: when one agent misclassified an intent at the top of the chain, the downstream agents executed the wrong workflow with full confidence.</p><p><strong>The engineering response.</strong> Explicit completion signals, not implicit ones. Each agent declares what success looks like for its step before the orchestrator advances. Validated handoffs: the orchestrator inspects the output schema and the confidence signal, and routes to a human if either is below threshold. Central state instead of distributed state, so the audit trail is one trace, not five overlapping ones. Golden-dataset evaluation runs the whole pipeline, not just individual agents in isolation. In a multiagent system, the orchestrator's job is to be sceptical of every agent it manages.</p><h3>4. Governance gaps: the second-day problem</h3><p><strong>The failure mode.</strong> The pilot ships. It works. Six months later, an audit lands and the team can't answer basic questions: which model version produced this decision, what data did it see, who approved this prompt change, was there human oversight when the agent made this irreversible call. The agent didn't fail technically. It failed the governance test, and the result is the same — pulled from production, programme cancelled, trust lost.</p><p><strong>Where we saw it.</strong> Regulated industries make this failure mode obvious early. The <a href="/case-studies/ai-communications-assistant-uk-water-utility/">AI communications assistant we built for a UK water utility</a> operates in a sector where every customer-facing communication is auditable. We couldn't ship a working agent. We had to ship a working agent whose every decision was traceable to an input, a model version, a prompt, and a tool call, and where a human could intervene before any outbound communication landed.</p><p><strong>The engineering response.</strong> Build governance into the agent, not next to it. Tool allowlists declared at config time and enforced at runtime. Output schemas the LLM has to fit into before any outbound action. Golden-dataset evaluation in CI on every prompt or model change. Full trace capture through OpenTelemetry into Langfuse or an equivalent. An agent registry the team can query: which agents are running, what tools they can call, who owns them. Audit logs designed from the start to satisfy SOC 2 and the EU AI Act's <a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai">August 2026 high-risk obligations</a>. The EU AI Act is not adding a new failure mode; it's making an existing one legally enforceable.</p><h3>5. Context engineering: the silent half of agent reliability</h3><p><strong>The failure mode.</strong> The agent has the right tools, governance, and data, and still produces wrong answers because the context it sees at decision time is incomplete, stale, or contradictory. Conversation history pollutes the prompt. Memory leaks across user sessions. Integration tests pass; the failure is in what the agent <em>knows</em> at the moment it acts.</p><p><strong>Where we saw it.</strong> The <a href="/case-studies/ai-training-platform-sana-hotels/">SANA Hotels AI training platform</a> personalises training delivery using AI avatars and Anthropic Claude. The agent has to track each learner's progress, adapt to their level, and avoid repeating content they've mastered. The early failure mode was not technical: the agent occasionally treated a previous learner's context as the current learner's, because session boundaries weren't enforced rigorously enough. The agent was working as designed; the design was wrong about what context belonged to whom.</p><p><strong>The engineering response.</strong> Context-isolated tasks: every interaction starts from an explicit, scoped context object, not from accumulated conversation. Managed conversation history with deliberate truncation, not whatever fits in the window. Prompt versioning so changes can be rolled back when a quality regression appears. And the part most teams skip: eval datasets that include adversarial cases for context contamination, so the test suite actually exercises the failure mode. This is where senior engineering judgment matters most, because the failure modes don't appear in unit tests.</p><h3>What the five failures have in common</h3><p>The engineering response converges on four properties. A production agent is <strong>observable</strong> (structured traces of every decision, tool call, and output), <strong>evaluated</strong> (golden datasets in CI plus online quality scoring plus human review), <strong>governed</strong> (tool allowlists, output schemas, audit logs, agent registry), and <strong>recoverable</strong> (retries, validation gates, no silently compounding errors). The failure modes vary by case study and industry; the engineering response does not. A team inheriting the system needs these patterns obvious, not buried.</p>]]></content:encoded>
            <author>hello@twistag.com (Twistag)</author>
            <enclosure url="https://twistag.com/images/cms/post-why-ai-agents-fail-in-production.jpg" length="0" type="image/jpg"/>
        </item>
    </channel>
</rss>