Framework · AI & Business Models

The Cement Curve: AI Coding and the End of the Managed Margin

In two earlier notes we argued that AI deflates capability and concentrates value onto proprietary substrates, and that the market's error was filing every software business under capability. The same framework produces a harder conclusion when applied to the compute stack. A cloud provider sells raw infrastructure and the managed services layered on top of it; the margin lives in the software. The premium of SageMaker over the EC2 instance underneath it, or DynamoDB over a database you run yourself, was always a labour arbitrage. AI coding collapses the labour side of that bargain. When an agent can assemble the orchestrator, inference server, IAM policy and patch cycle from open-source components, the managed margin deflates and what remains is a commodity business bounded by energy. The hyperscalers survive because they also sell liability, brand and a contract. The neoclouds are more exposed: they rent bare GPUs to customers who want the commodity itself. This is a framework note, not an initiation; the companies named are illustrations. No rating.

Larix Research · Framework · This note extends the arguments of The Loser's Game (11 July 2026) and The Winner's Share (24 July 2026). The businesses named below are large and well-covered; they are illustrations, not coverage. No rating, no target. Disclosures at the end.

The third leg of the stool

The first note in this series made the defensive argument: AI commoditises interfaces, so the disciplined software business protects the proprietary substrate the interface must pay to reach. The second made the offensive one: the owners of those substrates are not merely surviving the transition; they are receiving the value that AI displaces. The 2026 derating priced them as if they were capability sellers. Both notes were about application software. This one is about the layer underneath it, the cloud itself. A cloud service provider has the same structure the market has been mispricing all year: a physical substrate at the bottom, a capability layer on top, and a valuation that assumes the capability layer is where the margin permanently lives.

Strip a hyperscaler's product catalogue to its physics and there are two things for sale. The first is infrastructure: a virtual machine, a bucket, a rack of GPUs, a megawatt of energised capacity. It is scarce, capital-intensive, and bounded by power, silicon allocation, grid interconnection and land. The second is software: the managed database, Kubernetes control plane, machine-learning platform and inference endpoint. This layer is not physically scarce. Its scarcity came from the rare and expensive human labour needed to organise it well. The cloud industry's margin structure rests on that scarcity, which generative AI and AI coding are now deflating.

The managed margin

Amazon Web Services opened in 2006 with two products: S3 in March and EC2 in August. Raw storage, raw compute. The catalogue now runs to hundreds of services, but the move has remained the same: observe what customers build on top of the raw primitives, package it, and sell it back with a margin attached. DynamoDB in 2012 packaged the distributed database. SageMaker in 2017 packaged the machine-learning workflow. EKS in 2018 packaged Kubernetes, software Amazon did not write or own, because operating it yourself was painful enough to pay to avoid.

The markup is on the price list. SageMaker instances, the ml.-prefixed twins of ordinary EC2 instances running on identical hardware, have historically carried a premium of roughly 20 to 40 percent over the equivalent raw EC2 rate, averaging about 25 percent across comparable instance types. The premium buys OS patching, driver updates, endpoint orchestration and health checks. It buys the labour of platform engineering, productised. AWS's own total-cost-of-ownership materials make the arbitrage explicit: the managed service is priced against the engineering headcount it replaces, not against the compute it runs on.

That arbitrage built the most profitable infrastructure business in history. AWS's operating margin has run between 33 and 40 percent in recent years: 39.5 percent at the Q1 2025 peak and 34.6 percent in Q3 2025, on a segment generating $33 billion of quarterly revenue and growing at 20 percent. A commodity compute business does not produce those margins; energy, racks and depreciation see to that. The margin is in the software, the capability layer where the scarcity rents collect.

The managed layer always had a competitor. Inside every AWS account there was a standing alternative: run it yourself. SageMaker competed with a team loading its own model onto a raw g5 instance for three quarters of the price. EKS competed with self-managed Kubernetes on EC2. DynamoDB competed with Postgres, Cassandra or Valkey on instances the customer patched. The DIY path lost for twenty years on a simple calculation: the premium was cheaper than the payroll. A 25 percent markup on a compute bill is real money; a platform team that can safely operate distributed systems costs more, is scarcer, and takes longer to hire. The managed margin was the market price of a labour shortage.

The competitor that asks no permission

AI coding reprices that labour, and it reprices it from the buyer's side.

The current generation of coding models, including Claude Fable 5, Kimi K3 and GPT-5.6, does more than write snippets faster. These models can operate systems. The work that justified the managed premium was never the initial configuration; it was the continuing job of scoping and rotating IAM roles, patching CVEs, tuning autoscaling, and upgrading a cluster without taking it down. An agent can now be assigned that job, pulling Kubernetes, vLLM, Ray, Postgres, Ceph and the rest of the open-source stack together and then staying on shift. Open source does not charge a margin, and an agent does not draw a salary. The 25 percent premium is no longer priced against a platform team's payroll. It is priced against an inference bill that is falling faster than the premium ever has.

The open-source stack was always free, and hyperscaler customers always had access to it. What changed is the cost of operating it, which was the moat. CNCF's 2025 survey, published in January, found that 82 percent of container users run Kubernetes in production and two-thirds of organisations hosting generative-AI models use it for inference. The software is not exotic; the competence to run it well was. AI coding spreads that competence. The managed service's competitor used to be the hyperscaler's own cheaper tier, which the hyperscaler controlled. It is now the accumulated public work of infrastructure engineers, assembled and maintained by software the customer rents by the token.

The hyperscalers sit on both sides of this trade. The models that erode the managed margin are sold partly through the managed layer: Bedrock, Azure AI Foundry and Vertex AI. Every improvement in coding models makes orchestration software easier to replace, while every dollar of model revenue adds pressure to the platforms' highest-margin product line. Software incumbents faced the same choice: adopt the deflation or let a competitor do it first. Adoption is the correct strategy, but it does not repeal the arithmetic. The hyperscalers are racing to replace their own margin with a cheaper one before someone else does, leading to a lower-margin industry.

Cement

When the orchestration premium deflates, the physical layer remains. Its economics resemble an industry that has lived with energy constraints for a century.

Cement is the canonical energy-bounded commodity. The product is undifferentiated and heavy, the technology is mature, and energy accounts for 20 to 40 percent of production cost in the US EPA's formulation. Current industry benchmarks put fuel for the kiln at 30 to 40 percent and electricity for grinding at 20 to 25 percent. Nobody pays a cement producer for software. You pay for the quarry, kiln and energy contract. Returns depend on utilisation against a depreciating capital stock, and pricing power lasts only while capacity is scarce. Amrize, the North American cement business Holcim spun out in June 2025, lists energy as "an important part of our cost structure" in its own risk factors. It is a fine business, but not one that earns or is valued on a 35 percent operating margin.

The GPU rental market fits that template. An H100 is an H100 whoever racks it, and it is already priced like a commodity. Rental rates that peaked above $8 an hour during the 2023 scramble had fallen to roughly $2.85 to $3.50 at the budget tier by 2026, a decline of 64 to 75 percent. Each new NVIDIA generation compresses the prior generation's pricing. The cost structure is also energy-bounded. Amazon added 3.8 gigawatts of capacity in twelve months and still described power as the tightest constraint on AWS; grid operators warn that datacentre plans are straining load forecasts; converted bitcoin miners entered the neocloud business because they owned energised land. Returns depend on utilisation against a depreciating asset. One standard industry calculation has a debt-financed, 1,024-GPU cluster breaking even at roughly 70 percent utilisation, losing about $330,000 a month at 55 percent and earning about $340,000 at 85 percent. Swap "kiln" for "cluster" and the sentence needs no other edits.

The depreciation debate reaches the same conclusion through the accounts. In November Michael Burry argued that hyperscalers were understating depreciation by some $176 billion from 2026 to 2028 by extending GPU useful lives to five and six years. Meta moved from four-and-a-half to five-and-a-half years, Google from three to six, Oracle to six, and CoreWeave from four to six in 2023. His formulation was aggressive, but the structural point holds: the asset at the bottom of the AI stack wastes at a rate set by silicon progress and energy economics. Satya Nadella put the concern plainly on a podcast when he said he did not want to "get stuck with four or five years of depreciation on one generation." CoreWeave offers a real counterpoint: a customer re-contracted 2022-vintage H100s within 5 percent of the original price, and older chips can cascade into inference work. Yet this only shows that the commodity retains value while capacity is tight. It says nothing about the margin on the software above it. Cement plants can have good decades when the building cycle is up. They are still cement plants.

The neocloud squeeze

The neoclouds are where this framework matters most.

The neocloud category, including CoreWeave, Nebius, Lambda, Crusoe and the converted miners behind them, crossed $25 billion of revenue in 2025. Fourth-quarter revenue rose 223 percent year on year, and Synergy Research projects nearly $400 billion by 2031. The equity story embedded in that trajectory is about margins: these are cloud companies, and cloud companies earn software margins on hardware. CoreWeave markets an integrated platform, Kubernetes service, observability layer and ClusterMAX Platinum rating. Nebius makes the same claim through Token Factory, its inference product launched in November 2025.

The filings describe a bare-metal business. CoreWeave's FY2025 10-K reported revenue of $5.1 billion, up 168 percent, against a net loss of $1.2 billion. Remaining performance obligations rose to $60.7 billion from $15.1 billion a year earlier, with a weighted-average contract duration of roughly five years. Microsoft supplied approximately 67 percent of revenue; OpenAI committed up to $6.5 billion through 2031 and Meta up to $14.2 billion. The infrastructure is financed, in the company's own formulation, "primarily through asset-level debt supported by take-or-pay customer contracts." Capex ran near $23 billion in 2025, with $30 billion or more guided for 2026. S&P puts the debt load at roughly $21 to $30 billion depending on what is included, and the 2031 bonds have traded at double-digit yields. The revenue backlog approached $100 billion in the Q1 2026 print. The watch items for Q2 are utilisation, interest expense, and how much incremental revenue comes from the platform rather than the metal.

The customers signing these contracts are frontier labs and the hyperscalers themselves. Microsoft alone has struck commitments on the order of $60 billion across CoreWeave, Nebius and Nscale; Meta has signed for up to $27 billion with Nebius; IREN's November 2025 Microsoft contract is $9.7 billion for GB300 capacity across 200 megawatts in Childress, Texas, at a projected 85 percent EBITDA margin. These buyers employ, or are, the best infrastructure engineers in the industry. They are purchasing energised capacity they cannot build fast enough, not orchestration software. That is why the contracts are take-or-pay, GPU-collateralised and measured in megawatts. Microsoft is renting 200 megawatts from IREN while its own capacity is built. The commitment tapers when construction completes, and it is priced as capacity rather than software.

Pressure also comes from smaller customers. The neocloud's hoped-for margin layer, including hosted inference, fine-tuning platforms, RL tooling and Token Factory products, sells to mid-market AI companies and enterprises that cannot employ a frontier infrastructure team. Those are the customers to whom AI coding hands that team's competence. A two-person RL shop in 2026 does not need a hosted RL platform; it needs bare GPUs, an agent and the open-source training stack. Nebius's disclosures show the gap. Token Factory launched with the inference narrative in November, but receives no mention in the Q1 2026 6-K. Revenue appears as one undifferentiated AI-cloud line, compute rental, up 684 percent year on year to $399 million. The growth is real and enormous, but it is arriving in the commodity layer rather than the software layer. For the neoclouds, inference-as-software remains narrative ahead of financials.

None of this makes the neoclouds bad businesses. It changes how they should be valued. Bears who model GPU depreciation and customer concentration are analysing a commodity producer with a software multiple; bulls who model $400 billion of 2031 revenue are analysing a software company with a commodity cost structure. Both are pricing the wrong asset. A merchant-producer model is more useful: utilisation, energy, capacity discipline and the residual value of the asset when the next generation ships.

The IBM clause

The hyperscalers are not equally exposed. The difference is the authority channel the previous note identified in legal and medical corpora: a large enterprise buys accountability along with the software.

No one gets fired for buying IBM. The aphorism is older than the cloud, and it survives because it describes a real product. When the payment system fails at 3 a.m., there is a counterparty with a balance sheet, an SLA, an indemnification clause, a shared-responsibility model audited against SOC 2 and ISO, and a named human accountable for the failure. An agent-assembled stack has no such counterparty. It has a git history, which may be enough for a startup training a model. For a bank, hospital system, sovereign or insurer, the compliance certifications, audited IAM boundary and legal liability are the product. The managed premium is the price of someone to blame. Generative AI does not deflate that premium any more than it deflated Wolters Kluwer's corpus. A flood of agent-assembled infrastructure may even raise the value of the certified, audited and indemnified alternative.

The boundary matters. The authority premium defends regulated, audited and liability-bearing buyers at the top of the customer pyramid. It does not defend buyers with no compliance department and no one to be fired. The managed margin will not disappear; it will compress upward from the entry tier, workload by workload, starting with those carrying no liability: development environments, batch training, internal tooling and startup inference fleets. Consolidated segment margin will be the last place the compression shows. The earlier marker is the attach rate, or the share of new compute sold with the managed layer rather than bare. A sustained decline in that ratio would support the framework.

The hyperscalers hold a structural advantage the neoclouds do not: they sit on every side of the trade. They own silicon roadmaps such as Trainium and TPU, energy relationships and enterprise contracts. Through their capacity commitments, they are also the neoclouds' largest customers, renting capacity while their own datacentres are built and booking it as operating rather than capital expenditure. A hyperscaler owns every layer the commodity business depends on while also acting as the merchant producer's biggest buyer. It is temporarily renting someone else's curve. For hyperscalers, the predicted compression is a mix problem: managed-services growth slows relative to raw compute, and segment margin moves towards the infrastructure businesses it contains. That is serious but survivable. For a pure-play neocloud, the same compression is the whole business.

What would prove this wrong

Three developments would break this framework, and all three are observable in reported numbers.

The attach rate holds. The core prediction is that the managed premium deflates as agent-operated open source becomes the default alternative. If Bedrock, SageMaker, Azure AI Foundry and Vertex keep growing faster than the raw compute beneath them, the premium is holding and the framework is early or wrong. The markers are disclosed quarterly in hyperscaler commentary that splits managed AI services from infrastructure, and in the SageMaker premium itself. AWS cut P4-instance SageMaker pricing by up to 45 percent in June 2025. One data point supports the managed-services case: even CoreWeave, an archetype of build-it-yourself infrastructure competence, signed a $335 million storage agreement with Backblaze rather than operate that layer itself. Managed services survive where a workload is peripheral and the operator's accumulated expertise is real. The claim is not that the premium vanishes everywhere. It is that cheaper operating labour reduces the number of workloads for which the premium is worth paying.

The compliance boundary holds the pyramid. The IBM-clause argument assumes enterprise procurement continues to require the vendor's audited IAM, shared-responsibility model and certifications. If agent-operated infrastructure develops its own compliance primitives, including attestation frameworks, audited agent change-logs or insurance for self-managed stacks, the boundary moves. Compression would then reach regulated tiers faster than this note assumes. Procurement language matters more than benchmark scores: approval of an agent-operated control plane by a regulator or Big Four auditor would start to erode the authority premium.

The cascade saves the neoclouds. The commodity conclusion rests on GPU rental pricing following the commodity curve. If older silicon holds its rental value as CoreWeave's re-contracted H100s suggest, utilisation stays above the breakeven band through the next two NVIDIA generations, and the inference mix lifts pricing power, then the depreciation schedules hold and take-or-pay contracts convert to cash. The better neoclouds could earn through the cycle as merchant producers in a structurally short market. The decisive marker is the software layer they are building above the metal. If hosted inference and RL products such as Nebius's Token Factory appear as disclosed, material revenue lines rather than earnings-call narrative, the classification changes. Tonight's CoreWeave print will not settle that question. The next six quarters may.

The discipline is the one this series started with. Ask whether each layer sells capability or substrate, then ask whose labour the capability premium is priced against. In application software, the answer separated substrate owners from capability sellers that the market had filed together. In the compute stack, the capability layer is the margin, the labour it is priced against is being automated, and the layer underneath is bounded by energy, the one input no model generates more of. The cloud spent twenty years climbing from cement to software. AI coding is the escalator back down, moving fastest for businesses that never owned anything above the metal.

Sources and method

This is a framework note, not an initiation of coverage, and contains no recommendation, rating, or price target. It extends the arguments of "The Loser's Game: Why Software Survives AI by Not Losing" (Larix Research, 11 July 2026) and "The Winner's Share: Who AI Actually Pays in European Software" (Larix Research, 24 July 2026). Companies named (Amazon/AWS, Microsoft, Google, Meta, CoreWeave, Nebius, IREN, Lambda, Crusoe, Backblaze, Holcim/Amrize) appear as illustrations of a structural argument drawn from public reporting and filings; several are covered by sell-side research and all sit outside the undercovered universe this firm exists to examine. We hold no position in the securities mentioned and received no compensation from any party in connection with this note.

AWS history and pricing: S3 (March 2006) and EC2 (August 2006) launch dates, DynamoDB (January 2012), SageMaker (November 2017) and EKS (June 2018) launches per AWS's published service histories. SageMaker premium over equivalent EC2 (ml.-prefixed instances on identical hardware) of roughly 20–40%, ~25% average across comparable instance types — TrueFoundry pricing analysis, 13 February 2026; CloudZero SageMaker pricing guide, verified April 2026; Finout SageMaker cost guide, 2026; instance-by-instance EC2/SageMaker ratio table (~1.25x average) via Dev59 community compilation. AWS TCO claim (54% lower 3-year TCO vs self-managed EC2) as quoted in CloudZero. SageMaker P4-instance price cuts of up to 45%, effective June 2025, per Checkthat pricing coverage, 22 April 2026. AWS segment figures — Q3 2025 AWS net sales $33.0bn, +20%; operating margin 34.6% (Q3 2025), 39.5% (Q1 2025 peak), 35.9% TTM; more than 3.8 GW of power capacity added in the trailing twelve months — Amazon Q3 2025 earnings release, 30 October 2025 (SEC 8-K exhibit). "Power as the tightest constraint" — CFO commentary on the Q2 2025 call, via contemporaneous coverage, 26 August 2025.

Kubernetes and the open-source stack: production adoption of 82% of container users and two-thirds of organisations hosting generative-AI models using Kubernetes for inference — CNCF Annual Cloud Native Survey 2025, published January 2026, via CNCF and secondary summaries.

GPU rental pricing and neocloud unit economics: H100 budget-tier rental of roughly $2.85–3.50/hour in 2026, down 64–75% from the 2023 peak above $8/hour; B200 ~$6.50/hour; GB200 rack-scale ~$17.85/hour — Silicon Data pricing, via Moduledge, "How the Neocloud Business Works," 7 June 2026. Breakeven at ~70% utilisation; ~$330k/month loss at 55% versus ~$340k/month profit at 85% on a 1,024-GPU H100 cluster; gross margins of 55–65% before depreciation — American Compute, "Neocloud Business Model and Unit Economics," 5 March 2026. Neocloud category revenue exceeding $25bn in 2025, Q4 2025 revenue of $9bn (+223% YoY), forecast near $400bn by 2031 at a ~58% CAGR — Synergy Research Group, 2 April 2026. SemiAnalysis tier definitions per "AI Neocloud Playbook and Anatomy" (2024), as cited therein.

CoreWeave: FY2025 10-K (filed March 2026; SEC) — revenue $5.1bn/$1.9bn/$229m (2025/2024/2023); net losses $1.2bn/$863m/$594m; RPO $60.7bn versus $15.1bn; weighted-average contract duration ~5 years; ~67% of 2025 revenue from Microsoft; OpenAI commitment up to ~$6.5bn through May 2031; Meta initially up to ~$14.2bn through December 2031; financing "primarily through asset-level debt supported by take-or-pay customer contracts." Capex near $23bn in 2025 and $30bn+ guided for 2026; debt in the ~$21–30bn range; outlook revised to positive — S&P Global Ratings research update, 9 April 2026; Q4 2025 capex of $8.2bn and FY2025 adjusted EBITDA of $3.09bn per Q4 2025 results coverage. Revenue backlog approaching $100bn and active capacity above 1 GW — Q1 2026 results (8-K, 7 May 2026), via contemporaneous coverage. Kerrisdale Capital short thesis (cash burn, leverage) published September 2025. Q2 2026 results scheduled for release after the close on 11 August 2026 — company IR announcement, 28 July 2026. Backblaze–CoreWeave $335m strategic storage agreement — Backblaze Q2 FY2026 results and Futurum Group coverage, 5 August 2026. H100 re-contracting within 5% of original pricing and "demand remains robust across generations" — CEO Michael Intrator on the Q3 2025 call, quoted in Morningstar/MarketWatch, 11 November 2025.

Microsoft/neocloud commitments: ~$60bn across CoreWeave, Nebius and Nscale — Trending Topics, 6 July 2026. Meta up to $27bn with Nebius; Nebius contracted capacity past 3.5 GW and 2026 capex guidance of $20–25bn — Digital Applied, 18 May 2026. IREN–Microsoft: ~$9.7bn multi-year contract announced November 2025, NVIDIA GB300 across 200 MW at Childress, five-year average term, 20% upfront prepayment, ~$1.94bn annualised recurring revenue at ~85% project EBITDA margin; IREN targeting $3.4bn AI Cloud ARR by end-2026 — Converge Digest, 28 July 2026.

Nebius: Q4 2025 consolidated revenue $227.7m (+547% YoY), core AI cloud ARR of $1.2bn at year-end 2025, 170 MW active power — Q4 2025 results, via Converge Digest, 28 July 2026. Q1 2026 revenue $399.0m (+684% YoY) as a single AI-cloud line with inference not disaggregated, and zero Token Factory mentions in the 6-K — Nebius Group 6-K (SEC, filed April 2026), as documented in Jimmy Research, "Token Economics in the AI Era," 10 May 2026.

GPU depreciation debate: Burry's November 2025 critique — $176bn estimated understatement of depreciation across 2026–2028; Oracle earnings overstated ~27%, Meta ~21% on his calculations; useful-life extensions at Meta (to 5.5 years, January 2025), Google (3 to 6), Oracle (to 6) — Morningstar/MarketWatch, 11 November 2025; Fortune and InvestorPlace contemporaneous coverage. CoreWeave's extension from four to six years in 2023 and the spectrum of schedules (AWS nearer four years; Microsoft, Google, Oracle in the four-to-five-to-six range) — Yardeni Research Morning Briefing, 17 November 2025. "The $4 trillion accounting puzzle" — The Economist, as quoted in the November 2025 coverage. Nadella's "four or five years of depreciation on one generation" remark — via Stanley Laman, "Why GPU Useful Life Is the Most Misunderstood Variable in AI Economics," 21 November 2025. Counterpoints: Bernstein's Stacy Rasgon defending six-year schedules — Barron's, 17 November 2025.

Cement and Amrize: energy at 20–40% of cement production costs — US EPA, "Energy Efficiency Improvement and Cost Saving Opportunities for Cement Making" (S); fuel at 30–40% and electricity at 20–25% of production cost — iFactory industry cost breakdown, 8 July 2026, and the Cement Institute via Imubit, 13 March 2026. Amrize: spin-off completed 23 June 2025 (Holcim media release); 18 plants and 25M mt/year of cement capacity (S&P Global Commodity Insights, 2 June 2025); energy as "an important part of our cost structure" — Amrize Form 10 information statement (SEC, 2025); more than half of new capex targeting infrastructure, reshoring and hyperscale datacentres — contemporaneous spin-off coverage, June 2025, and Amrize Q4/FY2025 earnings presentation, 18 February 2026.

All market and financial figures were verified against primary filings or high-authority secondary sources on 11 August 2026. Where a figure rests on secondary coverage rather than the primary document (the Q1 2026 CoreWeave backlog, the Microsoft aggregate commitments, the Nadella quotation), it is attributed as such above.