| |
|
A daily bite-size selection of top business content.
PM edition. Issue number 1369
Latest 10 stories. Click the button for more.
|
| |
|
"Something has gone completely wrong. The basic view among enterprises in this country is: 'I am going to chillax and waste my time with tokens. I am going to get no value, and they are going to get my IP.'" - Alex Karp - Palantir CEO
Enterprise AI adoption is colliding with a harsh reality: many large organisations feel they are paying heavily for experimental systems while surrendering control over their most valuable asset, their intellectual property, to external model providers. In the US corporate environment Alex Karp describes on CNBC, the prevailing mood among sophisticated buyers is not excitement about cutting-edge models but frustration, distrust, and a growing sense that the economic bargain being offered by frontier AI labs is structurally misaligned with enterprise interests. That dislocation between value delivered and IP risk is the underlying tension driving both the quote and the strategic repositioning now under way across the AI ecosystem.
From Model Hype To Enterprise Disillusion
The immediate factual context is Karp's appearance on CNBC's "Squawk Box" in July 2026, where he argues that leading AI labs have "completely" mis-sold AI to enterprises by focusing on token-based access to frontier models rather than on controlled, outcome-focused deployments. In his account, many US enterprises have trialled generative AI services priced on usage tokens, only to discover that pilots rarely progress into production systems that materially improve manufacturing, logistics, pharmacological research, or other complex operations. Instead, budgets are consumed by experimentation while the providers accumulate fine-tuning data, usage patterns, and business process know-how which can be re-embedded into their own models, effectively harvesting the clients' alpha-the distinctive decision logic and competitive edge embedded in their data and workflows. This perceived asymmetry-"no value" for the buyer, strategic data for the seller-is what Karp frames as something having gone "completely wrong".
Token-based commercial models were originally marketed as democratising access: pay per token, experiment rapidly, avoid the capital expenditure of owning clusters or building internal model pipelines. Yet for complex enterprises operating in regulated environments or on the battlefield, capability without sovereignty quickly becomes a liability. They may obtain powerful generative capabilities but lack negotiated guarantees about where prompts are cached, how outputs are logged, and whether fine-tuning pipelines will internalise their proprietary domain knowledge. As concerns about model training data, reinforcement learning from human feedback, and long-term retention policies have intensified, token pricing has ceased to look like operational flexibility and instead resembles an opaque rent charged on access to infrastructure that may ultimately compete with the client.
IP Sovereignty And The Fear Of "Alpha Theft"
Central to the statement is a specific fear: that external labs will not only see sensitive data, but also infer and appropriate the core decision-making patterns that constitute an enterprise's alpha. In financial language, alpha denotes excess return above a benchmark; in Karp's usage, it extends to any proprietary operational edge-manufacturing recipes, logistics heuristics, risk scoring rules, or targeting doctrines-that can be implicitly reconstructed from usage data. The worry is not merely that sensitive documents or customer records might leak, but that the lab, by observing queries and feedback at scale, can build a generalised representation of how the client makes high-value decisions, then reuse or productise that representation for other customers or its own ventures. In this view, enterprise AI based on external closed models risks functioning as a one-way knowledge transfer: the client pays for tokens and experimentation, the provider quietly accumulates a distilled map of the client's decision landscape.
These concerns intersect with a broader security discourse on AI systems. Research on emerging AI security risks notes that large-scale models and agentic systems introduce new attack surfaces, including prompt injection, model theft, and indirect prompt attacks that can exfiltrate or reconstruct sensitive patterns from interaction histories. When enterprises are unsure who controls the weights, where model states are stored, or how fine-tuning data is segregated, the threat is not confined to adversarial actors; it includes the legitimate provider using aggregated insights to strengthen its own competitive position. Recorded analyses of AI security risk emphasise the need for zero-trust principles and strict governance for agent identities and model access paths. Karp's intervention translates that abstract security language into a blunt commercial accusation: some labs are effectively imposing a "wealth tax"-charging high fees while appropriating data-driven alpha that rightly belongs to the enterprise.
The Strategic Role Of Ontology And The Application Layer
Technically, Karp positions his organisation's ontology and application layer as the antidote to this perceived mis-selling. An ontology in modern enterprise AI is a structured, machine-readable representation of the entities, relationships, and actions that define a business domain. It encodes not merely metrics but the semantic and operational logic of an organisation: what a "shipment" is, how it relates to "warehouse", "route", and "risk score", and what should happen when a shipment is delayed or a risk threshold is exceeded. Sources on ontology describe it as the central nervous system of an enterprise AI stack, integrating data, business logic, and stateful decision processes into a unified model that both humans and AI agents can query and update. Palantir's documentation explicitly frames its Ontology system as modelling decisions through the integrated representation of data, logic, action, and security.
By situating the large language model behind such an ontology-driven application layer, Karp argues that enterprises can make frontier or open models "safe, useful and precise" without exposing underlying data or decision logic to uncontrolled caching or replication. The ontology constrains interactions: the model does not directly traverse raw databases or ungoverned knowledge graphs, but operates within a guardrailed semantic context where each read and write is mediated by the ontology's policies and security rules. This architecture allows enterprises to swap models-closed or open-while retaining ownership of the domain representation and controlling which signals are fed back into training pipelines. In effect, sovereignty is pushed up a layer: instead of negotiating at the level of tokens and prompts, enterprises assert control through a persistent context model that defines the vocabulary, relationships, and permissible actions of their AI systems.
Independent analysis of ontologies and semantic layers reinforces this framing. Multiple sources distinguish semantic layers (which standardise metrics and calculations) from ontologies (which model domain entities and relationships to support reasoning and multi-agent workflows). In more advanced architectures, an enterprise context layer extends the ontology with policy and judgement, enabling agentic AI to act with appropriate authority while remaining grounded in governed context. This layered approach speaks directly to Karp's critique: token-based access without such context mechanisms leaves enterprises reliant on system prompts and ad hoc safeguards, whereas moderated access through an ontology enables granular control over what the model sees, what it can infer, and how its outputs propagate through operational systems.
Tokens, Pricing, And The True Cost Of Enterprise AI
Behind the rhetoric lies a concrete economic dispute about how AI is priced and measured. Token-based billing makes sense for consumer chatbots or simple developer experiments, but it maps poorly onto the complex cost structure of enterprise AI deployments. Detailed breakdowns of total cost of ownership for enterprise AI highlight multiple categories beyond mere inference: infrastructure, data engineering, specialised talent, model maintenance, compliance, and integration can each account for substantial portions of spend. For example, GPU clusters and multi-cloud infrastructure can run from 200 000 to well over 2 000 000 annually; data engineering and pipeline maintenance often consume 25 to 40% of the total budget; and version control, monitoring, and retraining add further overhead. Against that backdrop, token fees imposed by labs are only one component of cost, yet they are often the most visible and least justified from a business outcome perspective.
Karp's point about "bad financials and growth while losing money" is that many frontier-lab-centric enterprises are stuck in a pattern where token usage rises, experimental deployments proliferate, but revenue-impacting applications fail to reach scale because clients will not pay the "true cost" for hardened systems integrated into core operations. Independent CIO analysis supports this diagnosis: most enterprise AI programmes struggle not because models are inadequate but because operating models, data governance, and decision workflows are not restructured to exploit intelligence effectively. Pilots remain isolated, metrics are vague, and AI doesn't become a "core business capability" driving measurable outcomes. In that environment, token spend looks like speculative experimentation rather than capital investment; when combined with fears about IP appropriation, the result is the disillusionment Karp reports-organisations feel they are "chillaxing" with models that burn compute and budget while shifting long-term advantage towards the providers.
Trust, Governance, And Who Owns The Risk
Layered through the backstory is a governance argument: enterprises are demanding clear answers to basic ownership questions that many labs have deferred or obscured. Who owns the data once it flows through prompts and logs? Where is it cached, and under what retention and deletion policies? Are fine-tuned weights shared, siloed, or reused across customers? Who is accountable if model behaviour evolves into risky territory due to accumulated training signals? Governance frameworks for AI risk increasingly emphasise cross-functional responsibility-security, legal, ethics, and operations convened in AI risk committees or similar bodies-to manage these questions systematically. Yet Karp suggests that labs have tried to bypass this institutional maturity by appealing to trust and speed, effectively asking enterprises to accept "I have never lied" narratives rather than auditable guarantees.
Studies on AI accountability in enterprises show that ownership is often fragmented: surveys find that a significant share of organisations cannot identify a single C-suite executive accountable for AI-related risks. Regulatory shifts, such as US federal guidance on Chief AI Officers, are slowly pushing towards clearer responsibility structures, but many enterprises are still negotiating basic guardrails. In that milieu, labs that offer generic indemnities or vague privacy commitments are increasingly out of step with the expectations of regulated industries and national security customers. Recorded Future and other security-focused analyses argue for continuous AI governance, validation, and monitoring, moving beyond traditional detection-first security towards active control over model access, prompt engineering, and behavioural auditing. Karp's narrative fits within that emerging consensus: he channels the frustration of enterprises that feel their governance questions are brushed aside in favour of token counters and marketing demos.
Open Models, Compute Control, And The Battle Over Means Of Production
A key strategic pivot Karp advocates is towards open-weight models, owned compute, and internal control over the "means of production" in AI. Rather than relying on external frontier labs to provide both models and infrastructure, he argues that sophisticated enterprises-and especially defence customers-should own their GPUs, run open-source or internal models, and maintain direct control over weights. Analysts observing his CNBC interview note that he positions this as an AI sovereignty agenda: enterprises should not depend on consensus views in Silicon Valley to run battlefields or critical infrastructure. This resonates with broader trends: open models from US, European, and Chinese providers, as well as model ecosystems associated with major hardware vendors, are increasingly attractive because they allow weight-level control and deployment into sovereign environments.
The partnership he describes with NVIDIA is emblematic of this shift. Rather than simply consuming NVIDIA's hardware indirectly through cloud platforms, the arrangement is framed as a way to build custom AI systems where enterprises own their compute, models, data stack, and alpha. Industry commentary highlights that technical customers now ask first-order questions about whether they can switch models, retain weights, and localise training within their own security perimeter. Open-weight models, combined with ontology-driven application layers, offer a pathway: enterprises can assemble model-plus-context stacks where the underlying infrastructure is under their control, and external labs become optional rather than central. Karp's blunt criticism of "deploy-co" structures-entities that merely deploy tokens while transferring alpha to third parties-captures his belief that the old model of centralised AI provision is being superseded by a more federated, sovereign architecture.
The Political And Geopolitical Dimension
Although the quote focuses on enterprise sentiment, it sits within a broader political and geopolitical frame. Karp warns that overselling AI to enterprises while refusing secure, controllable deployments for departments of defence or war is "effing insane" from a national security standpoint. He juxtaposes the willingness of labs to release powerful models to global adversaries with their reluctance to provide weight control and data sovereignty to allied governments, calling into question the alignment between commercial AI strategies and Western security priorities. Geopolitical analysis identifies America, China, and Israel as the key tech centres in this domain, with China acting as a peer adversary and building its own AI ecosystems without the same internal political and cultural frictions. In that context, he argues that Western debates about restricting AI for domestic governments while leaking capability to adversaries are strategically incoherent.
This geopolitical lens reinforces the enterprise IP concerns. If frontier labs spread powerful models widely, train them on cross-enterprise data, and retain unilateral control over their evolution, they accumulate a transnational reservoir of operational knowledge that may be difficult to govern or align with any particular state's interests. Policy analysts increasingly call for AI sovereignty: nations and major enterprises should maintain clear lines of control over critical models, training data, and deployment pipelines. Ontology-based architectures and sovereign compute fit this agenda, providing mechanisms to confine sensitive decision logic and operational knowledge within national or organisational boundaries while still exploiting shared model innovation where appropriate. Karp's argument surfaces the uncomfortable possibility that without such measures, enterprises are unwittingly contributing to a shared AI commons dominated by private labs whose strategic aims diverge from their own.
Debates, Objections, And Why It Matters
There are, of course, objections to Karp's framing. Supporters of frontier labs argue that token-based models enable rapid innovation, allow smaller firms to access capabilities they could never build themselves, and that strict privacy and data segregation policies prevent the kind of alpha appropriation he describes. They point out that most labs publish privacy guarantees, offer enterprise-grade instances, and in some cases commit not to train core models on customer data without explicit consent. Furthermore, independent commentators note that Karp is simultaneously warning about mis-selling and promoting his own stack as the solution, raising the question of how much of his critique is principled and how much is commercial positioning.
Yet even critics acknowledge that his intervention crystallises real anxieties in the market. CIO surveys repeatedly show that enterprise AI programmes struggle to progress beyond pilots, and that business stakeholders demand clearer ownership, risk management, and measurable outcomes. Security analysts document new attack vectors and emphasise the need for specialised AI governance rather than generic cybersecurity controls. Ontology and context-layer specialists underline that without a structured domain model, large language models will remain brittle, hallucination-prone, and difficult to integrate safely into mission-critical systems. In that sense, Karp's colourful language functions as a signal: the easy phase of model hype and token experimentation is ending, replaced by a more demanding era where enterprises insist on value, sovereignty, and trustable architecture.
Why it matters is straightforward. If large organisations continue to see AI as a high-cost, low-trust experiment, adoption will stall, and the transformative potential of integrated AI decision systems will be realised only in pockets-by those who build sovereign stacks with strong ontologies and controlled models. The competitive gap between enterprises that own their alpha and those that bleed it into shared clouds will widen. National security doctrines will be shaped by whether governments can deploy agentic AI that respects classified boundaries while reasoning effectively. And the structure of the AI industry itself will be determined largely by how this dispute over tokens, IP, and ownership is resolved: either frontier labs maintain a centralised, rent-extracting role, or the ecosystem rebalances towards open-weight, ontology-driven, sovereign architectures where labs are important but not dominant. Karp's remark is a snapshot of that inflection point, capturing a moment when the market is re-evaluating what it is willing to pay for-and what it is no longer willing to give away.

|
| |
| |
|
"AI compute refers to the raw processing power, hardware (like GPUs or TPUs), and computational resources required to build, train, and run machine learning models. It is the physical and electrical engine that makes artificial intelligence possible." - Compute - Artificial intelligence
The limiting factor in modern artificial intelligence is increasingly neither algorithms nor data, but the availability, efficiency, and governance of the underlying computational power that drives every training run and inference call . This constraint shapes which models can realistically be built, who can build them, and how they can be deployed in practice, turning technical capacity into a strategic economic and geopolitical resource . As models grow larger and more capable, the marginal gains from better architectures are often gated by access to sufficiently dense and affordable processing, memory, and interconnect, making the structure of computational resources central to the future trajectory of AI .
From abstract computations to physical infrastructure
Discussions of computational requirements for AI often blur three distinct but related layers: the number of mathematical operations needed to train or run a model, the performance of the hardware capable of executing those operations, and the physical infrastructure that supplies power, cooling, and connectivity . At the most abstract level, one can speak of the total number of floating point operations needed to complete a task such as training a large language model; this is a property of the model architecture, dataset size, and optimisation schedule . At the performance level, the relevant quantity is how many such operations a chip or cluster can execute per second, typically expressed as floating point operations per second, or FLOP/s, and scaled to tera-, peta-, or exa-levels for modern accelerators . Finally, the physical realisation includes racks of GPUs, TPUs, or other accelerators, backed by power distribution, cooling, networking, and storage, all of which determine whether theoretical performance can be sustained in practice .
This layered view matters because it separates the algorithmic compute demand from the hardware supply and the infrastructure that binds them. A model that mathematically requires floating point operations to train might in principle run on any hardware, but in practice only facilities with sufficiently many accelerators, reliable power, and high-bandwidth interconnect will complete the job within useful time and cost constraints . Conversely, a highly capable data centre with petascale compute capacity may be underutilised if software is poorly optimised or if algorithms do not parallelise efficiently across its architecture . This interplay between the abstract workload and its physical instantiation is where many of the practical and policy debates about AI compute now reside .
Substantive meaning: what compute encompasses
In operational terms, AI compute encompasses the processors, memory, storage, and interconnect needed to execute the numerical linear algebra at the core of contemporary machine learning . Processors include general-purpose CPUs and, increasingly, specialised accelerators such as GPUs, TPUs, NPUs, LPUs, and other AI-specific chips that are optimised for dense matrix multiplications and tensor operations . Memory covers both fast on-chip resources used to hold activations and parameters during computation, and off-chip system memory required for larger models and datasets . Storage and networking ensure that training data and model checkpoints can be moved, retrieved, and synchronised across nodes at sufficient speed to avoid idle accelerators .
This combination forms a stack in which hardware, software frameworks, and data centre infrastructure jointly determine the effective compute available for AI workloads . At the hardware level, GPUs and TPUs provide massively parallel arithmetic units; at the software level, frameworks such as TensorFlow, PyTorch, and JAX map high-level model descriptions into efficient kernels and collective operations; at the infrastructure level, orchestration systems schedule jobs, allocate accelerators, and manage contention and failures . When practitioners talk about scaling compute, they typically mean increasing one or more of these layers: adding more accelerators, improving software kernels and compilation strategies, or deploying in larger or more specialised data centres .
Training versus inference: distinct compute regimes
AI workloads impose very different computational profiles depending on whether the system is being trained or used for inference. Training deep models involves repeated forward and backward passes over large datasets, requiring extremely high aggregate throughput, long uninterrupted training runs, and careful coordination of parameter updates across many devices . This regime favours clusters of accelerators with high-bandwidth interconnects, large memory, and sophisticated parallelism strategies such as data, tensor, and pipeline parallelism to distribute the compute load .
Inference, by contrast, typically operates on single inputs or small batches but may need to respond within milliseconds at large scale, so latency and cost per query become the binding constraints . For many applications, the objective is to deliver acceptable quality with minimal compute per request, which drives interest in model compression, quantisation, distillation, and specialised inference chips . This divergence explains why training clusters may use general-purpose GPUs or TPUs capable of handling diverse operations, while inference at scale increasingly relies on highly specialised accelerators like LPUs optimised for deterministic, low-latency execution of large language models .
Quantifying AI compute: FLOPs and FLOP/s
To reason rigorously about computational requirements, AI research and policy communities have converged on two related quantities: the total number of floating point operations required by a workload, and the rate at which hardware can execute them . The total work for a training run can be represented as , where is an estimate of operations per example (a function of the model architecture), is the number of examples, and is the number of training epochs. This describes the abstract compute demand independent of any particular hardware implementation .
Hardware capability is characterised by its peak or sustained floating point operations per second, often written as for a given chip or cluster. In simplified terms, the minimum wall-clock time to complete a workload with total operations on a system with effective performance is , ignoring parallelisation overheads and communication costs . In practice, the realised is significantly lower than the theoretical peak due to memory bottlenecks, load imbalance, and suboptimal kernel use . Hence, much of the art of large-scale AI engineering lies in closing this gap through software optimisation, mixed-precision arithmetic, efficient batch sizing, and distributed training strategies that maintain high utilisation of available compute .
Key hardware paradigms: GPU, TPU, LPU and beyond
Modern AI compute is dominated by accelerator classes designed around the patterns of matrix multiplication and vector operations that underpin neural networks. GPUs began as graphics processors but evolved into highly parallel general-purpose accelerators capable of executing thousands of concurrent threads, making them the default platform for training and many inference workloads . Their strength lies in flexibility: they support a wide range of workloads, frameworks, and numerical precisions, and can be deployed in consumer devices, edge systems, on-premises clusters, and hyperscale cloud environments .
TPUs represent a more specialised design, using systolic arrays and custom data paths to accelerate dense tensor operations for deep learning, particularly in large-scale data centre deployments . By sacrificing some generality in favour of fixed-function matrix units and tightly integrated memory hierarchies, TPUs can deliver higher performance-per-watt on well-matched workloads, though they are closely tied to specific software ecosystems and cloud platforms . LPUs, as emerging accelerators targeted at language model inference, push specialisation further: architectures such as Groq's chip use deterministic, compiler-scheduled pipelines with thousands of arithmetic units and explicit dataflow to guarantee predictable latency and maximise throughput for sequential token generation . Alongside these, NPUs, IPUs, and other AI-specific processors explore different trade-offs in programmability, sparsity support, and on-chip memory to better align hardware with the computational structure of modern models .
Compute as a stack: hardware, software, and infrastructure
Thinking of AI compute as a stack highlights that raw processing units are only one component of a larger system that must be jointly engineered. At the base are the chips themselves, which embed microarchitectural choices about arithmetic precision, memory bandwidth, and interconnect topology . Above this sits the systems software layer, including device drivers, runtime libraries, compilers, and distributed training frameworks that translate model graphs into high-performance kernels and collective operations across many devices . At the top lies the infrastructure of data centres, including servers, racks, power delivery, cooling systems, and wide-area networking that enable reliable operation at scale .
A change at any layer can materially alter effective compute. Introducing more efficient kernel implementations or mixed-precision routines can reduce the total operations needed for a given level of model quality, effectively lowering in the workload equation . Upgrading interconnect from standard Ethernet to specialised fabrics can increase the fraction of peak FLOP/s that distributed training sustains by reducing communication overheads, thereby increasing realised . Investing in denser racks and advanced cooling allows more accelerators per square metre and per unit of power, expanding physical compute capacity without new algorithms or chips . This interdependence explains why companies and research institutions consider the full stack when planning AI investments, not just the nominal teraFLOP rating of individual accelerators .
Resource allocation, scheduling, and virtualisation
Because accelerator resources are scarce and expensive, managing their allocation across teams and workloads is a central operational concern. In large environments, compute is abstracted into schedulable units that can be requested and assigned to jobs, often via Kubernetes-based orchestration and higher-level platforms . Templates or profiles describe combinations of GPU count, type, memory allocation, and associated CPU and storage resources so that practitioners can submit workloads without micromanaging individual devices . The scheduler then matches these requests to available nodes, attempting to maximise utilisation while honouring constraints on memory, isolation, and performance .
Techniques such as GPU fractioning, where a single physical accelerator is partitioned among multiple workloads, further complicate the picture by enabling more granular sharing at the cost of potential interference and reduced per-job performance . Virtualisation and containerisation provide environment isolation, but also add layers that must be tuned to avoid bottlenecks in data loading or kernel launch overhead . As a result, the effective compute seen by an individual project depends not only on the data centre's headline capacity but also on organisational policies, queueing disciplines, and the sophistication of resource management tooling .
Schools of thought: compute-centric versus algorithm-centric views
Within the AI community, one can distinguish several positions on the role of compute in driving progress. A compute-centric view emphasises empirical scaling laws suggesting that model performance improves predictably with increased model size, dataset size, and computational budget, provided algorithms are reasonably well-chosen . On this view, access to ever larger compute budgets is a primary determinant of frontier capability, and thus controlling, forecasting, and prioritising compute becomes central to strategy and governance . Proponents often argue that even modest algorithmic innovations are amplified when combined with orders-of-magnitude increases in compute, as seen in the evolution of large language models and multimodal systems .
An algorithm-centric perspective stresses that improvements in architectures, optimisation methods, and data curation can yield substantial performance gains without proportional increases in compute. Advocates point to advances such as more efficient attention mechanisms, sparsity exploitation, or better training curricula that reduce the total operations needed for a given level of performance, effectively moving workloads to a lower for the same outcome . A third, more integrated stance treats compute, algorithms, and data as jointly constraining factors, where progress depends on simultaneously optimising all three. Under this hybrid view, investments in compute must be matched by research into more efficient methods and by strategies for high-quality dataset construction, else returns on additional FLOPs diminish .
Strategic and geopolitical dimensions
As training runs for state-of-the-art models require vast compute budgets, often aggregated in specialised AI supercomputers composed of thousands of accelerators, computational power acquires properties of a strategic resource . Such capacity is scarce, capital-intensive, and geographically concentrated in a small number of cloud providers and research labs, leading to concerns about market power, dependency, and unequal access . Governments and international organisations increasingly view domestic AI compute capacity as analogous to critical infrastructure, similar in strategic significance to energy supplies or advanced manufacturing bases .
This strategic lens raises questions about export controls on advanced chips, incentives for domestic data centre construction, and international coordination on the environmental and security implications of large compute clusters . Nations with limited access to leading-edge hardware may face barriers not only to competing at the frontier of AI capabilities, but also to deploying models tailored to local languages and contexts, potentially exacerbating digital divides . Conversely, concentration of compute in a few jurisdictions and firms creates levers for regulatory oversight, as controlling access to large-scale compute can act as an instrument for managing the pace and direction of powerful AI development .
Environmental and physical constraints
The physicality of AI compute carries environmental and infrastructure implications that are no longer peripheral. High-density accelerator clusters demand substantial electrical power, often measured in tens of megawatts for a single facility, and sophisticated cooling systems to keep chips within safe operating temperatures . As models and training runs scale, the cumulative energy consumption and associated carbon emissions of AI workloads have prompted scrutiny from regulators, researchers, and the public, particularly where power generation mixes are carbon-intensive .
Data centre operators respond with more efficient cooling designs, such as liquid cooling and hot-aisle containment, and with workload scheduling that shifts some computation to periods of lower grid stress or higher renewable availability . Hardware designers contribute by introducing more energy-efficient architectures, lowering the joules per FLOP for both training and inference . Nonetheless, because algorithmic and scale ambitions tend to expand to fill available capacity, there is an ongoing tension between efficiency gains and overall growth in compute demand, making governance of AI compute an environmental as well as a technological issue .
Why AI compute still matters and how it is evolving
Despite periodic claims that algorithmic breakthroughs might decouple progress from brute-force computation, current trends indicate that access to large-scale compute remains a central determinant of who can build and deploy advanced AI systems . Emerging modalities such as large multimodal models, long-context language models, and agentic systems often require significantly greater training and inference budgets than their predecessors, even when architectures are more efficient on a per-parameter basis . At the same time, edge deployments in mobile devices, vehicles, and industrial sensors demand increasingly capable inference under tight power and latency constraints, pushing innovation in specialised low-power accelerators and on-device optimisation techniques .
Looking ahead, the concept of AI compute is likely to become even more nuanced. Architecturally, heterogeneous systems combining different accelerator types may become standard, matching workloads to the most suitable chips within a single cluster . At the software level, advances in compilers, auto-parallelisation, and neural architecture search could make the mapping from high-level models to hardware more automated and efficient, narrowing the gap between theoretical and effective FLOP/s . At the governance level, discussions of responsible AI are steadily incorporating compute audits, reporting of training budgets, and assessments of energy and security implications, embedding computational power into broader frameworks for AI oversight . Far from being a background technical detail, AI compute has become a central lens through which the capabilities, risks, and opportunities of artificial intelligence are understood and contested.

|
| |
| |
|
Read the full brief at the link
Headlines for the last 24hrs
- AI Data Center Expansion and Extreme Weather Strain Power Grids, Prompting Emergency Measures
- Anticipated Shift in US Policy Suggests Deregulation of Artificial Intelligence
- AI Developers Pivot to Direct Drug Discovery, Transforming the Biotech Landscape
- DRAM and Memory Price Hikes Signal Continued Hardware Cost Pressures Despite AI Market Volatility
- Tech Giants and Startups Accelerate Shift Toward Autonomous AI Agents
- Nuclear Startups Hit Milestones as Tech Sector Seeks Zero-Carbon Power for AI
- Digital Content Deletions Highlight the Fragility of Modern Digital Ownership
- Media and Infrastructure Sectors Brace for M&A Wave Amid Take-Private and Consolidation Pressures
- Creative Industries Establish First-of-Its-Kind Partnerships to Navigate AI Integration
- Corporate Culture, Not Generous Leave Policies, Identified as Key Driver of Employee Burnout and Retention
Time window: 2026-07-03T05:00:33.070Z to 2026-07-04T05:00:33.070Z
|
| |
| |
|
"I'm simultaneously extremely AI-pilled and very bullish on humans." - Dan Shipper - Every CEO
Executives face a dual pressure that is unusually stark: frontier AI systems are improving at a pace that threatens to commoditise large swathes of white-collar capability, while competitive dynamics are simultaneously raising the premium on distinctly human judgment, taste, and leadership . For leaders trying to allocate capital, design organisations, and hire talent over the next decade, the risk is not simply underestimating AI. It is misreading how automation and human expertise interact, and therefore backing the wrong organisational bet .
The paradox of automation: more capability, more human work
The dominant narrative around AI has long oscillated between two poles: technology as a mass job destroyer, and technology as a benign assistive tool. Recent model performance data reinforces the first fear. On demanding reasoning benchmarks, top frontier models have moved from low single-digit scores to roughly 44 percent in about a year, and on GDPval - a test of how well models perform economically valuable tasks compared to humans - scores have jumped to around 85 percent in the same window . These figures underpin predictions that up to half of entry-level white-collar roles could be automated away by advanced systems .
Yet the same empirical trajectory points towards a subtler outcome: automation is altering the composition of work rather than deleting it. Current language models are trained on what has been called the visible residue of human competence - code, prose, images, tickets, specs - and they package this into a commodity layer available to anyone with an internet connection . Tasks that used to require highly paid experts, such as writing a decent newsletter draft, producing a YouTube thumbnail, or wiring up a simple product analytics pipeline, are now within reach of a much wider population .
Once these outputs become cheap and ubiquitous, their marginal value collapses. The strategic problem for firms is no longer how to acquire baseline competence, but how to differentiate in a world where the default AI-generated answer is everywhere and mostly acceptable. That push towards differentiation channels demand back towards human specialists who can decide which problems are worth solving, what quality bar matters, and how to assemble AI components into something coherent and distinctive .
In this framing, AI does not remove expert human knowledge work; it expands the volume of work and shifts value towards the hidden layer of human judgment that makes automation economically useful . The more automation arrives, the more there is to specify, critique, orchestrate, and refine - all activities that remain stubbornly human, even as models edge closer to artificial general intelligence .
Human judgment as the smuggled layer of intelligence
A key idea behind the current AI wave is that much of the "intelligence" attributed to large models is actually borrowed from humans. Any impressive system performance typically contains smuggled intelligence: the prompts, evaluations, and iterative corrections supplied by people who understand both the task and the model's limitations . When a model achieves a strong result on a benchmark like GDPval, it is not acting as an independent employee. It is participating in a tightly managed workflow where expert users define goals, tune prompts, interpret edge cases, and stitch multiple outputs together .
That additional layer of human input is rarely visible on a spreadsheet, yet it is essential to value creation. In organisational terms, it means that every supposedly autonomous agent still needs a person on top to ensure alignment with business objectives, manage failure modes, and adjudicate trade-offs . Even where AI agents are given wide latitude - running multi-step tasks across tools, data sources, and systems - they remain a means to a human-specified end . The billions now being invested by model vendors and toolmakers largely aim to make these systems more reliable executors of goals we give them, not independent actors who generate their own strategic intent .
The resulting paradox is clear. As models get more capable, the technical barrier to deploying them falls, but the organisational barrier rises. Every deployment decision requires sharper human thinking about risk, ethics, customer experience, and economic trade-offs. The more ambitious the automation target - job flows, pricing, content, customer support - the more crucial human overseers become, precisely because the cost of subtle errors scales with reach.
Dan Shipper's vantage point: AI-native work as a laboratory
Few observers have tested these dynamics as directly inside their own organisations as Dan Shipper. As co-founder and CEO of a media and software company built around exploring AI and the future of work, he has turned his business into a live laboratory for AI-native operations . Everyone in his organisation, including non-technical staff, is expected to be an early adopter of tools such as Codex, Co-work, and Claude-based environments for coding, writing, editing, and product development .
Under his leadership, the company grew from a handful of people to around 30 after the emergence of modern large language models, while simultaneously automating large portions of routine work . That trajectory contradicts the simplistic assumption that automation primarily shrinks headcount. Instead, AI adoption created more scope for new roles and specialisations, particularly in product management, design, and operational experiment running . His teams use models not only to produce content but to build lightweight applications, internal tools, and data-informed workflows that support a hybrid of human and agent activity .
Shipper's background matters here. Before focusing fully on AI, he worked at the intersection of writing, product management, and entrepreneurship, which gave him an unusually integrated view of how tools, narrative, and organisational incentives intersect . That combination of disciplines shapes a distinctive stance: rather than treating AI as either a pure threat to jobs or a harmless gadget, he sees it as a force that rewires how creative and knowledge work is structured, while opening new space for human differentiation .
Codex, Claude Code, and the "operating system for knowledge work"
One of Shipper's core predictions, shared in conversation with Lenny Rachitsky, is that the future of work will largely unfold inside environments such as Codex or Claude Code - integrated spaces where humans and AI agents co-create artefacts and workflows . Rather than treating models as isolated chatbots or API endpoints, he envisages them as part of an operating system for knowledge work, where documents, code, interfaces, and data pipelines are all constructed in partnership with AI tools .
In this view, traditional command-line interfaces (CLIs) fade because the affordances of natural language interaction and agent orchestration become more powerful and accessible . Engineers, designers, and PMs interact with systems through conversational commands, structured prompts, and high-level intentions, while the underlying environment translates these into code and configuration. Work ceases to be split neatly between "people who write specs" and "people who implement". Instead, the boundary between writing and building blurs, as writers create applications and app builders produce narrative content .
This reconfiguration has direct implications for organisational design. When the environment itself embeds generative capabilities, the value of siloed departments and rigid role definitions goes down. Shipper argues that org charts based on strict functional separation look increasingly out of step with a world where many employees can rapidly prototype, test, and deploy ideas with AI assistance . The emergent pattern is more fluid: small, cross-functional units that rely on a shared AI "backplane" to experiment cheaply and often.
Super-agents and the new centre of the company
Another of Shipper's predictions is that every company will come to rely on a single, highly capable "super-agent" integrated into its core collaboration environment, such as Slack . Instead of dozens of narrow bots, there will be one central agent that employees query, brief, and collaborate with on a daily basis. This agent will have access to company data, systems, and workflows, and will act as a universal interface for operational and analytic tasks .
The existence of such an agent introduces a new role: a human profile responsible for ensuring that the agent truly works for the whole company . That person, or team, designs the agent's objectives, constrains its behaviour, tunes its prompts, defines its access levels, and measures its impact. In a very concrete sense, they translate leadership intent into agent capabilities. This role cannot be meaningfully automated because it sits at the junction of strategy, data governance, security, and organisational culture.
From a market perspective, super-agents turn data and process quality into compounded advantages. Firms that have invested in clear documentation, robust analytics, and sane permissions can empower their agent to carry out more sophisticated tasks safely. Those that have neglected these foundations will find their agents hamstrung, regardless of how advanced the underlying model is. As a result, human-led improvements in information architecture and workflow design become key determinants of AI ROI.
Why Shipper rejects the "SaaSpocalypse" narrative
Against a backdrop of predictions that AI will hollow out the software-as-a-service sector, Shipper has taken a contrarian stance: he would buy SaaS stocks rather than abandon them . The logic rests on how AI economics alter the SaaS model. Instead of vendors absorbing all inference costs inside the application, he expects users to bring their own AI tokens, effectively subsidising part of the compute needed to run advanced features . That shift improves margins and makes it more viable to ship deeply AI-enhanced experiences.
More importantly, the rise of AI does not remove the need for well-designed software. Even if models can generate code and interfaces, the hard work of understanding users, orchestrating workflows, ensuring reliability, and building trust remains. Product managers and designers become more valuable under this regime, not less, because they are the ones who can frame problems that AI can help solve, decide which affordances to expose, and ensure that human agency is preserved in critical flows . When every product can bolt on a generic model, differentiation shifts towards the quality of human-led product thinking.
This is why Shipper predicts that PMs will thrive and full-stack designers will become "superheroes" in the AI era . Their leverage increases because they can use AI as a multiplier on their ability to prototype, test, and iterate, but the underlying value still comes from human taste and judgment about what should exist, not just what can be built cheaply.
Debates, objections, and the spectre of job loss
Critics of Shipper's stance argue that his optimism about human work may be biased by the nature of his organisation: a highly skilled, AI-forward company operating in sectors - content, software, training - that benefit directly from increased knowledge work . They point to projections from AI safety researchers and economists that warn of substantial displacement, particularly among routine white-collar roles such as junior analysts, paralegals, and entry-level coders . In those domains, the argument goes, automation will genuinely reduce headcount rather than just shift people to more creative tasks.
There is also concern that not all firms will invest in the human-layer capabilities Shipper emphasises. Some may pursue aggressive cost-cutting, replacing people wholesale with agents wherever regulators and customers allow. If many organisations follow that path, the aggregate demand for human expertise could still fall, even if the surviving roles become more interesting. Additionally, structural labour market frictions - retraining costs, geographic immobility, credential barriers - make it hard to move displaced workers into new expert positions quickly.
Shipper's counterpoint is not that displacement will never occur, but that the direction of travel for frontier firms shows a different pattern. His own company automated large numbers of internal processes and nonetheless expanded its team, precisely because automation exposed bottlenecks that only humans could clear: deciding which experiments to run, how to interpret ambiguous results, and how to build offerings that integrated AI in compelling ways . He also stresses that any credible AI deployment must be evaluated like any other business investment: leaders need to demonstrate that tools drive sales, increase margins, reduce risk, or build enterprise value, rather than simply cutting visible labour costs .
Another objection concerns overconfidence in human uniqueness. If models continue their trajectory, some argue, they will eventually match or exceed human performance not only in execution but in higher-order strategy and creativity. Shipper responds by pointing to how models are currently grounded: they learn from human-generated data and are shaped by human-defined objectives. As long as humans control the training regimes, safety constraints, and reward signals, human norms will remain embedded in AI behaviour - a fact that both reassures and demands responsibility .
Why this stance matters for CEOs and builders
The practical importance of Shipper's position lies in how it affects executive decision-making. A leader who believes that automation simply substitutes for people will design very different strategies to one who believes that automation amplifies and redirects human work. The former is likely to chase short-term cost reductions, restructuring teams around minimal human oversight. The latter will prioritise building AI fluency at leadership levels, redesigning workflows to integrate agents, and measuring outcomes rather than raw tool adoption .
Shipper's view suggests several concrete implications. First, AI literacy should become a core leadership expectation, not an isolated IT concern . Executives who do not understand what current models can and cannot do will misallocate capital, either overinvesting in speculative automation or underinvesting in leverage points where AI can dramatically improve speed and quality of execution. Secondly, organisations need to build roles explicitly responsible for agent effectiveness, data quality, and workflow integration - the "super-agent owner" becomes as central as traditional operations managers .
Thirdly, hiring strategies should recognise that some roles gain in importance as AI diffuses. Product managers, designers, and hybrid technical-creative profiles stand to gain disproportionate leverage because they are best placed to define the problems AI should attack and to interpret the messy, probabilistic outputs models generate . Finally, firms should expect that their most distinctive advantages will increasingly reside in human-shaped intangibles: culture, brand, narrative, and the quality of collective judgment. Models can replicate patterns in data; they cannot easily replicate the lived experience of teams that have navigated complex markets and honed a particular way of deciding.
For builders and operators, the backstory behind Shipper's stance offers a pragmatic way to resolve the tension between enthusiasm for AI's capabilities and commitment to human-centred organisations. Treat models as powerful tools that commoditise yesterday's expertise but expand the frontier of what is possible. Recognise that each agent, however impressive, still requires a human at the top to define success, manage risk, and connect outcomes to strategy. And design companies where the most valuable work is precisely the kind that models cannot fully absorb: deciding what to build, why it matters, and how to make it meaningfully different in a world where everyone has access to powerful AI.

|
| |
| |
|
"In the context of AI models, "nerfed" (often misspelled as "neerfed") is a slang term borrowed from gaming. It means a model's capabilities, intelligence, or responsiveness have been intentionally or accidentally reduced. Users often observe this when a previously stellar model starts giving shorter answers, making simpler mistakes, or refusing complex tasks." - Nerfed - AI slang
Frustration with declining model behaviour is one of the most persistent themes in day-to-day AI use: long, careful prompts suddenly yield shallow replies; complex coding tasks start failing on edge cases; and models that once felt razor-sharp begin hedging on straightforward questions. Users reach for a compact way to describe this experience, and increasingly settle on a gaming loanword: the system has been "nerfed". Behind the slang lies a substantive set of technical, economic, and governance issues about how frontier AI models are managed after release.
From game balance to model governance
In online games, "nerfing" describes the deliberate weakening of an overpowered weapon, character, or mechanic in the name of balance. Developers reduce damage, slow movement, or adjust abilities so the overall system remains fair . In that context, a nerf is an explicit design intervention, usually documented in patch notes and often debated fiercely by players.
When the term migrated into AI culture, the core intuition survived but the mechanics became less transparent. Users applying "nerfed" to AI are not merely saying "it feels worse"; they are implicitly alleging a change in the underlying model parameters, training, or deployment stack that reduces useful capability relative to an earlier baseline. In other words, they treat AI model performance like a live-game balance problem, except the changelog is often invisible.
A concise modern glossator offers an informal definition tailored to AI: a model has been "quietly degraded in capability, intelligence, or vibes - often without acknowledgment" . That phrasing captures both the technical claim (degradation) and the social accusation (lack of transparency). The term has thus shifted from purely mechanical balance to a critique of how AI providers exercise control over systems that users increasingly experience as critical infrastructure.
What users mean by a "nerfed" model
Practical usage of "nerfed" in AI settings tends to converge on a few behavioural signatures:
- Shallower outputs. Long, multi-step reasoning chains are replaced with short, generic answers, even when users explicitly request detailed analysis.
- Reduced risk tolerance. The model declines tasks it previously handled, citing safety or policy concerns more frequently or more broadly than before.
- Degraded agentic behaviour. Developers report that multi-file code edits, tool-using agents, or complex workflows become less reliable, more hesitant, or more prone to partial completions .
- New mistakes on familiar tasks. Benchmarks, test suites, or anecdotal tasks that once passed reliably begin failing in consistent ways, suggesting a systematic change rather than random fluctuation.
These observations form the empirical basis for claims that a model has been nerfed, especially when they appear quickly after a provider-side update or coincident with public safety announcements. The term operates as a compressed narrative: something changed upstream, that change was not fully disclosed, and users perceive the net effect as a reduction in value.
Why AI models get "worse" over time
The most obvious explanation is deliberate capability reduction: a provider may decide that certain behaviours are too risky, too costly, or too commercially sensitive, and adjust the model accordingly. However, the technical pathways from intention to user experience are more nuanced than a simple slider labelled "intelligence".
One well-documented case involves a safety fine-tuning update that introduced unintended regressions in complex tasks. A provider rolled out changes designed to tighten behaviour around a specific harm category. The update succeeded on those safety goals, but also generalised more broadly, making the system more conservative in long instruction chains, ambiguous requests, and multi-step autonomous workflows . Developers observed incomplete implementations, shorter docstrings, and hesitant multi-file edits, all in workflows that previously performed well .
In technical terms, post-training modifications alter the effective policy the model uses to map prompts to outputs. Safety-oriented fine-tuning or reinforcement learning from human feedback can shift the decision boundary between acceptable and unacceptable responses. If those updates are not carefully constrained, they can inadvertently suppress beneficial behaviours, especially in tasks that superficially resemble risky ones. Users interpret the resulting performance drop as a nerf, whether or not reduced capability was an intended outcome.
Other mechanisms can also produce a "nerfed" feeling without any explicit intent to degrade capability:
- Quantisation and compression. Reducing numerical precision or compressing weights for deployment efficiency can save compute but sometimes harms performance on edge cases or long-context reasoning .
- Instruction-following biases. Updated preference models might prioritise brevity, politeness, or caution over exhaustive analysis, leading to shorter, less incisive responses.
- Guardrail expansion. Newly tightened content policies or broader classification of "unsafe" topics can result in refusals where detailed answers were previously given, even for legitimate professional uses.
- Distribution shifts in usage. As models scale to new user bases and tasks, providers may tune them for average user satisfaction, diluting performance for specialised workflows that early adopters relied on.
From the user's vantage point, all of these pathways look similar: a previously reliable system now appears anaemic. The slang term compresses diverse technical causes into a single accusatory label.
Mathematical view: capability, alignment, and the "alignment tax"
Although "nerfed" is not a formal scientific term, it maps onto recognisable trade-offs in modern model design, especially the relationship between raw capability and safety alignment. One useful way to think about this is in terms of objective functions and post-training constraints.
During initial training, a large language model approximates a function from token sequences to probability distributions over next tokens. At a high level, we can think of the base model as learning a parametrised conditional distribution , where is the input context, the next token, and the parameter vector. Training seeks to minimise validation loss, a scalar summary of prediction error on held-out data, with lower values indicating better language modelling performance .
Post-training, providers introduce additional objectives. For example, they might define a reward function that measures how "aligned" a response to prompt is with safety, helpfulness, or brand tone. Reinforcement learning from human feedback then adjusts the policy towards responses with higher expected reward. Informally, we trade off pure predictive accuracy against alignment objectives, creating an "alignment tax": some high-capability behaviours are suppressed because they correlate with undesirable outputs .
We can sketch this trade-off by imagining two scalar metrics: base capability and deployed capability . Post-training alignment aims to satisfy constraints for various safety conditions . If the feasible set defined by these constraints excludes some of the high-capability behaviours the base model learned, then along certain dimensions. Users experience that gap as nerfing, particularly if they valued the excluded behaviour and do not see the safety benefit in their own use cases.
From this perspective, nerfing is not simply "making the model worse"; it is a re-optimisation of objectives. The controversy arises because users and providers weight different terms in the implicit objective function, and because the parameters of that optimisation are rarely disclosed.
Economic and platform incentives behind perceived nerfs
Beyond safety, economic pressures shape how models evolve. Providers operate under constraints of compute cost, monetisation, and regulatory exposure. Adjustments that appear as nerfs to power users can be rational responses to those constraints.
Cost pressures may encourage more aggressive quantisation, smaller context windows, or throttling of expensive behaviours such as long chain-of-thought reasoning. Safety and compliance pressures may drive broader guardrails around legally sensitive domains. Commercial strategy may prioritise "lite" configurations that feel responsive for casual users while reserving full capability for higher-priced tiers.
Commentary in public forums has argued that apparent degradation is less about technical limits and more about "management for profit": models are tuned, restricted, or tiered to fit business objectives, making them appear weaker compared with early, less constrained iterations . Whether one accepts that framing, it explains why nerfing is often discussed not merely as an engineering choice but as a form of platform governance.
Transparency, trust, and "silent nerfs"
The most contentious cases are "silent" changes: mid-cycle updates that alter behaviour without a clear, detailed changelog. Developers relying on stable behaviour for production workloads can find that their agent workflows or code generation pipelines start failing, with no straightforward way to attribute regressions .
In one analysis, this pattern was framed as "silent manipulation", arguing that unannounced behavioural shifts bake misalignment and erode user trust . When the only observable evidence is that a model feels different, users fill the gap with narrative: the model has been nerfed. The narrative may not match the provider's intention, but it reflects a rational response to opacity.
Where providers later clarify that changes were safety-related and that regressions were unintended spillovers, the underlying grievance persists: the absence of proactive transparency. Developers pressed for acknowledgement when model behaviour changed mid-cycle without documentation, and only then did they receive partial explanations and remediation commitments . In this context, "nerfed" functions as a rallying point for demands that AI systems be governed like critical software, with versioning, release notes, and regression tracking.
Competing interpretations and debates
Not all reports of nerfing are borne out by careful measurement. Some changes in user experience reflect adaptation to new norms, survivorship bias in anecdotal tasks, or shifts in prompt style. In fast-moving ecosystems, users may extrapolate from a handful of bad interactions and generalise too quickly to claims of systematic degradation.
There is thus a tension between subjective and objective assessments. On one side, highly engaged users and developers emphasise lived experience: they see specific workflows break, track sample tasks over time, and share detailed comparative logs. On the other side, providers may present internal benchmarks showing equal or better scores on standard evaluations, arguing that perceived nerfs are either local regressions or artefacts of changed prompts.
This disagreement is sharpened by the opaqueness of training and evaluation data. If the only published metrics are generic benchmarks, they may not capture the agentic or long-horizon tasks that matter most to certain user communities. Slang terms like "nerfed" arise precisely because formal documentation is insufficiently granular to explain real-world shifts.
Relationship to adjacent AI slang and concepts
The AI ecosystem has generated a growing lexicon of informal terms to describe phenomena that formal literature does not yet cover. Glossaries now document expressions such as "slop" for low-quality, repetitive AI-generated content , or "jailbreaking" for attempts to circumvent safety guardrails . "Nerfed" sits alongside these as a user-centric descriptor of perceived capability changes.
Unlike "hallucination" - the industry term for models generating incorrect information - nerfing is not about errors per se, but about systematic reduction of useful capacity. Where hallucinations signal problems in base model learning or inference-time reasoning, nerfing points to deliberate or incidental post-training changes. The term thus fills a conceptual gap: users needed a way to talk about models getting worse in ways that are not straightforward bugs.
Why the concept still matters
As frontier models become embedded in professional workflows, education, research, and everyday tools, the stakes of post-release changes rise. Whether or not providers embrace the slang, the underlying concerns it encodes are durable:
- Stability for production use. Organisations integrating models into software or processes require predictable behaviour. Unannounced shifts can carry direct operational and financial costs.
- Accountability for safety trade-offs. If alignment updates introduce capability regressions, users deserve a clear explanation of what was changed, why, and how performance will be restored or compensated.
- Trust in platform governance. Perceptions of "silent nerfs" erode confidence that providers will be candid about significant behavioural modifications, particularly when those changes reflect commercial or regulatory pressures.
- Democratic oversight of powerful systems. When models are increasingly central to information ecosystems, changes to their behaviour become a matter of public interest, not just product management. Slang terms become vehicles for broader critiques of power and control.
Even as technical language around post-training, alignment, and deployment matures, "nerfed" is likely to remain in circulation because it expresses, in one compact word, a mix of empirical observation and normative complaint. Users are not simply describing a weaker system; they are questioning why it became weaker, who authorised the change, and whose interests the new configuration serves.
Major schools of thought on nerfing
Debate around nerfing in AI roughly falls into three broad positions:
- Safety-first justification. Proponents argue that some reduction in apparent capability is a necessary price for preventing misuse, reducing harmful outputs, and complying with emerging regulation. From this view, complaints about nerfs reflect users undervaluing collective risk management.
- Transparency-critical stance. A second group accepts the need for safety updates but insists they be documented, measured, and reversible when regressions occur. They treat nerfing as a problem of governance rather than an inevitable technical reality .
- Suspicion of commercial motives. A more adversarial position claims that providers intentionally degrade free or lower-tier models to push users towards paid offerings, or to manage infrastructure load, and use safety as rhetorical cover .
These positions are not mutually exclusive. A model update might simultaneously satisfy genuine safety concerns, reduce costs, and affect commercial strategy. The nerfing discourse persists because most users lack visibility into how these motivations are balanced.
Practical implications for users and developers
For practitioners, taking nerfing seriously means treating models as dynamic, versioned dependencies rather than static commodities. Concrete responses include:
- Regression harnesses. Maintaining task suites, benchmarks, and evaluation scripts that can quickly detect behavioural changes when providers update models mid-cycle.
- Model diversity. Avoiding over-reliance on a single provider or model family, so that perceived nerfs can be mitigated by switching or ensemble strategies.
- Prompt and workflow robustness. Designing prompts and agent architectures that are resilient to minor preference shifts, while recognising that large safety updates may still require adaptation.
- Advocacy for changelogs. Pressing providers to publish detailed behavioural release notes and to gate updates behind new version identifiers, rather than silently modifying existing endpoints.
In this landscape, the slang term "nerfed" functions as both diagnosis and signal. It alerts communities to potential regressions and calls attention to the need for more disciplined model lifecycle management. While the word itself is playful, the issues it condenses - capability trade-offs, silent governance, and shifting platform incentives - are central to the future of AI deployment.

|
| |
| |
|
Read the full brief at the link
Headlines for the last 24hrs
- OpenAI Proposes Donating 5% Equity Stake to a US Sovereign Wealth Fund
- AI Infrastructure Demands Strain Power Grids, Triggering Emergency Regulatory Interventions
- Meta CEO Mark Zuckerberg Signals Slowdown in AI Agent Development Timelines
- Global Chip and Tech Stocks Sell Off Amid Growing AI Valuation and Capex Jitters
- Microsoft Launches $2.5 Billion AI Deployment Unit to Embed Engineers Directly in Enterprises
- EV Market Rebounds as Tesla and Rivian Deliver Strong Quarterly Beats
- US Labor Market Cools with Slower June Job Growth and Falling Participation Rates
- Liquidity Pressures Mount in Private Credit as Major Funds Impose Redemption Caps
- Google Loses Final Appeal Against Record $4.7 Billion European Union Antitrust Fine
- Anthropic Restricts Chinese Access to Claude and Explores Custom Chip Partnership with Samsung
Time window: 2026-07-02T05:00:33.072Z to 2026-07-03T05:00:33.072Z
|
| |
| |
|
"Don't stick your head in the sand and say, 'I hate all of this stuff.' That gives you a great feeling of moral superiority and you can go on Bluesky and shout at everybody about how evil AI is. Great, I'm happy for you, but that's not going to help. What helps is you diving into this and coming out understanding what you can do with it today." - Benedict Evans - Independent analyst
The real tension is not between optimism and pessimism about AI, but between posture and practice. Evans is pushing against a kind of performative refusal: the emotionally satisfying move of treating AI as inherently corrupt, then converting that stance into public identity, while leaving the underlying technology untouched. His argument is that rejection may feel morally clean, but it does not change the fact that the tools are already being built, deployed, and woven into workflows across software, media, knowledge work, and administration .
That matters because AI is not arriving as a single finished product with a neat set of social outcomes. It is arriving as a general-purpose capability whose effects depend on how it is embedded. In his broader work, Evans has repeatedly framed AI less as a monolithic intelligence and more as an enabling layer, similar to earlier platform shifts such as the internet and mobile, with the important caveat that the scale is large without being mystical . The practical question is therefore not whether one likes the technology in the abstract, but which tasks it can already improve, which ones it cannot, where it creates leverage, and where it introduces fresh failure modes .
The strategic problem behind the rhetoric
The rhetoric around AI often collapses three separate debates into one: whether the technology is impressive, whether its deployment is socially desirable, and whether its adoption can be slowed or redirected by denunciation. Evans separates those layers. His argument, as reflected in the source material, is that public moral outrage can become a substitute for analysis, especially on social platforms where opposition is rewarded with status signals and shared grievance . That dynamic is attractive because it offers immediate emotional certainty, but it is strategically weak because it does not answer the harder question of adaptation.
The underlying strategic problem is that AI changes the economics of knowledge work by lowering the cost of many low- and medium-complexity tasks, while leaving judgment, integration, and accountability in human hands for longer than the hype cycle suggests. Evans has said in multiple discussions that the most interesting use cases are not abstract claims about artificial general intelligence, but concrete applications where people can ask a model to do useful work today: analyse cancellations, parse documents, guide forms, summarise information, or assist in coding and software development . That is a more awkward story than the grand narrative of replacement, but it is the story that shapes budgets, hiring, and competitive advantage.
Why refusal is a weak strategy
One reason the refusal posture is weak is that it confuses moral evaluation with operational response. A company, profession, or individual may decide that certain AI uses are unacceptable, but that decision does not remove competitive pressure from the market. If rivals are using AI to shorten response times, reduce support costs, or speed up software delivery, then abstention becomes a strategic choice with measurable trade-offs rather than a pure ethical position . The same logic applies to workers: a blanket refusal to learn the tools may preserve a sense of coherence, yet it can also reduce employability in environments where AI becomes part of the expected toolkit.
Evans is not arguing that criticism is illegitimate. His point is that criticism becomes more credible when it is paired with operational understanding. That distinction matters because the most consequential debates about AI are not about whether it is good or bad in a totalising sense. They are about where it can be used safely, where it fails, how it should be governed, and how institutions should absorb it without degrading quality . A refusal to engage leaves those questions to people who are more interested in shipping products than in reflecting on their broader consequences.
The practical value of engagement
Engagement does not mean na?ve enthusiasm. It means acquiring enough familiarity to distinguish between capabilities that are real and those that are theatrical. Evans's earlier writing suggests that the strongest AI thesis is not that one model magically solves every problem, but that a model can become a flexible interface over many tasks, provided the surrounding workflow is designed correctly . That is a subtle but important shift. It moves AI away from the fantasy of autonomous replacement and towards a reality of assisted production, in which the human operator remains responsible for framing the task, checking output, and deciding whether the result is acceptable.
This is also why coding has emerged as a breakout use case. Software is already digital, the feedback loops are fast, and the domain has enough structure for models to provide obvious productivity gains without requiring the full complexity of the physical world . Once a technology proves itself in a high-velocity environment like software development, it becomes easier for businesses to imagine extensions into support, operations, finance, legal workflows, and internal knowledge systems. The significance is not that every task becomes automated. It is that the marginal cost of trying changes, and that alone can alter organisational behaviour.
The moral superiority trap
The language of moral superiority is central because it captures a recurring feature of technology debates: opposition can become a social performance detached from operational reality. Evans's target is not principled disagreement but the shortcut whereby people enjoy the identity benefits of resistance while avoiding the effort of understanding the tool they are criticising . That shortcut is especially tempting in AI, because the technology sits at the intersection of labour, creativity, power, and status. It therefore invites symbolic positioning, not just technical assessment.
But symbolic positioning has limits. If the technology is becoming embedded in search, office software, customer service, design, programming, and content production, then the relevant questions become more granular: which roles are exposed first, what forms of human oversight remain necessary, how quality assurance changes, and what kinds of failures are tolerable . Moral language is often too blunt to answer those questions. It can identify anxiety, but it rarely produces a usable operating model.
What the debate is really about
The deeper disagreement is over pace and shape. Critics often assume that if a technology is important, then its consequences must be immediate, total, and visible. Evans's view is closer to the historical pattern of earlier platform shifts: transformative technologies usually matter first in narrow, messy, and unequal ways before they become generalised . That means the early evidence can look underwhelming to people expecting dramatic collapse, while still being sufficient to change incentives inside firms.
There is also a genuine debate about whether AI is primarily a labour-saving tool or a new layer of infrastructure. If the models themselves become the value-capture point, then the economics may look concentrated. If, instead, the models become a commodity layer on which applications are built, then the value may move upward or outward into workflow software, services, and domain-specific products . Evans's work leans towards the idea that the most enduring effects may sit in how AI is applied rather than in the models alone . That makes the near-term competitive field more complex than the public conversation suggests.
Why this matters now
The urgency comes from the mismatch between how people talk about AI and how organisations actually adopt technology. Public debate rewards certainty, but deployment rewards iteration. A firm cannot manage AI risk by declaring it bad; it must decide where to allow it, where to block it, how to audit it, and how to train staff to use it responsibly. Likewise, an individual cannot assess the impact of AI on their career from slogans alone. They need exposure to the tools, because real displacement or augmentation will be mediated by tasks rather than ideology .
That is the most durable implication of Evans's position: understanding is a form of leverage. The people who stand outside the technology and denounce it may preserve a sense of purity, but they surrender influence over how the technology is used. The people who enter the system, test it, and learn its failure modes are better placed to shape policy, product design, and workplace norms . In a field moving as quickly as AI, that difference is not philosophical. It is commercial, organisational, and increasingly personal.

|
| |
| |
|
"A digital twin is a highly accurate, dynamic virtual replica of a physical object, person, system, or process. When combined with artificial intelligence (AI), the twin moves beyond just replicating real-time data. AI enables the twin to analyse vast datasets, predict future states, and autonomously optimise operations without human intervention." - Digital Twin - Artificial Intelligence
Operational systems increasingly need to be observed, predicted, and controlled as tightly coupled cyber-physical entities, not as isolated machines or spreadsheets of KPIs. The underlying issue is how to turn streaming telemetry, historical records, and domain knowledge into a continuously updated, decision-ready model that can both interpret current behaviour and test future interventions. Digital twins enhanced with artificial intelligence address this by binding data, models, and control logic into a dynamic virtual counterpart that can experiment safely while the physical system keeps running.
From passive monitoring to active optimisation
Traditional monitoring tools collect data and display dashboards, leaving humans to interpret patterns and decide what to do next. Conventional simulations, meanwhile, are static models that run offline on assumed conditions. The practical challenge is that complex assets and processes - an aircraft engine, a hospital, a city-scale energy grid - change faster than static models can be updated and exhibit interactions too intricate for simple rules of thumb. Digital twins resolve this tension by maintaining a high-fidelity virtual representation that is synchronised in real time with the physical counterpart via sensors and data platforms .
When artificial intelligence is added, the role of the twin shifts from passively mirroring reality to actively exploring and optimising it. Machine learning models embedded in the twin learn from time-series telemetry and historical incidents, forecasting likely future states and detecting anomalies that human operators would miss . Optimisation algorithms and AI planners then evaluate alternative actions - maintenance schedules, configuration changes, routing decisions - and recommend or autonomously trigger the option expected to best meet objectives such as cost, safety, throughput, or energy use . This turns the twin into a decision engine rather than a visualisation tool.
Core mechanisms and data flows
The defining mechanism is a two-way data connection between the physical entity and its virtual counterpart. Sensors and operational systems stream telemetry - temperatures, vibrations, pressures, positions, inventory levels, patient metrics - into the twin, often via IoT networks and industrial data platforms . The twin's models process this input to maintain a live state and perform analyses. In the opposite direction, the twin produces recommendations or control signals that influence the physical system, such as adjusting setpoints, rescheduling jobs, or changing treatment plans . This closed loop allows the combination of data and AI to be applied continuously rather than in occasional studies.
Data flows typically include three layers. First, raw telemetry, events, and contextual information (environmental conditions, schedules, configuration metadata) are ingested and harmonised into a common schema. Second, analytics and AI models transform this data into derived features - health indicators, risk scores, predicted demand - and scenario evaluations. Third, decision logic maps these model outputs into specific actions or alerts, factoring in business rules, regulatory constraints, and operator preferences . Over time, feedback from executed actions feeds back into model training, creating a learning loop.
Practical meaning in industrial and service contexts
In manufacturing, the practical consequence is that production engineers gain a live "digital sandbox" of the factory or process. They can simulate different batch sizes, routing rules, or layout changes inside the twin, observe expected utilisation, bottlenecks, and energy consumption, and only roll out changes that perform well in the virtual environment . Predictive models embedded in the twin anticipate machine failures or quality drift, enabling maintenance to be scheduled at convenient times and reducing unplanned downtime . This can translate directly into lower scrap rates, reduced maintenance labour, and higher throughput per line.
In energy networks, twins of turbines, substations, or entire grids enable operators to model the effect of weather patterns, load surges, or equipment outages before they occur. AI components learn relationships between conditions and failures, recommending reconfigurations or dispatch strategies to maintain stability and minimise losses . For built environments - offices, campuses, or smart cities - twins provide a unified view of occupancy, HVAC performance, lighting, and security systems. Optimisation algorithms can then adjust settings to balance comfort with energy efficiency, often cutting consumption without major capital investment .
In healthcare, patient-specific digital twins combine clinical records, imaging, sensor data, and biomechanical models to create virtual avatars that approximate an individual's physiology . AI-supported twins can simulate disease progression under different treatment options, estimate the risk of flares, and help design personalised rehabilitation regimes or dosage schedules . This reduces reliance on trial-and-error in the clinic and enables safer testing of innovative interventions before applying them to the patient.
Conceptual structure and variants
Although implementations vary by sector, most digital twin architectures share three conceptual layers. The first is the descriptive layer: a virtual model that represents geometry, topology, and configuration of the physical entity. This could range from a detailed CAD model of a component to a process flow model of a production line or a systems model of a hospital. The second is the analytical layer: algorithms and AI models that interpret data, run simulations, and compute KPIs. The third is the prescriptive and control layer: the set of rules, planners, and interfaces that turn analytical outputs into operational decisions .
Within this structure, practitioners often distinguish between asset-level twins and system-level twins. An asset twin focuses on a single physical item, such as an engine or a robot, providing detailed health monitoring and failure prediction. A system twin models interactions across multiple assets and processes - for example, an entire factory, a logistics network, or a hospital ward - allowing optimisation of flows, capacities, and resource allocation . As AI capabilities evolve, there is growing interest in person-level twins, particularly in healthcare and elite sport, where the physical counterpart is an individual human rather than a machine .
Mathematical specification and AI integration
Mathematically, a digital twin can be seen as a stateful model that maps observed variables and inputs into a latent system state and predicted outcomes. Let the physical system's state at time be represented by , observed through telemetry , and controllable inputs . The twin maintains an estimate via a state estimation function , such that , where are model parameters learned from data. Predictive components compute future trajectories for steps ahead. AI models implement and as machine learning architectures (e.g. recurrent networks, transformers, probabilistic models) trained on historical sequences and updated with streaming data .
Optimisation is often formalised as selecting control actions to minimise an expected cost function over a horizon . For example, given a cost , where penalises downtime, energy use, or safety risk, the AI planner searches over candidate sequences to find approximately optimal policies subject to constraints representing physical or regulatory limits. In complex settings, this is solved using reinforcement learning, model predictive control, or heuristic optimisation within the twin, leveraging its fast simulation capabilities .
For predictive maintenance, a simple probabilistic formulation might model time to failure as a random variable with distribution inferred from telemetry history. A commonly used approach is to assume a parametric distribution such as or a survival model and update and with incoming data; the twin then recommends maintenance when exceeds a threshold for some lead time . Here, AI improves the fidelity of by capturing non-linear relationships between operating conditions and degradation .
Key parameters and modelling choices
Several parameters critically shape the behaviour and value of an AI-enabled twin. Data latency determines how close the virtual state is to real time; high-latency data can make predictions stale, while ultra-low latency streams demand more robust engineering and edge processing. Spatial and temporal resolution specify how fine-grained the model is - whether it tracks component-level temperatures every second or aggregate performance every hour. Model fidelity reflects how accurately the twin captures physics, control logic, and interactions; higher fidelity typically requires more data and computational effort.
On the AI side, choices include the type of models used for different tasks: time-series forecasting models for demand and degradation, anomaly detection algorithms for spotting unusual behaviour, and causal or counterfactual models for estimating the impact of interventions. Hyperparameters such as learning rates, regularisation strengths, and window sizes influence stability and responsiveness. In autonomous scenarios, policy parameters in reinforcement learning or optimisation weights in multi-objective formulations determine how the twin trades off competing goals like cost versus reliability or speed versus quality .
Major schools of thought
One strand of thinking emphasises engineering roots, treating digital twins primarily as advanced simulation and control models grounded in physics-based representations. Here, the focus is on rigorous system identification, numerical methods, and safety-critical verification. AI plays a supporting role, augmenting classical models with data-driven corrections where physics is incomplete or too complex to model analytically . This perspective is common in aerospace, automotive, and energy sectors with strong engineering cultures.
A second school is data-centric, viewing twins as overlays on large data platforms whose main function is to orchestrate analytics and machine learning at scale. In this view, the twin is less a detailed physical model and more a semantic layer that ties together data from IoT devices, enterprise systems, and simulations into a coherent operational picture . AI becomes central, from predictive models to generative techniques that synthesise missing data and scenario variants. This orientation dominates narratives within Industry 4.0 and AI-first organisations.
A third perspective emerges in healthcare and personalised medicine, where the twin is conceived as a patient or human avatar integrating multi-modal data - clinical, behavioural, genomic - to support prediction and decision-making . The debate here centres on how far a twin can approximate individual physiology and whether such models should be used for high-stakes decisions. AI is essential, but it is constrained by ethical, legal, and interpretability requirements, leading to hybrid models that combine mechanistic descriptions with machine learning .
Tensions, risks, and debates
The first tension concerns autonomy. As AI-enhanced twins become capable of recommending and even executing actions, organisations must decide how much decision-making to delegate. In industrial environments, fully autonomous optimisation might conflict with operator intuition, safety protocols, or union rules. In healthcare, autonomy raises ethical questions about consent, accountability, and the acceptability of algorithmic treatment decisions. The debate is not only technical but institutional: how to structure governance and oversight for systems that continuously act based on learned patterns .
Another debate centres on model realism versus practicality. High-fidelity twins with detailed physics and rich AI components can be expensive to build and maintain, especially when telemetry is noisy or incomplete. Some practitioners argue for minimal viable twins focused on specific high-value use cases - such as predictive maintenance for a handful of critical assets - rather than grand, all-encompassing models of entire organisations. Others view comprehensive twins as necessary infrastructure for long-term digital transformation and systemic optimisation . The choice affects investment profiles, required skills, and expectations about ROI.
Data quality, privacy, and security form a third area of tension. Twins depend on continuous data flows, often from sensitive sources such as medical devices, worker wearables, or critical infrastructure sensors. Poor data quality can mislead AI models, while breaches or misuse can have serious consequences. In healthcare, building patient twins raises concerns about re-identification, surveillance, and secondary use of data. As AI models become more complex, explaining their recommendations to regulators and stakeholders becomes harder, increasing calls for transparency, validation, and robust risk management frameworks .
Why AI-driven digital twins still matter
Despite hype cycles and competing technologies, AI-enabled digital twins continue to matter because they align closely with structural shifts in how organisations operate. As systems become more interconnected and data-rich, the ability to reason about them in a unified, dynamic model - and to test changes safely before implementation - is a durable need. Twins provide a way to convert fragmented telemetry into coherent situational awareness and to use AI not just for offline analytics but for continuous, operational decision support .
In sectors facing tight resource constraints and sustainability pressures, the capability to optimise energy use, reduce waste, and extend asset life via predictive and prescriptive analytics embedded in twins can have material economic and environmental impact . In healthcare, the prospect of more personalised, proactive care based on AI-supported patient twins offers a route to managing chronic disease burdens and complex treatment decisions . As generative AI and agent-based architectures mature, twins are likely to become increasingly conversational and autonomous, orchestrating actions across diverse systems while remaining anchored to the physical realities they represent .
The continuing importance lies not in the label but in the underlying capability: maintaining a living, data-driven model of critical assets, systems, or people that can learn, predict, and act. Artificial intelligence expands this capability from faithful mirroring to intelligent stewardship, turning digital twins into central infrastructure for organisations seeking to manage complexity, risk, and performance in real time.

|
| |
| |
|
"It is called the lump-of-labor fallacy for a reason. Who knew when the internet was born that the internet was going to create a million and a half jobs as Uber drivers? We are in the first or second inning of [the AI] revolution. This is a big paradigm shift, both for the conduct of our policy and for our economies. I think the jobs will be greater and prosperity will be stronger." - Kevin Warsh - Chair of the Board of Governors of the Federal Reserve, CNBC policy panel at the ECB Forum on Central Banking 1 July 2026
Anxiety about artificial intelligence and jobs reflects a deeper confusion about how modern economies generate work, absorb technological shocks, and turn productivity gains into living standards. In public debate, the prospect of widespread automation is often treated as a zero-sum contest in which every task performed by a machine implies a human job lost forever. That framing misdiagnoses both the nature of labour demand and the historical record of past technological upheavals, from electrification and mass manufacturing to the internet and platform economies. It also risks steering policy towards defensive attempts to freeze existing job structures rather than towards the institutional and macroeconomic adjustments needed to harness a new wave of productivity growth.
The lump-of-labour intuition and why it keeps returning
The underlying belief that fuels much contemporary fear is the notion that there is a fixed quantity of work to be done in an economy. If technology allows one group of workers, or machines, to perform more tasks, the remaining workers are assumed to be surplus to requirements. Economists label this the lump-of-labour fallacy: the mistaken idea that the aggregate demand for labour is constant, so any addition of workers or machines simply redistributes a finite stock of jobs.
The doctrine has intuitive appeal. Individuals experience work as slots: a firm has a certain number of positions; if a robot takes one of those positions, a person is displaced. At the level of a single factory or office, that description may be accurate in the short term. But labour demand at the macroeconomic level is a derived demand: it is determined by the demand for goods and services that labour helps produce. When productivity increases, unit costs tend to fall, prices are reduced or quality improves, and new customers enter the market. The resulting expansion of output often raises, rather than reduces, total labour demand, even if the mix of occupations changes.
Historically, each major general-purpose technology has triggered a similar wave of fear. Mechanisation provoked nineteenth-century worries about permanent unemployment among textile workers; electrification reshaped manufacturing employment; computers were forecast to end clerical work. Yet in advanced economies, employment as a share of the adult population rose over long periods, especially for women, while working hours declined. The fallacy persists because transitional pain is real: job losses in specific sectors occur before new roles are visible, and those losses are concentrated in communities and occupations with limited mobility.
Internet platforms and the emergence of unforeseen occupations
The experience of the commercial internet illustrates how technological change can generate new forms of work that would have been difficult to anticipate during its early development. In the 1990s, policymakers discussed the web primarily as an information medium and a communication tool. Few foresaw a global ecosystem of search engine optimisation consultants, social media managers, app developers, and content creators. Ride-hailing services provide a vivid example: by combining smartphones, mapping, digital payments and algorithmic matching, platforms such as Uber enabled millions of people to earn income through quasi-formal driving work, often on a flexible basis that blurred the boundary between employment and self-employment.
From an aggregate labour-market perspective, this was not a straightforward net addition of jobs. Platform driving substituted for some traditional taxi work and for other forms of casual labour. But the key analytical point is that the technology did not merely redistribute a pre-existing lump of tasks. It lowered transaction costs in urban transport, brought new riders into the market, and turned underutilised time and assets into productive inputs. In standard economic terms, a fall in the effective price of rides expanded demand; as demand rose, so did the derived demand for driving services, even if the nature and stability of those roles raised separate regulatory questions.
For central banks and macroeconomic policymakers, the internet boom posed an early version of today's challenge: how to interpret a technology-driven surge in productivity that coexisted with structural labour-market change. Alan Greenspan's decision to allow the US economy to "run hot" in the late 1990s rested on a belief that higher measured and unmeasured productivity growth would allow faster output expansion without igniting inflation. That judgement reflected a willingness to treat technological progress as a disinflationary force, at least while labour supply and institutional frameworks could accommodate workers moving into new sectors.
AI as a new general-purpose technology
Artificial intelligence now occupies a similar, but more pervasive, role in macroeconomic debate. Advances in generative models, pattern recognition and optimisation promise simultaneous shocks to productivity across services, manufacturing, logistics, finance and creative industries. Senior policymakers increasingly frame AI as a general-purpose technology with the potential to reshape both the structure of the economy and the transmission of monetary policy.
Kevin Warsh has argued in speeches and testimony that AI is rapidly reshaping the US economy and that the Federal Reserve will need substantial changes in its models to account for the technology's impact. His view emphasises AI as a significant disinflationary force: by enabling workers and firms to produce more output with the same or fewer inputs, AI should lower the cost of goods and services and relieve some of the upward pressure on prices. That perspective connects directly to the historical experience of the internet and earlier waves of digitisation, in which higher productivity helped central banks maintain or lower interest rates even as economies expanded.
Conceptually, an AI-driven productivity shock can be represented in standard growth accounting. Output depends on capital , labour , and total factor productivity . A simple formulation is . AI, in this framing, operates primarily through : better algorithms, decision-making and automation raise , allowing higher for given and . If monetary policy holds nominal demand roughly stable, higher can lead to downward pressure on prices rather than reduced employment, provided that labour and capital reallocate towards expanding sectors.
The macroeconomic policy tension: productivity vs unemployment risk
The debate among central bankers and economists centres on how quickly and smoothly that reallocation occurs, and whether AI's productivity effects could be so extreme that they overwhelm traditional adjustment mechanisms. At the European Central Bank's annual forum, participants highlighted both the opportunity for efficiency gains and the risk that AI could accelerate bubbles, transmit shocks through financial markets, and destabilise labour demand in ways that conventional models may not capture.
One scenario discussed in policy circles imagines AI-driven automation replacing large segments of routine and cognitive work, from call centres and paralegal tasks to driving and warehousing. If the resulting productivity gains are huge, the cost of these services might plummet. Consumers would have more disposable income, which they could spend on new goods and experiences, creating fresh jobs elsewhere. This is the canonical anti-lump-of-labour mechanism. However, if automation outpaces the creation of new high-demand sectors, or if capital owners capture the bulk of the gains without translating them into broader spending, the adjustment could be slower, leading to pockets of long-term unemployment and downward pressure on wages.
Critics of an overly optimistic view argue that standard economic reasoning relies on assumptions about demand elasticity and redistribution that may not hold in extreme AI scenarios. If machines perform an ever-widening set of tasks, including creative and managerial functions, the share of income going to labour could fall significantly. In that case, the circular flow described in textbooks-workers earn wages, spend them on goods and services, businesses hire more workers-may be disrupted. The lump-of-labour fallacy, they contend, is not a fallacy in every conceivable technological regime; rather, it is an empirical claim about how far substitution can proceed before the structure of income and demand fundamentally changes.
Warsh's stance within the emerging policy landscape
Warsh's position sits within a broader spectrum of central bank thinking. He acknowledges labour concerns but insists that the labour market will evolve over time, with new roles emerging even as some existing jobs are displaced. In his public remarks, he stresses that AI does not repeal the basic logic of competitive markets or the derived nature of labour demand. Instead, he views AI as an accelerant of trends already familiar from earlier technological shifts, such as the internet: greater productivity, lower costs, and new categories of employment that are difficult to forecast in advance.
For the Federal Reserve, such a stance has concrete policy implications. If AI is expected to be a disinflationary force, the central bank may judge that the neutral interest rate-consistent with stable inflation and full employment-is lower than previously assumed. That could justify a more accommodative stance than headline inflation data might suggest, especially if AI-related productivity gains are initially undercounted. However, Warsh has also emphasised that near-term inflation risks have only recently declined and that the Fed must still work to manage elevated prices. AI therefore becomes part of a medium-term story about potential growth and the equilibrium real interest rate, rather than a justification for immediate aggressive easing.
Furthermore, Warsh's comments about updating the Fed's models indicate an institutional recognition that standard forecasting tools-such as Phillips curve-based relationships between unemployment and inflation-may not fully capture AI-induced shifts in labour-market behaviour. If workers can augment their output dramatically using AI tools, measures of slack and productivity may become more volatile. Monetary policy would need to incorporate richer data on technology adoption, sectoral reallocation, and wage dispersion to avoid misreading the signals.
Why the lump-of-labour dispute matters for AI governance
The argument over whether fears of job loss are a fallacious carryover from earlier eras, or a rational response to a uniquely powerful technology, shapes not just interest-rate decisions but the broader governance of AI. If policymakers are persuaded that labour demand is fundamentally elastic and that new jobs will arise to absorb displaced workers, they are more likely to focus on transitional support: retraining schemes, portable benefits, and regional adjustment policies. If, by contrast, they believe that AI could induce a structural decline in labour's share of income and a persistent shortage of suitable jobs, they may consider more radical measures such as universal basic income, public job guarantees, or aggressive regulation of automation.
Economic history offers both reassurance and caution. On the reassuring side, the simple circular-flow reasoning supported by empirical studies suggests that as workers gain new jobs and income, the additional spending increases demand for goods and services and thereby for labour. Labour is not a fixed resource; workers move from declining sectors to expanding ones, and the "economic pie" grows over time. On the cautionary side, globalisation and digitalisation have already produced regions and cohorts experiencing years of stagnation or decline, even while aggregate indicators improved. Institutional context-education systems, social insurance, labour-market flexibility, competition policy-determines how effectively economies reallocate workers.
In the AI context, the distributional question looms particularly large. If AI makes low-wage jobs more productive, allowing the same output with fewer workers, the proportion of middle- and high-wage roles could rise, potentially reducing income inequality by expanding access to high-value services for lower-income households. Yet that benign scenario depends on complementary policies that foster broad AI access, prevent excessive concentration of market power, and mitigate the risks of algorithmic discrimination or exclusion.
Debates, objections and political constraints
Warsh's optimism about jobs and prosperity attracts several lines of criticism. First, some economists argue that general-purpose technologies may differ in their labour impact depending on which tasks they primarily automate. The internet and earlier computing waves largely enhanced information access and communications, creating new industries around digital advertising, e-commerce and user-generated content. AI systems that can handle complex reasoning, code generation and professional services may threaten a wider range of occupations, including many that were previously insulated.
Second, sceptics question the pace at which new roles appear. Historically, job-creating industries have sometimes lagged behind job-destroying innovations, leaving multi-decade stretches of adjustment. If AI accelerates both destruction and creation, but institutional mechanisms for retraining and relocation remain slow, communities dependent on vulnerable sectors could face prolonged stress. The lump-of-labour fallacy, they suggest, may be less a logical error and more a reflection of practical bottlenecks in matching displaced workers to emerging opportunities.
Third, political economy considerations complicate the picture. Large firms investing heavily in AI may have incentives to lobby for regulatory frameworks that favour capital-intensive solutions over labour-intensive ones. If those firms capture outsize productivity gains and return them primarily to shareholders, the macroeconomic feedback loop from wages to demand could weaken. Central banks might then confront a world in which headline productivity is strong, inflation is subdued, but labour participation and wage growth stagnate, challenging traditional mandates focused on maximum employment and price stability.
Finally, there is an epistemic objection: the claim that we simply do not yet know enough about AI's trajectory to make confident statements about its long-run labour impact. Warsh himself acknowledges that key effects remain uncertain and that the Fed must update its frameworks as evidence accumulates. This humility contrasts with stronger pronouncements from both techno-optimists and AI pessimists, underscoring that the argument about the lump-of-labour fallacy is partly a dispute about the reliability of historical analogies in unprecedented circumstances.
Why the argument matters for future prosperity
Despite disagreement on specifics, the discussion reveals a common thread: the prosperity of advanced economies over the coming decades will hinge on how effectively they translate AI-driven productivity into broad-based gains in employment, income and welfare. A policy stance that overestimates the risk of permanent job loss may discourage innovation and lead to defensive regulation that freezes outdated production structures. Conversely, an overly sanguine belief that jobs will always appear could justify neglect of transitional support and structural reforms.
Warsh's framing places central banks in an active role. By recognising AI as both an opportunity and a source of model uncertainty, monetary authorities are encouraged to reconsider assumptions about equilibrium interest rates, inflation dynamics and labour-market slack. At the same time, their mandate does not extend to redistributive policy, education or industrial strategy. That division of responsibilities means that even if AI turns out to be strongly disinflationary and job-creating at the aggregate level, failings in other parts of the policy system could still leave many workers behind.
The unresolved tension between lump-of-labour intuitions and dynamic labour-demand models therefore matters far beyond academic economics. It shapes public expectations, electoral debates and the legitimacy of institutions tasked with managing technological transitions. If societies accept that the amount of work is not fixed but remain sceptical that new work will be accessible and decent, they may demand more explicit guarantees of inclusion. The way central bankers, including Warsh, articulate the relationship between AI, jobs and policy will influence whether those demands are channelled into constructive reforms or into resistance to technological change itself.
Ultimately, the question is not whether AI will eliminate a certain share of current occupations, but whether economic systems can be steered so that higher productivity expands the range of meaningful, well-paid work rather than narrowing it. Contesting the lump-of-labour fallacy is one part of that steering: it insists that labour demand can grow with innovation. The harder task lies in building the institutional scaffolding-skills, mobility, safety nets, competition frameworks-that ensures the expansion of jobs and prosperity that optimistic policymakers anticipate becomes a lived reality rather than a theoretical promise.
!["It is called the lump-of-labor fallacy for a reason. Who knew when the internet was born that the internet was going to create a million and a half jobs as Uber drivers? We are in the first or second inning of [the AI] revolution. This is a big paradigm shift, both for the conduct of our policy and for our economies. I think the jobs will be greater and prosperity will be stronger." - Quote: Kevin Warsh - Chair of the Board of Governors of the Federal Reserve, CNBC policy panel at the ECB Forum on Central Banking 1 July 2026](https://globaladvisors.biz/wp-content/uploads/2026/07/20260701_17h45_GlobalAdvisors_Marketing_Quote_KevinWarsh_GAQ.png)
|
| |
| |
|
Read the full brief at the link
Headlines for the last 24hrs
- Geopolitical Tug-of-War Over AI Intensifies as Anthropic's Export Controls are Lifted and OpenAI Proposes Trump Administration Stake
- Bending Spoons' Blockbuster $18B+ Nasdaq Debut Defies SaaS Slump
- Meta and Tech Giants Pivot to Selling Excess AI Compute and Cloud Capacity
- The Escalating Energy and Infrastructure Crisis Driven by AI Data Centers
- Monetary Policy and Inflation Concerns Dominate Global Central Bank Agenda
- Google Hit with Landmark $1.5B+ Antitrust Damages Awarded to Klarna
- Consolidation in Retail and Logistics: Kroger's $1.65B Giant Eagle Acquisition & FedEx's $1.4B Supply Chain Sale
- GLP-1 Weight Loss Drugs Enter Mass-Market Era with Medicare Coverage and Price Drops
- Donald Trump's Crypto Windfall Highlights Growing Intersection of Politics and Digital Assets
- Supply Chain Geopolitics: Polestar's US Ban and Apple's Chinese Chip Lobbying
Time window: 2026-07-01T05:00:33.070Z to 2026-07-02T05:00:33.070Z
|
| |
|