This is the second part of the series. Part 1 showed how AI infrastructure maps evolving models onto evolving hardware. This part explains why that mapping is fast-moving and challenging: the workload, hardware, and useful optimization frontier all change at once. Both the work inside one goal and the number of machine-generated steps behind it can grow.
Inference demand still outruns one accelerator
Key insightThe public full-context inference envelope has expanded much faster than the dense compute available from one accelerator.
Inference demand versus one accelerator
Normalized to BERT-Large request work and V100 compute in 2018; logarithmic scale
For a dense Transformer, a useful first-order proxy for one request is active parameters multiplied by processed tokens: (W_{\text{request}}\propto P_{\text{active}}T). This intentionally leaves out attention, KV-cache traffic, output decoding, repeated samples, and agent steps; all of those widen the real systems envelope.
Using public model specifications gives an estimated envelope. BERT-Large exposed 340 million parameters and 512-token sequences in 2018.1 GPT-2 increased the model to 1.5 billion parameters in 2019, and GPT-3 reached 175 billion parameters with a 2,048-token context in 2020.23 PaLM was a dense 540-billion-parameter model, while Llama 3.1 combined 405 billion parameters with a 128K context window.45
The 2026 open-model frontier changes how that curve should be read. GLM-5.2 has a 753-billion-parameter checkpoint, DeepSeek-V4-Pro reports 1.6 trillion total parameters, and Kimi K3 reports 2.8 trillion.678 All three are sparse MoE models with roughly 40–50 billion parameters active per token and about a one-million-token window. Total parameters therefore describe weight capacity, memory footprint, expert placement, and communication pressure; active parameters are the better first-order input to per-token arithmetic.
These are not average production requests; they show how much inference work a single request can ask the system to carry. The hardware curve uses NVIDIA’s published dense FP16/BF16 figures, with C870 FP32 as a legacy proxy and Rubin marked preliminary.9
Show the normalized inference-demand assumptions
| public model example | year | total parameters | active parameters | advertised token window | total-parameter index | request-work index |
|---|---|---|---|---|---|---|
| BERT-Large | 2018 | 340M | 340M | 512 | 1x | 1x |
| GPT-2 | 2019 | 1.5B | 1.5B | 1,024 | 4.4x | 8.8x |
| GPT-3 | 2020 | 175B | 175B | 2,048 | 515x | 2,059x |
| PaLM | 2022 | 540B | 540B | 2,048 | 1,588x | 6,353x |
| Llama 3.1 405B | 2024 | 405B | 405B | 128K | 1,191x | 297,794x |
| GLM-5.2 | 2026 | 753B | approximately 40B* | 1M | 2,215x | approximately 229,779x |
| DeepSeek-V4-Pro | 2026 | 1.6T | 49B | 1M | 4,706x | 281,480x |
| Kimi K3 | 2026 | 2.8T | approximately 50B* | 1M | 8,235x | approximately 287,224x |
The request-work index divides (P_{\text{active}}T_{\max}) by the BERT-Large baseline. It is a deliberately simple compute envelope, not a latency benchmark. Maximum context is not average context; prefill and decode have different execution shapes; and memory capacity, bandwidth, attention, batching, quantization, and KV-cache reuse determine how much of peak arithmetic becomes useful throughput. GLM-5.2 and Kimi K3 do not publish an explicit active-parameter total in their launch material; the approximate values are derived from their disclosed MoE routing and checkpoint architecture, so they are shown as estimates rather than measurements.
OpenAI publishes GPT-5.6 Sol’s 1.05M-token context but not its parameter count.10 A 2–4T total-parameter range is sometimes proposed as an external estimate; the chart shows that range as an unverified annotation and excludes it from both trend lines and index calculations.
This is the space AI infrastructure has to close. It increases effective hardware through lower precision, kernels, batching, parallelism, and communication overlap. It reduces demand through caching, sparsity, smaller routed models, speculative decoding, retrieval, and fewer wasted tokens or retries. The middle layer does not repeal the gap; it decides how much useful product can fit inside it.
The workload multiplies
Key insightA traditional application pauses for the next human decision; an agentic application can turn one delegated goal into many machine-paced branches, steps, and retries.
Pre-LLM consumer workloads usually keep most of these factors bounded. A person clicks, watches, scrolls, types, or plays; the user remains in the loop and paces expensive work. A video frame, feed request, or game tick can be optimized aggressively, but its shape does not expand because a model gained parameters or a context window doubled.
Agentic systems remove that pacing limit. One goal can launch several agents. Each agent can issue model calls, retrieve data, invoke tools, run tests, retry failures, and verify results for hours. At the same time, larger models, longer context, longer outputs, multimodal inputs, and inference-time reasoning increase (T) and (q). Brown et al. showed that repeated sampling with verification can improve measured solution coverage across large increases in sample count.11
Strong and weak scaling are the systems analogy
Key insightTraditional infrastructure is usually asked to finish bounded, human-triggered work faster; agentic products spend added capacity on more work inside each goal.
Strong scaling
Fixed total work; communication creates a floor
Weak scaling
Fixed work per processor; total work grows with resources
The analogy becomes clearer when comparing the control planes. Traditional Dev Infra and ML Infra automate large jobs, but humans usually decide when a build, test, deployment, or experiment begins. Agentic Infra also automates the creation of new work inside the goal: the next model call, tool action, retry, verification pass, or concurrent subtask.
| dimension | traditional Dev / ML Infra | current Agentic Infra |
|---|---|---|
| trigger | engineer submits a build, test, deployment, or experiment | user delegates a goal once |
| unit of work | bounded pipeline or predefined job graph | evolving task tree of model calls, tools, and checks |
| pacing loop | system finishes and waits for human feedback | machine decides and launches the next step |
| work per trigger | mostly fixed after submission | compounds across agents, steps, tokens, retries, and tools |
| concurrency | bounded by teams, commits, and planned experiments | many agents and branches can run for one user concurrently |
| natural limiter | human attention and decision latency | compute budget, policy, quality threshold, and infrastructure capacity |
| optimization target | turnaround time, reliability, reproducibility, utilization | useful completed work per dollar, watt, second, and user goal |
Traditional Dev/ML demand can still grow substantially, and one training experiment can be enormous. The important bound is the request-generation loop: new jobs usually enter at human speed. Agentic demand can grow much faster because one human decision creates a machine-paced loop. Its work envelope multiplies as (A\cdot S\cdot T\cdot q), so increases in autonomy, model size, context, reasoning, and concurrency can produce an exponential-looking demand curve without requiring exponentially more users.
Strong and weak scaling have precise meanings in parallel computing.
For strong scaling, total work (W) is fixed and more parallel resources (P) reduce completion time:
For weak scaling, work grows with resources so that work per processor stays approximately constant:
Dev/ML and agentic systems are not literally parallel-computing benchmarks, but the analogy is useful. Traditional infrastructure mostly asks the system to finish a bounded, human-submitted job faster: a strong-scaling-shaped objective. Agentic products repeatedly spend new capacity on larger models, longer contexts, more reasoning, more branches, and more autonomous steps. The workload expands with the available system, which makes it weak-scaling-shaped.
The bottleneck therefore moves instead of disappearing:
- More GPUs expose collective communication, topology, and straggler costs.
- Longer context turns KV-cache capacity and bandwidth into serving constraints.
- Larger batches improve throughput while increasing latency and memory pressure.
- MoE routing reduces active compute while adding placement and load-balancing problems.
- Test-time compute improves quality while making per-goal cost less predictable.
Why every efficiency point is magnified
Key insightA fixed percentage improvement saves more absolute money as agents multiply branches, steps, tokens, and compute per token.
If an optimization leaves a fraction (r) of the original unit cost, the percentage is ordinary; the multiplying workload base is not.
As (A), (S), (T), and (q) grow, the same kernel, compiler, cache, quantization, batching, or routing improvement acts on more work. This is the economic meaning of “the more you buy, the more you save”: not that scale makes waste acceptable, but that each percentage point of efficiency converts into a larger absolute saving.
Good infrastructure improves both sides of the equation:
- Increase (\eta): execute each unit of work more efficiently.
- Reduce (W): finish the goal with fewer tokens, tool calls, retries, or model passes.
The AI infra control loop
Key insightAI infrastructure repeatedly measures the workload, finds the binding hardware constraint, and changes models, kernels, compilers, or scheduling to recover useful capacity.
Hardware no longer provides a uniform free lunch. Single-thread performance stopped scaling automatically; dark silicon made power limits explicit; and data movement can cost orders of magnitude more energy than arithmetic.121314
Modern accelerators answer with HBM, larger SRAM, tensor cores, lower precision, faster fabrics, and advanced packaging. Infrastructure turns those features into useful work through a recurring loop:
- Measure the real workload: prefill/decode mix, cache residency, batch distribution, communication, and kernel hotspots.
- Change model architecture or serving policy when the bottleneck is structural.
- Change kernels, compiler lowering, layouts, and runtime scheduling when execution is the bottleneck.
- Feed the remaining constraints back into model and hardware design.
That is why AI infrastructure has more room and more responsibility than the conventional middle layer. It can change not only how cheaply a product runs, but what the product can afford to attempt. Part 3 examines the supply side: compute, memory, communication, and edge hardware, and what their different trajectories imply for the future stack.
References
-
Jacob Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, 2018; Google Research, BERT model repository, documenting the 340M-parameter BERT-Large model and 512-token sequence length. ↩
-
OpenAI, Better Language Models and Their Implications, 2019; Alec Radford et al., Language Models are Unsupervised Multitask Learners, documenting GPT-2’s 1.5B parameters and 1,024-token context. ↩
-
Tom Brown et al., Language Models are Few-Shot Learners, 2020, documenting 175B parameters and a 2,048-token context. ↩
-
Aakanksha Chowdhery et al., PaLM: Scaling Language Modeling with Pathways, 2022; Reiner Pope et al., Efficiently Scaling Transformer Inference, documenting dense 540B-parameter inference with a 2,048-token context. ↩
-
Meta, Introducing Llama 3.1, 2024, documenting the 405B model and 128K context window. ↩
-
Z.ai, GLM-5.2: Built for Long-Horizon Tasks, 2026, documenting the 1M-token context and sparse-attention architecture; the official GLM-5.2 Hugging Face repository reports a 753B-parameter checkpoint, while its configuration exposes 256 routed experts with eight selected per token. The approximately 40B active estimate uses the disclosed architecture and the GLM-5 technical report, which reports 40B active parameters for the closely related 744B checkpoint. ↩
-
DeepSeek, DeepSeek V4 Preview Release, 2026, documenting 1.6T total and 49B active parameters for V4-Pro with a 1M-token context. ↩
-
Moonshot AI, Kimi K3 Tech Blog, 2026, documenting 2.8T total parameters, a 1M-token context, and 16 selected experts among 896. The approximately 50B active value is an architecture-derived estimate rather than an official published total. ↩
-
NVIDIA, Tesla technical brief, P100 architecture, V100 architecture, A100 architecture, H100 specifications, DGX B200, and preliminary Vera Rubin NVL72 specifications. The normalized hardware series uses 0.518, 21.2, 125, 312, 989, 2,250, and preliminary 4,000 TFLOP/s for C870, P100, V100, A100, H100, B200, and Rubin respectively; C870 uses FP32 as a pre-native-FP16 proxy. ↩
-
OpenAI, GPT-5.6 Sol model documentation, 2026, documenting a 1.05M-token context window. OpenAI’s release material does not disclose total or active parameter counts. ↩
-
Bradley Brown et al., Large Language Monkeys: Scaling Inference Compute with Repeated Sampling, 2024. ↩
-
Herb Sutter, The Free Lunch Is Over, 2005. ↩
-
Hadi Esmaeilzadeh et al., Dark Silicon and the End of Multicore Scaling, 2011. ↩
-
Mark Horowitz, Computing’s Energy Problem, ISSCC 2014. ↩
Discussion
Comments