Skip to content
AI Infra Weekly2026 · 08 · 30

NVIDIA Posts $96.2B Quarter, Guides to 70% Growth, AWS Adds Two Million GPUs, GLM-5.3-Flash Matches Opus 4.8 at 1/40 the Price

Explore NVIDIA’s $96.2B quarter and 70% growth outlook, AWS’s two-million-GPU expansion, OpenAI’s inference ASIC, and GLM-5.3-Flash’s price challenge to Claude Opus 4.8.

Over the past week, the inference business was priced by three parties at once. On August 26, NVIDIA disclosed quarterly revenue of $96.2 billion with $89.0 billion from Data Center, and took the unusual step of guiding next fiscal year's growth to 70% — a guide that explicitly assumes zero contribution from Data Center compute in China. The same day, AWS announced it would deploy two million more NVIDIA GPUs across 2027–2028 and bring NVIDIA's custom high-bandwidth memory into its own Trainium. At Hot Chips, held August 23–25, OpenAI disclosed the full architecture and third-party measured curves for Jalapeño, NVIDIA announced that Groq 3 LPX had entered full production, Intel introduced an inference card carrying up to 480GB of memory, and Google split its eighth-generation TPU into separate training and serving lines. On August 26, Zhipu open-sourced GLM-5.3-Flash, which reaches a composite intelligence score level with Claude Opus 4.8 on 320B total and 18B active parameters, priced at $0.15 per million input tokens and $0.50 per million output tokens, with public traffic running on domestic chip clusters. In the same week, prosecutors in Keelung, Taiwan indicted nine people over 74 B300 servers rerouted to mainland China along three paths, and MCP published a new roadmap listing agent identity and enterprise-ready security among its five priority areas.

AgentsFlare is the enterprise AI control plane — as models, clouds, and agents keep fragmenting, it keeps routing, cost attribution, and access audit under your control.

AI Infra Weekly is AgentsFlare's strategic column for enterprise teams, tracking the pivotal shifts across the global AI infrastructure layer. By design, a control plane backs no single model or cloud — so we read the structural shifts in models, compute, and regulation without a stake in who wins, and chart the direction before the landscape hardens.

$96.2 Billion, and One Line That Reaches Into Next Year

On August 26, NVIDIA reported results for the second quarter of fiscal 2027, ended July 26: revenue of $96.2 billion, up 106% year over year and 18% quarter over quarter, above the $92.2 billion consensus; Data Center revenue of $89.0 billion, up 117% year over year, also above the roughly $85.7 billion analysts expected. Q3 guidance is $108 billion plus or minus 2%, a range of $105.84 billion to $110.16 billion. GAAP and non-GAAP gross margins are both guided to 74.0%, plus or minus 50 basis points.

What is rare here is that annual outlook: NVIDIA gave a figure of roughly 70% revenue growth for the next fiscal year, against a prior analyst consensus of 44%. Chip companies normally guide only one quarter ahead; putting a full year's growth rate into public communication amounts to putting order visibility on display. In the results statement, Jensen Huang attributed the acceleration to expanding adoption of agentic applications, demand for the Blackwell architecture, and the first production shipments of the next-generation Vera Rubin platform.

The guidance carries one further qualifying condition: the $108 billion Q3 expectation assumes zero Data Center compute revenue from China. That sentence moves geopolitics from the variable column to the constant column. If export policy loosens at some future point, that revenue is pure upside; if it tightens further, the guidance is unaffected. For buyers, this means the tightness in production scheduling and lead times will not ease in the near term because of a policy change, since the supply allocation assumption never counted the China market in the first place.

Revenue and the growth outlook both beat expectations, yet what moved the stock that day was the 74% gross margin figure. It closed down 1.59% at $209.66 before the release and rose 4.71% to $219.53 after hours, with the earlier decline coming from a gross margin guide below expectations. That level is high by semiconductor standards, but the step down from prior levels points to one thing: memory and advanced packaging costs are rising, and supply chain price increases are working their way into the financials of the most upstream vendor itself. Where that cost ultimately lands depends on the price elasticity of end demand, and no elasticity is visible at present.

AWS Orders Two Million More GPUs and Brings NVIDIA's Memory Into Its Own Silicon

On the same day, AWS and NVIDIA announced the deployment of two million additional Blackwell Ultra, Rubin and Rubin Ultra GPUs across AWS global infrastructure in 2027–2028. For comparison, the plan AWS announced at GTC this year was more than one million starting in 2026, and the announcement's wording is that demand exceeded those expectations. The two-million scale itself indicates that hyperscaler purchasing has not slowed in response to rising prices.

The deepest change in the same announcement is to memory. NVIDIA introduced NVHBM that day, moving the memory controller from the accelerator compute die into the HBM base die. The stated gains against standard HBM4E are up to 30% more bandwidth, 15% lower HBM power, and 25% more usable area on the compute die. The first named collaborator is AWS's Annapurna Labs, whose next-generation Trainium will use both NVIDIA's interconnect and its custom memory through NVLink Fusion.

A cloud provider that designs its own accelerator handing both the interconnect standard and the memory specification to its largest competitor to define carries more weight than an order for two million GPUs. One of the original goals of custom silicon was to reduce dependence on a single supplier; once the custom chip's memory subsystem is designed to the other party's specification and the in-rack interconnect runs that party's protocol, the dependency moves back from the procurement layer to the architecture layer. For enterprise customers, the near-term benefit is that Trainium and NVIDIA GPUs can be scheduled within the same rack-scale architecture, lowering the engineering difficulty of heterogeneous placement; the longer-term point to watch is that custom cloud silicon carries less weight as a bargaining chip, and the price gaps on the rate card may no longer be negotiable the way they were over the past two years.

The remaining items sit across CPUs, government orders, model distribution and data processing: the Vera CPU will come to AWS; the two plan to build AI factories for the U.S. government, including 100,000 GPUs on secure infrastructure capable of carrying federal and national-security workloads at Impact Level 6 and above; NVIDIA's Nemotron family of open-weight models remains available on Amazon Bedrock and SageMaker; GPU-accelerated data processing on Amazon EMR using cuDF is stated to run up to 3.7 times faster with 30% better price performance; and offloading vector index construction on OpenSearch to dedicated GPUs is said to deliver up to 9 times faster indexing at a quarter of the cost. All of these figures are vendor-stated, with no third-party reproduction yet.

Three Inference-Specific Chips on the Same Stage at Hot Chips

Hot Chips ran August 23–25 at Stanford. This year's theme was tightly focused: what was a roadmap last year is silicon this year, and nearly every part of it was designed for agentic workloads.

OpenAI disclosed Jalapeño's full architecture and measured data for the first time. The inference ASIC, co-developed with Broadcom, carries 216GB of HBM4 with up to 15.4 TB/s of bandwidth and delivers 3.4 MXFP8 PFLOPS and 13.4 MXFP4 PFLOPS at 700W, currently running at 1.70GHz in silicon with an engineering plan to reach 1.80GHz. Architecturally it uses a NUMA-style memory-sliced design in which each of 64 compute slices is paired with its own HBM slice, with a dedicated high-bandwidth, low-latency collective network moving data between slices and a separate general-purpose on-chip network — deliberately weaker than convention — handling less common accesses, the intent being to keep performance-critical traffic off the general path. On scaling, 128 chips per rack are interconnected over Ethernet at 600 GB/s, and a 16-rack pod holds 2,048 chips at 200 GB/s each, with the complete system providing 27 EFLOPS of MXFP4 compute, 432TB of HBM4 and 32 PB/s of aggregate memory bandwidth; the switches are Broadcom Tomahawk 6 and the hardware is built by Celestica.

OpenAI did not compete on peak compute. It used SemiAnalysis's InferenceX benchmark to run full latency-versus-throughput curves and normalized by package power — 700W for Jalapeño, 1,200W for GB200, 1,400W for GB300. The conclusion is 1.5 to 1.9 times higher peak throughput per watt, 1.7 to 3.6 times lower end-to-end latency, and 2.1 to 4.1 times lower minimum time between tokens. The 104.3x figure in the announcement needs to be read carefully: it refers to the throughput multiple Jalapeño can sustain at GB300's lowest-latency operating point, a methodology clearly favorable to low-latency scenarios that does not represent an overall performance gap.

The nine-month development pace is another point of information. The first line of RTL was written in February 2025, tapeout came that November, first silicon arrived in May 2026, and Codex was running on it that same month. More than half of the core was written in the XLS hardware language, with OpenAI's own models searching for improvements in power, performance and area; against human baselines, the BF16 multiplier improved 56%, the FP4 dot-product block 21% and the FP32 accumulator 10%, with matrix and SIMD unit area shrinking 10% and 8% respectively. The second generation is close to tapeout and the third is in development. Using models to design chips that run models compresses the custom silicon cycle from years to nine months, and once more vendors copy that path, the number of accelerator varieties will be far larger than in recent years.

On August 24, NVIDIA announced that Groq 3 LPX had entered full production. Fabricated by Samsung, the chip integrates 500MB of on-die SRAM to bypass the memory bandwidth bottleneck and is positioned as a decode-phase extension of the Vera Rubin platform, at up to 256 chips per rack. Third-party measurement by Artificial Analysis recorded 3,431 output tokens per second on a 100K-context Gemma 4 31B reasoning workload, roughly four times the 870 tokens per second of the next-fastest public endpoint. The product line comes from a $20 billion transaction NVIDIA closed in December 2025, structured as a non-exclusive IP license plus the hiring of the team; Groq the company continues to operate independently and closed a $350 million round on August 17 while pivoting to building NVIDIA clouds.

Intel, for its part, introduced Crescent Island: 32 Xe3P cores, 256 XMX engines, 32MB of unified L2 cache and a PCIe Gen5 x16 interface in a 350W air-cooled PCIe add-in card, with Intel's own card carrying 160GB of LPDDR5X and ODM versions going up to 480GB. Using LPDDR5X rather than HBM is an explicit trade-off: bandwidth falls short of HBM, but per-card capacity reaches the highest level among current accelerators and the cost structure is entirely different. Inference scenarios with long context and resident large-model weights have high capacity requirements, and this card targets precisely that class of workload. Intel also disclosed the 256-core next-generation Xeon, Diamond Rapids, at the same event. Google split its eighth-generation TPU into two lines, one biased toward training and one toward serving, explicitly partitioned along agentic-era workloads. Meta's MTIA, Arm's server CPU and AMD's Instinct MI400 series were all on the same agenda.

Within one week, the available silicon forms on the inference side went from two or three to seven or eight, and their programming models differ from one another. Jalapeño uses a low-level environment within the open-source Triton ecosystem, where software must explicitly specify where tensors sit physically on the chip and how they are distributed; once the next generation changes the core count or local memory organization, the existing placement and scheduling scheme has to be redone. By contrast, code written for Blackwell runs on Rubin and successors without modification. That difference determines how much of the new inference supply can actually be converted into bargaining leverage — whether you can switch depends on the software layer, not the hardware layer. What the enterprise side can do is keep that difference out of application code: on AgentsFlare, the same calls are routed across different backends by task type, cost budget and SLA, so when the underlying chip or deployment location changes, what changes is configuration rather than the code in each business line.

GLM-5.3-Flash: Opus 4.8's Score at One Fortieth the Price

On August 26, Zhipu released and open-sourced GLM-5.3-Flash on HuggingFace under the MIT license, with 320B total and 18B active parameters, the first natively multimodal model in the GLM-5 series. It scores 57 on the Artificial Analysis Intelligence Index, above GLM-5.2 and level with Claude Opus 4.8; on Zhipu's own Z.ai Code Bench, coding performance is described as comparable to Opus 4.8, the latter being a company self-assessment. The context window is 1M and maximum output is 128K.

The architectural changes all point at the unit cost of long context. The model uses a hybrid of linear and sparse attention, with linear attention capturing local dependencies through a recurrent mechanism and sparse attention recalling global context through a lightweight indexer, and a weighted pooling module compressing the indexer's four cache vectors into one. The stated comparison against GLM-5.3 is a 3.01x reduction in attention compute and a 4.44x reduction in KV cache size. Both figures map directly onto the unit cost of serving long context, and they explain the gap on the rate card.

The list price is $0.15 per million input tokens, $0.03 for cached input and $0.50 for output; through 24:00 on September 9 it is half price, at $0.075, $0.015 and $0.25. The official framing is one tenth of GLM-5.3 and one fortieth of Claude Opus 4.8. For lateral comparison: Claude Opus 5 is $5 and $25, Fable 5 is $10 and $50, and Sonnet 5's promotional $2 and $10 expire on August 31, reverting to $3 and $15 on September 1. At the same composite intelligence score, rate cards differ by one to two orders of magnitude, a gap now too large to explain by fine differences in model capability.

The hardware serving this traffic is a domestic chip cluster. Zhipu states in its technical write-up that over the past week it served large-scale traffic on domestic chip clusters for the first time, with chips connected by a self-developed high-bandwidth interconnect. To work around limits in per-card compute and memory capacity, the team built its own inference engine on top of SGLang, using intra-node tensor parallelism, W8A8 quantization, mixed INT8/FP8/BF16 cache quantization and layer splitting for memory optimization, and adopted a production architecture at the cluster level that separates encode, prefill and decode into independently scalable worker pools. The company states that end-to-end serving performance improved threefold over the initial baseline on the same hardware, and that hardware efficiency and per-token cost have reached parity with mainstream NVIDIA GPUs. Press reports put the cluster at roughly 100,000 domestic accelerator cards; that figure comes from a single exclusive report and has not been formally disclosed by the company. Moore Threads announced same-day adaptation on its MTT S5000.

For the enterprise side, start with price. A model with a composite intelligence score at closed-source flagship level, with downloadable weights that can be privately deployed and an output price compressed to one fortieth of a flagship, will directly rewrite the cost model for high-volume, low-sensitivity workloads such as batch processing, document generation and code completion — route by task tier, move that class of workload to the low-price tier, and the bill changes by multiples. On AgentsFlare, per-project token budgets and four-dimensional cost attribution reflect the actual effect of such a switch as it happens, rather than surfacing the difference only when the month-end bill arrives. Then supply. The claim that per-token cost on domestic cards matches NVIDIA, if it can later be verified by third parties, would mean that inference supply on the two sides becomes comparable on cost for the first time, while the compliance requirements, data residency rules and available model lists on each side are not the same.

Seventy-Four B300 Servers Reached Mainland China Along Three Routes

On August 24, prosecutors at the Keelung District Prosecutors Office in Taiwan indicted nine people, alleging they rerouted NVIDIA B300 servers to buyers in mainland China. The defendants include a manager at NVIDIA's Taiwan office and two employees of Super Micro's Taiwan branch; prosecutors identified the NVIDIA manager as the key figure, alleging he was responsible for authorizing the release of the B300 GPUs.

The routes laid out in the indictment are highly specific. The 130 approved B300 units were released in three tranches of 2, 64 and 64; of the 74 units already traced, three simultaneous routes were used — 16 shipped directly to mainland China in January 2026, 50 went first to Indonesia and were transshipped, and 8 went to a Japanese entity controlled by the defendants before passing through Hong Kong into mainland China. Some defendants are alleged to have created fake websites and falsified information to evade export controls. Prosecutors are seeking the statutory maximum of five years for four of the nine.

The failure point for export controls along this chain lies not in the regulatory text but in the authorization step. A person inside the vendor with release authority, a transshipper within the distribution system, and a shell entity in a third country are together enough to move restricted equipment out. The more such cases are prosecuted, the more predictable the follow-on: stricter channel verification by vendors, more detailed end-user statements, more pre-delivery review steps, and correspondingly longer procurement cycles for cross-border deployment. Enterprises operating across multiple jurisdictions have one more layer to account for: the same hardware differs in availability and compliance treatment by region, so available model lists and compute availability must both be set per region rather than covered by one global configuration.

MCP Puts Agent Identity on the Priority List

On August 22, the Model Context Protocol published a new roadmap developed by the Core Maintainers together with the Working Groups, with five priority areas: agentic messaging primitives, HTTP-native transport unification and hardening, agent identity and enterprise-ready security, improved primitives, and improved SDK developer experience.

The July 28 specification had already landed a batch of changes. Protocol-level sessions and the initialization handshake were removed, so a server can scale horizontally without holding state; clients can call server/discover to learn a server's supported versions and capabilities before doing anything else; list results are cacheable; Tasks were moved into an official extension; and a new Multi Round-Trip Requests pattern replaced server-initiated requests so that flows such as elicitation work on stateless servers as well. Authorization changes included issuer validation, issuer-bound client credentials, and Client ID Metadata Documents as the preferred registration path, while Enterprise-Managed Authorization is now a stable extension.

The item in the new roadmap with the highest enterprise relevance is identity. MCP authorization today is built around a person approving access in a browser, while more and more callers are agents running as cloud workloads with their own identity, acting on behalf of a user who is not present, or delegating narrower authority to sub-agents. The work ahead is to finalize DPoP and drive its adoption, and to define an opinionated path for agent identity and delegation through Workload Identity Federation, the ID-JAG grant behind Enterprise-Managed Authorization, and standard token exchange, while continuing engagement with the IETF OAuth and WIMSE working groups. Another item is progressive tool discovery: connecting to a server carrying a hundred tools means the model pays for that entire surface before the user has asked a single question, and tool selection accuracy degrades as the list grows.

The protocol layer putting these issues on the schedule does not mean the enterprise side can wait. Specification finalization, SDK follow-through and implementation across MCP servers is a timeline measured in quarters or more, while agents are already calling tools in production. In the meantime, someone has to answer the same set of questions: which real user and which team is behind this call, which tools and data it is allowed to reach, and whether the whole chain can be pulled up afterward for review. These questions have to line up with the enterprise's own identity system and compliance definitions, which the protocol itself cannot answer. AgentsFlare applies identity isolation and request-level audit to MCP access, and enforces those rules uniformly across hybrid multi-cloud environments without binding to any single vendor's agent stack.

Upstream: Memory Moves Into the Chip, Land Gets Locked by Local Councils

The most substantive upstream change over the past week is who defines memory. NVHBM moves the memory controller from the accelerator compute die into the HBM base die, with stated gains of up to 30% more bandwidth, 15% lower HBM power and 25% more usable area on the compute die. HBM is the high-speed memory stacked beside the compute die specifically to feed it data, and inference speed in many scenarios is capped by its bandwidth rather than by compute. Relocating the controller means the memory interface specification is defined by the accelerator vendor rather than the memory vendor, moving the three memory suppliers from selling standard parts toward building to someone else's drawings. It will take several quarters for this change to show up in procurement, but the direction is clear: the more customized the part, the less substitutable it is, and the spot market for memory no longer applies to the customized portion.

Memory price increases are already visible in NVIDIA's own financials. Its 74% gross margin guide came in below expectations, with the step down driven mainly by memory and advanced packaging costs; this is consistent with a market in which third-quarter server memory contract prices continue to rise and suppliers compete for capacity on 16-high HBM4 orders from NVIDIA. The implication for enterprises is direct: upstream cost pressure has not reversed, the assumption that model rate cards will shift down across the board in the second half does not hold, and the price differential will show up between tiers — the low tier keeps falling while the high tier stays elevated along with monitoring overhead and memory costs.

Another constraint comes from land and power. As of August 19, 533 local-government moratoria or restrictions on data centers were in force across the United States, with another tracker counting 242 under a stricter definition; state-level moratorium bills were introduced in 11 states this year, and New York became the first state with a statewide moratorium in force under Executive Order 62. Two August cases: the Louisville Metro Council voted 24 to 1 on August 13 for a six-month moratorium, and the Indianapolis City-County Council voted 23 to 1 on August 10 to bar new data center construction in Marion County through the end of 2027. In Texas, the governor has required electricity and utility regulators to complete a full audit of data center power and water usage before grid connections proceed, covering roughly 250 to 300 projects representing about 200 gigawatts of requested demand — more than twice ERCOT's all-time peak — with the audit expected to take months.

These constraints show up in the rental market as a near-threefold spread on the same chip. In August's public quotes, B200 ranges from $5.99 to $16.11 per GPU-hour: Hyperbolic at $5.99, Lambda at $6.69, RunPod at $6.79, Nebius at $7.15 and CoreWeave at $8.60, rising to Oracle at $14.00, AWS at $14.24 and Google Cloud at $16.11 among hyperscalers; H100 spans $1.49 to $6.98. That spread indicates that what the market lacks is not the chip itself but powered, networked slots available for immediate delivery. The more local moratoria there are, the more slowly that spread narrows.

The Week That Re-Sorted the Cost Table

The same thing showed up once at each of three layers over the past seven days. NVIDIA used an annual guide reaching into next year to demonstrate order visibility, AWS used an incremental order of two million GPUs to show that hyperscalers have not slowed, and the three companies putting inference-specific chips on the same stage at Hot Chips showed that identical workloads are being split across an increasing number of hardware forms. All three point to one conclusion: inference-side supply is broadening quickly, and every new source of supply brings its own programming model, memory specification and delivery timeline.

Price, meanwhile, pulled apart in two directions in the same week. GLM-5.3-Flash reached a composite intelligence score level with Opus 4.8 at one fortieth the price, in the same week that Sonnet 5's promotional pricing expires and returns to list, and NVIDIA's gross margin stepped down on memory costs. The low tier is falling, the high tier stays elevated, and the gap between them has widened past the point of being ignorable — the old practice of drafting an annual budget around a single supplier will keep drifting out of true under this price structure.

The capability enterprises actually need to build this year is one that turns all of the above into configuration changes. Which chip, which model, which region will all change repeatedly over the next two years; what should not change is one set of admission rules, one definition of cost attribution, and one request-level audit chain. Hold those three yourself, and how the rate card changes and where supply comes from become parameters.