AI Infra Weekly

NVIDIA Leases a Texas Data Center for $50 Billion, Hyperscalers Raise 2026 Capex to $725 Billion, and Qwen Prices Multimodal AI at $0.03/$0.13

Author

AgentsFlare Research

Date Published

This week saw a dense cluster of pricing moves. On July 30, OpenAI cut the price of GPT-5.6 Luna, its lowest-cost tier, by approximately 80%, and reduced Terra, its mid-tier model, by around 20%. Alibaba’s Qwen 3.7 Flash, released on July 27, brought a multimodal model with a 1-million-token context window down to $0.03/$0.13 per million tokens. Vendors in both China and the United States are simultaneously pushing prices toward the cost-sensitive end of the market. Yet the more consequential activity took place at the physical layer of compute. On July 28, NVIDIA signed a Texas data center lease worth up to $50 billion, meaning a chipmaker is now directly taking on the role of a long-term compute tenant. In the same week, the four largest hyperscalers raised their combined 2026 capital expenditure plans to approximately $725 billion, up 77% year over year. For the first time in this cycle, capital markets voted collectively with their feet: Alphabet, Amazon, Microsoft, and Meta all saw their shares decline after earnings. On the industry-alignment front, NVIDIA assembled 37 members into an open AI security alliance on July 27, while all four frontier-model developers—OpenAI, Anthropic, Google, and Meta—were absent. On July 30, the European Union launched a €30 billion AI Gigafactory tender, writing compute sovereignty directly into its procurement documents. Chipmakers, cloud providers, and governments all moved in the same week to secure the physical layer, pushing the activity at the model layer into the background.

Executive summary

If you read nothing else this week:

  • NVIDIA signed a long-term lease with data center developer Hut 8 for the Beacon Point data center in Texas. The base 15-year term is worth approximately $19.6 billion and could reach $50.2 billion if all renewal options are exercised. The site can support up to 1 GW of power, marking the first time a chipmaker has directly secured physical compute capacity as a tenant (7/28)
  • AMD committed to a strategic investment of up to $5 billion in Anthropic, while Anthropic will deploy up to 2 GW of AMD Instinct MI450 Series GPUs, with the first 1 GW scheduled to come online in the first half of 2027. A second chip pole beyond NVIDIA took shape in the same week (7/22; AMD Helios rack-scale solution launched 7/23)
  • The four hyperscalers raised their combined 2026 capital expenditure guidance to approximately $725 billion, up 77% year over year: Microsoft around $190 billion, Amazon around $220 billion, Google $180–190 billion, and Meta $130–145 billion. Alphabet shares fell approximately 7% after earnings, while Amazon, Microsoft, and Meta also declined during the period (late-July Q2 earnings season)
  • OpenAI reduced the prices of two lower-cost GPT-5.6 tiers: Luna fell from $1/$6 to $0.20/$1.20 per million tokens, a reduction of approximately 80%, while Terra fell from $2.50/$15 to $2/$12, a reduction of approximately 20%. Flagship Sol remained unchanged at $5/$30, with the company explicitly positioning the changes for cost-sensitive customers (7/30)
  • Alibaba released Qwen 3.7 Flash at $0.03/$0.13 per million tokens, with a 1-million-token context window and vision support, making it the lowest-priced multimodal model with a million-token context window currently available (7/27)
  • NVIDIA and 37 members including Microsoft, Palantir, IBM, Dell, and CrowdStrike launched the Open Secure AI Alliance, while all four frontier-model developers—OpenAI, Anthropic, Google, and Meta—were absent (7/27)
  • The European Commission launched a tender for seven AI Gigafactories: €1 billion in public funding intended to mobilise €2 billion in private investment, for a total of approximately €30 billion. Bids close on 11/12, with the stated aim of reducing Europe’s dependence on U.S. and Asian compute capacity (7/30)
  • Upstream compute signal: TrendForce forecasts that server DRAM contract prices will rise 13–18% quarter over quarter in Q3, shifting pricing pressure toward customers without long-term contracts, while HBM contract prices are expected to double in 2027. At the same time, on-demand rental rates for H100 and B200 GPUs are at historical lows (July)

AgentsFlare is the enterprise AI control plane—preserving enterprise control over request routing, cost attribution, and access auditing in an environment where models, clouds, and agents continue to fragment. AI Infra Weekly is a strategic column for enterprise users produced by AgentsFlare. It independently tracks key shifts across the global AI infrastructure layer, helping enterprises understand the direction of the market before the structure becomes fixed.


Major Events

NVIDIA Leases a Texas Data Center for $50 Billion: The Chipmaker Starts Acting Like a Compute Landlord

On July 28, NVIDIA was confirmed as the tenant of data center developer Hut 8’s Beacon Point facility in Texas. The lease has a base term of 15 years and is worth approximately $19.6 billion; if all renewal options are exercised, the total could reach $50.2 billion. The facility can support up to 1 GW of power. For a company that historically sold chips while leaving data center construction to customers, securing physical compute capacity through a 15-year tenant agreement represents a substantive shift in position. NVIDIA’s role in the supply chain was previously that of an upstream component supplier. It is now also a capital-intensive compute tenant bearing operational complexity and regulatory exposure. The market has widely read this lease alongside separate reports that NVIDIA may be providing up to $250 billion in financial guarantees for OpenAI’s large-scale data centers. The latter has not been officially confirmed, but taken together, the two stories point in the same direction: NVIDIA is using its own balance sheet to underwrite downstream customers’ compute expansion.

This was not the only transaction this week pushing a chipmaker into the physical layer. On July 22, AMD announced a strategic investment of up to $5 billion in Anthropic. Anthropic will deploy up to 2 GW of AMD Instinct MI450 Series GPUs, with the first 1 GW scheduled to come online in the first half of 2027. The hardware form factor is the AMD Helios rack-scale solution, launched on July 23, which integrates GPUs, CPUs, networking, and storage at the rack level. A single rack provides approximately 2.9 exaFLOPS of inference compute and is priced at $5–5.5 million. Microsoft and OpenAI also stated during the same week that they would deploy Helios at scale. Taken together, the transactions indicate that a second chip pole beyond NVIDIA took shape in the same week. The common feature of both poles is that chipmakers are no longer merely selling hardware; they are embedding themselves into customers’ compute stacks through equity investment and direct leasing. For enterprises, this means supply-chain concentration is moving downward from the model layer into hardware and capital. When the roles of chip manufacturer, model provider, and data center builder begin to overlap across the same group of companies, vendor lock-in is no longer visible only in the API bill; it is embedded in the ownership structure of physical compute. The value of preserving scheduling and migration capability across a dual hardware stack—NVIDIA and AMD racks, networking, and software ecosystems that are not mutually compatible—is greater than it was a year ago.

Hyperscalers Raise 2026 Capex to $725 Billion, and the Market Votes Against the Race for the First Time

During the late-July Q2 earnings season, the four largest hyperscalers again raised their 2026 capital expenditure guidance, bringing the combined total to approximately $725 billion, up around 77% from $410 billion in 2025. Microsoft set full-year capital expenditure at approximately $190 billion, well above analyst expectations of around $152 billion. Amazon raised its plan to approximately $220 billion. Google maintained a range of $180–190 billion, while Meta raised the upper end of its range to $145 billion. Demand-side figures remained strong: Google Cloud quarterly revenue increased by approximately 63% year over year to around $20 billion, significantly faster than the market had expected.

The real inflection point was the market reaction. Alphabet shares fell approximately 7% in a single day after earnings, and Amazon, Microsoft, and Meta declined over the following days. For the first time, capital markets collectively expressed doubt about the current capital expenditure race. Investors are concerned that cash reserves are being consumed continuously while the certainty of returns has not kept pace. For enterprise customers, the significance lies not in the share prices themselves, but in how this pressure may change hyperscalers’ future pricing and capacity-allocation behaviour. If capital markets continue to demand returns on capital expenditure, cloud providers will have stronger incentives to pass expensive infrastructure costs into inference pricing and to prioritise dedicated capacity for their own models and largest customers. When an enterprise binds its AI workloads to a single hyperscale cloud, it is simultaneously betting on that provider’s capital discipline and its future ability to pass costs through to customers. Both assumptions have now entered a period of uncertainty.

OpenAI Cuts GPT-5.6’s Low-Cost Tier by 80%, While Qwen Prices Multimodal AI at $0.03/$0.13: The Price War Moves Toward Cost-Sensitive Buyers

The most concentrated activity at the model layer this week occurred in pricing, with vendors in China and the United States pushing prices down simultaneously. On July 30, OpenAI cut the prices of two lower-cost GPT-5.6 tiers. Luna, the lowest-cost option, fell from $1/$6 to $0.20/$1.20 per million tokens, a reduction of approximately 80%. Mid-tier Terra fell from $2.50/$15 to $2/$12, a reduction of approximately 20%. Flagship Sol remained unchanged at $5/$30, while cached input retained a 90% discount. OpenAI attributed the price reductions to end-to-end software-stack optimisation—improved hardware routing, faster inference software, and more intelligent context-caching algorithms that reduce repeated work by agents. The timing is also significant: the cuts came only around three weeks after the three GPT-5.6 tiers became generally available. The company’s public framing directly identified an increasingly cost-sensitive group of enterprise customers unwilling to deploy expensive models before returns are clear. This is the model-layer expression of the same return pressure hyperscalers faced from capital markets this week.

At the other end of the same week, Alibaba released Qwen 3.7 Flash on July 27 at $0.03 per million input tokens and $0.13 per million output tokens. The model supports a 1-million-token context window, visual input, tool use, and function calling, making it the lowest-priced multimodal model with a million-token context window currently on the market. Within Alibaba’s own portfolio, it sits below Qwen 3.7 Plus at $0.28/$0.86 and Qwen 3.7 Max at $2.50/$7.50. One qualification is necessary: Alibaba has not released a technical report, benchmark suite, or architecture description for this variant, and independent third-party evaluation data is not yet available. Its actual capability will need to be tested by the community; aggressive pricing does not by itself establish performance parity. At the opposite end of the pricing spectrum, frontier pricing remains firm. Three days earlier, on July 24, Anthropic released Claude Opus 5 at $5/$25, approximately half the price of Fable 5, while still maintaining double-digit output pricing for a frontier proprietary flagship.

Taken together, the three moves show that the price spread in 2026 continues to widen, but the force behind that widening has expanded beyond Chinese vendors cutting prices. U.S. frontier-model developers are now proactively lowering their entry-level tiers to capture cost-sensitive volume. For enterprises, the direct implication is that the best model for different stages of the same workflow may come from entirely different price bands. High-volume multimodal preprocessing, classification, and extraction can be handled by models priced in fractions of a cent or a few cents, while only the reasoning steps that truly require frontier capability justify double-digit pricing. Routing requests by task type and cost budget, and assigning token quotas by team and project, is the prerequisite for converting the widening price spectrum into actual cost savings. As different tiers from the same provider and comparable models across providers continue to reprice frequently, dynamic routing based on current price lists—and cost attribution by model and supplier—captures more of the benefit than locking into a single model.

NVIDIA Brings Together 37 Members in an Open Security Alliance, While All Four Frontier-Model Developers Are Absent

On July 27, NVIDIA and 37 members launched the Open Secure AI Alliance to develop and share open-source tools for AI cybersecurity and agent security. Members include Microsoft, Palantir, IBM, Dell, CrowdStrike, Snowflake, Databricks, Cisco, Palo Alto Networks, Red Hat, Hugging Face, and the Linux Foundation. One of the stated catalysts was an earlier cyberattack involving OpenAI. The most notable feature is the list of absentees: all four frontier-model developers—OpenAI, Anthropic, Google, and Meta—were missing from the founding membership.

This configuration brings an industry fault line into the open. The alliance brings together the infrastructure camp, which controls chips, servers, data stacks, and security capabilities; absent are the model developers that control proprietary model APIs. A representative judgement voiced by Palantir at the event was that enterprises are paying for tokens while watching their proprietary insights flow toward model laboratories through those same calls, and that this exchange is beginning to provoke a backlash. For enterprise decision-makers, the divide elevates an issue that has long existed but is easy to overlook. When model capability is concentrated among a small number of proprietary providers, while influence over security and data governance remains distributed across the infrastructure ecosystem, enterprises need an access and audit layer under their own control. That layer must ensure that sensitive data and proprietary context do not accumulate uncontrollably on the model-provider side whenever models are invoked. For enterprises running AI traffic through AgentsFlare, request-level audit trails, call records by user and endpoint, and tool-access controls operate within the enterprise’s own control plane. Models can be replaced or used across multiple providers, while data sovereignty remains with the enterprise rather than shifting with the terms of any single model vendor.

EU Launches a €30 Billion AI Gigafactory Tender: Compute Sovereignty Enters the Procurement Documents

On July 30, the European Commission formally launched a tender for seven AI Gigafactories, using €1 billion in public funding to mobilise €2 billion in private investment, for a targeted total of approximately €30 billion. The facilities are intended to provide startups, small and medium-sized enterprises, and industry with compute for training, inference, and fine-tuning of frontier models. The programme is divided into two tiers: four smaller facilities, each equipped with at least 75,000 AI chips and eligible for up to €500 million in public funding, and three larger facilities, each equipped with at least 100,000 chips and eligible for up to €1 billion. Bids close on November 12, with awards expected in early 2027 and construction beginning later that year. The explicit objective is to reduce Europe’s dependence on U.S. and Asian AI compute capacity.

Read alongside the preceding events, a common move emerges: physical compute is becoming a strategic asset, and enterprises are no longer the only actors competing for it. In the United States, chipmakers are using their balance sheets to secure data centers. In Europe, governments are using public capital to build sovereign compute pools. For enterprises operating in the European Union, this means that over the next two to three years, a new set of local compute options may emerge that are not tied to a single U.S. hyperscaler, expanding the available choices for regional deployment and data residency. But the AI Gigafactories will not begin construction until 2027, so Europe’s near-term compute-supply structure will not change. The practical step today is to incorporate this development into medium- and long-term regional deployment planning rather than current procurement decisions.

Upstream Compute Signal: Memory Prices Are Rising While Rental Rates Are Falling, and the Two Curves Diverge for the First Time

The most important upstream signal this week is the movement of two price curves in opposite directions. The first is memory. TrendForce forecasts that server DRAM contract prices will rise another 13–18% quarter over quarter in Q3. Because several U.S. cloud providers had previously signed multi-year contracts that capped pricing, the current round of increases will fall primarily on customers without long-term agreements and on incremental purchases above the contracted quotas of customers that do have them. Contract prices for HBM—high-bandwidth memory, in which multiple memory dies are vertically stacked close to the processor package and which has become one of the performance bottlenecks for AI chips—are expected to double in 2027, while HBM4 is set to become mainstream in the second half of 2026. The second curve is GPU rental pricing. July market tracking shows H100 on-demand rental rates at historical lows, with lower-cost channels offering prices of $1.38–1.49 per hour. B200 on-demand pricing is around $6 per hour, with spot rates as low as $2.12.

The transmission path suggested by these market signals is as follows: memory pricing is constrained by production capacity and long-term contract structures, making it more likely to rise than fall, particularly for small and medium-sized enterprises and new customers without bargaining power. GPU rental rates, by contrast, are declining as providers continue to add capacity and channel competition intensifies, lowering the marginal cost of renting mainstream accelerators for inference. Taken together, the two curves indicate that the current elasticity in inference costs comes primarily from compute-rental channels rather than the underlying hardware. Switching providers and placing workloads with lower-priced capacity can generate greater savings than waiting for the next generation of chips. But this requires workloads to be engineered for migration across suppliers. Workloads locked into one provider’s instance formats or software stack cannot capture this benefit. The memory curve is instead a variable for enterprises building their own infrastructure or negotiating long-term contracts. Incremental procurement costs without long-term price protection will be materially higher after Q3, making the timing and volume of such commitments worth recalculating.

Overall Assessment

Taken together, this week’s developments reveal a deeper theme precisely because the model layer was relatively quiet: the physical layer of compute is being redistributed, and all three groups now competing for it—chipmakers, hyperscale clouds, and sovereign governments—did not previously hold this layer directly in the same way. NVIDIA is becoming a long-term tenant, AMD is using equity investment to embed itself into Anthropic, and the European Union is using public funding to build a sovereign compute pool. The forms differ, but they all point to the same conclusion: whoever controls physical compute will occupy a stronger position in the next phase of bargaining. At the same time, capital markets have collectively expressed doubt about the capital expenditure race for the first time, meaning this redistribution is no longer accelerating without conditions. Return pressure will begin to flow backward into cloud-provider pricing and quota allocation.

For enterprises, these developments narrow into one conclusion: concentration in the AI supply chain is moving downward from the model layer into hardware and capital, and the risk of single-vendor lock-in is likewise moving from API bills into the ownership structure of physical compute. Maintaining a multi-provider model strategy, preserving scheduling flexibility across the two hardware stacks, and keeping data sovereignty within a layer controlled by the enterprise are the most direct hedges against this structural shift. The divide between the infrastructure camp and proprietary model developers visible in NVIDIA’s security alliance this week merely brought an existing problem to the surface earlier than expected.

Sources

Anthropic / Alibaba / OpenAI model releases and pricing: Best AI Models of July 2026, ThursdAI July 2026 Releases, LLM Updates — 2026-07

Qwen 3.7 Flash release and pricing: Qwen3.7 Flash — OpenRouter, Qwen 3.7 Flash Complete Guide — 2026-07-27

GPT-5.6 lower-tier price cuts: OpenAI cuts GPT-5.6 Luna and Terra prices by up to 80% — Quartz, OpenAI cuts prices for two of its GPT-5.6 AI models — CNBC, AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% — VentureBeat — 2026-07-30

NVIDIA’s Hut 8 Texas data center lease: Nvidia Signs $50 Billion Lease with Hut 8 — GuruFocus, Nvidia behind $50 billion lease, FT reports — Yahoo Finance — 2026-07-28

AMD × Anthropic $5 billion investment and 2 GW of MI450 GPUs: AMD and Anthropic Announce Strategic Partnership — AMD IR, CNBC — 2026-07-22

AMD Helios rack-scale solution launch: AMD Launches Helios — AMD Newsroom, DCD — 2026-07-23

Hyperscaler 2026 capital expenditure and earnings: Google/Microsoft/Meta/Amazon capex to hit $725B — Yahoo Finance, Hyperscalers face higher capex scrutiny — CNBC — 2026-07

NVIDIA Open Secure AI Alliance: Industry Leaders Join Open Secure AI Alliance — NVIDIA Blog, Nvidia forms 37-member AI security alliance — CoinDesk — 2026-07-27

EU AI Gigafactory tender: EU Pledges €10 Billion — Bloomberg, EU opens applications for €10bn AI gigafactory scheme — Sifted — 2026-07-30

Upstream memory and GPU rental pricing: Server DRAM Contract Prices to Rise 13-18% QoQ in 3Q26 — TrendForce, NVIDIA B200 Pricing (July 2026) — Thunder Compute, AI GPU Rental Market Trends (July 2026) — Thunder Compute — 2026-07