The two most important AI infrastructure stories of Q1 2026 landed within 24 hours of each other. On April 15, Snap restructured its entire workforce around AI productivity. On April 16, TSMC reported its fourth consecutive record quarter — $18.2 billion in net profit, up 58% year-over-year. And framing both was a deal announced back in January that keeps getting more significant the longer you look at it: OpenAI committed more than $10 billion to Cerebras over a multi-year agreement for 750 megawatts of dedicated compute.
Taken together, these data points tell a story about where the real competition in AI is happening. Not at the model layer — everyone is converging on capable models. The competition is at the infrastructure layer: who controls the chips, the fabs, and the compute delivery pipelines that make real-time AI possible at scale.
We think this matters far beyond the hyperscalers and semiconductor giants. If you’re building AI products — or planning to — the infrastructure decisions being made right now will determine what you can access, at what cost, and at what latency, for the next five years. Let’s walk through what we’re actually looking at.
The Cerebras Deal: What $10 Billion Buys You
The OpenAI-Cerebras agreement, announced publicly in January 2026, isn’t a research partnership or a pilot. It’s a production infrastructure commitment: $10 billion-plus over the 2026–2028 window, delivering 750 megawatts of compute capacity. That’s enough to power a small city, and it’s all dedicated to AI inference.
The technical case for Cerebras centers on their WSE-3 chip — a wafer-scale engine that uses nearly an entire 300mm semiconductor wafer as a single chip. At 46,225 square millimeters, it packs 4 trillion transistors and 900,000 AI cores into one unit. The claimed result is 15x faster inference than GPU-based systems. Cerebras CEO Andrew Feldman put it simply: “Just as broadband transformed the internet, real-time inference will transform AI.”
For OpenAI, the strategic logic is equally clear. Sachin Katti, VP of Research Infrastructure, described Cerebras as adding “a dedicated low-latency inference solution” that enables faster responses and more natural interactions. Translation: OpenAI wants inference that feels instant, not just fast. That’s a UX requirement as much as a technical one.
There’s also a diversification play here. Cerebras had been dangerously dependent on a single customer — G42 in the UAE accounted for 87% of revenue in the first half of 2024. The OpenAI deal changes the company’s risk profile entirely and gives it the revenue visibility to pursue a future capital raise or relaunched IPO. Both parties needed this deal. Both parties got something from it.
TSMC’s Record Quarter and What It Signals
TSMC’s Q1 2026 numbers aren’t just impressive — they’re a real-time indicator of where AI investment is flowing. Net profit of NT$572.48 billion (roughly $18.2 billion USD) on revenue of $35.9 billion represents TSMC’s fourth consecutive record quarter. Gross margin hit 66.2%. Operating margin hit 58.1%. These are numbers that suggest a company operating at the center of a structural, not cyclical, demand wave.
The composition of that revenue tells the deeper story. Seventy-four percent came from advanced chips at 7nm or below. Twenty-five percent came from sub-3nm processes alone. AI and high-performance computing now account for 61% of TSMC’s total revenue — up from roughly 45% in 2024. For the full year 2026, TSMC is guiding to 30%+ revenue growth in USD terms, with Q2 expected at $39–40.2 billion.
What this means practically: every major AI chip — Nvidia’s H200, Cerebras’s WSE-3, Apple’s M-series, custom silicon from Meta and Amazon — runs through TSMC’s fabs. The company isn’t just a supplier. It’s the shared infrastructure layer beneath the entire AI industry. When you see TSMC post four consecutive record quarters, you’re seeing the financial signature of an industry-wide bet on AI compute.
The Bifurcation: Training Stays Nvidia, Inference Is Contested
Here’s the dynamic we find most interesting, and most relevant to anyone building AI products right now: the infrastructure market is splitting in two.
For model training, Nvidia’s dominance is essentially intact. Training large foundation models requires massive parallelism across thousands of GPUs, deep ecosystem integration (CUDA, cuDNN, the full software stack), and the kind of multi-year hardware-software co-development that Nvidia has been running since 2012. No credible challenger has emerged for training at frontier scale.
Inference is a different market with different constraints. Speed matters more than raw throughput. Latency is a product requirement, not just a performance benchmark. And the economics are different — inference happens continuously at scale, not in concentrated training runs. This is exactly where Cerebras’s wafer-scale architecture has a structural advantage: by eliminating the inter-chip communication bottlenecks that slow down GPU clusters, the WSE-3 can deliver dramatically lower latency for workloads that fit its architecture.
The OpenAI deal is OpenAI explicitly betting on this bifurcation. They’re keeping Nvidia for training. They’re adding Cerebras for inference. As we’ve explored in our coverage of AI deployment strategies, the companies making smart infrastructure choices now are building competitive advantages that will compound over the next three to five years.
How Mid-Market Companies Can Still Access High-Performance Compute
Reading about $10 billion chip deals and record semiconductor profits can feel alienating if you’re building AI products without hyperscaler budgets. We get that. But the infrastructure arms race at the top of the market has real downstream benefits for everyone else — and you don’t need to buy a Cerebras chip to benefit from the dynamics it’s creating.
First, competition at the infrastructure layer drives inference costs down broadly. When OpenAI has access to 15x faster inference through Cerebras, it can offer faster API responses to every developer building on its platform — including the startups and mid-market companies that power their products through the API. You benefit from the infrastructure investment without paying for it directly.
Second, the cloud providers are all racing to offer specialized inference capacity as a service. AWS Trainium and Inferentia, Google TPUs, Azure’s expanding AI infrastructure — these are all attempts to capture the inference market that Cerebras is targeting at the frontier. Competition means better options and more competitive pricing at every tier.
Third — and this is where strategy matters most — the companies that will win in the next cycle aren’t necessarily the ones with the most compute. They’re the ones that route workloads intelligently: using the right compute for the right task, optimizing for latency where it affects user experience, and managing cost where it doesn’t. That’s an architectural and operational discipline, not a capital one.
What This Means If You’re Building AI Products
The Cerebras deal and the TSMC earnings together point to a world where AI infrastructure is maturing fast, getting more specialized, and becoming a genuine competitive differentiator. Here’s what we think that means in practice for anyone building right now.
Latency is becoming a product spec. The reason OpenAI is paying a premium for Cerebras inference is that users can feel the difference between a 100ms response and a 500ms response. If your product depends on AI-generated output in a user-facing context, inference speed is a UX decision — treat it like one.
Model selection and infrastructure selection are increasingly separate decisions. You might choose a model based on capability and then choose an inference provider based on speed, cost, and reliability. These choices don’t have to be bundled. The infrastructure layer is disaggregating, which gives you more control and more complexity to manage.
The build-vs-buy calculus for AI infrastructure is shifting toward buy — but buy from whom matters more than ever. TSMC’s margins tell you that the fab layer is concentrated and stable. The inference layer, where Cerebras is competing with Nvidia and cloud-specific silicon, is still being decided. The companies that watch this space closely and make deliberate provider choices will have advantages that are hard to replicate.
We track these infrastructure shifts closely because they shape what’s possible for the products we help build. The gap between the companies that understand the compute layer and the ones that treat it as a black box is widening — and it’s widening fast. You can read more about how we think about this in our AI infrastructure coverage.
If you want to talk through how these infrastructure shifts apply to what you’re building — whether that’s model selection, inference architecture, or just figuring out where to start — we’d welcome that conversation.



