<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Chipstrat]]></title><description><![CDATA[Semiconductors, AI, and business strategy. Read by tech leaders and investors. Sits between SemiAnalysis and Stratechery.]]></description><link>https://newsletter.chipstrat.com</link><image><url>https://substackcdn.com/image/fetch/$s_!_zdi!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ed93e6-778c-47de-a3f9-607ed3ed93ce_1024x1024.png</url><title>Chipstrat</title><link>https://newsletter.chipstrat.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 10 Oct 2026 03:24:50 GMT</lastBuildDate><atom:link href="https://newsletter.chipstrat.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Austin Lyons]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[chipstrat@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[chipstrat@substack.com]]></itunes:email><itunes:name><![CDATA[Austin Lyons]]></itunes:name></itunes:owner><itunes:author><![CDATA[Austin Lyons]]></itunes:author><googleplay:owner><![CDATA[chipstrat@substack.com]]></googleplay:owner><googleplay:email><![CDATA[chipstrat@substack.com]]></googleplay:email><googleplay:author><![CDATA[Austin Lyons]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[🎙️ The Compute 100: Startups vs. Incumbents in Semis, Intel Foundry, Physical AI, and more]]></title><description><![CDATA[Watch now (66 mins) | Intel&#8217;s EMIB vs. TSMC&#8217;s CoWoS, prefill and decode disaggregation, 576 GPUs in a rack, edge compute for end-to-end transformers, AI-native EDA, and more]]></description><link>https://newsletter.chipstrat.com/p/the-compute-100-startups-vs-incumbents</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/the-compute-100-startups-vs-incumbents</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Wed, 07 Oct 2026 20:44:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4450aeca-d948-4f85-84d8-5e77915f20b3_1280x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I joined my friends Brian Schechter and Gaby Lorenzi, investors at Primary, on The Compute 100 podcast. It was a great chat and I wanted to share it with you. You can also watch it on X: </p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/thecompute100/status/2107532903080952220&quot;,&quot;full_text&quot;:&quot;Today's guest is <span class=\&quot;tweet-fake-link\&quot;>@austinsemis</span>, the writer behind <span class=\&quot;tweet-fake-link\&quot;>@chipstrat</span>.\n\nAustin started his career as an engineer at Intel before moving into startups and independent analysis. Today, he writes about semiconductors, optics, networking, and physical AI, bringing an engineer's perspective to &#8230;&quot;,&quot;username&quot;:&quot;thecompute100&quot;,&quot;name&quot;:&quot;The Compute 100&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2055322120930963456/owFMqZ3B_normal.jpg&quot;,&quot;date&quot;:&quot;2026-10-06T18:06:14.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!1G58!,w_1028,c_limit,f_auto,q_auto:best,fl_progressive:steep/l_play_button_usfui2,w_88,e_colorize:0/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F__ss-rehost__tw-video-preview-13_2107523219997544448.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/ilYmsPIZ3r&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:1,&quot;retweet_count&quot;:6,&quot;like_count&quot;:50,&quot;impression_count&quot;:34768,&quot;expanded_url&quot;:null,&quot;video_url&quot;:&quot;https://video.twimg.com/amplify_video/2107523219997544448/vid/avc1/1280x720/tYlgNN8Zu3ZdIr_C.mp4&quot;,&quot;video_preview_media_key&quot;:&quot;13_2107523219997544448&quot;,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p><strong>Things we cover:</strong></p><ul><li><p>Why the CUDA moat shrank once LLM inference became the biggest market</p></li><li><p>Intel Foundry&#8217;s culture change, and EMIB as an alternative when CoWoS capacity is short</p></li><li><p>Advanced packaging, from OSATs to 3D stacking</p></li><li><p>Prefill and decode disaggregation, and why it becomes a rack and data center design problem</p></li><li><p>Edge compute for physical AI, including Rivian&#8217;s RAP1</p></li><li><p>AI-native EDA startups vs. Cadence, Synopsys, and Siemens</p></li><li><p>Rapid fire on TSMC, Anthropic, OpenAI, Hock Tan, and quantum</p></li></ul><p>Also, Primary is hiring:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/gabyllorenzi/status/2107578223655833977&quot;,&quot;full_text&quot;:&quot;Big news! I&#8217;m heading west to SF with <span class=\&quot;tweet-fake-link\&quot;>@PrimaryVC</span>. And we're hiring three new roles for our compute team to work alongside <span class=\&quot;tweet-fake-link\&quot;>@BSchech</span> and me.\n\nI believe compute is the most compelling, and challenging, market opportunity in my lifetime. Visionary founders are building toward a&#8230;&quot;,&quot;username&quot;:&quot;gabyllorenzi&quot;,&quot;name&quot;:&quot;gaby lorenzi&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1696241161315168256/5xxE9ard_normal.jpg&quot;,&quot;date&quot;:&quot;2026-10-06T21:06:20.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;&quot;,&quot;username&quot;:&quot;BSchech&quot;,&quot;name&quot;:&quot;Brian Schechter&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1437450948197105666/WOqTHMM2_normal.jpg&quot;},&quot;reply_count&quot;:21,&quot;retweet_count&quot;:10,&quot;like_count&quot;:160,&quot;impression_count&quot;:12036,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p><em>This podcast is lightly edited for clarity.</em></p><h2>Introducing Austin</h2><p><strong>Gaby:</strong> Today&#8217;s guest is Austin Lyons. He is the writer behind Chipstrat. When I was getting up to speed on the compute market, one of the first places I looked for information was Substacks. All these independent analysts were writing incredible pieces at the intersection of technology and business, and Chipstrat was one of the earliest ones that I found. So getting to talk with him and hear his journey, not only building the media engine around Chipstrat, but also being in conversation with some of the most senior leaders in compute, has been really fun. What did you think?</p><p><strong>Brian:</strong> One of the things about Austin, one, he&#8217;s just a really nice guy, but he has this set of experiences, from his MBA to startups to working in industry, that has now given him some credibility to predict what are the most important themes coming, whether it&#8217;s emerging ASICs or the reemergence of Intel as a really exciting force in the semiconductor industry. Just an exciting conversation to hear what he&#8217;s thinking about these days of what&#8217;s coming.</p><p>The Compute 100 is a podcast about the most consequential market in AI.</p><p><strong>Gaby:</strong> These are the conversations with builders and investors about the technology and the capital powering the greatest infrastructure buildout ever.</p><p><strong>Brian:</strong> I&#8217;m Brian Schechter.</p><p><strong>Gaby:</strong> I&#8217;m Gaby Lorenzi.</p><p><strong>Brian:</strong> We&#8217;re investors at Primary.</p><p><strong>Gaby:</strong> Let&#8217;s dive in to The Compute 100. This podcast is for informational purposes only and does not constitute investment advice or an offer to buy or sell any securities.</p><h2>From Intel Engineer to Chipstrat</h2><p><strong>Gaby:</strong> Austin Lyons, thank you so much for joining us on The Compute 100 podcast. We&#8217;re super excited for the conversation.</p><p><strong>Austin:</strong> Awesome. Yes, thank you for having me.</p><p><strong>Gaby:</strong> You sit at one of probably the most interesting places in the content-meets-semiconductor world over the past five years. I&#8217;m sure your career has gone through a complete evolution as this market has caught up to the world that you&#8217;re really interested in. I&#8217;d love to maybe start 10 years back, even your first career post-grad, and hear your story up to what you&#8217;re working on today at Chipstrat.</p><p><strong>Austin:</strong> Sure, totally. I&#8217;m older than I look, so it might be longer than people will think. I got trained as an electrical engineer in undergrad and in grad school. Did my master&#8217;s at the University of Illinois. And interestingly, when I was at the University of Illinois, I was actually in the PhD program, but I started a startup with some friends, some other grad students. This was back in the Kickstarter days.</p><p>We launched something on Kickstarter, had a lot of fun with that, and it really opened my aperture for what&#8217;s possible. So I was like, oh man, I&#8217;ve got to go get into industry and maybe get into startups. I actually went straight to Intel out of grad school, down in Austin, Texas, and was working on a startup at night, not in semiconductors. This was more of the IoT, iPhone apps, B2B SaaS early days era.</p><p>So I was doing hardware engineering by day, I was doing software engineering by night. And that set me down a path of thinking beyond just engineering, but more about who is the customer? What do they really want? Is this useful? Ultimately, we raised some money, we built a product, and it turned out customers didn&#8217;t really want it. So you learn early on the hard way. Fast forward, I ended up eventually moving into product management and getting an MBA.</p><p>I was really impacted by those entrepreneurial experiences. So I sit at this nexus of, I understand the deep technology, I can think like a product manager and ask the questions about roadmap, value prop, pricing, that kind of thing. And then I have a bit of an entrepreneurial spirit myself. So of course, I always love to study startups as well.</p><p><strong>Gaby:</strong> And so all of this was phenomenally well-timed with what is today the compute revolution. Take us through the last five years of deciding to go out on your own with Chipstrat and really build around the ecosystem that is education and technical deep dives around the compute ecosystem.</p><p><strong>Austin:</strong> Yes. So I was working as a product manager by day and of course the world was changing. ChatGPT was launched and all of a sudden semiconductors, which was my original first love, became super interesting. Early on, of course, it was just from the perspective of logic, but now it&#8217;s so much more. It&#8217;s logic, memory, optics, power, foundry, all the things. I was spending a lot of time consuming media.</p><p>I was training for marathons, so consuming podcasts, reading people&#8217;s things, trying to pay attention. And I just saw a gap in the market from an education content perspective. There&#8217;s either very technical things or very non-technical, thin, FinTwit-type things. And I just thought, we&#8217;re missing a Ben Thompson for the semiconductor industry here.</p><p>I felt like my background gave me a particular take on certain things. So I just decided to start writing, because it was as accessible as ever. And then it snowballed from there. You start writing and essentially you start putting your public thoughts out there in the court of public opinion. It&#8217;s a little intimidating at first, but it turns out most people are really nice and really interested, and you find the other nerds like you who want to talk about this stuff.</p><p>And then, as you guys know how it goes, it just snowballs from there.</p><h2>Early Takes: AI Accelerator Startups and AMD</h2><p><strong>Brian:</strong> So what is your particular take? I know that there&#8217;s a lot in that, but walk us through some of it. Ben Thompson has some core ideas that animate his writing. Curious to hear what are some of your core ideas right now. I&#8217;m sure they&#8217;re constantly evolving, but as some of them become solidified, what are they?</p><p><strong>Austin:</strong> Sure. So some of my early takes. Very early on, there&#8217;s an AI accelerator startup, Etched, and I started looking at them. Early on, it was just the Nvidia gravy train. Everyone realized, holy cow, GPUs are the future of AI, and Nvidia is obviously dominating in training, and we weren&#8217;t even into the inference era yet. And I saw that Etched was out there and I started asking questions like, could there be companies in this space other than Nvidia?</p><p>So Etched was one of them that got me into thinking about, what would it take for a startup to compete with Nvidia? Would the model labs trust and buy from a startup? Would the hyperscalers trust and buy from a startup? And so looking at first principles of, what&#8217;s wrong with the GPU? And what&#8217;s right with the GPU? I started pulling on threads and realizing, hey, GPUs are super flexible.</p><p>That&#8217;s great for training, that&#8217;s great for inference. But at the same time, once you start to fast forward, if transformer-based inference is the defining workload of our time, which has turned out to be true, then it would make a lot of sense for hyperscalers who actually know their workloads to invest in something that might be a little bit less flexible, but as a trade-off, it&#8217;s hyper-designed for the workload at hand, for attention and LLMs and KV cache.</p><p>So early on I started saying, hey, I think there&#8217;s something here. I think these AI accelerator companies can compete. And then similarly, early on, I was actually pretty bullish AMD. I think that was one of my first posts, about a year in, that started to get a lot of traction. Everyone at the time was talking about this CUDA moat.</p><p>Part of the CUDA moat early on was the fact that Nvidia had written software so that anyone in almost any domain could accelerate their workload, from scientific computing to oil and gas to quantitative finance and so on and so forth. So anyone who&#8217;s trying to compete, say AMD was trying to compete, people would say, oh, well, you have to write libraries in all of these different domains and it&#8217;s going to take you hundreds of thousands of man-years.</p><p>But once it became that the TAM for LLM inference was going to be way bigger than any of these other high-performance computing things, all of a sudden that CUDA moat kind of shrank. To compete for all these dollars, maybe AMD doesn&#8217;t have to even worry about all those other markets. Maybe if they just focus on making their Instinct chip and their ROCm software competitive for LLM inference and start there, they could actually compete.</p><p>So those are actually two very similar takes, which is, I think startups can compete in this market and I think that AMD can compete. I&#8217;ll just pause there, but it was very focused on logic early on.</p><h2>Intel: Fix CPUs and Foundry First</h2><p><strong>Brian:</strong> I also enjoyed some of your articulation around Intel and where you see their place in the market right now. Would love to hear you go on that. Maybe we could set the table of Austin&#8217;s takes on various dimensions of the compute world and then we&#8217;ll double click wherever seems most interesting.</p><p><strong>Gaby:</strong> I think, Austin, we&#8217;ve had probably the most ex-Intel people on this podcast. So we&#8217;re happy to have another one and just excited to get your take as someone who was there.</p><p><strong>Austin:</strong> Yes. Well, I&#8217;m wearing my Intel blue today. So, Intel. Early on, once Nvidia started selling all the GPUs and making a ton of money, everyone was saying, where&#8217;s Intel? Where&#8217;s Intel&#8217;s GPU? At the time, I said, hold the phone, Intel is a CPU company. And frankly, at the time, this was back in the Pat Gelsinger days, when they were even trying to figure out, are we going to have an external foundry or not? What is our strategy? And Intel was behind.</p><p>Basically, the advantage that Intel always had was they could do their own manufacturing. For many, many years, that gave them a performance benefit compared to everyone else, because they could get to the next node from a manufacturing process faster, and therefore they could reap that and bring more performant, more power-performant products to the market. Again, mostly CPUs, both for client and for server.</p><p>Then there was a period of time around 10 nanometer, 7 nanometer, where Intel actually fell behind for various reasons. One being not getting on EUV fast enough. Basically you can think of Intel as manufacturing and as product. They&#8217;re an integrated device manufacturer. They do both. They don&#8217;t use TSMC. Manufacturing used to really benefit the product team. Once manufacturing fell behind, all of a sudden that really started to drag down the product team and expose, if you have the same technology as everyone else or worse transistor technology, are your designs actually as performant and do they compete? So manufacturing started to fall down and then the product team started to fall down.</p><p>Meanwhile, everyone&#8217;s starting to say, okay, Intel&#8217;s struggling, but oh, they just need to make GPUs to catch up. And my question was, should they make GPUs? Or should they not make GPUs, and not make all these other networking chips and stuff, and actually just focus on fixing both of those things? Make better CPUs like they used to. Of course, AMD was starting to eat their lunch, Arm was starting to eat their lunch. And figure out the manufacturing thing.</p><p>Not only get back to being competitive, but then of course the problem with manufacturing. Everyone used to manufacture their own chips, because there&#8217;s actually a benefit if you can co-design and design for manufacturing. If you can do that together, not just design the chip but also design the manufacturing, you can optimize up and down the stack. But with every generation of foundry, it becomes more and more expensive, to the point where most companies didn&#8217;t have the volume to fully utilize and amortize all of those huge CapEx investments. So people kept falling off, but Intel had enough volume. Even they were starting to get to the limits. So it started to beg the question, if their manufacturing could catch back up and if they could take on external companies, could they get back into the game and compete with TSMC, both on performance but also financially? Could it actually be competitive from a cost perspective?</p><p>So early on, my Intel hot take was simply, they shouldn&#8217;t focus on GPUs. They should actually just right the ship, make sure their CPUs are performant, and double down on foundry. And now, I will say, that has paid off for Intel. CPUs have exploded with agentic AI. I didn&#8217;t necessarily expect that early on, but it has worked out for them in that their CPUs have been in high demand and can help continue to fund the business as they&#8217;re starting to bring foundry up.</p><p>And of course, within the past six months, we&#8217;ve started to hear a lot about Intel Foundry and external customers. So they are on stable ground with CPUs and with manufacturing. It still begs the question, what is their AI strategy? What is their GPU strategy? But I think now Lip-Bu has put enough sturdy ground underneath their business that they have earned the right to expand into GPUs.</p><p><strong>Gaby:</strong> Was your belief that investing in manufacturing is what will enable Intel to be a better partner to external customers, hence using that spare capacity to deliver for whoever they might be selling to in the next decade?</p><p><strong>Austin:</strong> Yes. So the question for Intel is, why didn&#8217;t people just raise their hand when everyone wanted to make GPUs and TSMC said, we are supply limited, and it depends on where you are in line, and you get allocated so many wafers? And Intel said, hey, we&#8217;re going to get into the game. Why didn&#8217;t people rush over to Intel?</p><p>Obviously, there&#8217;s a big cost in saying, now we have to redesign everything for a different foundry. But ultimately it was cultural. When you&#8217;re an IDM, an integrated device manufacturer, your foundry ultimately serves your product team, and you do have this internal groupthink of, this is the way we do things and this is the best. That is not a customer-centric mindset. Your customer is the product team, but you&#8217;re all the same company. So when Intel said that they were going to become an external foundry and they started working even with small customers, it did give them an opportunity to start trying to build that muscle.</p><p>And probably the best thing that has happened at Intel is they&#8217;ve brought in external hires who&#8217;ve worked at other foundries, whether it was GlobalFoundries or even Micron or Samsung, memory companies, who already have that correct orientation of, I work in a service business. Because that&#8217;s what manufacturing is.</p><p>So yes, I think Intel has had to undergo a lot of cultural change, but they seem to have worked through this. And ultimately it will end up benefiting Intel products, because the Intel manufacturing team now has to make the decision, who should get the capacity? Who should get the wafers? Should it be our internal products team or should it be our external customers? And that&#8217;s actually going to be great for both of them, because now the products team can&#8217;t be sloppy and just be like, ah, we screwed up here, do a hot run and run a couple more wafers for us. It&#8217;s like, no, no, no, you have to behave like an external customer. You need to plan in advance, and there might be times where you have to go to the end of the line and wait your turn.</p><p><strong>Gaby:</strong> Yeah, it&#8217;s interesting. It&#8217;s almost similar to what&#8217;s playing out in some of the hyperscalers that are designing their own chips now, and the way that they have to think about resource allocation between their own internal teams using their systems for the development of AI versus selling those chips, which are ultimately becoming really valuable product lines. So it is going to be interesting to see how that ultimately affects Intel and what their business looks like in the future. Very, very curious to see.</p><p><strong>Austin:</strong> Yes, totally. It&#8217;s a brand new world. And of course, again, none of us could have predicted all the geopolitical stuff that&#8217;s going on, which also indirectly benefits Intel Foundry as customers are incentivized to say, I&#8217;ll give you a chance.</p><h2>Advanced Packaging</h2><p><strong>Gaby:</strong> Totally. So one area that you spoke to that Intel has done a lot of really great innovation in is packaging. Something that we&#8217;ve been tracking from a startup perspective is how much value is going to be accruing to the fabs themselves. As we start to see systems like power become more integrated into the package, cooling becoming more integrated, so much value is going to be going to packaging and who&#8217;s doing that. So I&#8217;d love to hear a little bit more about what you saw and those learnings from your deep dives into the Intel space, and then do a little bit of a segue into what you&#8217;re tracking in packaging these days.</p><p><strong>Austin:</strong> So packaging used to be a fairly boring space, low value, low margin. If you go way back in time, packaging is literally taking discrete transistors and wiring them together. And then it was taking discrete chips on a board, printed circuit boards, and just big copper traces in between them.</p><p>But nowadays, we&#8217;ve moved into advanced packaging, where we&#8217;re taking chiplets or dies and trying to connect them in three dimensions. And that&#8217;s way more sophisticated from the perspective of how do you route everything electrically. It&#8217;s complicated from both a design and a manufacturing perspective. From a design perspective, you&#8217;re thinking about crosstalk and noise and electrically routing all these signals to the right places.</p><p><strong>Brian:</strong> Austin, maybe if I could break it down even more simply and just check for understanding, because the 3D thing was confusing for me for a while. One way that I think about advanced packaging right now is, we used to have city streets where every building had at least alleys between them, if not roads.</p><p>And we now have begun building the buildings to be touching each other and have shared plumbing and electrical connectivity between the two of them underneath. And so that is now the 3D element, where it&#8217;s no longer just everything laid out on a board assuming that it&#8217;s its own piece. Is that what we mean by 3D?</p><p><strong>Austin:</strong> Yeah, that&#8217;s a good analogy. Historically it was just, you had a chip and that was it. You&#8217;ve got a house and it&#8217;s in the country, whatever. We&#8217;re not even dealing with roads. And then eventually there&#8217;s this 2.5D, where it was, yes, let&#8217;s start to put buildings near each other, and now we have to put roads and alleys and connections between them.</p><p>And then as we&#8217;re headed into 3D, for example, high bandwidth memory is 3D-stacked DRAM chips. You&#8217;re literally stacking them up and you have to think about the internal plumbing. If you&#8217;re talking about a high-rise building, how do I get heating and cooling to all of these, get the waste out, electricity?</p><p>And we will continue. In the future, there are actually companies who are working on stacking logic chips. You could either have logic chips with memory on top or logic chips stacked on logic chips. So yes, now we do have to think about the infrastructure in three dimensions.</p><p>It used to be OSATs that did this packaging, and that stands for outsourced semiconductor assembly and test. It says outsourced. Right there you can tell it&#8217;s low value. But now it&#8217;s obviously getting higher and higher value and getting more and more complicated. At the end of the day, it&#8217;s very similar to what foundries do when they&#8217;re fabricating logic chips or memory chips. It&#8217;s that level of deposition and etch and patterning.</p><p>Now the OSATs are trying to move into that space, because they want to accrue the value at the OSAT layer and they want to say, no, no, we can do CoWoS now, so if you can&#8217;t get it from TSMC, we can do that too. But of course it&#8217;s also an opportunity for TSMC and Intel to expand from just making logic chips to doing more of the full-on logic, dicing, packaging, everything.</p><p><strong>Gaby:</strong> And so when you think about it, maybe selfishly from a startup perspective, thinking about where there&#8217;s new technologies that are interesting to invest in, as so much more value is going to be accruing to people who are doing packaging or TSMC as a fab, how do you think about the startup landscape evolving? Is it largely going to be startups that are building around new processes or startups that are enabling new types of packaging techniques?</p><p>We&#8217;ve seen companies that want to be a next-gen Amkor or a next-gen OSAT more broadly. But would love to just hear how you think this affects the newcomers and how they build something that&#8217;s really valuable.</p><p><strong>Austin:</strong> Yeah, this is tough. One of the hard parts when you first look at the semiconductor industry is it feels like everywhere in the value chain, there&#8217;s two or three or four companies and they&#8217;re huge. So you have to ask, okay, if I am a PhD student and I&#8217;m studying some really interesting new material or something, how do I take on Applied Materials or someone like that?</p><p>I don&#8217;t have scale, but what could I do? Obviously, one thing that you&#8217;re thinking about right away is, could I build some interesting technology that they would want to acquire? But if you want to be a full-on company, let&#8217;s talk advanced packaging. As this is getting smaller and smaller and the power consumed by these dies continues to increase, there&#8217;s really interesting questions around how do you deliver power there?</p><p>Power used to come in from way off to the side. Are there ways to integrate it much closer on the package, or even as a die, as part of this 3D stack? So you can think about, are there power delivery opportunities? Are there power dissipation opportunities? Interestingly, I think what probably doesn&#8217;t get appreciated enough is from an EDA or a design tool perspective.</p><p>There used to be people that would design and think about the electrical characteristics, and maybe different mechanical engineer type people thinking about the thermal characteristics. And now we live in a world where you have to always be doing both and thinking about both, because it&#8217;s just such small places with such hot, hot dies that are interconnected in three dimensions.</p><p>So I think there&#8217;s opportunities to ask, are there ways that we could help solve this or create a new user experience from a design perspective, where the electrical design and the thermal design is taken into consideration? There&#8217;s definitely opportunities. It&#8217;s just about that seam of how do you take your innovation and bring it to market and get in here.</p><h2>EUV, EMIB, and CoWoS</h2><p><strong>Brian:</strong> Can we actually spend another minute on that? Austin, take us further into that mistake of not investing in EUV, and then also the technical bets that they did make that people are now starting to say look like they&#8217;re starting to be really promising.</p><p><strong>Austin:</strong> Both of those, because both are important. From the logic perspective, one thing that people don&#8217;t appreciate about Intel is they actually do a lot of the transistor research and development that enables future generations that the whole industry uses, TSMC included.</p><p>Intel&#8217;s always been really good at that early R&amp;D. They have innovations from the early days like strained silicon and using different materials as oxides and things like this. So they&#8217;ve always really pushed forward on stuff. But there was a period where they took some bets on some materials that didn&#8217;t work, and they ended up delaying jumping to EUV. I can&#8217;t remember off the top of my head exactly why, if it was a cost concern.</p><p><strong>Brian:</strong> I think it was a cost concern. I think the cost was too big.</p><p><strong>Austin:</strong> That actually aligns with the idea that they had some more financial-oriented folks who were driving the ship during those times and not engineering-centric folks. So Intel, yeah, they slowed down there. Basically they were stuck at a particular node, 10, 7, while TSMC was able to catch up and keep going. And then of course everyone, AMD for example, is using TSMC and they&#8217;re getting on smaller and smaller transistors, and Intel was essentially stuck.</p><p>At the same time, Intel has always been innovating on packaging, and EMIB is one of their packaging processes that they use. The idea is, once you start to make chips that are really big and you hit the reticle limit, you might want to make two of them and tie them together.</p><p>TSMC, their approach is CoWoS, chip on wafer on substrate. And right there you can hear it. There&#8217;s three levels there, chip, wafer, substrate. So there&#8217;s this interposer where the connections happen. Intel has historically had EMIB, where basically just in the substrate itself, they embed these little bridges. So they just put the connections in between the chips right where you need them. The rest of it is a substrate.</p><p>And it hasn&#8217;t gotten much attention at all until very recently. But one of the benefits of being an integrated device manufacturer is, when your technology works and you have the scale of Intel&#8217;s product team making CPUs for tons of laptops and every company worldwide, you have the opportunity to take your packaging and put it in production and run it at scale. So now Intel Foundry is talking to potential customers who either just can&#8217;t get enough CoWoS capacity, ultimately they might be gated by the amount of wafers they can produce at TSMC, or, there&#8217;s different flavors of CoWoS, but some of the early flavors that were adopted have particular trade-offs.</p><p>And even as TSMC has come up with something called CoWoS-L, which is local silicon interconnect, and it&#8217;s fairly similar to EMIB, there&#8217;s still the three layers versus two. But there are some scaling problems with the way that CoWoS is implemented. And Intel Foundry can essentially stand on the sideline and say, hey guys, we&#8217;ve got this thing called EMIB, it scales really well. We are already making it in volume.</p><p>We&#8217;re servicing all of our Intel products team and all of their products that they&#8217;re shipping. So we know that it works and we have spare capacity. So yes, as we&#8217;ve moved to these AI accelerators that want to just be as big as freaking possible and have all this HBM that needs to be interconnected, it&#8217;s the perfect timing for EMIB and Intel Foundry.</p><h2>Information Sources and the Master Plan</h2><p><strong>Brian:</strong> Where are you getting your information from? We listen to various thought leaders in the compute space and they all have their information edge. Where do you get your information? And also, to the extent that you&#8217;re open to it, what are your goals? What&#8217;s the Austin Lyons master plan?</p><p><strong>Austin:</strong> Yeah, these are awesome questions. So where do I get my information? I of course try to consume all the public information that I can, earnings reports, whatever. When I can get my hands on expert interviews, those are interesting. You always have to take those with a grain of salt. And then I have the opportunity in my day job with Creative Strategies as an analyst to talk to a lot of these merchant silicon companies and talk with their executives.</p><p>Sometimes I&#8217;ll bring them on and publicly interview them, or other times I&#8217;ll get to talk to them behind closed doors and just ask them what&#8217;s top of mind. So that is an awesome privilege. But I also am always trying to talk to startups, see what people are saying on X. Again, a lot of it is public conversation. If there&#8217;s one thing I could do more of, it&#8217;s continue to try to talk to people who are in the trenches as much as possible.</p><p>Going to OCP or conferences is a really good place to meet people and have the opportunity to just ask them. I try to balance, okay, what is the market narrative around something? What is the FinTwit narrative? What are the thought leaders saying? What is Gavin Baker saying? But what am I hearing from companies? What are they saying?</p><p>And then thinking about questions where it&#8217;s like, well, the company said this, but did they actually mean this? Or they specifically didn&#8217;t talk about this. So trying to find those gaps and noodle on some of those. And then I will say, one of my advantages, I think, is that I don&#8217;t live in Silicon Valley. I&#8217;m outside of the thought bubble of, everyone&#8217;s thinking this or saying this. I live in the Midwest, and so we have a particular perspective on how the world works. And as you guys are experienced with as well, I&#8217;ve worked in industries outside of this.</p><p>That lived experience also gives you an ability to think critically and differently about certain things. And then to Brian&#8217;s question of the master plan. When I was a product manager, what was awesome was, ah, I like to think about strategy, I like to think about technology, this is so fun. I get to make a plan. What&#8217;s our one-year plan? What&#8217;s our three-year plan?</p><p>And then you think about it for a little bit and then you just spend the rest of your life implementing it. So it&#8217;s a little bit of strategy and a lot of execution. What I like about my current role, working as an analyst, writing Chipstrat, and then I also host a podcast called Semi Doped with Vik Sekar, is I just like to be on the forefront as new information comes out and think strategically. Oh, what does this mean?</p><p>China&#8217;s going to shut down access to indium phosphide. Oh no, is that real? What does it mean? So a lot of just thinking strategically. And then fortunately I&#8217;ve been able to build up relationships with executives and people at these companies to start to ask, I think this means this for your company. What is your answer to that?</p><p><strong>Brian:</strong> Yeah, it&#8217;s really interesting, because what it points to, for the hardware designer, whether it&#8217;s a startup or an incumbent or a hyperscaler, is having a point of view around the TAM associated with different types of value that AI can create. Because ultimately the TAM needs to be there, and the margin profile will need to be there as well.</p><h2>Disaggregated Compute</h2><p><strong>Gaby:</strong> I&#8217;d love to spend more time on disaggregated compute. One of the areas that people are spending time thinking about is, in response to disaggregation, a lot of new areas emerge, or maybe not new, but the downfalls of disaggregation emerge, like networking or how close your power supply is. So walk through your framework. If we move to this disaggregated world where a single vendor has all of these different compute offerings for different parts of workloads, how do you actually bring that together in a way that you can still enable efficient serving of inference, or maybe it ends up being for training, but largely for inference, when you have to deal more with networking, power, memory, et cetera?</p><p><strong>Austin:</strong> Yes. So when it comes to disaggregation, we&#8217;ll start with prefill-decode disaggregation, because everyone knows that. It actually makes a lot of sense. There&#8217;s these two subtasks and they have two different memory constraints or profiles, how much memory capacity and memory bandwidth they need. Prefill doesn&#8217;t need much and decode needs a lot.</p><p>So right there you&#8217;re like, wait a minute, we could have two different hardware SKUs, and then we could run part of the workload on one and part of the workload on the other. But then to your point, Gaby, now that means you do add some complexity. You need to orchestrate across them, you need to send information across them. So right away that starts to point to, we&#8217;re not talking about chip design, we&#8217;re talking about rack design and ultimately data center design. You start to ask questions about, what is the most power-efficient, most cost-effective, fastest way to send information back and forth?</p><p>You could go so far as to ask, how should we even architect the data center? Do you have racks of prefill chips, racks of decode chips, or do you have a rack that has prefill and decode in it so they can talk back and forth? And that even goes to co-design with the model companies. How are they taking their mixture of experts and splitting it across? We&#8217;ve got these layers on this GPU and these layers on that GPU, this expert on this GPU, this expert on that GPU.</p><p>I do think that&#8217;s why it pushes Nvidia to just say, oh, I want to try to get 576 GPUs in the same rack, even if the power&#8217;s crazy, if the cooling&#8217;s crazy, if the manufacturing&#8217;s crazy. You&#8217;re essentially trying to keep as much of the compute as close as possible to make it easier to send everything over copper if you can. But it does start to become a problem, and it opens up opportunities again to rethink. Oh yeah, but what about long context size and KV caching, and a cache miss versus a cache hit?</p><p>Should we get clever and interesting with having a rack right next door that&#8217;s just a ton of storage, so you can store the KV caches as close as possible? So we&#8217;re obviously disaggregating things and saying, how can we have the proper SKU for the proper workload need? But we&#8217;re also realizing, oh crap, there&#8217;s trade-offs, and how can we also aggregate it as much as possible and get them as close as possible?</p><p>Time will tell where other interesting innovations happen, from networking to memory to cooling, to continue to let us get more done faster but not waste as much power sending bits all over the place.</p><h2>Physical AI and Edge Compute</h2><p><strong>Brian:</strong> As we&#8217;re talking about this, a lot of the reason for all of this comes back to the size of the models and the size of the context windows that they benefit from in terms of developing more and more useful forms of intelligence. This is a change of topic some, but what are you most excited about in terms of AI and what capabilities are coming online? And what are you most hungry to see the semiconductor industry do to help enable those types of possibilities?</p><p><strong>Austin:</strong> I&#8217;m most excited about physical AI, about embodied AI. If you think about it, yes, in the last 30 years, an increasing amount of our world&#8217;s GDP is impacted by digital things, Netflix and Facebook and the internet and Riverside.fm and what have you. But so much of the world is still physical. Everything that you buy, everything that&#8217;s manufactured, Apple, Nvidia, it&#8217;s all supply chains, it&#8217;s all physical hardware, it&#8217;s all getting moved, it&#8217;s all getting shipped.</p><p>Of course, I live in places where people are growing corn and soybeans and hogs and things. Anyone who eats anything, that had to get physically moved. And most of that stuff is still not touched by AI. It&#8217;s barely touched by digital. There&#8217;s some logistics and things that go on. But literally, I was at a large grocery retailer and we were using mainframes still, and this was not a decade ago.</p><p>And they will continue to. There&#8217;s places where people are still using pen and paper. So it&#8217;s crazy to think, obviously just getting digital is one thing, but man, how can physical AI just unlock even more productivity? And it&#8217;s not just making what you&#8217;re doing more productive, like saying, here&#8217;s the way we work today, how can we be more productive? One of the cool things with anything digital that&#8217;s AI-enabled is it enables some solopreneur to do so much more than they ever could before.</p><p>Oh, hey, I&#8217;m a domain expert. I&#8217;m a dentist and I&#8217;ve always thought our software sucks, but now with AI, I can just write better dental software, because I&#8217;m the domain expert and I can just do it, send Fable off for a week. My oldest is in middle school and he&#8217;s been spending this summer coding an indie game and it&#8217;s super cool. He knows how to code but not that good, but Fable can just go off and do it.</p><p>And then he can spend a lot of time thinking about the story and the graphics and the art and stuff like this. Now imagine if we could take that and put it in the physical world. What if I could realize that I hate this particular part of my car? Of course I could 3D print it and stuff, but what if I could actually manufacture and assemble it and be like, hey, I&#8217;m going to make this bespoke add-on piece to cars and start selling it Etsy-style or something?</p><p>But it&#8217;s actually something physical. In Iowa where I live, lots of small towns are supported by a local manufacturer. It might be snowmobiles or basketball hoops or whatever. But they suffer from brain drain like everyone else does, which is everyone going off to the big city, and they don&#8217;t want to come back and live there and work on it. What if you could use robotics and physical AI innovations somehow to make it so the same amount of people could get more physical work done, or they could graduate to managing fleets of humanoids instead of being the person doing the thing?</p><p>It&#8217;s obviously very fuzzy, but what I&#8217;m excited about is, man, it feels like there&#8217;s so much human productivity that&#8217;s not unlocked yet, and there&#8217;s so many interesting, capable people that have very interesting mechanical or physical ideas, but we just don&#8217;t have the tools for them to just build.</p><p><strong>Brian:</strong> And so what hardware, compute-specific unlocks do you think will lead to breakthroughs in physical AI? Or do you think that&#8217;s part of the equation?</p><p><strong>Austin:</strong> Yeah, good question. Historically, you&#8217;ve just got embedded compute, and it&#8217;s little CPUs and not much memory and it&#8217;s all pretty wimpy. Nvidia has had Orin and Thor. They have a small portfolio, and it has certain power requirements, and it fits perfectly for certain use cases. But if you go talk to people who are doing AI at the edge, you&#8217;ll find out there are customers who are like, oh yeah, we just have to take three Nvidia SOMs and bolt them together so that we have enough compute to do what we want to do.</p><p>But that requires a ton of power. There&#8217;s all these trade-offs today. So I just don&#8217;t think we&#8217;re yet to a point where we&#8217;ve got enough different compute options for the edge, whether it&#8217;s more FLOPs or more memory or less power or just more different configurations. We are in the early days and it is a little bit of, you get what you get. Think about it. With a humanoid, it&#8217;s got a brain, it&#8217;s got sensing, it&#8217;s got all these little sensors at the fingertips, and you could have distributed computing within it.</p><p>I don&#8217;t think we even know what is the right hardware architecture there. There&#8217;s CPUs, there&#8217;s GPUs, there&#8217;s FPGAs. But are we designing the systems we build today based on the hardware that&#8217;s given to us, or are we building the right hardware for where these workloads are going?</p><p><strong>Gaby:</strong> We talk a good bit about the hardware lottery, and this idea that what is working and the research ideas that win are just what is suitable to what exists at the time. If you think about it, cloud was enabled by CPUs, AI enabled by GPUs, and there&#8217;s likely another iteration of this going to happen in the physical AI world, a little bit more than just what a Jetson offers, if we use that as the best edge compute that exists today. But it almost makes it hard to think about what are the use cases that a next-gen Jetson, or the next best piece of edge compute, could enable.</p><p>We&#8217;ve spent most of our time thinking about how edge will impact personal computers, or the way that we&#8217;re running more models locally, and how that eventually will trickle down into the economics of the AI economy. But of course the physical AI world is going to be, hopefully, an explosive customer of edge compute. In terms of the conversations you&#8217;re having with people around this market, where are the real pull and pain points of what exists today in the physical AI space? We hear a lot about power, a lot about weight, about latency. But I&#8217;d be curious, in the conversations that you&#8217;re in, where you feel like physical AI&#8217;s compute is not stacking up today.</p><p><strong>Austin:</strong> I think what Rivian is doing with their RAP1, the Rivian Autonomy Processor, is pretty interesting, because they are designing their own platform. They said what they needed off the shelf wasn&#8217;t there. And it was very specific. For personal vehicle autonomy, we need a real-time operating system and some real-time safety that has hard latency budgets and everything. But we also have cameras, we have lidar, so all this sensor data input, lots of data, it&#8217;s moving around. They actually created their own communication protocol. They called it RivLink, which is like NVLink or something.</p><p>They made a lot of trade-offs for their particular portfolio of vehicles, and I think that they&#8217;re also extending it to Mind Robotics potentially, which is, hey, can we also have humanoids that help make our vehicles someday? I thought it was just an interesting example of needing certain amounts of compute, a certain amount of interconnect bandwidth to send all this information around, and yet still having particular power constraints. Now, of course, at the end of the day, they&#8217;re connected to a big battery in your car. They don&#8217;t want to drain your battery, but they have more power.</p><p>Whether it&#8217;s humanoids or, what&#8217;s interesting is what Meta&#8217;s doing with those Orion glasses. Those were going to be crazy expensive, of course for some of the optics reasons, but still at the end of the day, it&#8217;s battery, compute. And if you think about it, most of the edge compute that we&#8217;re using today was probably designed two or three or four or five years ago. What was the AI that they were designing for at the time? It was convolutional neural networks and more traditional vision AI and ML.</p><p>So I actually don&#8217;t know if anyone has designed edge compute around end-to-end transformers yet, other than, again, maybe Rivian, maybe Tesla. That would probably be a big opportunity. If you started from first principles and you said, I want to run a transformer end-to-end at the edge, what is the right compute to support that?</p><h2>Where Startups Can Get a Foothold</h2><p><strong>Brian:</strong> What are you most excited about from a startup or early-stage company creation standpoint?</p><p><strong>Austin:</strong> I always love hearing about the technology, but what I&#8217;m most interested in is, where can a startup actually gain a foothold against a very large incumbent? In networking, what always seems to happen over the last 20 years is a startup, it&#8217;s usually people with a lot of experience, spin out of a company and then they go start something and then they get acquired by Cisco or Marvell or Broadcom. And it&#8217;s amazing, it works. If you want to start a startup and get acquired, go to networking. But it&#8217;s a little bit less interesting, because you&#8217;re like, oh, I know the playbook here.</p><p>What&#8217;s more interesting is the people who are playing in the space of ASML or foundry, like we talked about, where everyone on the internet&#8217;s just like, oh, you can never do that. That&#8217;s the dumbest thing I&#8217;ve ever heard of. How could you ever compete with ASML? There&#8217;s some startups, Substrate, xLight, where they&#8217;re saying, hey, what if we use X-ray lithography, or what if we use electron beams? What if we just think really crazy from first principles? What could that unlock?</p><p>I&#8217;m excited, from a strategy and entrepreneur perspective, to see people with chutzpah who say, no, no, no, we actually think there&#8217;s a way to pull this off. And I just like to watch it. One, I want to understand their crazy. Tell me more. You think this is possible. Why is it possible? I want to learn from you. And then two, I just want to see how it plays out.</p><p><strong>Brian:</strong> One of my passions as a VC is thinking about people and what they want to do and where they&#8217;re going, and then the questions associated, like, how do you get there? One thing that leads me to wonder about for you is, what is the question that you think people are asking that maybe they don&#8217;t even know how to fully frame, that comes up again and again, maybe in the context of consulting or maybe in the context of people who are engaging with your content that are hungry to come to you? What are people really trying to find out?</p><p><strong>Gaby:</strong> Can I give a guess and then you can tell us if it&#8217;s right? My guess is that people want to know how to trade in the stock market and are trying to get some intel.</p><p><strong>Austin:</strong> Yeah, totally. Ultimately that&#8217;s what it boils down to, and I was trying to abstract it in my head. The simplest thing people ask is essentially, is there an interesting insight that you have that I could profit off of that is not obvious to other people? If I abstract that, even when I work with tech companies, I think what the executives want to know is, what am I not taking seriously enough that I should be? It could be a new idea.</p><p>One of the things that I put out into the world before a lot of people saw it coming was, hey, I do think CPU demand is going to shoot up. I didn&#8217;t see it as early as I should have, but once you start to think it through, you&#8217;re like, oh, wait, I&#8217;m just going to have a bunch of interns doing my bidding. What are they going to do? They&#8217;re going to use Microsoft Excel or whatever, Google Drive, whatever. And then I had tech executives say, oh, we read your piece and we hadn&#8217;t really thought this through, but this could impact our supply chain decisions if we&#8217;re going to need more CPUs than we realized.</p><p>So abstracted a little bit, it&#8217;s, because I&#8217;m afforded the time and I have the experience, how can I traverse the frontier of knowledge and just think through interesting things and find one that has an implication that&#8217;s going to really matter? Could matter for institutional investors, could matter for tech executives.</p><p>And actually, I was in a PhD program. I just stopped after my master&#8217;s. But when I was doing research, it was actually very, very similar. I was doing research in graphene and carbon nanotubes, and what you would do is you would survey the landscape. I&#8217;d work with my professor and he&#8217;d be like, go read all these papers as a new student and then see what people are doing, and what haven&#8217;t they done that should be done?</p><p>It&#8217;s that same muscle. What is Nvidia doing? What is AMD doing? What are the problems that startups are working on? Because that can tip you off into problems that the big companies aren&#8217;t solving. And then, what&#8217;s an interesting space that people just haven&#8217;t really thought through much? Let me quick go think about it. Are there any interesting implications? And then I could publish it or I could sell it to clients or what have you.</p><p>That&#8217;s what also gets me up, because you don&#8217;t know what that&#8217;s going to be tomorrow. You&#8217;re like, well, I did this yesterday and I came up with something interesting, but what about next week? I should go put pen to paper and come up with, what are my theories, and try to name them and come back to them.</p><p>But for me, it does usually end up at, okay, this is the way the industry works. There&#8217;s various incentives for two or three players to exist. Yet there&#8217;s always technological revolutions, even on a cadence. If you look at optics and networking, it&#8217;s 400G, then it&#8217;s 800G, then it&#8217;s 1.6T. So there&#8217;s the status quo, and then there&#8217;s inflection points that are going to happen.</p><p>So there&#8217;s questions of, can incumbents take advantage of those inflection points, or could startups come in and take advantage? I guess some of my rules of thumb are, just because you&#8217;re the incumbent doesn&#8217;t mean you will continue to be the incumbent. And just because you&#8217;re a startup doesn&#8217;t mean you can&#8217;t compete here. Usually you just have to be clever about routes to market and whatnot.</p><p>Even with transistors, there&#8217;s a company, imec, that&#8217;s over in Belgium, and they research this stuff and they publish a roadmap out to 2040. So it&#8217;s not crazy for a grad student to just be like, well, let me go think about some of that 2040 stuff. Is there a company to build there?</p><h2>AI-Native Chip Design</h2><p><strong>Gaby:</strong> It&#8217;s fascinating, because as they&#8217;re publishing roadmaps till 2040, we&#8217;re also living in this world of just accelerated, compressed development times, research periods, all of this. What Dario calls the compressed 21st century, where all of research, whether it be across pharma, semis, whatever, is just going to be compressed. I imagine you&#8217;re an avid user, and it sounds like your son is an avid user of AI. I&#8217;m curious how you&#8217;re seeing the strongest teams actually applying AI to this industry.</p><p>We&#8217;re investors in a company that&#8217;s doing AI for materials science with a focus on compute. We see that as a really valuable way to be identifying areas where there&#8217;s new opportunities for materials to drive innovation. But obviously chip design has been a big area, verification, manufacturing. Would love to hear about where you&#8217;re actually seeing AI applied to this industry from a software perspective.</p><p><strong>Austin:</strong> I think this is going to be the Harvard case studies of our time, when we look at this period where you&#8217;ve got companies that weren&#8217;t built AI-native, and then there&#8217;s going to be startups that are built AI-native. So even AI-native chip design, let&#8217;s talk it through. Right now there&#8217;s three big EDA companies, Cadence, Synopsys, Siemens.</p><p>They have all the customers. They have been qualified with TSMC with their PDKs, process design kits. So there&#8217;s all of this momentum, institutional knowledge, and everything that&#8217;s built up. On the other hand, you&#8217;ve got startups who have the freedom to think from first principles and build AI-native chip design. So the question is going to be, who&#8217;s going to win?</p><p>Is it the big guys, who have actually solved a lot of really hard non-technical problems, but have to overcome their old way of thinking and figure out, do they bolt AI into EDA? Or the startups, who are going to build EDA for AI from the start, but then they have to go convince TSMC to work with them and they have to go convince customers to use them?</p><p>I think that&#8217;s what&#8217;s going to be fascinating to watch play out. As far as companies using AI, I&#8217;m surprised when I go talk to even some of my friends in Iowa. They&#8217;ll be working at power companies, traditional power companies. And even they are saying, Austin, all we talk about is AI, and are we doing enough and are we thinking enough? And hey, we&#8217;re doing on-premises. We&#8217;re buying huge racks of GPUs.</p><p>It&#8217;s totally new to us, but we have to either sink or swim, and we&#8217;ve got to learn how to swim quickly.</p><p><strong>Gaby:</strong> Yeah, it&#8217;s interesting. On the AI EDA side, we saw, not Chipstrat, we saw ChipStack get acquired by Cadence. And I think a lot of people are interested to see how Synopsys is going to evolve. But like you said, a lot of these duopolies, or areas where there&#8217;s two to three major players, whether it be in EDA or other areas, are thinking about, how can we integrate AI into our stack?</p><p>And a lot of the question, like you mentioned, is going to be, do we bolt it on or does someone new come in entirely? Part of the question is going to be how relationship-driven and vendor-connected these industries are. Beyond EDA, and the power thing you mentioned around everybody thinking about AI at the forefront, are there other areas in the semiconductor world or compute landscape more broadly where you feel like AI hasn&#8217;t been applied enough? One area I&#8217;ve been learning a little bit more about has been the defect detection world.</p><p>If you think about defect detection, maybe a company like KLA or Lam Research is building technology that can use computer vision and really powerful AI models to detect whether a chip has been produced correctly. That&#8217;s one area I&#8217;m spending time in and learning more about, but I&#8217;d be curious if there&#8217;s areas you feel like are underserved.</p><p><strong>Austin:</strong> Interesting. Yeah, that one&#8217;s cool. AMD just bought a company, I think they&#8217;re called MEXT, and they were doing something really cool, which is applying AI, it is more traditional AI, but to memory. Basically trying to predict cache hits and cache misses and trying to predict access patterns. I thought that was a pretty cool use case. Even something as simple as what you study in computer science 201 or something, when do you write something to a cache, and when there&#8217;s a cache miss, there&#8217;s a write-through, some of this very simple stuff.</p><p>It&#8217;s probably actually still some very simple heuristics or algorithms. What if you applied AI to that? So I do think there&#8217;s probably an opportunity to literally go up and down the stack, at the operating system level, down to how we access memory, and there&#8217;s probably a lot at the networking and data sending layer. Could you predict how you send information even more intelligently than we do?</p><p>I think there&#8217;s a lot of interesting opportunity to explore where to apply AI next.</p><h2>What&#8217;s Next at Chipstrat</h2><p><strong>Gaby:</strong> Perfect. So Austin, tell us about what&#8217;s on the docket. What are you working on? What&#8217;s coming out soon in the Chipstrat world?</p><p><strong>Austin:</strong> What I&#8217;ve been starting to do lately is zoom out from individual companies and look at the markets as a whole, because there&#8217;s a lot of new dynamics happening because of the huge demand-supply imbalance. So memory, zooming out and saying, oh, there&#8217;s all these LTAs, these long-term agreements. How should we think about the dynamics of memory markets, which were traditionally more of a commodity, when they have an LTA in place? They&#8217;ve got agreed commitments from people to buy capacity for five years.</p><p>How does that impact pricing? How does it impact whether they should spin up more capacity or not? There&#8217;s a lot of interesting game theory things where it&#8217;s like, oh, the rules just changed a little bit here. Trying to look at that across optics and networking. And then of course, I think foundry is super interesting again. For all intents and purposes, TSMC was an earned monopolist, which means if they wanted to, they could charge whatever prices they want, but they hadn&#8217;t historically. Then they started leaning into selling their value.</p><p>But now all of a sudden, as Intel Foundry and Samsung Foundry come into play, what does that mean for how quickly they raise their prices, and also how quickly they expand capacity or don&#8217;t? Because expanding capacity is super expensive and risky for someone like TSMC. Again, there&#8217;s these dynamics I&#8217;m trying to dive into, which is, how are they calculating how much they could spend and how much capacity they are leaving on the table?</p><p>And this is also so interesting because if you don&#8217;t build it today, that impacts you in four years. And if you don&#8217;t build it today, you know that it&#8217;s going to go to Intel or Samsung potentially. So are you lifting them up? Or is TSMC believing, no, we don&#8217;t think they&#8217;ll be able to fulfill it? There&#8217;s a lot of gamesmanship. It&#8217;s fun to look across memory, look across foundry, look across all these markets and figure out, how are people thinking about this world where demand is so crazy?</p><p>And then we still don&#8217;t know. Companies are ramping up capacity for 2028, 2029. And even just predicting, is the rate of growth going to slow or is it going to stay? Because of course we see free cash flows going towards zero and there&#8217;s all these financing questions. So I&#8217;m just trying to also project a couple years and think three, four years down the road.</p><p><strong>Gaby:</strong> Awesome. Very excited to be reading whatever market you&#8217;re digging into next. Brian, what are your rapid fire questions for today?</p><h2>Rapid Fire</h2><p><strong>Brian:</strong> So this is a new experiment for us on The Compute 100 podcast. Gaby, we&#8217;ll ping pong. And Austin, you just need to answer quickly. All right. You have to put your entire net worth into one company in the semiconductor industry. Which one would it be?</p><p><strong>Austin:</strong> Oh man. TSMC.</p><p><strong>Gaby:</strong> What pending IPO are you most excited for?</p><p><strong>Austin:</strong> Anthropic.</p><p><strong>Brian:</strong> If one marquee name in the AI revolution, public or private, were to take a true nosedive within the next six months, which would you guess it would be and why?</p><p><strong>Austin:</strong> Off the cuff, my brain goes straight to OpenAI. They&#8217;re awesome, but it&#8217;s just their path dependency of how they got to where they were, essentially at the core being researchers and then thinking, oh yeah, we will go to consumer and we&#8217;ll make a bunch of money. And then it&#8217;s like, dude, people only want to pay $20 a month. Meanwhile, Anthropic&#8217;s like, I know where it&#8217;s at. It&#8217;s charging enterprises hundreds of thousands of dollars for tokens.</p><p>Now, of course, to Sam Altman&#8217;s credit and the OpenAI team, they have done great with Codex and they&#8217;re chasing that, so I think actually they&#8217;re fine. But if they fell off, it could be them. Obviously there&#8217;s distractions around hardware and lots of other vectors it could be tempting to go into. So if they&#8217;re trying to prevent a nosedive, I would just say, dude, focus on enterprises and big customers.</p><p><strong>Gaby:</strong> Who is your favorite CEO on earnings calls?</p><p><strong>Austin:</strong> Favorite CEO on earnings calls. Oh, Hock Tan. Hock Tan from Broadcom. I love him because he is very wise, very experienced, very old, also has a lot of power in the industry, and sometimes he doesn&#8217;t really give a you-know-what. And I love how that comes through.</p><p><strong>Brian:</strong> Shoot, I had one for you. That was a great one, Gaby. Who would you like to hear us interview on a podcast who is not someone you typically hear?</p><p><strong>Austin:</strong> Interview Richard Ho from OpenAI. He&#8217;s their head of hardware. Of course, they&#8217;re building their own hardware, which we&#8217;ve heard about recently with the Jalape&#241;o chip and Broadcom stuff. But they&#8217;re also building the AI that helps them to build the hardware. So what a cool perspective that he has, building the AI, building the hardware. Amazing.</p><p><strong>Gaby:</strong> If you weren&#8217;t doing this, what would you be doing?</p><p><strong>Austin:</strong> Oh man. If I wasn&#8217;t doing this, I would be trying to create the best lifestyle I could, and using my brain. How I make career decisions is just, am I going to get to use my brain more or less? Of course, you&#8217;re balancing the trade-offs there too. When I was in grad school, you have total intellectual freedom and you also don&#8217;t make any money. So at some point you&#8217;re like, well, I have to feed kids.</p><p><strong>Gaby:</strong> The thing that is exciting to your brain is also the biggest market ever, so you&#8217;ve coincided in a good way.</p><p><strong>Brian:</strong> All right, let&#8217;s keep going. Again, I had one, but this is fun. Gaby, your questions are good and I want to get Austin&#8217;s answers.</p><p><strong>Gaby:</strong> What is your favorite book in the semiconductor world?</p><p><strong>Austin:</strong> <em>Focus: The ASML Way</em> is really good. It&#8217;s about ASML. I thought that one was cool. I thought Tae Kim&#8217;s book on Nvidia was good too when it came out. And of course everyone&#8217;s going to say <em>Chip War</em>, Chris Miller. That&#8217;s a great one too. There&#8217;s actually lots of good ones out there.</p><p><strong>Brian:</strong> All right, maybe one more. What&#8217;s a question for us that you&#8217;d be curious to hear about before we wrap up?</p><p><strong>Austin:</strong> What are you guys most excited about, looking three or four years out, not three months out, into the future about compute?</p><p><strong>Brian:</strong> That reminds me of one of my questions for you, which was actually, quantum relevant within the next five years, yes or no?</p><p><strong>Austin:</strong> Quantum companies, please convince me yes, but I say no.</p><p><strong>Brian:</strong> Cool, great. Most excited. Gaby and I spend a lot of time thinking about the rising generation of entrepreneurs and operators who are AI-native and want to build the future of compute for AI. Watching them do seemingly impossible things, and transforming the industry along the lines of what we&#8217;ve been talking about, is a thing of constant fascination.</p><p><strong>Austin:</strong> Yes, I love that. It&#8217;s like, dude, AI is magic. It&#8217;s like Harry Potter before he went to Hogwarts and after he went to Hogwarts. Think about these kids that just have this magic wand. It&#8217;s amazing.</p><p><strong>Gaby:</strong> It&#8217;s crazy. It is a different world spending time on college campuses. I feel like they are learning in their freshman year what I learned in all of college. And that&#8217;s not actually a curriculum thing. That is just the speed at which they can keep up with the world around them and get exposure to things, and a lot of that is obviously because of AI. It&#8217;s just so fascinating. So the new generation of talent here is just really exciting.</p><p><strong>Austin:</strong> Love it. That&#8217;s exciting.</p><p><strong>Gaby:</strong> Amazing. Well, Austin, thank you so much for joining us. We really appreciate the conversation, and thank you for coming on the podcast.</p>]]></content:encoded></item><item><title><![CDATA[Is EDA Dead?]]></title><description><![CDATA[GPT-Synopsys and whether model labs can replace EDA tools]]></description><link>https://newsletter.chipstrat.com/p/is-eda-dead</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/is-eda-dead</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Mon, 05 Oct 2026 18:54:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5w0g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Synopsys had a lot to say about the future of AI and EDA at last week&#8217;s Investor Day in NYC. <em>And I have a lot of thoughts on it myself after attending!</em></p><p>As <a href="https://x.com/StephenNellis">Stephen Nellis</a>  <a href="https://www.reuters.com/business/synopsys-openai-strike-deal-develop-ai-model-chip-design-work-2026-09-30/">reported for Reuters</a>:</p><blockquote><p>Synopsys, a maker of software used in designing computing chips, said on Wednesday that it has struck a deal to share revenue with OpenAI as the two develop an AI model for the chip business.</p><p>GPT-Synopsys, as the model will be called, will be specifically trained and tuned to take on chip design tasks, which start with a description of the chip&#8217;s circuits in a code-like language and span to determining how to lay out billions of transistors on a small square of silicon...</p><p>Ghazi told Reuters in an interview that OpenAI will pay Synopsys a training subscription fee for learning to use Synopsys tools. When Synopsys customers use the product, Synopsys and OpenAI will share revenue based on how well the model improves the design of a chip.</p><p>&#8220;We structured the agreement in a way that it will not be cannibalizing our business,&#8221; Ghazi said. &#8220;It will be an upside to our business given we&#8217;re delivering more value to the customer.&#8221;</p></blockquote><p>Hmmm... is EDA <em>not</em> dead? I could have sworn everyone was saying the model labs were going to eat EDA&#8217;s lunch, because EDA is &#8220;just software&#8221; and OpenAI&#8217;s Jalape&#241;o showed the death of EDA? <em>Right?</em></p><p>Enough FUD, let&#8217;s work through this ourselves.</p><p>First, let&#8217;s discuss how Synopsys positions agentic AI with respect to Synopsys&#8217; EDA tools. Here&#8217;s the portfolio slide from Investor Day:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iZVq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iZVq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png 424w, https://substackcdn.com/image/fetch/$s_!iZVq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png 848w, https://substackcdn.com/image/fetch/$s_!iZVq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!iZVq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iZVq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png" width="1958" height="1106" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1106,&quot;width&quot;:1958,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1003508,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!iZVq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png 424w, https://substackcdn.com/image/fetch/$s_!iZVq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png 848w, https://substackcdn.com/image/fetch/$s_!iZVq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!iZVq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7128cfa8-fdc0-40af-805c-1e9bbf05e22e_1958x1106.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Importantly, the foundation is the <em>tool layer</em>, i.e. the existing EDA tools like VCS/Verdi, RTL-to-GDS, SPICE, Mask/TCAD, and Ansys physics simulation tools on the S&amp;A side.</p><p>One layer above this are <em>task agents</em>, i.e. narrow jobs like Coverage Closure agent, PPA Closure agent, Analog Design agent, and so on.</p><p>Atop this are long-horizon agents which take an objective, decompose it, and orchestrate task agents + tools.</p><blockquote><p>&#8220;A verification agent engineer is able to take an outcome or an objective, like here is a spec 200, 300, 500 pages long, and give me a verified RTL and test benches compared to the spec. Or here is a design that I&#8217;m trying to improve the coverage on, and I have about 30%, 40% coverage. Take this and make this into a 90% coverage situation.&#8221;</p><p><em>Shankar Krishnamoorthy, Chief Product Development Officer</em></p></blockquote><p>As an analogy, if you use Claude Code or Codex for software vibe-coding, the tool layer is using git for version control or curl to hit an API; these are the lower level tools.</p><p>One layer up are skills that package the instructions and scripts for a particular job, like a skill to &#8220;bump this dependency and fix whatever breaks&#8221; or &#8220;write a database migration&#8221;. The skill knows which commands to run and how to check the output.</p><p>And at the top is the long-running agent that takes an outcome like &#8220;upgrade this app to the new version of its web framework&#8221;, and it plans the steps, calls the skills and tools, runs the test suite, fixes what fails, and keeps going until the tests pass.</p><p>Synopsys&#8217; stack is the hardware equivalent concept.</p><p>Of course there are differences&#8230; including the business model. git and curl are free, whereas Synopsys&#8217; tools are licensed. <em>Are there incentives to create free open-source tool alternatives in that foundation layer? Hold that thought.</em></p><p>Synopsys pointed out that agents enable parallel exploration. Why not have several task agents each using a tool to explore the state space in parallel? Better outcomes for the customer (explore more solutions), and more tool use (license revenue) for Synopsys:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ljuX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ljuX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png 424w, https://substackcdn.com/image/fetch/$s_!ljuX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png 848w, https://substackcdn.com/image/fetch/$s_!ljuX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png 1272w, https://substackcdn.com/image/fetch/$s_!ljuX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ljuX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png" width="1974" height="1082" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1082,&quot;width&quot;:1974,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:748504,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ljuX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png 424w, https://substackcdn.com/image/fetch/$s_!ljuX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png 848w, https://substackcdn.com/image/fetch/$s_!ljuX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png 1272w, https://substackcdn.com/image/fetch/$s_!ljuX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8110b83-f432-475c-bb6a-56dc8dddf1df_1974x1082.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p>&#8220;The total work that can be accomplished is gated by the amount of human cycles that are available... you&#8217;re only gated by the compute on the right-hand side. The more compute you have, the more exploration, the more agents that can run in parallel&#8221;</p><p><em>Shankar Krishnamoorthy</em></p></blockquote><p>That&#8217;s pretty awesome. Solution space exploration is definitely limited by the number of engineers on a team, which can&#8217;t easily scale up; but agents can!</p><p>But wait... what about OpenAI and Jalape&#241;o using AI? Recall that OpenAI says <a href="https://openai.com/index/jalapeno-first-results/">AI helped</a> take Jalape&#241;o, its first custom inference chip, &#8220;from initial design to tapeout in nine months&#8221;&#8230; was that replacing Synopsys? Or was OpenAI scaling up it&#8217;s use of Synopsys?</p><p>At Investor Day, Synopsys CEO Sassine Ghazi pointed out Jalape&#241;o still ran through Synopsys EDA, then Broadcom as back-end ASIC partner:</p><blockquote><p>&#8220;when Jalape&#241;o was announced, there was this simplistic extrapolation that if a model can build software, therefore the model can build a Fusion Compiler or a PrimeSim or VCS, and EDA is doomed... That simplistic extrapolation cannot be more far off than reality&#8221; &#8212; <em>Sassine Ghazi</em></p></blockquote><blockquote><p>&#8220;The model needs these guardrails in order to check the physics&#8221;<br>&#8212; <em>Sassine Ghazi, to Reuters</em></p></blockquote><p>Ah ok, so OpenAI was running AI during front-end design, and mostly at the task/agent layer, not the tool layer and seemingly not during back-end design (Broadcom). </p><p>But could the tool layer be replaced by AI too? If not, why not? <em>Hold that thought too, we&#8217;ll get there.</em></p><p>OK then, EDA vendors are the tool layer. </p><p>But how much of that agentic exploration layer value will accrue to EDA vendors vs the model labs?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uE2S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uE2S!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png 424w, https://substackcdn.com/image/fetch/$s_!uE2S!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png 848w, https://substackcdn.com/image/fetch/$s_!uE2S!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png 1272w, https://substackcdn.com/image/fetch/$s_!uE2S!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uE2S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png" width="1962" height="1094" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1094,&quot;width&quot;:1962,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1125965,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uE2S!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png 424w, https://substackcdn.com/image/fetch/$s_!uE2S!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png 848w, https://substackcdn.com/image/fetch/$s_!uE2S!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png 1272w, https://substackcdn.com/image/fetch/$s_!uE2S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde2be1f6-bab3-4a11-8081-1de4d7ef75bb_1962x1094.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Might model labs take over the agentic layers entirely?</figcaption></figure></div><p>The bear case is model labs will own the agent layers circled above, and EDA vendors are limited to the tool layer. </p><p>Synopsys acknowledged these agent layers are something customers are exploring owning themselves, but suggested seen three different paths forward with customers so far:</p><p><strong>1) Use Synopsys agentic layers</strong></p><p><strong>2) Roll your own</strong></p><p><strong>3) Use a frontier lab offering</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bs26!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bs26!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png 424w, https://substackcdn.com/image/fetch/$s_!bs26!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png 848w, https://substackcdn.com/image/fetch/$s_!bs26!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!bs26!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bs26!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png" width="1946" height="1080" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1080,&quot;width&quot;:1946,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1437672,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bs26!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png 424w, https://substackcdn.com/image/fetch/$s_!bs26!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png 848w, https://substackcdn.com/image/fetch/$s_!bs26!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!bs26!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6baf0d6-51e8-4615-b094-1d30408614de_1946x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Digging into each:</p><p><strong>1) Synopsys full stack</strong></p><p>The customer takes everything from Synopsys. The value prop of this turnkey offering is speed-to-market for the customer, and trust that Synopsys is the domain expert here and can unlock token efficiency that others can&#8217;t:</p><blockquote><p>&#8220;We have access to the guts of the tool. We have deep API that they&#8217;re not available in lane two or three.&#8221; &#8212; <em>Sassine Ghazi</em></p></blockquote><p>Obviously, this provides the most revenue opportunities for Synopsys:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q7l4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q7l4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png 424w, https://substackcdn.com/image/fetch/$s_!q7l4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png 848w, https://substackcdn.com/image/fetch/$s_!q7l4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!q7l4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q7l4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png" width="1932" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1932,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:727267,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!q7l4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png 424w, https://substackcdn.com/image/fetch/$s_!q7l4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png 848w, https://substackcdn.com/image/fetch/$s_!q7l4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!q7l4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8977b135-2668-4c69-9dae-95b27c84ce0a_1932x1086.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Pay for the platform, the tool licenses, and agents</figcaption></figure></div><p><strong>2) Roll your own (customer-owned platform + agents)</strong></p><p>Synopys said most early adopters are exploring this lane today; their agents still use Synopsys tool licenses, but they want to own the agentic layer.</p><blockquote><p>&#8220;I have my own special sauce. I want to build my own agent... I do not want to reinvent a debug agent.&#8221;    &#8212;<em>Sassine Ghazi, describing lane-two customers</em></p></blockquote><p>I&#8217;m not surprised most customers would land here, as customers retain optionality (can explore using different models, innovate or optimize the harness, etc), and can limit dependence on Synopsys (and reduce spend with Synopsys). <em>Granted, you&#8217;re paying your own engineer salaries for development and maintenance instead. No free lunch.</em></p><p>OK, so now the surprising announcement, which I was there for :) </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5w0g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5w0g!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg 424w, https://substackcdn.com/image/fetch/$s_!5w0g!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg 848w, https://substackcdn.com/image/fetch/$s_!5w0g!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!5w0g!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5w0g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg" width="2400" height="1728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1728,&quot;width&quot;:2400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:622747,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5w0g!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg 424w, https://substackcdn.com/image/fetch/$s_!5w0g!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg 848w, https://substackcdn.com/image/fetch/$s_!5w0g!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!5w0g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b9add99-9397-43e3-8bd6-f39222e41864_2400x1728.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Paid subscribers get my thoughts on</p><ul><li><p><strong>GPT-Synopsys</strong></p></li><li><p><strong>Labs owning the agentic layer. Bear case? Or not?</strong></p></li><li><p><strong>Outcome-based pricing</strong></p></li><li><p><strong>Can AI replace the tool layer?</strong></p></li><li><p><strong>Retraining</strong></p></li><li><p><strong>Open-source EDA</strong></p></li><li><p><strong>AI simulation shortcuts</strong></p></li><li><p><strong>OpenAI&#8217;s commitment</strong></p></li></ul><p>Is EDA dead?</p><p>Let&#8217;s get into it.</p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/is-eda-dead">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[AI With No Decode (Jev), and What It Means for AI Infrastructure]]></title><description><![CDATA[Jev from TypeSafe answers in one pass instead of writing text token by token. Why LLMs are hard to use inside software, how Jev fixes it, and what a model with no decode phase means for AI Infra]]></description><link>https://newsletter.chipstrat.com/p/ai-with-no-decode-jev-and-what-it</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/ai-with-no-decode-jev-and-what-it</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Sat, 26 Sep 2026 16:41:24 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8ac3b906-16a8-4441-9325-5eb656d1be1f_1734x907.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>LLMs are amazing at writing software. <em>Deterministic software.</em> </p><p>But actually <em>using</em> LLMs within software.. well, that still sucks.</p><p>Which is too bad, because LLMs are great at making sense of unstructured data. They can make judgment calls like &#8220;read this customer complaint and figure out what to do about it.&#8221; </p><p>That use case is so freaking valuable. Just ask <a href="https://sierra.ai/">Sierra</a>, with its $15B valuation and long list of well-known customers.</p><p>Why do LLM calls embedded within deterministic software suck? It&#8217;s not just that it&#8217;s slow and expensive. <em>which is true...</em> </p><p>The big concern is the type signature. An LLM call is a function typed <code>String -&gt; String</code>. If you&#8217;ve ever worked with a statically typed language, this is like the least informative signature in software. <em>For example, if the LLM is expected to return a number, you&#8217;d rather the output signature be Integer</em>. </p><p>When the function just says &#8220;I&#8217;ll take anything and will return anything&#8221;&#8230; lol. It&#8217;s very difficult for the caller to reason about it, so they end up having to write a runtime validator to parse the output and check the schema.</p><p>Sure, you can use <a href="https://openai.com/index/introducing-structured-outputs-in-the-api/">Structured Outputs</a>, which makes sure model-generated outputs match JSON schemas provided by the developer. That&#8217;s an improvement. Before, asking &#8220;Is this customer asking for a refund? Reply yes or no&#8221; to the LLM might give back &#8220;Yes, it looks like they want a refund, although...&#8221; And your code then has to parse that string and hope to extract the correct &#8220;yes&#8221;. </p><p>But with structured outputs, you send a schema like</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;json&quot;,&quot;nodeId&quot;:&quot;93ce84a0-d151-444b-bc1c-efe7bea5a384&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-json">{ &#8220;refund_requested&#8221;: &#8220;boolean&#8221;, &#8220;category&#8221;: &#8220;billing | technical | account&#8221; }</code></pre></div><p>and the reply is guaranteed to be valid JSON in that shape, e.g. </p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;json&quot;,&quot;nodeId&quot;:&quot;ee4935db-2f98-4e43-8248-75643e399121&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-json">{&#8221;refund_requested&#8221;: true, &#8220;category&#8221;: &#8220;billing&#8221;}</code></pre></div><p>This is obviously much better for the developer. </p><p>Yet under the hood it works through constrained decoding, which is still slow. The model still generates the answer token by token, and at each step the server blocks any token that would break the schema. So if the next valid character has to be <code>t</code> or <code>f</code> for a boolean, every other token gets zeroed out. <em>This works, but it&#8217;s still slow and expensive...</em></p><p><strong>There&#8217;s also no confidence signal.</strong> The input might be a long, rambling paragraph from a customer. How confident is the LLM that the customer actually asked for a refund? <em>Shouldn&#8217;t the developer act differently if the confidence is 95% versus 52%?</em></p><p>Yes, you can ask the LLM to include a confidence field in its response, but that number is just another generated token.</p><p>Then there&#8217;s hallucination. What if the answer that comes back was made up? It looks fine to the engineer, and it may even parse cleanly. But it&#8217;s not real! A trustworthy confidence level would help here too.</p><p>So what do we need to call an LLM over and over, deep in the bowels of a program, with confidence? A fixed set of possible answers, a guaranteed response time, and an honest confidence attached to every answer.</p><p>In other words, <strong>intelligence with a function signature</strong>.</p><p>With <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev">Jev</a> from Typesafe AI, we get just that. A frontier-intelligence function call. Unstructured state in, typed probabilistic decisions out. AND FAST:</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;cbada6bd-8e36-4a89-98ad-f4b175fecdd2&quot;,&quot;duration&quot;:null}"></div><p><em>Same 27 questions, one request each, started at the same time. Jev returns all 27 typed answers with probabilities in 0.114 seconds for $0.000081. GPT-5.6 Terra is still generating its answer token by token. Source: <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev">TypeSafe</a>.</em></p><p>With Jev, you send the state and a set of questions, each with a type. A Noul is a yes/no question, and it returns the probability the answer is yes. A Choice picks one option from a list you define. A Score places the input on a scale you define, like calm to very angry:a</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;json&quot;,&quot;nodeId&quot;:&quot;311b2bc1-02e3-4d00-bf42-279da341c0af&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-json">respon se = client.system_one(
    state={
        "ticket": "My flight was cancelled. Can I get a refund?",
        "refund_policy": "Cancelled flights are eligible for a full refund.",
    },
    questions={
        "refund_requested": Noul(instructions="Does the customer request a refund?"),
        "request_type": Choice(
            instructions="What is the main request?",
            criteria={"refund": "...", "rebooking": "...", "information": "..."},
        ),
        "frustration": Score(
            instructions="How frustrated does the customer appear?",
            criteria=["Calm and neutral", "Concerned but civil", "Very angry"],
        ),
    },
)
</code></pre></div><p>What comes back from the model is typed values. <code>refund_requested</code> comes back as a probability, say 0.97. <code>request_type</code> comes back as &#8220;refund&#8221;, along with a probability for every option and a confidence score. <code>frustration</code> comes back as a position on your scale, say 0.4. </p><p>Then your code decides what to do. <em>You focus on the business logic:</em></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;041b0f39-a770-4764-824f-de51d1a67f70&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">a = response.answers

if a["refund_requested"].noul &gt; 0.9 and a["request_type"].choice == "refund":
    issue_refund(ticket_id)

elif a["frustration"].score &gt; 1.5:
    route_to_human(ticket_id)</code></pre></div><p><em>Dude! So simple!</em></p><p>TypeSafe calls it a &#8220;smart if-statement&#8221;, and in their manifesto they say</p><blockquote><p>&#8220;Computers can do so much by just branching on bits, imagine if they could also branch on common sense, understanding, and intent.&#8221;</p></blockquote><p>Jev also can&#8217;t hallucinate an answer, because it can only answer with the options you gave it! Jev can still be wrong, of course. It can pick rebooking when the customer wanted a refund. But it comes with a confidence signal too, which should help clue you in.</p><p>TypeSafe trains Jev with a new post-training method it calls RLCD, reinforcement learning for calibrated decisions. RLHF trained models to write answers people prefer. RLVR trained models to get checkable answers right, like math. RLCD trains the model&#8217;s probabilities to be honest.</p><p>Calibrated means that when Jev says 0.8, it&#8217;s right about 80% of the time. That&#8217;s across many answers, not a promise about any one of them.</p><p><strong>Calibrated probabilities are something you can program against.</strong></p><p>And Jev is fast. </p><p>Why? </p><p><strong>It answers every question in a single pass through the model.</strong> An LLM writes its answer one token at a time, and each token waits on the one before it. Jev has no token-by-token loop! <em>That&#8217;s why output tokens are free btw.</em></p><p>What is Jev good for? Quick judgment calls. Classifying, routing, scoring, extracting, verifying, and ranking. </p><p>Just think of Jev as a tool for developers. <em>Soo&#8230; for agents? :) </em> </p><p>It&#8217;s meant to put fast and cheap intelligence into software.</p><p><strong>And it has implications for AI infrastructure!</strong></p><p>Let&#8217;s dig into some questions you might have:</p><ul><li><p>Decode is the reason AI accelerators have HBM. What does a model with no decode phase need instead?</p></li><li><p>Is this a slice of today&#8217;s LLM inference market, or a new one?</p></li><li><p>What happens to fast decode solutions like Groq and Cerebras?</p></li><li><p>Who wins?</p></li><li><p>Can a humanoid or a car run frontier-class judgment on an edge SoC? </p></li><li><p>What about a phone?</p></li></ul><p>Let&#8217;s get into it.</p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/ai-with-no-decode-jev-and-what-it">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[🎙️The Race to Build the Next Trillion-Dollar AI Chip Company ]]></title><description><![CDATA[TechSurge: Deep Tech VC Podcast]]></description><link>https://newsletter.chipstrat.com/p/the-race-to-build-the-next-trillion</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/the-race-to-build-the-next-trillion</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Thu, 17 Sep 2026 19:29:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/bXNJU3RhW1I" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I recently joined David Goldman, a partner at Celesta Capital, on the TechSurge podcast to talk about the race to build the next trillion-dollar AI chip company. </p><p>We debated things like: </p><blockquote><p><em>What if you could get 10x more tokens, at a fixed interactivity, for a particular model, out of your megawatts than you could buying something off the shelf? </em></p></blockquote><p>It was a great conversation; here&#8217;s the video and transcript.</p><p><em>Highly recommend reading / listening / AI&#8217;ing.</em></p><div id="youtube2-bXNJU3RhW1I" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;bXNJU3RhW1I&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/bXNJU3RhW1I?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><em>This transcript is lightly edited for clarity.</em></p><h2>Why AI sells as systems</h2><p><strong>David Goldman:</strong> Austin, welcome to the show.</p><p><strong>Austin Lyons:</strong> Thank you for having me.</p><p><strong>David Goldman:</strong> AI infrastructure is in the middle of one of the biggest capital buildouts in history, but under the surface, even though the numbers are going up, there&#8217;s been a pretty big change as the dollar spend on inference has gotten bigger than the spending on training. This is a trend that you&#8217;ve covered a lot in your writing and something you have a very unique perspective on. So I want to spend a little bit of time unpacking the implications of this shift.</p><p>Let&#8217;s start at the top with how these systems get sold and what system selling is. Historically, when cloud buyers went out to the semiconductor industry, they were trying to commoditize server hardware as much as possible. They would work with industry groups, they would try to get standard designs done so that they could drive down pricing and negotiate with each part of that design. But increasingly over the last few years, they&#8217;ve been buying whole systems. So maybe let&#8217;s start at the top and just say what goes in a system, and why is this trend happening?</p><p><strong>Austin Lyons:</strong> Yes. I like the idea of starting with systems when we&#8217;re talking about AI, because that&#8217;s really what it&#8217;s all about these days. It&#8217;s rack scale, all the way out to full data center clusters. The compute at the heart of these systems is ultimately GPUs, and that&#8217;s what gets the airtime &#8212; or AI accelerators &#8212; but it&#8217;s actually a full system.</p><p>So stepping back, the question is how did we get here, and why are people buying a full system from one vendor and not commoditizing it? The workload that we all care about today is LLM inference. If you&#8217;re running at the frontier and you&#8217;re wanting to use the latest and greatest OpenAI model or Anthropic model, these are models that are two trillion or more parameters. What that means is, even as you quantize these and try to run them without using as much memory as possible, you might still need something like a terabyte, or two terabytes, or several terabytes of memory just for the model. Because with these models, as we saw with the scaling laws over late 2022, 2023, 2024 &#8212; the bigger the models and the more compute you have, the better answers you get.</p><p>Way back in the day, we might have taken an AI model and put it on just one GPU. But if you think about today&#8217;s frontier models, you need maybe terabytes of memory, and yet a single GPU can only have, for example, 288 gigabytes of HBM on it to store these model weights. So quickly you say, wait a minute, one model can&#8217;t fit in a single GPU&#8217;s memory. So what do you have to do? You need several GPUs. You might say, okay, take that model and split it across four GPUs or eight GPUs. And that&#8217;s how we started to work our way into &#8220;this isn&#8217;t just a single chip anymore, this is a whole system.&#8221; That&#8217;s how we got to, for example, the Grace Blackwell NVL72 rack &#8212; a rack with 72 GPUs in it.</p><p>We&#8217;ve gotten to a place where we&#8217;re talking about AI systems even just for inference. Even if you were only to have one rack, you&#8217;ve got GPUs, CPUs, networking, power, and cooling. So back to the question of why the industry isn&#8217;t commoditizing it: a big piece of where we&#8217;re at today is that Nvidia did a really good job of getting to this rack scale first. They said, hey, we can help design the whole system and make sure it works across 72 GPUs and 36 CPUs, and that they can all talk to each other. And by the way, it also takes software and compilers so that you, the model developer, can write your model and actually run it across all this. So even though the industry always wants commoditization, competition, and multiple suppliers, <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">Nvidia did a good job of frankly getting there first and making this complicated system turnkey</mark>, so that for the software developer it&#8217;s as easy as writing their model, deploying it, and it &#8220; just works&#8221;.</p><p><strong>David Goldman:</strong> So you need a very big, complex system to make these LLMs work in a cloud. But if I&#8217;m a cloud buyer &#8212; Nvidia has famously high margins, and they charge that on all of the different parts of the system, not just on the GPU. If you go out in the Valley, there are all sorts of companies offering one piece of this puzzle. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">If I&#8217;m a cloud buyer, why would I be willing to give so much margin to Nvidia instead of trying to piece each piece together</mark> &#8212; buy my GPU from Nvidia maybe, my networking from someone else, my software from someone else, my cooling from someone else? What kind of value do you get from getting it all at once? Is it around speed? Is it around simplicity? What are the trade-offs that people think about?</p><p><strong>Austin Lyons:</strong> Yeah. <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">We are not yet in the era where people are trying to squeeze down their costs as much. We&#8217;ve still been in the era of speed to market.</mark> We saw that with Elon Musk and xAI, where he&#8217;s been able to stand up data centers very quickly using Nvidia systems. I don&#8217;t think Elon&#8217;s goal was to say &#8220;how can we do this as cheaply as possible,&#8221; but more &#8220;how can we stand up this compute and get productive tokens out of it as quickly as possible?&#8221; So even though a company can potentially pull the components from many different vendors, there&#8217;s just a desire to buy it and stand it up quickly.</p><p>And to your other point, there&#8217;s definitely simplicity there. We see even AMD, with their Helios rack that they&#8217;re bringing to market, seeing that same pull from customers: I want to be able to buy the full rack, get it, stand it up as quickly as possible, and get software running as quickly as possible.</p><p><strong>David Goldman:</strong> Each of Nvidia and AMD have had decades to put together this puzzle, and tons of financial firepower to do M&amp;A. Nvidia acquired Mellanox; AMD has done a bunch of acquisitions to build out these system-selling strategies. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Is it possible for a startup to compete now that you have to sell systems?</mark></p><p><strong>Austin Lyons:</strong> That is a tough and very interesting question. To your point, if you&#8217;re a startup now, you can&#8217;t necessarily just come in and say &#8220;I&#8217;m going to build a better GPU, a better AI ASIC.&#8221; People are going to say, okay, great &#8212; now what do I do with that? Are you going to make your end customers piece together the whole system? We just said they want to stand it up and deploy it as quickly as possible. Or are you, as the startup, now going to have to take on building the rest of the system?</p><p>I think there are probably different approaches here. And by the way, this is why we see that it takes hundreds of millions of dollars now for a chip startup, when maybe back in the day you used to do several rounds of just a couple million dollars to prove out your little proof of concept. <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">So can a startup compete here? Yes, but they have to be very smart</mark>. It could also be death by a thousand paper cuts if you&#8217;re trying to invent a new AI accelerator and you&#8217;re also trying to invent your own proprietary networking and maybe your own proprietary cooling.</p><p>Actually, for example, Cerebras took a very interesting approach. They&#8217;re an AI accelerator startup &#8212; they&#8217;ve gone public &#8212; and they make what they call wafer-scale engines. For those listeners who aren&#8217;t as familiar: instead of taking a whole wafer of chips and dicing them up into individual chips, like two GPUs and a CPU that you package, they said, why don&#8217;t we just leave it at the wafer scale and make the whole wafer our compute engine? But then they had to invent how you make them all communicate, how you cool all of that, how you deliver power to all of that. They basically had to reinvent all of it, and that takes a lot of time. So I think any other AI ASIC company that followed could look at an example like Cerebras and say: <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">wow, they did very innovative technical things here, but they had to invent everything, and that&#8217;s time-consuming and expensive.</mark> How should we do this differently?</p><p><strong>David Goldman:</strong> Yeah. I think you see some parallels with the buildout of internet infrastructure in the &#8216;90s and early 2000s. You start with proprietary, then as open standards proliferate, there&#8217;s more capability for more people to offer things, and you can get a little bit more mix and match. I know you&#8217;ve seen SambaNova, for example, has partnered with Intel to offer x86 CPUs alongside their system. So there seem to be some trends, even though incumbents clearly have an advantage in system selling.</p><h2>Prefill, decode, and the split inference workload</h2><p><strong>David Goldman:</strong> One of those trends, which I think you&#8217;ve touched on a lot in your writing: as we move into this inference era, you&#8217;re starting to see the inference workload itself get split, into what&#8217;s called prefill and decode. Can you tell us what those two things are, and why you might want to split that workload onto different types of silicon?</p><p><strong>Austin Lyons:</strong> Sure. At the end of the day, if you just think of simply talking to an AI chatbot &#8212; you write a paragraph of explanation, &#8220;hey, I want you to go research this and give me an answer&#8221; &#8212; the first thing the AI system, the model under the hood, needs to do is read through the whole prompt that you gave it. That&#8217;s the prefill stage.</p><p>What&#8217;s happening is, in parallel &#8212; one of the key innovations of the transformer model is this idea of attention &#8212; you work through everything in the long English paragraph, and you ask which of these words are connected to other words in the sentence. That way you can piece together the context of the whole paragraph and what each word is referring to. And you can do all of this in parallel. If there are a hundred words, you can look at all hundred words in parallel and compare them to the other hundred words and do these calculations to figure out whether they&#8217;re related or not. That&#8217;s prefill. You&#8217;re ultimately doing a bunch of linear algebra, a bunch of matrix multiplication.</p><p>Now, when it&#8217;s time to give you an answer &#8212; especially if you remember back to when ChatGPT first launched and was a lot slower &#8212; you&#8217;d see the answer come out word by word. That&#8217;s because every word that I&#8217;m about to say depends on the words that I just said. And when I say that word, the next word also depends on those other words. This is decode. This is where you&#8217;re predicting the next tokens, or the next words if we&#8217;re talking about English. And that is sequential. You can&#8217;t do it in parallel.</p><p>People running this inference at scale started to realize: hey, wait a minute. In prefill, I&#8217;ve got all these GPUs doing all of this parallel computation, and it actually doesn&#8217;t require a huge amount of memory. I&#8217;ve paid for this high-bandwidth memory &#8212; which is very expensive, and getting more expensive by the day &#8212; and in prefill, that HBM is just sitting there being underutilized. Interesting. Okay, now I&#8217;ve got these other GPUs that are running decode, and they&#8217;re actually not fully utilizing all of their FLOPS, all their units of compute, because decode is what&#8217;s called memory-bound. It&#8217;s a lot of &#8220;I made a prediction, and I need to get the weights and whatever I need from the KV cache, then do a little calculation, then go back and forth and back and forth,&#8221; one at a time. You&#8217;re not doing things in parallel; you&#8217;re really just waiting on memory.</p><p>So at the highest level &#8212; even Nvidia led the way with this &#8212; they said: we&#8217;ve got a bunch of GPUs in the prefill phase being heavily utilized for compute with their memory underutilized, and a bunch of GPUs in the decode phase basically underutilizing their compute and totally utilizing their memory. Naturally, any engineer is going to look at that and say: huh, maybe we should disaggregate these. Maybe prefill should run on systems that have lots of compute but don&#8217;t necessarily need all that memory. And maybe on the decode side, we should really emphasize memory bandwidth &#8212; how quickly information can get shuttled around &#8212; and maybe it doesn&#8217;t even need quite as much compute; we should essentially overindex on the memory part.</p><p>So naturally, we started to move into a world where Nvidia shipped a software layer for their systems called Dynamo, which helps orchestrate this across Nvidia GPUs. And that actually gave rise to the Groqs and the Cerebrases &#8212; these AI ASIC startups that were actually started even before transformer-based LLMs were the defining workload of our era. They had made architectural choices where they used a lot of this really fast on-chip memory called SRAM &#8212; where you use transistors to store the memory, instead of DRAM, which is capacitors and transistors, which is what HBM is made of. We won&#8217;t go way down into those memory details, but basically these early startups had made a bet on having really high memory bandwidth. And once the workload got separated into prefill and decode, they could raise their hand and go: oh wait, we&#8217;re actually really good at decode. In fact, we can go even faster than GPUs.</p><p>This took us from the inference era where everything was on GPUs to saying: what if you could slot in one of these AI ASICs, heavily built on SRAM, that can maybe unlock a thousand tokens a second? Whereas a GPU running decode &#8212; just the way GPUs are more general-purpose in their design and their memory hierarchy decisions &#8212; maybe could only run at a fraction of that.</p><p><strong>David Goldman:</strong> Is that something you should always do? <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Do we always need to split up prefill and decode? </mark>Or are there just certain applications where it&#8217;s really good to have speed &#8212; I&#8217;m willing to pay a premium for speed, therefore I&#8217;m willing to go through the hassle of splitting these things up, having orchestration software, having different types of silicon and all the things associated with that &#8212; versus just running it on the GPUs or the system I&#8217;ve already bought?</p><p><strong>Austin Lyons:</strong> Yeah, there are so many different nuances here. Take the hyperscalers deploying tens of billions of dollars of GPUs. Early on, GPUs are very general-purpose, very flexible &#8212; they give you the freedom to change your workloads. But let&#8217;s say you&#8217;re OpenAI and you say: no, these are our specific models that we know we&#8217;re iterating on &#8212; the frontier one, the medium-sized one, the small one. You start to say, hey, we should really cater to this workload&#8217;s needs, and therefore it would make sense to deal with the complexities you pointed out &#8212; splitting up prefill and decode and orchestrating that. It might be worth it. Maybe that unlocks, for example, being able to sell an ultra-premium tier where you get really, really fast inference, and maybe there&#8217;s a small subset of users who would pay 10x more to get tokens that are 5x faster. So I think t<mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">here are certain model labs and hyperscalers at scale that are saying the complexity is totally worth it</mark> &#8212; we&#8217;re willing to deal with different SKUs, different chips.</p><p>On the other hand, I think there are going to be tons of enterprises, and arguably the long tail of consumers, who of course want inference, and they&#8217;re thinking about cost and speed, but they aren&#8217;t going to want to manage all of that complexity. For them, where we&#8217;re at today, they&#8217;re not going to need it. Of course, everyone always wants faster inference &#8212; so if there&#8217;s a way for them to get that outcome without dealing with all the complexity, they&#8217;re going to want it.</p><p>When I&#8217;m trying to think about this space and ask whether it&#8217;s one-size-fits-all GPU or many different chips &#8212; where is this going to go in the end &#8212; I look to CPUs. When you look at any CPU vendor, or any cloud like Google Cloud, they don&#8217;t just deploy one CPU, even for customers who want to rent them from their cloud. They have a portfolio: hey, this one has a lot of memory in case you&#8217;re running a database; oh, this one&#8217;s actually pretty vanilla and it&#8217;s nice and cheap if you&#8217;re just running an API server. There are still shapes &#8212; this family of chip, that family of chip. I do think there&#8217;s a world where we get to a couple of different shapes of AI accelerators.</p><p>Where we are now: we had this training era, then this inference era &#8212; those were all on GPUs. Then we&#8217;ve got this next era, which is a GPU plus a specific decode chip that&#8217;s made a good memory-hierarchy trade-off so you can have high interactivity, as they call it &#8212; really fast tokens. But where we are today, those are multi-vendor: Nvidia plus Groq &#8212; of course, Nvidia acquired Groq, so they&#8217;re trying to bring that all in-house &#8212; or someone else plus Cerebras &#8212; AMD plus Cerebras, or Trainium plus Cerebras, or whatever. <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">My thought is that we will ultimately go to those SKUs living inside the same silicon vendor, because it starts to get complicated when your route to market depends on another company</mark> &#8212; I sell a GPU and they sell a decode-specific thing. But it is working right now. There&#8217;s definitely demand for it.</p><p><strong>David Goldman:</strong> It feels like there&#8217;s this inherent tension in the market. You want to buy systems &#8212; that&#8217;s what customers are saying, they want someone to do the work of putting all of these things together. But at the same time, they also want the right silicon for the right job. They want specialized things for decode, and they want the right proportion of CPUs for the workload. So it&#8217;s not a one-size-fits-all system. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Ultimately, do you think you get more and more fragmentation here?</mark> Or are customers going to say: &#8220;Nvidia, please solve this problem for me &#8212; buy Groq, buy the next company, buy the next company, keep selling me systems. AMD &#8212; buy Talis [sp?], work with Cerebras, figure it out for me.&#8221;</p><p>Because, to give you the counter side: there are separate companies for CPUs and GPUs. We&#8217;ve decided those are separate enough markets that they can have standalone companies. We can figure out the system with CPUs and GPUs from different vendors &#8212; that&#8217;s been a solved problem for a long time. So maybe that could be an end state, where you have decode silicon that is a completely different market, and we give it a catchy name like DPU or something. Well, not that one, because it&#8217;s already been used.</p><p><strong>Austin Lyons:</strong> Right, totally. Zooming way out, <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">when you look at the semiconductor industry in the long run, there always seem to be three or four vendors </mark>in a certain thing. Whether you look at wafer fab equipment, foundries &#8212; it&#8217;s complicated, but we&#8217;re getting back to maybe having two or three &#8212; CPU vendors, GPU vendors. We are of course in an era, as happens whenever there&#8217;s drastic innovation, where a ton of competitors have popped up. So I think it would not be crazy to zoom out and say: in the grand arc, maybe only three or four people will shake out, and therefore there&#8217;ll be some sort of consolidation.</p><p>But I personally think this isn&#8217;t as simple as &#8220;Nvidia buys [that startup], AMD buys them,&#8221; even though we&#8217;re seeing some of that &#8212; and of course there are regulatory things to talk about there. If you look up and down the stack, there&#8217;s actually a lot of competition. </p><p>For example, at the neocloud layer: a neocloud might be incentivized to say, hey, I can tell the customer wants the right silicon for the right job, but they also don&#8217;t want to deal with the complexity. I could buy a bunch of different silicon, deal with all the complexity myself, sell them tokens as a service, and try to give them the speed or the cost they&#8217;re looking for &#8212; and differentiate from other neoclouds that way.</p><p>So I do think there are still routes to market where startups can come in today and say: hey, I am the best decode solution, you should try me. Or maybe I&#8217;m a prefill solution. And figure out a way to get to customers while letting the end developer not deal with all that complexity. Then you keep pulling on that thread and play it forward: could an AI ASIC startup ever merge with a neocloud? I don&#8217;t know. Maybe it won&#8217;t just be Nvidia or AMD buying all of these companies.</p><h2>The rise of the neoclouds</h2><p><strong>David Goldman:</strong> You touch on an interesting thing here, because I think neoclouds are really under-discussed when people talk about AI infrastructure. We tend to focus a lot on semiconductor companies and systems companies, and not as much on the cloud layer. And I sort of have a pet theory that it&#8217;s because most of the traditional VCs missed out on those as investments, and so we don&#8217;t like to spend too much time giving credit where we don&#8217;t get to claim any.</p><p><strong>Austin Lyons:</strong> Yeah.</p><p><strong>David Goldman:</strong> But the reality is that if you look at the handful of neoclouds that came up in this first wave, they&#8217;ve created a lot more equity value. The top three are worth something like $125 billion in public markets, which is a lot more than what we&#8217;ve been talking about with some of these chip startups. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">So why do you think investors had a lot of trouble understanding the first wave of neoclouds and missed out on those investments?</mark></p><p><strong>Austin Lyons:</strong> Yeah. For listeners, by &#8220;neocloud&#8221; we mean a cloud company that started by just renting GPUs. I&#8217;ll talk through my hesitation when I first saw the idea of neoclouds, and maybe I&#8217;m a fair proxy for some investors &#8212; maybe they thought this way too.</p><p>Let&#8217;s back up. Who&#8217;s renting from these neoclouds? Well, it turns out a lot of it is hyperscalers, which is ultimately driven by demand from the biggest model labs. So the question is: wait a minute, you&#8217;re saying that OpenAI is running on GPUs that Microsoft is renting from some neocloud? Doesn&#8217;t Microsoft just have their own data centers? And the answer is: of course they do, and they&#8217;re trying to build more, but at the end of the day we are limited by access to power. There are financial reasons to rent versus buy &#8212; maybe even go off balance sheet as the capex has increased year over year over year. So there are very legitimate reasons why even someone as sophisticated as Microsoft Azure might say: actually, I want to rent GPUs from someone, and I&#8217;m also building other data centers, but maybe it&#8217;s a stopgap.</p><p>Looking at that, I said: okay, so you&#8217;re going to have a neocloud company &#8212; maybe they have access to power, and maybe they have a good way to raise capital against assets like this. Bitcoin miners, for example, historically have had this experience, and maybe they&#8217;re shifting into GPUs. And so I thought: huh, okay. They were doing Bitcoin mining, now GPUs are hot, so they&#8217;re going to do GPUs, but they&#8217;re just going to rent them to Azure. But Azure is also standing up their own data centers. S<mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">o how is this sustainable? They&#8217;ve got huge customer concentration &#8212; literally maybe one customer</mark>. But to their credit, that&#8217;s how you get the financing: you say, I&#8217;ve got Microsoft who&#8217;s going to rent these from me, and so people will lend against that and believe in that.</p><p>That&#8217;s part of why I missed it as these companies were popping up. There&#8217;s a legitimate need in the marketplace for people who have access to power, who can get financing, who can manage this, and who can move quickly &#8212; to stand up GPUs, run models on them, and rent it out as bare metal or maybe at a higher level of abstraction. <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">And even though they have serious customer concentration risk, so does everyone else in the semiconductor industry right now.</mark> Who are Nvidia&#8217;s end customers? Even Nvidia, the biggest and best, is selling a lot of GPUs to a small set of customers.</p><p><strong>David Goldman:</strong> Yeah, it&#8217;s not uncommon to see someone go public with 90% customer concentration these days.</p><p><strong>Austin Lyons:</strong> Totally. That is the name of the game. Look at any component supplier &#8212; even in the interconnect space or the switching space. Look at Credo, who made these active electrical cables, a really awesome invention. Same thing: even when they went public, they were selling to a handful of customers. The name of the game for this era is watching how this unfolds and seeing how some of the people who win one big customer ultimately win a couple more, build out more of a portfolio, and reduce a little bit of that risk. But that&#8217;s just the way the industry works right now.</p><p><strong>David Goldman:</strong> If you wind the clock back five years, before these neoclouds got big &#8212; AWS is an incredible business; Amazon and Microsoft have fortress balance sheets, great relationships with Nvidia and all the semiconductor companies, expertise in setting up data centers, software, customer relationships. It seems to me they could have done this, and certainly most people thought they would, which is why so many missed out on these investments. Was it a strategic decision, where they thought &#8220;maybe this isn&#8217;t going to be a big enough market, I&#8217;m not sure I want to spend all the money and take the risk&#8221;? Or was there some special sauce in what these neoclouds were able to do &#8212; setting up quicker, being more creative in financing, converting old Bitcoin data centers? Were they doing something different from what Microsoft might have done, or was Microsoft ceding market share to them, or some other factor? Not to pick on Microsoft &#8212; insert anyone here.</p><p><strong>Austin Lyons:</strong> Maybe a little bit of everything. At the end of the day, this is risky business, because the investments might be $40 billion this year, $80 billion next year, a hundred billion the year after that. I don&#8217;t think it&#8217;s that the existing players didn&#8217;t believe the future we&#8217;re in now would manifest &#8212; I think it&#8217;s all about timing. And they&#8217;re probably not incentivized to just sprint out and stand up all these data centers. It&#8217;s a huge investment.</p><p>The neoclouds were able to move faster and take on more risk. They can kind of go for broke: all right, cool, let&#8217;s convert, let&#8217;s buy a bunch of GPUs, let&#8217;s get moving. It&#8217;s a bit like the innovator&#8217;s dilemma. There are already three big clouds, they already have all the customers, and I think they were probably looking at each other &#8212; AWS and Google and Azure &#8212; saying: are we all believing this future is coming and investing in GPUs at the same rate as each other? But that could still not be enough supply to meet demand.</p><p>Now, I do think demand went higher faster than everyone expected, especially once we got reasoning models, and now of course the agentic age. So even if they all looked around and said, &#8220;we think if we grow supply like this &#8212; it&#8217;s a little risky and these numbers feel really big, but we think the demand will be there&#8221; &#8212; demand skyrocketed past that, and it gave neoclouds an opportunity to come in and say: yeah, we&#8217;ll fill that gap.</p><p><strong>David Goldman:</strong> I think there was also maybe a difference in interest level. You heard some of the hyperscalers make comments about how these bare-metal GPU instances are low-margin and sort of commodity, and therefore they didn&#8217;t want to support them as much. Whereas the neoclouds read them as revenue, and therefore good.</p><p><strong>Austin Lyons:</strong> Yeah, you make a fair point. Anyone whose business was renting CPUs, or selling services on top of CPUs &#8212; the cost structure there is much better than GPUs. So you could see CFOs saying: wait a minute, we&#8217;re going to spend a ton of money and our margins are going to go down, even if our margin dollars go up. There are still conversations to be had that might make you slow down or hesitate a little bit.</p><p><strong>David Goldman:</strong> <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">So is this a durable state of affairs?</mark> Obviously this is a very fast-moving market, but &#8220;first mover,&#8221; &#8220;more willing to take risk,&#8221; &#8220;willing to take on lower margin&#8221; &#8212; those are not necessarily durable advantages that will last a decade or more. What do you see playing out with these neoclouds? Do they get acquired by hyperscalers? Do they consolidate into a neo-hyperscaler or something?</p><p><strong>Austin Lyons:</strong> Yeah, it&#8217;s very interesting. I&#8217;m not sure there will always be a need for a hundred neoclouds, but I do think it&#8217;s real demand, and it will remain real.<mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);"> I don&#8217;t think the big three clouds will always meet everyone&#8217;s needs indefinitely.</mark> Not only that &#8212; there&#8217;s always going to be innovation. Take World Labs, Fei-Fei Li&#8217;s company. They&#8217;re coming out with world models. Who&#8217;s going to make the bet there? What if those don&#8217;t run well on the exact shape of hardware all these big clouds have invested in? I think there will always be new workloads or new demands popping into existence where a neocloud can pop up quickly and say: I can meet your need, and I can innovate there. That will always exist. There may not be enough cutting-edge frontier demand for a hundred neoclouds to hop on, but I definitely think there will always be a need for these more nimble, smaller GPU and AI ASIC rental companies that can innovate a lot closer to where the model labs and AI-enabled software companies are innovating.</p><h2>Circular financing</h2><p><strong>David Goldman:</strong> I feel like we can&#8217;t talk about neoclouds without at least discussing circular financing risk. This is probably the thing I hear most from people who are skeptical about the durability of AI infrastructure. The argument goes: the neoclouds get equity investment from Nvidia in many cases. They use that equity to buy GPUs. They then use those GPUs as collateral to take on debt. A lot of that debt is sometimes backstopped by either their hyperscaler or by Nvidia itself. It all sort of perpetuates, and the revenue from the neocloud buying GPUs goes back to Nvidia, who invests in more neoclouds. People see a circular financing issue, like the vendor financing that happened in the dot-com bubble and was revealed to be an inflator of that bubble. On their last earnings call, Nvidia addressed this directly and said they see it differently. Do you agree with them, or do you think there are some concerns here?</p><p><strong>Austin Lyons:</strong> You know, I was definitely the type of person where, right away, I said: okay, this is different, this feels funny, I need to dig in and understand it and try to understand both sides. I can definitely see why it looks like circular financing &#8212; call it what you will.</p><p>But going back to thinking through demand, supply, and the cost of capital: it is a fact that there&#8217;s just insatiable demand, especially now with agentic AI, where the barrier to entry for software development has gone as close to zero as possible. I&#8217;ve got a son who made a 70,000-line video game this summer, and he didn&#8217;t write those lines himself &#8212; he used Codex to do it. That&#8217;s amazing, it&#8217;s unreal. And just wait until he&#8217;s in high school and college and beyond. He&#8217;s going to use AI so much more intelligently than me. He&#8217;s going to use way more tokens than me. I can&#8217;t even believe what it&#8217;s going to be like in the future. You can look at every industry and see lots of people writing software, doing interesting things they couldn&#8217;t do before. So the demand is totally real.</p><p>The supply is very fixed. At the end of the day, you might look at someone like TSMC &#8212; there are only so many wafers that come out, only so much CoWoS capacity. But even as GPUs get built, the question is who has the capital to buy them. Today we might be talking $5&#8211;10 million a rack, or more. Who has that kind of money lying around? So there is this cost-of-capital, financing thing that comes into play. If customers are saying &#8220;I just want inference, as fast as possible, as soon as possible, please make it happen&#8221; &#8212; who in the supply chain has the money to invest in standing up all these data centers and running them? It could be neoclouds. Okay &#8212; do they have access to the capital they need to make tens of billions of dollars of investment, or even a few billion? A lot of these might be early companies, or former Bitcoin miners. They might have some access to capital, but not enough.</p><p>If you&#8217;re Nvidia, sitting there with all this money as the world&#8217;s biggest company, and you&#8217;re saying demand is incredible and supply is what it is &#8212; but it&#8217;s not just building the GPUs, it&#8217;s getting them powered up, financed, stood up &#8212; of course it makes sense from their perspective to help get this stood up. The banks aren&#8217;t so sure they want to lend to the neoclouds, but if Microsoft or Google or AWS says &#8220;I&#8217;ll be the offtaker &#8212; don&#8217;t think about the neocloud when you&#8217;re lending, think about me,&#8221; and if Nvidia can also come in and say &#8220;we want to help make this happen, can we put our brand behind it, backstop it, whatever&#8221; &#8212; I can see why Nvidia would want to do that. Now, does that mean they also benefit from it? Of course, totally. It&#8217;s customers, and maybe even revenue sharing &#8212; a new source of revenue for Nvidia.</p><p><strong>David Goldman:</strong> Well, that&#8217;s capitalism. They&#8217;re not going to do it if they don&#8217;t benefit.</p><p><strong>Austin Lyons:</strong> Exactly. But as sort of a techno-optimist: if it&#8217;s going to feel funny, but there are ways to get more compute stood up faster so that more people around the world can do the awesome things they&#8217;re trying to do, then I&#8217;d say, all right, I can get behind that.</p><p><strong>David Goldman:</strong> I think it ultimately boils down to differences of opinion on the durability of the cash flow that comes from these assets. If you went out and said &#8220;I&#8217;m going to build a toll road,&#8221; you can get a lot of financing for that. You don&#8217;t have to put in a lot of equity, and you can get a lot of debt, because people know this road is going to have X number of cars, we know the traffic patterns, you&#8217;ll collect this amount of money. It&#8217;s very safe. You can raise lots of debt at very low rates for projects like that, even if you&#8217;re a new company. Obviously there&#8217;s a big difference between a toll road and a GPU-based AI factory, as some people are calling them. But from Nvidia&#8217;s perspective, and from some of the hyperscalers&#8217; perspectives, these are fairly safe assets that will be completely paid back in two or three years, and then there&#8217;s a stream of cash flows coming out of them. Even if demand goes down a little or doesn&#8217;t grow at the same rate, you&#8217;ll still get value out of them, and they&#8217;re not going to depreciate super quickly. Nvidia has that view and is willing to put their balance sheet behind it, the hyperscalers feel the same way &#8212; and not everyone else has the same view. That&#8217;s where the rubber is going to meet the road.</p><p>The flip side, though, is that you have the existing neocloud set who are now very embroiled with Nvidia &#8212; they&#8217;re in lockstep, they rely on them for financing. To the extent that customers are going to demand more and different silicon, that creates an opportunity for new neoclouds who can figure out how to make it all work together and choose the right silicon for what customers want. Do you think that&#8217;s effective counter-positioning? Will we see another wave of neoclouds like the first time, or will the existing set &#8212; the CoreWeaves, the Nebiuses, companies like that &#8212; figure this out and just start using Nvidia plus Groq, or AMD plus whoever?</p><p><strong>Austin Lyons:</strong> Right. I think both will exist, but you make a good point about the trade-off. It&#8217;s kind of like golden handcuffs. The trade-off for a neocloud is: hey, they&#8217;re backstopped by Nvidia, maybe they got financing, and they&#8217;re obviously getting allocation for GPUs. So those particular neoclouds might feel like, if there&#8217;s different silicon out there that&#8217;s very competitive, even for a subset of workloads, we may not feel like we can go buy it and offer it &#8212; because what if we don&#8217;t get as much allocation in the future? Don&#8217;t bite the hand that feeds you. So I do think there will be opportunities for other neoclouds to come in and say: we&#8217;ve got a bunch of different silicon, and maybe we can abstract it and run your workloads across it so you don&#8217;t need to worry about it. There will be opportunities for someone to come in and counter-position.</p><h2>Four conditions for the next trillion-dollar chip company</h2><p><strong>David Goldman:</strong> In the last couple of years there have really been two big outcomes &#8212; we&#8217;ve touched on both &#8212; in the semiconductor space: Groq, which sold to Nvidia, and Cerebras, which went public. But neither one really won a hyperscaler before they were able to do this, and in the case of Groq, they sold to Nvidia. You wrote an article about what you think the conditions are for the next trillion-dollar chip company, and you had four conditions: capable of running trillion-plus-parameter models; rack-scale chips, which we&#8217;ve touched on; beating an incumbent on a KPI; and landing a frontier anchor. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Points one and four &#8212; running trillion-parameter models and landing a frontier anchor &#8212; are correlated around the idea that the frontier matters most. Why do you think that&#8217;s the case, and a prerequisite to being the next big breakout company in silicon?</mark></p><p><strong>Austin Lyons:</strong> Very interesting question. <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">Ultimately, today, it&#8217;s the best models &#8212; especially if you can run them at fast enough speeds &#8212; where my belief is the outsized value will accrue.</mark> Yes, there are lots of use cases where you can use older models, smaller models, and you don&#8217;t have to run them as fast. I think that pie will continue to expand. I just don&#8217;t think people will pay a premium for it.</p><p>So if you&#8217;re an AI accelerator company and you&#8217;re trying to put as much muscle behind a few arrows as possible, you&#8217;d want to compete at the frontier. Today it&#8217;s software developers saying: yes, I will pay &#8212; not $200 a month; we&#8217;ll pay for tokens, thousands of dollars a month, tens of thousands a month &#8212; if we can get it fast, if we can get Claude Fable, for example, or the latest OpenAI model. And of course the frontier will always keep getting better.</p><p>It also feels like the frontier is where there will be the least competition, because GPUs, for example, can&#8217;t get there today on speed, and we know that. If you&#8217;re aiming at a 70-billion-parameter Llama 3 and going after all those workloads that are valuable but don&#8217;t need the highest intelligence, there&#8217;s going to be a lot of competition there too &#8212; and it could literally be old Hoppers or old Amperes from Nvidia.</p><p>When you look at that Pareto frontier curve we always see &#8212; I&#8217;ll describe it for people who are just listening: the x-axis is interactivity, which is how fast the tokens come, tokens per second per user; the y-axis is throughput. The slower you go, the more tokens you can generate concurrently; but way out there on the far right, going fast, even if you can&#8217;t serve as many users &#8212; that&#8217;s where the value is accruing today. And it&#8217;s hard to see a world where that changes.</p><p><strong>David Goldman:</strong> That graph gets shown a lot, particularly when Jensen or people from Nvidia talk about what they&#8217;re going to be able to do with Groq plus Nvidia. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">But is tokens per second per user &#8212; that interactivity KPI &#8212; still the right one for startups to think about?</mark> Or are there changing needs because of power constraints, cost constraints, new workloads like agentic coding? Is it still all about speed?</p><p><strong>Austin Lyons:</strong> It&#8217;s not all about speed. In my opinion, the way to compare people is at a fixed interactivity, for a given unit of power. We are power-constrained. Ultimately, if a neocloud gets access to 100 megawatts, they&#8217;re going to have to ask themselves: how can I get as much revenue as possible out of this 100 megawatts? They might say, I want to bet on allocating some of my megawatts to really fast tokens, because I think we can charge more for them. So if you&#8217;re going way to the right on the interactivity curve &#8212; let&#8217;s say they&#8217;re aiming for 800 tokens per second or higher, because they feel they can charge a premium for that, and they&#8217;ve only got so many megawatts &#8212; then they&#8217;re going to ask: how many concurrent users can I serve? What is my token throughput? <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">So I think it&#8217;s token throughput at a fixed interactivity, normalized by power.</mark> But again, not every workload needs that.</p><p><strong>David Goldman:</strong> I didn&#8217;t hear you say the word &#8220;cost.&#8221; One of the ways you get better interactivity is by using more expensive memory, using SRAM. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">So how much does cost play into that equation?</mark></p><p><strong>Austin Lyons:</strong> Cost plays into it. If you&#8217;re in that use case &#8212; you&#8217;re a neocloud, you&#8217;ve got 100 megawatts, you&#8217;re trying to generate as many tokens at a fixed interactivity as you can &#8212; you also have a fixed budget to spend on compute. If you buy an Nvidia Vera Rubin rack, the latest and greatest, plus nine accompanying Groq LPU racks, you might get really high on that interactivity, and it might be pretty good power-normalized &#8212; but you might have spent half your budget, or all of your budget, right there. So <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">cost absolutely comes into play.</mark> When I&#8217;m thinking about the user experience, I&#8217;m thinking about interactivity and how many people can be served. But if you&#8217;re a neocloud or any buyer of compute, you&#8217;re definitely thinking about cost. <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">And cost could mean: I can get the same performance out of two racks from this vendor versus eleven racks from that one. Or, for the same fixed cost, what if I could get eleven racks from this new competitor &#8212; and therefore five or ten times more tokens at that interactivity?</mark></p><h2>Enterprise and on-prem AI</h2><p><strong>David Goldman:</strong> This kind of buyer thinking is very emblematic of a cloud or hyperscaler with a huge instance they&#8217;re trying to spread over lots of users. I personally always struggle with holding two ideas in my head at the same time. On the one hand, all of the initial demand and value has been going to the frontier labs, served by a combination of hyperscalers and neoclouds. On the other hand, we and many other people believe AI is going to be something as big as the internet. It&#8217;s going to diffuse into businesses all over the world. Every company is going to have some AI element, in the same way every company has a website now &#8212; there are no more dot-com versions of companies; everyone&#8217;s got a website, everyone&#8217;s got an app. Soon everyone will have some AI element in their business. And in today&#8217;s internet world, more workloads exist on-prem than in the cloud. If I&#8217;m a coffee shop, I might have a server in my coffee shop; I&#8217;m probably not going to have an AWS account. So ultimately, <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">does enterprise actually become the big market here instead of cloud?</mark></p><p><strong>Austin Lyons:</strong> That is a very good question. <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">I think they&#8217;re both going to be massive. </mark>Let&#8217;s talk enterprise, because I don&#8217;t think people appreciate that enough. Even Nvidia is trying to get ahead of it &#8212; they changed their reporting and their business units so that for data center it&#8217;s essentially hyperscaler and non-hyperscaler. It gets a little fuzzy, but they&#8217;re saying that, by the way, the non-hyperscaler segment is growing faster than the hyperscaler one, and right now the revenue is pretty close on both &#8212; a little fuzzy because they put neoclouds in there.</p><p>So the question is: what workloads are going to go to the enterprise, and why? You&#8217;re totally right that there is going to be diffusion of generative AI across every industry, and I definitely don&#8217;t think we&#8217;re there yet. Marc Andreessen said fifteen years ago that software is eating the world &#8212; every company is going to be a software company. And to some extent that was right: logistics, manufacturing, healthcare &#8212; they&#8217;re all using software, even if it&#8217;s just internal tools. <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">I think we&#8217;re going to a world where agentic AI eats the world, and every company is going to be an agentic AI company.</mark> Like I said, pointing at my son: fast-forward fifteen years, and yes, they&#8217;re all going to be agentic.</p><p>Okay &#8212; so if agentic AI is core to how businesses run, are they all going to have hundreds of millions of dollars to spend on tokens every year? Totally not. There are going to be all sorts of reasons for companies to want to deploy workloads on premises. One: cost. Two: owning your own data, and figuring out how you even differentiate in a world like that. It won&#8217;t be all-or-nothing &#8212; go to the cloud for the frontier workloads, and do as much as you can of the older, smaller models on premises.</p><p>By the way, you could also read Nvidia&#8217;s hyperscaler/non-hyperscaler split as a proxy for frontier closed models versus open-source models, because if you&#8217;re running enterprise AI locally today &#8212; on premises, on a server, on your desktop &#8212; it&#8217;s got to be an open-source model. That&#8217;s a little bit why we&#8217;re in the world we&#8217;re in today: if you want the best model, you have to go to the cloud. And for everything else &#8212; <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">do you want to pay for tokens, or for token generators? I think a lot of people, if they can afford it, would rather own token generators.</mark></p><p>But fast-forward a little: what happens if open-source frontier models can keep up, or be good enough? I think there will continue to be a rise in the amount of workloads run on premises. There are all sorts of reasons. Of course you can point to regulated industries that will have to run on premises. But even if I&#8217;m not regulated: maybe I think the labeled data my humans generate is valuable. Say I&#8217;m an insurance company. We have all this agentic stuff doing claims processing, and my humans are going in and correcting it &#8212; that was good, that was good, that was wrong. Let&#8217;s keep that data internally and fine-tune our own model so it gets it right in the future. Maybe I want to run that locally because I&#8217;m in charge of the model, I&#8217;m in charge of the data, I keep it, it&#8217;s all my intellectual property. And maybe I see that as how I differentiate in the future &#8212; I&#8217;ve got better agents than my insurance competitor.</p><p>Can you do all this stuff in the cloud and feel like it&#8217;s secure? You totally can. But at the end of the day, when we&#8217;re talking about diffusion, we want every engineer at every company to be able to tinker with it, touch it, play with it, use it themselves. Sometimes when stuff&#8217;s in the cloud, you get the convenience, but you lose the ability to get under the hood &#8212; whether it&#8217;s a closed model or even an open model in the cloud.</p><p><strong>David Goldman:</strong> It&#8217;s funny you say that &#8212; <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">I almost feel the opposite way about enterprise. To me it seems gated a little bit by software</mark>. So many of these companies haven&#8217;t moved workloads even to the cloud, for data sovereignty, regulatory, privacy, IP protection reasons. They&#8217;re very concerned about stuff leaving their IT premises, and they would gladly do more in AI, but they don&#8217;t have engineers in-house who know how to post-train a model or tinker with this stuff. Those people are expensive, and they don&#8217;t want to hire them. So they want to do it on-prem in the enterprise, but no one&#8217;s quite figured out how to help them do that yet. That, to me, feels like the missing piece that would unlock a lot of enterprise hardware sales.</p><p><strong>Austin Lyons:</strong> I definitely agree with you &#8212; that&#8217;s where we are today. Back up eight years: everyone wanted to be a software company, but they didn&#8217;t have software engineers. I live in Iowa, so if you were a software engineer there and willing to work in insurance or ag or retail, you were a rock star. You could walk in and they&#8217;d say: yes, thank you, we need you, we didn&#8217;t have this capability before. Fast-forward to now, and anyone can vibe-code, which is actually pretty awesome, because now the domain experts &#8212; the person in insurance who knows insurance really well &#8212; can actually build the solution they want.</p><p>I think we&#8217;re now where we were eight years ago, but for fine-tuning: I&#8217;m an insurance expert and I can vibe-code a thing, and I&#8217;ve got some software people here, but none of us know how to fine-tune yet. That&#8217;s the education piece. <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">If my children had to go to college today and pick a major, I&#8217;d say: pick anything, plus machine learning, and learn how to fine-tune, because then you can understand a domain and also understand how to actually apply AI. So I think there will probably be a rise of AI engineers, if you will.</mark> Maybe agents will do this for you and bring that cost down to zero faster than agents brought software down to zero &#8212; which took, you know, forty years. But it is a pain point today that companies don&#8217;t have generative AI familiarity yet. I don&#8217;t think that pain point will be there forever. Maybe it&#8217;s five years, maybe more, I don&#8217;t know. But I don&#8217;t think it will always be a blocker. <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">If every company became a software company with software-literate people on staff, I think eventually everyone will have LLM-fine-tuning-literate people on staff too.</mark></p><p><strong>David Goldman:</strong> So when that day comes, is this just a huge unlock for Nvidia, and they get that much more revenue? Or do you think anyone else has a chance at that market?</p><p><strong>Austin Lyons:</strong> That&#8217;s a great question. What happened to IBM, you know? There are giants, and they&#8217;re first, and they ride a huge wave. And then, to all the points we&#8217;ve talked about &#8212; zooming out, there are three or four winners, and people want competition. The more people can tinker, the more this diffuses, the more opportunity there is. One company cannot meet everyone&#8217;s needs at the right price point, at the right speed. They can meet lots of people&#8217;s needs, but there will always be people trying to do some interesting bespoke thing who say the off-the-shelf stuff uses too much power &#8212; I know I&#8217;ve got this crazy setup, but I can&#8217;t do 130 kilowatts, I can only do 50. There will always be emerging workloads where the off-the-shelf stuff just doesn&#8217;t fit.</p><p>And the tough part is, when you&#8217;re Nvidia and you&#8217;ve got these huge hyperscalers, you&#8217;re not necessarily incentivized to go find those little people. You&#8217;re not interested in picking up pennies; you&#8217;re interested in picking up billion-dollar bills. So I think Nvidia will be totally fine &#8212; they have great solutions, and they&#8217;re always going to have customers coming to them for the latest and greatest, deployed quickly. But I do think there will continue to be new opportunities popping up, especially as this diffuses, where people can compete.</p><h2>Clean-sheet silicon for LLMs</h2><p><strong>David Goldman:</strong> Circling back to this idea of the next trillion-dollar company. There&#8217;s an explosion of opportunities in different workloads, and also, as we&#8217;re hearing, different markets. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Does that mean there&#8217;s an opportunity for a trillion-dollar company? Or are we going to get $500 billion companies?</mark> AMD is still not even a trillion-dollar company, and they&#8217;ve been around for a really long time and have a lot of pieces of this puzzle.</p><p><strong>Austin Lyons:</strong> I just think about the size of the market. Here&#8217;s the deal, and part of why I came to that conclusion. GPUs obviously have a history in graphics and doing things in parallel, and that has been shifting toward AI-centric. Nvidia is saying: yes, these data center GPUs &#8212; you&#8217;re not going to run Doom on them. They used to support FP64, and they still do, but we&#8217;re going to spend all of our transistors, as we do a node shrink, on FP4 and FP8, the lower precision that AI models really want. But at the same time, it was still a general-purpose GPU.</p><p>That&#8217;s why I said we went from training with GPUs, to inference with GPUs, to this next era where <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">LLMs are </mark><em><mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">the</mark></em><mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);"> workload &#8212; and we still haven&#8217;t had silicon designed specifically for LLMs</mark>. We had GPUs that morphed from their early roots to fit the shape of what we&#8217;re doing. Let&#8217;s take a GPU and slap on this SRAM thing. <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">But no one yet has really brought a chip to market that was designed specifically for LLMs. So the question is: what if you can be the first one that can stand up a gigawatt&#8217;s worth? </mark>Which is not simple.</p><p><strong>David Goldman:</strong> Well, to some degree, the Nvidia GPUs of the last couple of cycles <em>are</em> chips designed specifically for LLMs. It&#8217;s not like you can take Grace Blackwell and play a video game on it easily. It&#8217;s highly specialized for this type of workload. And in particular, when you get into these combination GPU-LPU systems, or the AMD Helios, these are really designed specifically for LLM workloads.</p><p><strong>Austin Lyons:</strong> They are <em>morphed</em> specifically for LLM workloads. They have legacy ways of doing the networking, legacy ways of thinking about the memory hierarchy. I definitely agree that they are iterating toward what is best for the inference workloads they serve. But the question is: what if you started with a blank sheet?</p><p>We actually saw this from OpenAI with their Jalape&#241;o chip, which they presented at Hot Chips recently. They said: we started with a blank design, and we&#8217;re thinking very differently about it, making very different architectural decisions. They said, instead of shipping all this KV cache around and having all this shared memory and all this contention &#8212; the data is never in the right place when we want it, and our compute is always sitting around &#8212; what if every accelerator had its own little HBM slice? What if we mapped the workload and rethought it so we don&#8217;t have all this contention and shipping data around?</p><p>Tensordyne is another example. They said: should it be matrix multiplication, or could we do log math? That would turn multiplies into adds, which are really fast in silicon. So I actually do think there are architectural knobs. </p><p>Nvidia&#8217;s next GPU still has to have backward support for the software &#8212; not exactly, but to some extent; they want to support all the workloads, so if it ran on Hopper they want it to mostly run on Blackwell. But what if you could start with a clean sheet and make different architectural decisions &#8212; about the way you do the compute, the memory hierarchy, the way you network it up?</p><p>Etched said: let&#8217;s do low-voltage inference. What would happen if we ran this at lower voltage, instead of putting in a lot of oomph so it can go really fast? If power is fixed, maybe there are benefits, even with trade-offs &#8212; it doesn&#8217;t run as fast, but it&#8217;s significantly lower power. </p><p>So I do think there&#8217;s opportunity to make clean-sheet designs, make different architectural decisions, and therefore unlock that KPI I talked about. <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">What if you could get 10x more tokens, at that fixed interactivity, for that particular model, out of your 100 megawatts, than you could buying something off the shelf from Nvidia or AMD?</mark></p><p><strong>David Goldman:</strong> You hit on another interesting tension. On the one hand, you have someone like OpenAI, who is a large customer of Nvidia but is also now designing their own chip. When you hear them talk about how they did that design, they claim it was very AI-optimized &#8212; that they were able to do it much quicker because of AI acceleration. And then you also have companies like Tensordyne, who have an entirely new paradigm, a whole new way of approaching this problem, that a customer probably wouldn&#8217;t have thought of on their own, because they&#8217;re more focused on their specific workload than on new ideas from first principles.<mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);"> How do you think about the trade-off between custom silicon companies using AI to create something specific for what they need &#8212; maybe working with a Broadcom or Marvell to help them finish and get it into production &#8212; versus these companies with entirely new approaches?</mark></p><p><strong>Austin Lyons:</strong> It&#8217;s so interesting to think about OpenAI buying compute from the big vendors, mostly GPUs, while also building their own silicon &#8212; with the advantage of their chip designers working hand-in-hand with their software team, co-designing for specific workloads. Off-the-shelf silicon vendors are not inventing in a vacuum; they&#8217;re talking with their biggest customers, asking where their roadmaps are going and how to make sure their silicon meets those needs. But that&#8217;s different from OpenAI internally having their ML team and their chip team working closely together and co-designing.</p><p>That said, they&#8217;re going to land on a particular set of trade-offs. Everything in engineering is about trade-offs. Do you want more HBM or more SRAM? It&#8217;s going to cost you something either way. Do you want more die size, or do you stack it higher and take the thermal trade-offs? <mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">The merchant vendors made a particular set of trade-offs &#8212; and they need to sell their chips to as many customers as possible, even if it&#8217;s only a handful these days. The internal teams will make their own set of trade-offs, given what they know about the workloads they&#8217;re running.</mark></p><p><mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">But to think those two sets of trade-offs are all you&#8217;ll ever need? I think there&#8217;s opportunity for someone like a Tensordyne to say: what if there are these totally crazy trade-offs, and we take the risk of doing the R&amp;D on them?</mark> The log math stuff &#8212; surely it&#8217;s crossed the minds of people at the merchant silicon vendors or the internal XPU teams; they&#8217;ve seen a paper. But they may not be incentivized to take the risk. If you&#8217;re OpenAI&#8217;s XPU team making your first chip, are you going to play around with log math? Or are you going to say: no, no, let&#8217;s make these other interesting but proven decisions, like HBM slices that have been used elsewhere in industry?</p><p><strong>David Goldman:</strong> Yeah. People forget there are people involved in these decisions who have career risk.</p><p><strong>Austin Lyons:</strong> Exactly.</p><p><strong>David Goldman:</strong> If you can make a decision that&#8217;s high-probability and still works, but maybe isn&#8217;t the 10x &#8212; that&#8217;s probably better for you if you&#8217;re at a big company, particularly delivering your first version of something. You don&#8217;t want to screw up.</p><p><strong>Austin Lyons:</strong> And it&#8217;s expensive, too. You&#8217;re going to tape it out, and you might be all-in a billion dollars or something. You don&#8217;t want to get that wrong, or have it get cancelled before you can stand it up. So yes, there&#8217;s this whole human side, these incentives. And again, that continues to create opportunities for startups to say: we&#8217;re going to take that risk. Or, in Tensordyne&#8217;s case: we&#8217;ve been taking that risk &#8212; we were trying the log math in a different market, and now we&#8217;re ready to bring it to this one. Could an OpenAI say &#8220;that&#8217;s super interesting, let&#8217;s try a couple racks&#8221;? Absolutely &#8212; these are very sophisticated buyers, and they&#8217;re always watching what else is out there, because they completely understand that designs made at a particular point in time may not fit exactly what they need a couple of years from now.</p><h2>AI-designed chips and the falling barrier to custom silicon</h2><p><strong>David Goldman:</strong> Are more companies looking at custom chips? OpenAI has a lot of money and a lot of really talented engineers &#8212; even without AI, they probably could have designed their own chip. But as AI makes designing a chip easier &#8212; it&#8217;s not at the level of your son creating a video game with Codex, but you talk to teams and they&#8217;re seeing lots of improvements &#8212; there&#8217;s some floor to how cheap it can get, because you ultimately need to tape out and do things in the physical world. <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Do you think we&#8217;ll see a big proliferation of companies that never would have tried custom silicon giving it a shot?</mark></p><p><strong>Austin Lyons:</strong> I definitely think so, and you can already look at examples where it&#8217;s happening. Why would you design your own silicon? If you know your workload really well, what you bought off the shelf &#8212; where design decisions were already made &#8212; might not map to your workload perfectly. That might be okay at first. But take Rivian, for example. They went through a couple of different merchant silicon vendors, mapping their workload to the parts, and it was fine enough. But if you&#8217;re running an electric vehicle and trying to do autonomous driving, you have very specific needs. You want to use as little power as possible, because otherwise you&#8217;re taking battery away from the customer being able to drive another couple of miles. On the other hand, you need to run as fast as possible, because if it takes too long to make a decision, that&#8217;s another 20 meters of braking distance while you were thinking. Now all of a sudden they&#8217;ve got lidar and cameras and all this data flowing around that they need to do inference on as fast as possible. And by the way, it used to be convolutional neural networks; now they&#8217;re doing end-to-end VLA &#8212; vision-language-action &#8212; models. The workload has been changing, and in this world of real-time inference on big models with power under control, plus a particular interconnect bandwidth and memory capacity and memory bandwidth &#8212; they&#8217;re saying: man, this stuff off the shelf just doesn&#8217;t fit our needs and our cost profile.</p><p>So the answer could be: design your own chip. Well, there&#8217;s a cost to that. You need engineers familiar with front-end design, back-end &#8212; maybe you can partner for testing &#8212; and at the end of the day, that might take you three years of development. You&#8217;re paying those engineers for years, there&#8217;s a roadmap, it&#8217;s a very big investment, and then there&#8217;s the cost to tape out and actually get it built. But what if, with the help of AI, it still takes a hundred people but instead of three years it takes one? Maybe your cost is cut by two-thirds. Where before, even though the performance and the headroom would let you do interesting things, your CFO was saying &#8220;we just can&#8217;t add another $2,000 to the bill of materials&#8221; &#8212; maybe now you come back and say it&#8217;s only going to be $700 on the BOM, and they say: okay, that&#8217;s really interesting, we think we can hack it.</p><p>So I think adding AI to chip design &#8212; speeding up time to market, doing more with the same number of people &#8212; will ultimately be net good, and will reduce that cost barrier to entry. Or it could be a talent thing: how do I go find a hundred people? Maybe you only need fifty. <mark data-color="rgb(255, 243, 160)" style="background-color: rgb(255, 243, 160); color: rgb(0, 0, 0);">I just think it will reduce barriers to entry, and we&#8217;ll see all sorts of cases where people who never thought about making their own chip will say: that&#8217;s actually something we could do.</mark></p><p><strong>David Goldman:</strong> It&#8217;s an interesting dynamic, because the same speed-up available to these companies is also available to the merchant silicon teams &#8212; who are probably even better positioned to use these tools, because an experienced engineer who understands the trade-offs is going to use the tool better than someone approaching their first chip design. So maybe they can start proliferating their number of SKUs and serve more customers with more semi-custom things within what is somewhat merchant. I just think it&#8217;s really interesting. I don&#8217;t know how it&#8217;s going to work.</p><p><strong>Austin Lyons:</strong> Yeah. I would hope that merchant silicon companies &#8212; the thing is, when you&#8217;re a merchant silicon company making a particular product, there has to be a big enough market to capture enough customers for it to be worth your time and investment. There&#8217;s probably a point where the ROI didn&#8217;t make sense: the market&#8217;s only this big, the SKU costs this much to develop &#8212; not worth it. But if the cost can come down by half, maybe it clears the internal rate of return you needed. So yes, I&#8217;d expect big companies to be able to create more innovations as well.</p><p><strong>David Goldman:</strong> Awesome. Well, I don&#8217;t think anyone really knows how it&#8217;s going to play out. This is such an exciting time. Thank you so much &#8212; this has been an incredible conversation. I&#8217;ve really enjoyed it.</p><p><strong>Austin Lyons:</strong> Yes, thank you. This was fun. Let&#8217;s do it again.</p><p><strong>David Goldman:</strong> Absolutely.</p><h2></h2>]]></content:encoded></item><item><title><![CDATA[A Billion Agents Need Somewhere to Run]]></title><description><![CDATA[Grok Bot, Meta Muse, and what consumer adoption means for CPUs, accelerators, memory, and advanced packaging.]]></description><link>https://newsletter.chipstrat.com/p/a-billion-agents-need-somewhere-to</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/a-billion-agents-need-somewhere-to</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Fri, 11 Sep 2026 21:31:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/537a45e4-fc6f-4c65-8b04-87d31acc2d98_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It finally seems like the average consumer can adopt agentic AI. Reasoning models were the first milestone, harnesses like <a href="https://www.anthropic.com/news/enabling-claude-code-to-work-more-autonomously">Claude Code</a> came next, and personal agents like OpenClaw unlocked useful work beyond just coding. Of course, running your own agent still required a willingness to tinker and spend money. <em>Go buy a Mac Mini! Spend $200/mo on Claude Code!</em></p><p>My hypothesis is that we&#8217;re on the verge of agentic AI crossing the chasm from early adopters to the early majority because product companies are turning these capabilities into simple, turnkey, yet useful services. <em>And the upfront cost is low or zero. For now&#8230;</em></p><p><a href="https://x.ai/news/introducing-grok-bot">Grok Bot</a> showed the way:</p><div id="youtube2-F1_0Lkp16Rc" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;F1_0Lkp16Rc&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/F1_0Lkp16Rc?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Free. Simple. </p><p>From the video:</p><blockquote><p><em>There wasn't anything to learn. It was just like bringing on a co-worker to my team</em></p></blockquote><p>The same video hints at the impact this could on the AI infra supply chain:</p><blockquote><p><em>It's not using your computer. So you can keep doing work in parallel while bot is working on another instance.</em></p></blockquote><p>Meta shipped <a href="https://ai.meta.com/muse/">Muse</a> recently; a personal agent with its own cloud computer, accessible through an app or WhatsApp:</p><div id="youtube2-wHn0hTjvFoo" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;wHn0hTjvFoo&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/wHn0hTjvFoo?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><em><strong>WhatsApp!!!</strong> </em></p><p>Meta&#8217;s distribution <em>(3.6B daily active users across the family of apps)</em> is where adoption could inflect. <em>If Meta does it right, they have a direct connection to billions of active users!! And AI infra demand too. Good thing they are building out all those datacenters, even if they keep getting punished by investors because of it&#8230; </em></p><p>Of course right now Muse is starting with a US rollout so those billions of users are <em>potential</em> reach, not necessarily day one Muse users. But obviously downloading an app is a familiar reflex for nearly everyone in the world; setting up a Mac Mini etc is not. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ufPP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ufPP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png 424w, https://substackcdn.com/image/fetch/$s_!ufPP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png 848w, https://substackcdn.com/image/fetch/$s_!ufPP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!ufPP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ufPP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Agentic AI Is About to Take Off&quot;,&quot;title&quot;:&quot;Agentic AI Is About to Take Off&quot;,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Agentic AI Is About to Take Off" title="Agentic AI Is About to Take Off" srcset="https://substackcdn.com/image/fetch/$s_!ufPP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png 424w, https://substackcdn.com/image/fetch/$s_!ufPP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png 848w, https://substackcdn.com/image/fetch/$s_!ufPP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!ufPP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d04427-e69b-4c97-811c-9a2f6f97e197_3072x2048.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The path from reasoning models to consumer agents, with innovations like Meta Muse opening the door to the early majority.</em></figcaption></figure></div><p><em>And btw, if you&#8217;re feeling brave, there&#8217;s also <a href="https://instinct.co/">Instinct</a>:</em></p><div id="youtube2-6Yeo6M8-W_8" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;6Yeo6M8-W_8&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/6Yeo6M8-W_8?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong>Consumer cloud agents will continue to accelerate demand for server CPUs.</strong> </p><p>Yes we have a lot of ChatGPT and Claude adoption already amongst normies, but not necessarily heavily agentic use from most that aren&#8217;t coding/vibe-coding. <em>Although Astra is showing more normie use cases like driving DaVinci for editing. </em></p><p>But Muse, Grok Bot, Instinct show simple user experiences have the potential to quickly drive adoption to the masses. And each person can generate more than a single agent run per day; Grok Bot&#8217;s docs say:</p><blockquote><p><em>Multiple Bots share one user-scoped computer and can run in parallel.</em></p></blockquote><p><strong>Could 100 million people run 1 billion daily agents?</strong> <em>Seems plausible.</em> </p><p>Of course, a billion daily agentic tasks doesn&#8217;t imply needing a billion CPU cores because it&#8217;s not like we&#8217;ll have a billion <em>simultaneously active</em> agents. That said, there will still be a heck of a lot of CPU cores running simultaneous agents, and those agents will each kick off some LLM inference and quite often also stress the existing cloud infra (fetching data from the web and so on). <em>And the question is whether Grok and Meta and Instinct have enough CPUs for this growth?</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wYpJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wYpJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg 424w, https://substackcdn.com/image/fetch/$s_!wYpJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg 848w, https://substackcdn.com/image/fetch/$s_!wYpJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!wYpJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wYpJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg" width="550" height="313" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:313,&quot;width&quot;:550,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The iconic line 'You're gonna need a bigger boat' from the movie Jaws&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The iconic line 'You're gonna need a bigger boat' from the movie Jaws" title="The iconic line 'You're gonna need a bigger boat' from the movie Jaws" srcset="https://substackcdn.com/image/fetch/$s_!wYpJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg 424w, https://substackcdn.com/image/fetch/$s_!wYpJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg 848w, https://substackcdn.com/image/fetch/$s_!wYpJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!wYpJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8606e730-8f89-4a67-a829-74eced8d2a87_550x313.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>And these are all ADDITIVE to OpenAI and Anthropic, who will continue to innovate and make it easier for folks to use agents to do the things they find useful at work or home.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.chipstrat.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.chipstrat.com/subscribe?"><span>Subscribe now</span></a></p><h2>What does this mean for the ecosystem?</h2><p><strong>For paid subscribers:</strong></p><ul><li><p><strong>CPUs:</strong> We obviously need more CPUs. But which ones?</p></li><li><p><strong>Accelerators:</strong> Which accelerators? Nvidia? AMD? XPUs? <em>Startups?</em></p></li><li><p><strong>HBM and DRAM:</strong> Agents need memory for the model and for the computer it drives. Which demand grows faster?</p></li><li><p><strong>NAND:</strong> Agents leave a lot of work behind. What does that mean for flash demand?</p></li><li><p><strong>CoWoS and EMIB:</strong> More accelerators means more advanced packaging. Who captures the growth?</p></li></ul>
      <p>
          <a href="https://newsletter.chipstrat.com/p/a-billion-agents-need-somewhere-to">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Oligopoly Equilibrium]]></title><description><![CDATA[Why most semiconductor markets settle at three players]]></description><link>https://newsletter.chipstrat.com/p/oligopoly-equilibrium</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/oligopoly-equilibrium</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Tue, 08 Sep 2026 17:11:23 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/65a4f810-987b-4d10-8cfc-99e064b5a0db_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Why do most semiconductor markets have ~3 main companies? &#8230; </strong><em>give or take, depending on the market.  </em>For example:</p><ul><li><p><strong>EDA:</strong> Synopsys, Cadence, Siemens</p></li><li><p><strong>Memory:</strong> Samsung, SK Hynix, Micron</p></li><li><p><strong>Leading-edge foundry</strong>: TSMC, Samsung, Intel</p></li><li><p><strong>GPU:</strong> Nvidia, AMD</p></li><li><p><strong>CPU</strong>: Intel, AMD, Arm</p></li><li><p><strong>AECs:</strong> Credo, Marvell, Astera</p></li><li><p><strong>Scale-up interconnect:</strong> NVLink, UALink, ESUN</p></li><li><p>Litho: ASML alone at EUV, only ASML and Nikon at immersion</p></li></ul><p><strong>Fixed costs are the main driver.</strong> The buy-in to compete in nearly all semiconductor markets is large and rises every generation. <em>This isn&#8217;t SaaS! Interestingly, model lab dynamics look similar; the costs to play the game are enormous and rising every generation.</em></p><p>Take leading-edge foundry as an illustration. Chip companies used to own their fabs. But the cost of a leading-edge fab never stops rising... $100M, then $1B, then $10B: </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gO9G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gO9G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png 424w, https://substackcdn.com/image/fetch/$s_!gO9G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png 848w, https://substackcdn.com/image/fetch/$s_!gO9G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!gO9G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gO9G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png" width="544" height="414.94068965517243" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1106,&quot;width&quot;:1450,&quot;resizeWidth&quot;:544,&quot;bytes&quot;:167033,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.chipstrat.com/i/214747618?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gO9G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png 424w, https://substackcdn.com/image/fetch/$s_!gO9G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png 848w, https://substackcdn.com/image/fetch/$s_!gO9G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!gO9G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6da2bd-8e81-4f2b-bfc7-3c790dd2c9b5_1450x1106.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><a href="https://www.construction-physics.com/p/how-to-build-a-20-billion-semiconductor">source</a></figcaption></figure></div><p><strong>Only companies with serious volume can amortize that.</strong></p><p>Most chip companies didn&#8217;t sell enough chips to keep a fab running at full capacity. So each generation, more of them sold off their fabs. <em>That&#8217;s where the fabless design industry and merchant foundries came from.</em></p><p><strong>So the number of players in leading-edge fabs keeps falling because the market share needed to break even keeps rising.</strong> We can turn that into a rough rule. Take the market size and divide it by the revenue one player needs to break even, and that gives us how many players the market can support. <em>As a first-order toy model&#8230;</em></p><div class="pullquote"><p><strong>N &#8776; market size / breakeven scale</strong></p></div><p>Let&#8217;s do some napkin math as an illustration. We can approximate the size of the leading-edge foundry market; TSMC did $122B in<a href="https://investor.tsmc.com/english/encrypt/files/encrypt_file/reports/2026-01/00fe50f72b38d74e6b9b066398f020f337cd4e9d/4Q25%20Management%20Report.pdf"> revenue in 2025</a>, and 74% of that was from &#8220;advanced nodes&#8221; of 7nm and below (call it ~$90B), so if TSMC has 90%+ of that market, we can call the market $100B+. Then let&#8217;s say a single sub-2nm fab is ~$20B-$30B. A $100B+ annual market with a $20-$30B buy-in supports about three players? <em>Seems reasonable. Of course if the market share is even 60%, 25%, 15%, you&#8217;re looking at less revenue in a year than the cost of a single fab&#8230; </em></p><p>Yes, you amortize a fab and don&#8217;t pay for the whole thing up front; TSMC depreciates equipment over five years, so a $20B-30B fab is roughly $6B a year of cost. <em>So that&#8217;s how one company can build and own many fabs at once.</em> But the CapEx treadmill never stops; you&#8217;re building the next fab while depreciating the last one, hence the $41B of annual CapEx. And that doesn&#8217;t account for R&amp;D costs. Staying on the roadmap cost TSMC <a href="https://investor.tsmc.com/english/encrypt/files/encrypt_file/reports/2026-01/00fe50f72b38d74e6b9b066398f020f337cd4e9d/4Q25%20Management%20Report.pdf">$41B of CapEx plus about $8B of R&amp;D in 2025</a>. <em>And I think <a href="http://chipstrat.com/p/tsmc-could-spend-10b-more-on-n2-capacity">they could spend even more</a>.</em></p><p>We can visualize the break-even volume for a credible program at each node, meaning a fab or two of capacity plus the process R&amp;D to stay on the roadmap. At 2nm that&#8217;s roughly two $28B fabs depreciating plus a few billion a year of R&amp;D, call it ~$14B a year of fixed cost:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aAXg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aAXg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png 424w, https://substackcdn.com/image/fetch/$s_!aAXg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png 848w, https://substackcdn.com/image/fetch/$s_!aAXg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png 1272w, https://substackcdn.com/image/fetch/$s_!aAXg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aAXg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png" width="1456" height="1028" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1028,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:190862,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.chipstrat.com/i/214747618?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aAXg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png 424w, https://substackcdn.com/image/fetch/$s_!aAXg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png 848w, https://substackcdn.com/image/fetch/$s_!aAXg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png 1272w, https://substackcdn.com/image/fetch/$s_!aAXg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42598627-7c58-4a82-9166-3fce750a7133_2040x1440.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Directional estimates. The dots show a breakeven point. Each curve is a credible program at that node (fab capacity plus process R&amp;D), not a single fab. Clearly each new node requires more wafer starts per month just to breakeven, even if ASPs are increased too. To the left of the dot, the program loses money on every wafer. Breakeven volume is why most IDMs transitioned to fabless.</figcaption></figure></div><p></p><p>Look at the hypothetical 2nm break-even of ~97K wafer starts per month, nearly 1.2M wafers a year. That&#8217;s an entire <em><a href="https://www.tsmc.com/english/dedicatedFoundry/manufacturing/gigafab">gigafab</a></em> worth of demand (TSMC&#8217;s own term for a 100K wspm site), running essentially full, all year, just to pay for the program.</p><p>Of course, fabs don&#8217;t start at that kind of volume. A new fab ramps for months at lower wafer starts while yield matures. So the buy-in isn&#8217;t just the fab. It&#8217;s the fab plus the underutilized ramp, losing money on every wafer until volume and yield catch up.</p><p><strong>So why do most semiconductor markets hold ~3 main players?</strong> </p><p><strong>Why not fewer? </strong></p><p><strong>Why not more?</strong> </p><p><strong>Do growing markets allow more entrants?</strong> </p><p>And sure, foundry has the highest fixed costs so it&#8217;s a great illustration, but <strong>why does this extend to other markets?</strong></p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/oligopoly-equilibrium">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Open Models & American Diffusion; Nvidia & Poolside]]></title><description><![CDATA[America needs agentic AI to diffuse. Diffusion needs competitive frontier open models.]]></description><link>https://newsletter.chipstrat.com/p/open-weights-and-american-diffusion</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/open-weights-and-american-diffusion</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Fri, 28 Aug 2026 14:41:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c2085bb3-b9c5-43e7-b648-151f560cb7f2_1456x802.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Nvidia is paying $6 billion for Poolside, a model lab. From the WSJ&#8217;s <a href="https://www.wsj.com/tech/ai/nvidia-is-spending-6-billion-to-build-a-powerful-u-s-alternative-to-chinese-ai-c51c38cc">Nvidia Is Spending $6 Billion to Build a Powerful U.S. Alternative to Chinese AI</a></p><blockquote><p>Nvidia is planning to use a $6 billion deal it struck this week to build one of the world&#8217;s most powerful open-weight AI models, one that would compete with Chinese heavyweights like DeepSeek and Kimi K3, according to people familiar with the matter.</p><p>The chip giant&#8217;s licensing deal with the AI startup Poolside is also set to present a direct challenge to frontier U.S. AI companies including OpenAI and Anthropic, since open-weight models are generally far cheaper to operate and allow easy customization.</p></blockquote><p>Nvidia has long been shipping the Nemotron family of open models, with the <a href="https://www.youtube.com/watch?v=_y9SEtn1lU8">stated goal</a> of giving every company a trusted base model to specialize for their bespoke needs:</p><blockquote><p><em>&#8220;Our goal with Nemotron is to raise the bar for general intelligence, but provide that intelligence in a way that allows businesses around the world to specialize it for their own problems&#8221;.  &#8212;Bryan Catanzaro, Nvidia VP</em> </p></blockquote><p>Nemotron is open and customizable. <a href="https://artificialanalysis.ai/articles/nvidia-nemotron-3-ultra-released">Artificial Analysis</a> called Ultra &#8220;the leading US open weights model&#8221;, while noting it sits behind &#8220;the Chinese-led open weights frontier&#8221;. <em>For now.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CXlw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CXlw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png 424w, https://substackcdn.com/image/fetch/$s_!CXlw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png 848w, https://substackcdn.com/image/fetch/$s_!CXlw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png 1272w, https://substackcdn.com/image/fetch/$s_!CXlw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CXlw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png" width="1200" height="505" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:505,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CXlw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png 424w, https://substackcdn.com/image/fetch/$s_!CXlw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png 848w, https://substackcdn.com/image/fetch/$s_!CXlw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png 1272w, https://substackcdn.com/image/fetch/$s_!CXlw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F190cb7ab-be2e-4823-84eb-b504133a14e4_1200x505.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Benchmarked back in June. Ultra sits in the middle here, see the little purple arrow. <a href="https://artificialanalysis.ai/articles/nvidia-nemotron-3-ultra-released">Source</a></figcaption></figure></div><p>Nvidia is working on it. Back in March, it launched the <a href="https://nvidianews.nvidia.com/news/nvidia-launches-nemotron-coalition-of-leading-global-ai-labs-to-advance-open-frontier-models">Nemotron Coalition</a> with partners like Mistral, Cursor, Perplexity, and Thinking Machines Lab to build competitive frontier open models. <em>Can a coalition move at the speed of light though?</em> </p><p><em>As we&#8217;ll see below, Poolside can help with iteration speed&#8230;</em></p><p>But first and more importantly, why do competitive frontier open models even matter?</p><h2><strong>Diffusion matters as much as innovation</strong></h2><p>A technology that can transform every sector is, by definition, a general-purpose technology. <em>General-purpose technology is also abbreviated GPT, but not that GPT... I know, slightly confusing.</em></p><p>Check out Jeff Ding&#8217;s book: <a href="https://press.princeton.edu/books/paperback/9780691260341/technology-and-the-rise-of-great-powers?srsltid=AfmBOopm_L8xGZxUaftNpp6S0E_-gVTN8UDi3GFl6X432PXZMdsQXPaG">Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition</a>. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uyFx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uyFx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg 424w, https://substackcdn.com/image/fetch/$s_!uyFx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg 848w, https://substackcdn.com/image/fetch/$s_!uyFx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!uyFx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uyFx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg" width="254" height="385.9560975609756" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:623,&quot;width&quot;:410,&quot;resizeWidth&quot;:254,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uyFx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg 424w, https://substackcdn.com/image/fetch/$s_!uyFx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg 848w, https://substackcdn.com/image/fetch/$s_!uyFx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!uyFx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2794dc9e-eeb0-414d-9e1d-f15cbd2a65fa_410x623.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The important takeaway is that the <strong>countries that win with GPTs aren&#8217;t the countries that </strong><em><strong>invent</strong></em><strong> the GPT, but the countries that </strong><em><strong>diffuse</strong></em><strong> the GPT.</strong></p><p><em>We idolize invention; diffusion sounds boring. But it&#8217;s super important!</em></p><p>Actually, let&#8217;s back up a bit. Marc Andreessen&#8217;s 15-year-old "<a href="https://a16z.com/why-software-is-eating-the-world/">Software Is Eating The World</a>" saw the diffusion of another general-purpose technology (software) and predicted it would upend every industry. <em>Even the software industry!</em> Every company would become a software company. Of course software wasn&#8217;t <em>new</em>, but thanks to various innovations like cloud computing and higher-level programming languages &amp; frameworks, the rate at which it could <em>diffuse</em> was increasing.</p><p>Notably, Andreesen said the constraint preventing companies and industries from software-enabled productivity gains was <em>people</em>.</p><blockquote><p><em>&#8220;Many people in the U.S. and around the world lack the education and skills required to participate in the great new companies coming out of the software revolution... There&#8217;s no way through this problem other than education, and we have a long way to go.&#8221; <br>&#8212;Marc Andreesen</em></p></blockquote><p>This is a <em>diffusion</em> argument. For software to eat the world, it needs to diffuse across every industry, and that&#8217;s predicated on having enough software-engineering-literate <em>people</em> in every industry.</p><p>Crazy enough, fifteen years after Andreesen&#8217;s article, we see traditional software development diffusing rapidly across every industry thanks to generative AI! Now, suddenly, domain experts can be literate software engineers. </p><p><em>Hey Marc, let me update it 15 years later:</em></p><blockquote><p><em>&#8230; There&#8217;s no way through this problem other than <s>education</s> vibe coding.  </em></p></blockquote><p>Jeff Ding&#8217;s <a href="https://www.youtube.com/watch?v=wXucwFa_7do">interview with Microsoft&#8217;s Brad Smith</a> makes the same skilling argument for Generative AI. Professor Ding argues that countries that see expanded GDP and increased productivity are those that can properly diffuse the general-purpose technology to the &#8220;average engineer&#8221;.</p><p>Should that occur, industries will get rebuilt not just around software, but around generative AI. </p><p>After all, agentic AI is a general-purpose technology that has already transformed software engineering and will eventually transform every industry. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HOeS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HOeS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HOeS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HOeS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HOeS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HOeS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg" width="443" height="354.24150268336314" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:447,&quot;width&quot;:559,&quot;resizeWidth&quot;:443,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HOeS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HOeS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HOeS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HOeS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b0d75a-acc5-49ca-a97b-63435e8442ee_559x447.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>No longer &#8220;every company will be a software company&#8221; but now &#8220;every company will be an agentic AI company&#8221;</em></figcaption></figure></div><p>The question that follows, then, is what will it take for agentic AI to diffuse broadly? Not just the diffusion of <em>software engineering</em> thanks to agentic AI, but the actual application of agentic AI across every industry.</p><p>Jensen argues that one important ingredient is the <strong>existence of frontier-class, competitive open-weight models.</strong></p><p>And I wholeheartedly agree. </p><p>From Jensen&#8217;s <a href="https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf">&#8220;Open Weights and American AI Leadership&#8221;</a> published on July 24, 2026:</p><blockquote><p>&#8220;Our AI leadership will be judged not by one frontier AI model, but by whether the United States builds a strong, open ecosystem that diffuses into every sector&#8221;</p></blockquote><p>Yes. Diffusion. The fastest diffusion is predicated on the existence of competitive frontier opens weights models. Again from Jensen:</p><blockquote><p>&#8220;Startups, established businesses, universities, and public institutions can build on advanced models without training one from scratch or paying frontier-model prices for every task. Open weights let every organization match the right model to the right job at the right cost, reserving frontier-scale capability for genuine frontier problems and running efficient, specialized models everywhere else.&#8221;</p></blockquote><p><em>Preach it Brother Huang!</em></p><p><em>And ideally fully <a href="https://opensource.org/ai/open-weights">open source</a>, i.e. open model, data, and weights &#8212; see <a href="https://openmdw.ai/">OpenMDW</a>. Also a nice explainer table from <a href="https://opensource.org/ai/open-weights">opensource.org</a>:</em></p><div class="captioned-image-container"><figure><div class="image-link image2" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CSKJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png 424w, https://substackcdn.com/image/fetch/$s_!CSKJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png 848w, https://substackcdn.com/image/fetch/$s_!CSKJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png 1272w, https://substackcdn.com/image/fetch/$s_!CSKJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CSKJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png" width="1456" height="345" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:345,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:121361,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.chipstrat.com/i/213018303?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CSKJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png 424w, https://substackcdn.com/image/fetch/$s_!CSKJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png 848w, https://substackcdn.com/image/fetch/$s_!CSKJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png 1272w, https://substackcdn.com/image/fetch/$s_!CSKJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18603330-a430-45de-a701-9e5b2aa6094e_2610x618.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></div><figcaption class="image-caption"><a href="https://opensource.org/ai/open-weights">Source</a></figcaption></figure></div><p>Why are open models what diffusion needs?</p><p>First, open weights means the <strong>average engineer across any industry can tinker with it</strong>. Quantize it to fit the hardware you use. Prune it. Fine-tune it on your proprietary data. Want more control over costs and performance? Host it yourself, choose batch size vs interactivity tradeoffs, tier the KV cache to improve the hit rate, and so on. All the while keeping your data private, with full transparency into the methods powering your models.</p><p>Don&#8217;t want to host it? Then don&#8217;t. The weights are portable, so you rent one host this quarter and a cheaper one next quarter. <em>Free market. </em></p><p>Second, open models spur competition on price. A handful of closed vendors who all monetize tokens have little reason to race each other toward pennies-per-task. But diffusion needs lower prices. <em>Jevon&#8217;s paradox, right?</em> Open models mean many hosts can serve the same model, pressuring token prices down toward the cost of compute.</p><p>One more. Open models help companies retain and make use of their proprietary data. <em>There&#8217;s the obvious sovereignty-and-regulated-industries angle, but that&#8217;s well-trodden.</em> What&#8217;s also interesting here is having more control over the feedback loop: fine-tune a model, use it, generate more labeled data as you use it, feed that back into the post-training, and so on. Easiest and fastest when you&#8217;re in full control.</p><h2><strong>Diffusion already happening, but can be sped up with competitive frontier open model</strong></h2><p>Diffusion of GenAI across American industries is already happening, thanks in part to open models.</p><p>It&#8217;s there in Nvidia&#8217;s numbers, although a bit fuzzy. From the <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581026000073/q2fy27cfocommentary.htm">Q2 FY27 CFO commentary</a> earlier this week, ACIE (AI clouds, industrial, and enterprise) was at about 45% of data center revenue, growing roughly twice as fast as hyperscale. </p><p>Of course, ACIE includes AI clouds (neoclouds), which mostly serve the model labs.  So ACIE alone isn&#8217;t a direct measure of pure enterprise diffusion, but it&#8217;s in there.</p><p>I turn to Dell for enterprise read-through. On the Q4 FY26 call, Dell said </p><blockquote><p><em>&#8220;the <strong>enterprise</strong> portion of our five-quarter pipeline grew and actually <strong>was the fastest-growing portion of the five-quarter pipeline</strong>.&#8221;</em></p></blockquote><p>Yes, it&#8217;s easier to grow small numbers. But the Q1 FY27 call said the AI customer count &#8220;surpassed five thousand&#8221; and up over 50% in six months. </p><p>And note what those customers buy. Dell&#8217;s mix is shifting to &#8220;x86 + Blackwell... driven primarily by enterprise deployment, air cooling being the number one consideration&#8221;. </p><p><em>Dell is referring to products like the air-cooled <a href="https://www.dell.com/en-us/shop/ipovw/poweredge-xe9780">PowerEdge server</a> on HGX B200/B300 baseboards with x86 CPUs, not the Arm-based GB200/GB300 NVL72 racks hyperscalers buy. One can argue this is a &#8220;diffusion SKU&#8221;.</em></p><p>Yet to accelerate the diffusion of GenAI across healthcare, insurance, agriculture, retail, manufacturing, defense, and so on, we need to close the gap between proprietary, frontier-leading models and American open-source models.</p><p>I previously thought that a frontier-capable open model was simply not possible given the amount of recurring investment needed to stay at the frontier. Nvidia is one of the few companies with that kind of money, and I assumed Nvidia was incentivized to train only &#8220;big enough&#8221; Nemotron models to test their hardware end-to-end. <em>i.e. <a href="https://en.wikipedia.org/wiki/Eating_your_own_dog_food">eating your own dogfood</a>.</em> </p><p>Catanzaro <a href="https://www.youtube.com/watch?v=_y9SEtn1lU8">said as much</a>:</p><blockquote><p><em>&#8220;Nemotron is not just a model, it&#8217;s all of our AI technology that helps us build the GPUs and systems for the future. There&#8217;s a lot of questions about how networking or the GPU should work in order to make AI progress, and we are able to answer those questions better because of building Nemotron&#8221;.</em></p></blockquote><p>Moreover, on the business strategy side, wouldn&#8217;t a competitive open frontier model compete with the closed labs, the biggest consumers of Nvidia compute?</p><p>Well, Nvidia is buying Poolside, and I&#8217;ve changed my mind. I think Jensen wants to build competitive frontier open models. <em>And it&#8217;s a good thing.</em></p><p>Sure, Jensen told us analysts at GTC in March something to the effect of, &#8220;we don&#8217;t have to be absolutely the world&#8217;s best model, but we have to be near it&#8221;. But with Poolside and <a href="https://www.theinformation.com/articles/nvidia-agrees-buy-open-source-model-repository-hugging-face-12-9-billion?rc=ij60ts">the rumored Hugging Face acquisition</a>,  Nvidia is spending ~$20B (<em>allegedly</em>) on the talent and distribution needed to succeed. And that $20B will pale in the long run compared to what Nvidia will spend on compute to train models. </p><p>You don&#8217;t spend that kind of money if you aren&#8217;t trying to compete. Jensen Huang is a winner. I think he now wants to be as close to the frontier as possible, especially on the jagged edges that matter, like coding. </p><p>My read on the acquisition of Poolside is that it gives Nvidia the iteration speed needed to catch up. </p><p>Yes, via talent, but also via a bunch of internal tools and IP. According to <a href="https://poolside.ai/blog/introducing-the-model-factory">Poolside&#8217;s blogs</a>, training experiments that &#8220;used to take weeks to organize and schedule are now handled by the orchestrator in under an hour, often as few as ten minutes&#8221;. Batch size and learning rate sweeps run in &#8220;fewer than ten minutes, as opposed to days&#8221;. Deduplication went from &#8220;over a week&#8217;s worth of engineering time to schedule and run&#8221; to &#8220;the click of a button&#8221;. </p><p>Jensen is adding speed to his team through tools and people.</p><p><em>So this is somewhat like Meta buying rockstars for Meta Superintelligence Lab, except Poolside is 100+ people that already work well together and have already built a bunch of tooling.</em></p><p>This makes sense for Nvidia and for the diffusion of agentic AI across America and the world. </p><h2><strong>Why this is a good thing</strong></h2><p>Competitive frontier open models mean GenAI diffuses more quickly into every industry. </p><p>Open models enable fungibility, which unlocks competition. Competition and portability give the end customer control. Control unlocks innovation and cost efficiency.</p><p>American and global productivity and GDP benefit accordingly.</p><p>Yes, this aligns with Nvidia&#8217;s business model. <em>As it should! </em>A lower token price means more tokens. Cheaper tokens mean more demand. More demand means more compute. Open models pulling forward diffusion also pulls forward Nvidia revenue.</p><p><em>Give the model away, monetize the hardware.</em> </p><p>This can be done. Existence proof one is autonomous driving with <a href="https://www.nvidia.com/en-us/solutions/autonomous-vehicles/alpamayo/">Alpamayo</a>, a family of open VLA models. Nvidia monetizes via <a href="https://blogs.nvidia.com/blog/three-computers-robotics/">three computers</a>: &#8220;training (NVIDIA DGX), simulation and validation (NVIDIA Omniverse and NVIDIA Cosmos on NVIDIA RTX), and in-vehicle computing (NVIDIA DRIVE AGX)&#8221;.</p><p>Existence proof two, robotics. <a href="https://www.nvidia.com/en-us/ai/cosmos/">Cosmos</a> world foundation model weights are open too. Monetize via same three computers.</p><p><a href="https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/">Nemotron</a> + enterprise AI looks similar. Nvidia trains the model, releases the weights, and sells the hardware needed for fine-tuning and deployment. <em>DGX Station and DGX Spark on the desk if you want, too. An important part of enterprise GenAI tinkering will be on local compute.</em></p><p>Zooming out, competitive open frontier models also lead to a healthier market structure. Dylan Patel on the <a href="https://www.dwarkesh.com/p/dylan-patel-3">Dwarkesh Podcast</a> recently said &#8220;Anthropic and OpenAI are taking as much as 40% to 50% of compute next year&#8221;. <strong>Is a world where two labs buy half the world&#8217;s compute the healthiest outcome?</strong> Or do we want the long tail of enterprises buying compute or renting from neoclouds and CSPs, so one lab missing on a quarter doesn&#8217;t punish the entire AI infrastructure ecosystem? <em>Like a local or state government putting the brakes on a datacenter buildout...</em></p><p>Put differently, more compute buyers is good for anyone selling optics, copper, power, cooling, and memory. <em>And everyone investing in it.</em></p><p>Now the big question: is this in direct competition with Nvidia&#8217;s biggest end-customers, OpenAI and Anthropic? Yes and no. <em>It&#8217;s nuanced.</em> </p><p>We have to ask on what plane of competition are we talking about frontier? <em>It&#8217;s jagged after all, right?</em> Having open models catch up at different points of different jagged spikes doesn&#8217;t mean death to Anthropic and OpenAI, and in fact it may help them focus on which spikes matter. </p><p>Whatever happened to Sora? Oh yeah, Codex. Competition pushed OAI to be better, and that&#8217;s a good thing. More competition is good.</p><p>This is technology expansion. It&#8217;s not zero-sum. We can have a healthy OAI and Anthropic, and competitive frontier open models too.</p><p></p>]]></content:encoded></item><item><title><![CDATA[NPO State of the Union]]></title><description><![CDATA[When NPO? Who NPO? Where NPO? Why NPO? Winners, losers, and more.]]></description><link>https://newsletter.chipstrat.com/p/npo-state-of-the-union</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/npo-state-of-the-union</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Fri, 21 Aug 2026 13:37:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!d965!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The scale-up optics trade used to focus on the question &#8220;when does CPO displace copper?&#8221; But the <a href="https://www.youtube.com/watch?v=90WD_ats6eE">times they are a-changin&#8217;</a>, and this earnings season the industry is clearly articulating NPO as an intermediate step to CPO.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!d965!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!d965!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png 424w, https://substackcdn.com/image/fetch/$s_!d965!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png 848w, https://substackcdn.com/image/fetch/$s_!d965!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png 1272w, https://substackcdn.com/image/fetch/$s_!d965!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!d965!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png" width="458" height="329.0302197802198" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1046,&quot;width&quot;:1456,&quot;resizeWidth&quot;:458,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;NPO State of the Union&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="NPO State of the Union" title="NPO State of the Union" srcset="https://substackcdn.com/image/fetch/$s_!d965!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png 424w, https://substackcdn.com/image/fetch/$s_!d965!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png 848w, https://substackcdn.com/image/fetch/$s_!d965!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png 1272w, https://substackcdn.com/image/fetch/$s_!d965!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22d2e059-e759-446f-ac92-f626151a9e52_2480x1782.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>It&#8217;s coming up a lot on earnings calls. From GlobalFoundries:</p><blockquote><p><strong>Karl Ackerman, BNP Paribas:</strong> For my follow-up, if I may, could you discuss what portion of those 7 customers on your SCALE platform are working on NPO or near-packaged optics -- and I guess how should we think about the timing of your NPO opportunity?</p><p><strong>Tim Breen, GF CEO</strong>: Yes. Maybe just to take a step back, and I think there&#8217;s obviously a year ago, the industry wasn&#8217;t talking a lot about NPO -- now it&#8217;s talking a lot about it.</p></blockquote><p><em>Let me remind newer folks quick about NPO.</em></p><p>Today&#8217;s <a href="https://www.chipstrat.com/p/optics-primer-part-1-traditional">pluggable transceiver</a> is the industry&#8217;s workhorse, but it burns a lot of power. The module connects at the switch faceplate (far from the switch chip), so a long copper trace on the board is needed. <a href="https://www.chipstrat.com/p/credo-aecs">High-frequency signals over copper</a> degrade, so a power-hungry digital signal processor (DSP) sits in the transceiver, cleaning up and retiming signals in both directions, transmit and receive.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!g5vo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!g5vo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png 424w, https://substackcdn.com/image/fetch/$s_!g5vo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png 848w, https://substackcdn.com/image/fetch/$s_!g5vo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png 1272w, https://substackcdn.com/image/fetch/$s_!g5vo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!g5vo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png" width="426" height="373.07272727272726" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1156,&quot;width&quot;:1320,&quot;resizeWidth&quot;:426,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!g5vo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png 424w, https://substackcdn.com/image/fetch/$s_!g5vo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png 848w, https://substackcdn.com/image/fetch/$s_!g5vo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png 1272w, https://substackcdn.com/image/fetch/$s_!g5vo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa165aa47-5320-4e84-adb8-b5328864e2fd_1320x1156.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><a href="https://www.chipstrat.com/p/linear-optics-trade-offs-lro-and">Source</a></figcaption></figure></div><p><a href="https://www.chipstrat.com/p/linear-optics-trade-offs-lro-and">Linear optics</a> such as LRO above give that cleanup work to the switch ASIC&#8217;s own <a href="https://www.chipstrat.com/p/serdes-matters">SerDes</a>. The transceiver can shed the DSP silicon (LRO sheds the receive half; LPO sheds all of it), and the power spent per transmitted bit drops.</p><p><a href="https://www.chipstrat.com/p/optics-primer-part-3-co-packaged">Co-packaged optics</a> take this to the logical conclusion and move the optical engine onto the chip&#8217;s own package.</p><p><strong>NPO is the middle step</strong>. The optical engine moves onto the board a few centimeters from the chip, into a socket, but remains outside the package:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OjOM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OjOM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png 424w, https://substackcdn.com/image/fetch/$s_!OjOM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png 848w, https://substackcdn.com/image/fetch/$s_!OjOM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png 1272w, https://substackcdn.com/image/fetch/$s_!OjOM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OjOM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png" width="1456" height="849" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:849,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The efficiency ladder&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The efficiency ladder" title="The efficiency ladder" srcset="https://substackcdn.com/image/fetch/$s_!OjOM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png 424w, https://substackcdn.com/image/fetch/$s_!OjOM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png 848w, https://substackcdn.com/image/fetch/$s_!OjOM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png 1272w, https://substackcdn.com/image/fetch/$s_!OjOM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a49090d-0eab-4505-9ac0-68d9be2ff536_2400x1400.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The closer the optical engine sits to the chip, the less energy each bit needs.</em></figcaption></figure></div><p>A nice two min video about this from my chat with Tom Barber:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/semidoped/status/2089468682019749935]&quot;,&quot;full_text&quot;:&quot;Why are data centers moving from pluggables to co-packaged optics (CPO)? It's a ladder of efficiency.\n\n&#8226; Pluggable: 20-25 pJ/bit (35dB loss)\n&#8226; NPO: ~10 pJ/bit (15-20dB loss)\n&#8226; CPO: &amp;lt;5 pJ/bit (&amp;lt;6dB loss)\n\nGlobalFoundries' Thomas Barber on the physics driving the change. &quot;,&quot;username&quot;:&quot;semidoped&quot;,&quot;name&quot;:&quot;Semi Doped&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2078487165185486848/S73FzrFG_normal.jpg&quot;,&quot;date&quot;:&quot;2026-08-17T21:45:28.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!mxaH!,w_1028,c_limit,f_auto,q_auto:best,fl_progressive:steep/l_play_button_usfui2,w_88,e_colorize:0/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F__ss-rehost__tw-video-preview-13_2089468591611555840.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/E5FL7JQTCO&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:1,&quot;retweet_count&quot;:2,&quot;like_count&quot;:18,&quot;impression_count&quot;:3609,&quot;expanded_url&quot;:null,&quot;video_url&quot;:&quot;https://video.twimg.com/amplify_video/2089468591611555840/vid/avc1/1280x720/oSQEkS00sRM1pBoA.mp4&quot;,&quot;video_preview_media_key&quot;:&quot;13_2089468591611555840&quot;,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>So <strong>&#8220;What is NPO?&#8221;</strong> is one of the <a href="https://comm.gatech.edu/resources/writers/5ws">five Ws</a>, and now you might now be wondering about the other four Ws:</p><ul><li><p><strong>WHEN NPO?</strong></p></li><li><p><strong>WHERE NPO?</strong></p></li><li><p><strong>WHO NPO?</strong></p></li><li><p><strong>WHY NPO?</strong></p></li></ul><p><em>Well, more context.</em></p><p>Remember that the CPO switches in the news like Nvidia&#8217;s Quantum-X and Broadcom&#8217;s Davisson are CPO used in <em>scale-out</em>. And a quick reminder that scale-out doesn&#8217;t simply mean &#8220;rack-to-rack&#8221; and scale-up doesn&#8217;t simply mean &#8220;within a rack&#8221; as lots of folks tend to say online. </p><p>GPUs/XPUs/AI ASICs want to access each other&#8217;s memory so fast that the remote memory feels like their own. Accelerators on the same &#8220;scale-up domain&#8221; can do just that. <em>They wish they had more HBM, but for now it&#8217;s just you get whatever is attached to your board and that&#8217;s it, so the hack is to access your neighbors as if it was your own. This landscape will change with HBM pooling btw, but that&#8217;s a topic for another day</em>. </p><p>If accelerators aren&#8217;t on the same scale-up domain, then they communicate over the <em>scale-out</em> network.</p><p>Scale-up has historically lived within a server node first (<em>e.g. just 8 GPUs in a scale-up domain</em>) then within a rack (<em>72 GPU domain</em>). But the scale-up domain can span neighboring racks too! <a href="https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/">NVL576</a> connects eight racks of 72 Rubin Ultra GPUs into a single NVLink domain. We even see Kyber rack NVL1152 taking it to eight racks of 144 GPUs, with <em>optical</em> links between racks.</p><p>So the first CPO products are <em>switches</em> on the scale-out network. <em>Not GPUs.</em> The optical engine moves onto the <em>switch</em> substrate, and nothing changes on the GPU side; the endpoints still run pluggables and copper. <em>This is just like the images above where the O/E is moving closer to switch ASIC</em>.</p><p><em>OK, with that in mind, back to what the industry is saying about when, where, who and why.</em></p><p>We can see a shift in language from various industry participants; for example, back in February, Macom CEO Stephen Daly positioned Macom&#8217;s products as &#8220;optimized for co-packaged and highly integrated architectures, like CPO and NPO&#8221;. <em>Ah ok, both NPO and CPO are a-comin&#8217;</em></p><p>In May the prepared remarks carried the same script line in different words, and it was the only CPO mention in the call. By August, the line read &#8220;highly integrated architectures like NPO and XPO&#8221;. <em>What? Where is CPO? And <a href="https://x.com/semidoped/status/2079320561264607543">XPO</a>?! And that call also mentioned NPO nine times.</em> Macom now talks about NPO and XPO, not CPO.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bfSC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bfSC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png 424w, https://substackcdn.com/image/fetch/$s_!bfSC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png 848w, https://substackcdn.com/image/fetch/$s_!bfSC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png 1272w, https://substackcdn.com/image/fetch/$s_!bfSC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bfSC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png" width="406" height="533.7115384615385" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1914,&quot;width&quot;:1456,&quot;resizeWidth&quot;:406,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Who says NPO, who says CPO&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Who says NPO, who says CPO" title="Who says NPO, who says CPO" srcset="https://substackcdn.com/image/fetch/$s_!bfSC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png 424w, https://substackcdn.com/image/fetch/$s_!bfSC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png 848w, https://substackcdn.com/image/fetch/$s_!bfSC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png 1272w, https://substackcdn.com/image/fetch/$s_!bfSC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ef1340f-5be7-47dc-ac63-2e023e2bd4a3_2160x2840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This raises follow-on questions.</p><ul><li><p><strong>When is NPO ramping?</strong> And why didn&#8217;t this NPO-before-CPO come up sooner?</p></li><li><p><strong>Does NPO push CPO back entirely, or only in certain sockets?</strong></p></li><li><p><strong>Who benefits from NPO?</strong> <strong>Who doesn&#8217;t?</strong></p></li><li><p><strong>How much of the &#8220;CPO&#8221; shipping today is actually NPO?</strong> Is marketing at hand here?</p></li><li><p><strong>Is NPO accretive or cannibalistic?</strong> For whom?</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QI74!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QI74!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png 424w, https://substackcdn.com/image/fetch/$s_!QI74!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png 848w, https://substackcdn.com/image/fetch/$s_!QI74!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png 1272w, https://substackcdn.com/image/fetch/$s_!QI74!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QI74!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png" width="1456" height="592" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:592,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Winners and losers in NPO and CPO&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Winners and losers in NPO and CPO" title="Winners and losers in NPO and CPO" srcset="https://substackcdn.com/image/fetch/$s_!QI74!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png 424w, https://substackcdn.com/image/fetch/$s_!QI74!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png 848w, https://substackcdn.com/image/fetch/$s_!QI74!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png 1272w, https://substackcdn.com/image/fetch/$s_!QI74!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b20baef-8589-4d75-81f4-71f445c11fcc_2560x1040.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We&#8217;ll answer those below for subscribers.</p><p>We&#8217;ll also unpack the related MSAs and what they tell us.</p><p>We&#8217;ll touch on many companies including <strong>GlobalFoundries, TSMC, Tower, Lumentum, Coherent, AAOI, Macom, Semtech, Credo, Astera Labs, Marvell, Nvidia, AMD, Broadcom, Arista, Ciena, ASE, Celestica, Fabrinet, InnoLight, Eoptolink, HG Genuine.</strong></p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/npo-state-of-the-union">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Build Your Own DRAM Supply Model]]></title><description><![CDATA[Wafer starts, die size, yield, HBM mix, node migration impacts, etc. Simple enough to follow, but detailed enough to learn from.]]></description><link>https://newsletter.chipstrat.com/p/build-your-own-dram-supply-model</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/build-your-own-dram-supply-model</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Fri, 14 Aug 2026 23:11:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Iggk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve been thinking about the pros and cons of long-term agreements in the memory market, for both customers and suppliers. As groundwork for explaining the implications, I thought it&#8217;d help to build toy supply and demand DRAM models. </p><p>Then we can ask whether LTAs change how we model demand, how we think about supply and price, and whether the incentives to bring new capacity online sooner change.</p><p>So two overarching questions we&#8217;ll look to address</p><ol><li><p>How is today&#8217;s memory supply responding to demand, and on what timeline?</p></li><li><p>Do incentives like LTAs change that in any direction?</p></li></ol><p>This article will begin by building our own simple supply model of the memory market. Not aiming for the best forecast out there, rather a simple enough one to be useful and teach us some important details along the way.</p><p>So a basic tutorial to get you started. <em>Teach a man to fish...</em></p><p>Let&#8217;s model supply. <em>In subsequent articles we&#8217;ll look at demand.</em></p><p>Today we&#8217;ll address</p><ul><li><p><strong>Wafer starts and total bits</strong></p></li><li><p><strong>Cell size and process nodes</strong></p></li><li><p><strong>Node generations</strong></p></li><li><p><strong>Node adoption across the industry</strong></p></li><li><p><strong>HBM impact on bit density</strong></p></li><li><p><strong>Industry totals</strong></p><ul><li><p>Checking the number two ways</p></li></ul></li><li><p><strong>Supply growth model</strong></p><ul><li><p>Node migration puts and takes</p></li><li><p>HBM wafer absorption</p></li><li><p>The conventional mix die-size penalty</p></li></ul></li><li><p><strong>Capacity coming online</strong></p></li><li><p><strong>Checking the toy model</strong></p></li></ul><p>Let&#8217;s dig in.</p><p>One might start with the market-level question &#8220;how many bits the industry will produce?&#8221; </p><p>We can estimate this on a company-by-company basis. If you know the number of wafer starts per month, and the die size, you can figure out how many die per wafer and how many gigabytes of DRAM per die. Then you can make assumptions about yield to get the number of good die per wafer.</p><p><code>bits per wafer = usable wafer area &#247; die area &#215; yield &#215; bits per die</code></p><p>Or more directly, if you have the bit density:</p><p><code>bits per wafer = usable wafer area &#215; bit density &#215; yield</code></p><p>For reference, <a href="https://www.techinsights.com/blog/memory/micron-1a-dram-technology">TechInsights measured</a> Micron&#8217;s 1&#945; DRAM at a die size of 25.41 mm&#178; for an 8Gb DDR4 part with a bit density of 0.315 Gb/mm&#178; and a cell size of 1,672 nm&#178;.</p><p>Plugging that in,</p><ul><li><p>300mm wafer = &#960; &#215; 150&#178; = 70,686 mm&#178;</p></li><li><p>With a 3mm edge exclusion, and subtracting the partial die that fall off the edge, about 63,800 mm&#178; is usable</p></li><li><p>63,800 &#247; 25.41 = <strong>2,511 total die</strong></p></li><li><p>At 85% yield, <strong>2,134 good die</strong></p></li><li><p>2,134 &#215; 8 Gb &#247; 8 = 2,134 GB, so <strong>about 2.1 TB per wafer</strong> at 1&#945;</p></li></ul><p>2.1 TB per wafer at the 1&#945; node. <em>Yes, one alpha is actually its name</em>.</p><div class="captioned-image-container"><figure><div class="image-link image2" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b4bL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png 424w, https://substackcdn.com/image/fetch/$s_!b4bL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png 848w, https://substackcdn.com/image/fetch/$s_!b4bL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png 1272w, https://substackcdn.com/image/fetch/$s_!b4bL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b4bL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png" width="1456" height="292" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:292,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!b4bL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png 424w, https://substackcdn.com/image/fetch/$s_!b4bL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png 848w, https://substackcdn.com/image/fetch/$s_!b4bL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png 1272w, https://substackcdn.com/image/fetch/$s_!b4bL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84090c80-d7d9-40b4-a1d7-5b41f3960441_3708x744.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></div><figcaption class="image-caption">Node names for memory are different than logic&#8230; Source: TechInsights</figcaption></figure></div><p>Note that not every company can squeeze the same amount of bits out of a given area. It depends on the process node that the company and the fab are on. A DRAM bit consists of one transistor and one capacitor. The area it takes on the die is roughly 6F&#178;, where F is the feature size set by the process node.</p><p>With each process node, the transistor and capacitor shrink linearly. Because the cell is 6F&#178;, the area is reduced by the square of the shrink. <em>i.e., going from say 16.5nm to 11.5nm is a 1.43x reduction in feature size and a 2.06x reduction in cell area. So process shrinks do matter.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Iggk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Iggk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png 424w, https://substackcdn.com/image/fetch/$s_!Iggk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png 848w, https://substackcdn.com/image/fetch/$s_!Iggk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png 1272w, https://substackcdn.com/image/fetch/$s_!Iggk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Iggk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png" width="1456" height="665" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:665,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!Iggk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png 424w, https://substackcdn.com/image/fetch/$s_!Iggk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png 848w, https://substackcdn.com/image/fetch/$s_!Iggk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png 1272w, https://substackcdn.com/image/fetch/$s_!Iggk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a7310a-fb7a-4daa-8620-085eb21cbefd_1720x786.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The Greek letters make it obvious that memory process nodes are different than logic nodes. But note that Micron uses Greek letters, while Samsung and SK hynix use Latin letters. Hence 1-gamma and 1c are the same generation. Micron explains why they switched:</p><blockquote><p>&#8220;We started with 1x, but as we continued to shrink and name the next nodes, we hit the end of the Roman alphabet. That&#8217;s why we switched to the Greek alphabet alpha, beta, gamma and so on.&#8221; &#8212; <a href="https://www.micron.com/about/blog/memory/dram/inside-1a-the-worlds-most-advanced-dram-process-technology">Micron</a></p></blockquote><p>Logic nodes are just marketing today, as there are no transistor dimensions that are actually 2 nanometers or 18 Angstroms. But DRAM generations do somewhat attempt to tie back to an actual dimension, the &#8220;half-pitch&#8221;.</p><p>SK hynix sometimes writes it &#8220;1cnm&#8221;, &#8220;1c-nanometer&#8221; and &#8220;1c nm&#8221;. <em>I don&#8217;t know why they include the nanometers bit.</em> The most explicit name is &#8220;the sixth-generation of the 10-nanometer technology&#8221;.</p><p>Here&#8217;s the memory process size and naming by generation:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SbOj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SbOj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png 424w, https://substackcdn.com/image/fetch/$s_!SbOj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png 848w, https://substackcdn.com/image/fetch/$s_!SbOj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png 1272w, https://substackcdn.com/image/fetch/$s_!SbOj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SbOj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png" width="520" height="321.218836565097" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:892,&quot;width&quot;:1444,&quot;resizeWidth&quot;:520,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SbOj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png 424w, https://substackcdn.com/image/fetch/$s_!SbOj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png 848w, https://substackcdn.com/image/fetch/$s_!SbOj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png 1272w, https://substackcdn.com/image/fetch/$s_!SbOj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7df929d-63c8-4bea-b261-8ea049201bf6_1444x892.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>1&#947; is the leading edge in production, two generations on from the 1&#945; we used in the example above. Bit density is roughly 0.55 Gb/mm&#178;, so a wafer capacity is ~3.7 TB. That&#8217;s a lot, and is obviously more than the 2.1 TB per wafer we estimated above. <em>So process shrinks matter.</em></p><p>The big three are roughly on the same roadmap, but they don&#8217;t get there at the same time. <em>Sort of like TSMC vs Intel vs Samsung Foundry for 18A and 14A logic nodes.</em></p><p>CXMT is behind on process nodes. Nomura&#8217;s CXMT initiation puts its mainstream node at 1x-1y, four generations back from the 1&#947;/1c leading edge. <em>This is an important detail! You can&#8217;t just compare wafer starts per month.</em></p><p>Anyway, we&#8217;re starting to build some understanding. We can work out the bits on one wafer. <em>2.1 TB at 1&#945;, 3.7 TB at 1&#947;.</em></p><p>What else do we need? Well, we aren&#8217;t accounting for HBM and how it decreases bits per wafer. Then we need to blend the process mix (not all fabs are on the same process), multiply by each producer&#8217;s wafer starts, and compare this total of our bottoms-up approach against industry estimates.</p><p>Note that gives us a good understanding of <em>today&#8217;s</em> supply, but doesn&#8217;t tell us about <em>future </em>supply. And recall we want to understand whether supply can respond to the increase in demand, and how LTAs impact that. </p><p>For that, we need to account for growth: what&#8217;s getting built, when does it actually produce bits, and any gains/losses from node transitions and HBM mix along the way. <em>All of that is below.</em></p><p>And we have to check all that work and see if we&#8217;re on the right track with industry estimates. If not, we should think about why.</p><p>Oh yeah, and we&#8217;ll pencil out how far CXMT is behind and whether it matters.</p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/build-your-own-dram-supply-model">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[AMD on Agentic CPUs, and Why There's No One Size Fits All]]></title><description><![CDATA[Listen now | Madhu Rangarajan, Corporate VP at AMD, on why the agent era is a CPU buildout, and why the workloads under the "agentic rack" label are too varied for one SKU]]></description><link>https://newsletter.chipstrat.com/p/amd-on-agentic-cpus-and-why-theres</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/amd-on-agentic-cpus-and-why-theres</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Wed, 12 Aug 2026 13:03:11 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/210824040/19b83e503d080d79dcba17dede3b6ecf.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><em>Madhu Rangarajan spent 26 years in servers, at Dell, then Intel, then Ampere. Now he is Corporate VP at AMD for Compute and Enterprise AI Products, which covers the EPYC CPUs and every GPU that isn&#8217;t rack-scale. I sat down with him for the latest Bit by Bit conversation.</em></p><p><em>His main argument: there is no one-size-fits-all for agentic infrastructure. The workloads under that label are too different from each other to land on one machine.</em></p><p>The industry has spent this year announcing agentic racks. Madhu thinks the label hides how different those machines actually are. Some of what gets sold as one is <em>&#8220;just dense liquid-cooled racks that take advantage of what was provisioned in the data center for GPUs already&#8221;. </em>But a gateway, a RAG database, a storage node, and a sandbox server each need different memory, storage, and CPU.</p><p>What is AMD&#8217;s agentic Venice part is tuned for? With 256 cores at roughly 400 watts, his framing for it is <strong>&#8220;maximizes threads per megawatt&#8221;</strong>. When your power constraint is a fixed amount of MWs, threads is the number that decides how much agent capacity you get.</p><p>The takeaways:</p><ul><li><p><strong>Agents turn tokens into CPU work.</strong> Tool calls, sandboxes, verification, then a loop back around when the answer isn&#8217;t good enough. The GPU produces the tokens. The CPU runs the code those tokens describe.</p></li><li><p><strong>CPU to GPU has reached about one to one</strong>, and Madhu expects it to keep shifting toward CPUs. Every gain in GPU token output creates more tool calls and more sandbox runs, and those land on CPUs.</p></li><li><p><strong>AMD now cuts the server market three ways.</strong> Host nodes feeding GPUs, general purpose, and a new tier for agent orchestration and sandboxes.</p></li><li><p><strong>Concurrency beats per-thread speed, for now.</strong> Coding agents are coarse-grained and want every thread they can get. <em>&#8220;What&#8217;s happening today with agents may not be what&#8217;s happening in two years.&#8221;</em></p></li><li><p><strong>Operators deploy as few SKUs as they can.</strong> Every extra SKU is another load-balancing layer and more stranded capacity. One trick is to buy the 256-core part and run 128 cores for the turbo headroom if needed.</p></li><li><p><strong>To date, agents mostly don&#8217;t run beside the GPUs.</strong> They run on laptops, on-prem boxes, or general-purpose scale-out, decoupled from the nodes serving the LLM.</p></li><li><p><strong>Tokenomics is a routing problem.</strong> Not every task needs the PhD-grade model when an intern-grade one will do. AMD&#8217;s own IT is testing whether an MI350P on-prem can cut token spend by around 40%.</p></li><li><p><strong>Openness wins the orchestration layer eventually.</strong> Bespoke now, standardized later, with Kubernetes as the template.</p></li></ul><p><em>Full transcript below (lightly edited for clarity).</em></p><p><strong>All right, hello everyone. We have a special guest today, Madhu Rangarajan from AMD. Welcome, Madhu. Tell us about your role at AMD.</strong></p><p><strong>MR:</strong> Hi. So, I&#8217;ve been at AMD for about three years now, and I run the compute and enterprise AI product team. So, that&#8217;s all of the EPYC CPUs and all of the GPUs that are not rack scale. So, think the PCIe card, think GPUs that are eight-way and so on. And I&#8217;ve been in the server industry for about 26 years now. Prior to this, I was at Ampere, prior to that at Intel, prior to that at Dell, predominantly in engineering and architecture roles, but more of a product role for the last few years.</p><p><strong>Okay, awesome, interesting. So, perfect. You&#8217;re the perfect guest today because I want to talk CPUs for agentic AI, and then I also want to hit on enterprise AI, especially like sub rack scale. So, we&#8217;ll try to get to those both. So, let&#8217;s get into the fun stuff. So, let&#8217;s talk about CPUs. So, like maybe to lay the groundwork for listeners, at advancing AI recently, Lisa&#8217;s kind of called out these three different workloads in to explain sort of how the server CPU market is shaking out. Can you walk us through those and maybe tell us like why AMD sees the world this particular way?</strong></p><p><strong>MR:</strong> Yeah. So, stepping back a little bit, right? The transition from the chatbot era to the agentic era, basically drives a lot more CPUs because instead of being a chatbot just answering questions that you ask it, you now have an entire workflow where based on the token outputs from the GPUs, the real work is being done, whether it&#8217;s coding agents or any other sort of agents. And what that&#8217;s driving is a much more complex end-to-end workflow that involves a lot of CPU. And this is where I think we have been decomposing this workflow all the way from there&#8217;s a gateway. There&#8217;s a orchestration step. There is rag in order to inject enterprise data into your results. There&#8217;s the reasoning itself, which has been the interesting part for the last few years and all the attention has been there, but now making the reasoning do some useful work is driving CPU demand again. And after the reasoning, you&#8217;re going to have tool use because, well, at the end of the day, you want to do real work, and that means you have to run some tools and the tools tend to run on the CPUs. Then you have to do some verification to see, hey, did the agent do what I asked it to correctly? And then you loop around until the verification model says, yeah, that&#8217;s satisfactory. And then a user gets a response saying, okay, the agent did its thing. And at that point, the user might go, well, I don&#8217;t like the answer. Can you go do it again and this time fix this? So, you now have a series of steps that are CPU bound and some that are GPU bound running in a loop, and that&#8217;s what&#8217;s driving this crazy CPU demand we are seeing in the market right now. And the workload itself is not one size fits all. A lot of this workload is very familiar. All of the things that we know and love from the past like databases and CRM and ERP and all of the stuff that agents calling those workloads instead of humans calling those workloads. And in addition to that, you have some new emerging tiers like agentic orchestration and agentic tools running, where you now have agents creating code on the fly and executing those at extreme concurrency. So that&#8217;s creating a new tier of servers. So based on all of that analysis, we kind of broke it up into three large buckets, right? There is the host node itself, which is the CPU that connects up to the GPU, where you want a lot of single threaded performance, where you want a lot of memory bandwidth and IO bandwidth. And the goal is to make sure you maximize GPU throughput and the GPU is never left waiting for the CPU. Right? So that&#8217;s one tier of servers that we&#8217;ve been focused on. You&#8217;ve seen us have the Turin high frequency part serving that and we&#8217;re going to continue that with Venice and Verona. The second bucket of CPUs is the tried and true general purpose because agentic drives demand for databases and all and agents may run EDA tools and agents are going to do ERP. So all of the CPUs across enterprises and cloud that have been doing those workloads, you&#8217;re going to need a lot more of them now that agents are trying to make use of those CPUs. And the third bucket is this agentic orchestration and tools, where now you have swarms of agents trying to do a lot of work. So, you have to now think about it at a fleet scale level, where you want to balance how many agents you can run concurrently and how quickly any given agent is able to complete its task. And that&#8217;s what data center operators would be optimizing all day long and that&#8217;s that third tier that we introduced a new lower power Venice for in advancing AI recently.</p><p><strong>Nice. Okay, this is helpful history. And I think like what I&#8217;m hearing you say if we kind of like go back in time was obviously like back in the cloud era, there&#8217;s these general purpose CPUs. You guys had different SKUs there to fit the different workload needs there, whether it&#8217;s a database or an API server, whatever. Then of course, fast forward to the early chatbot days 2023, 2022. All of a sudden there&#8217;s the need for the head node to feed the GPUs. And then of course with reasoning models and now agentic AI, it&#8217;s suddenly like, oh, we need this third bucket, the place where the agent&#8217;s going to live, do work, and a lot of that is CPUs and maybe ping-ponging back and forth with the GPU, but a lot is actually happening on those CPUs.</strong></p><p><strong>MR:</strong> Yeah, for example, right, the you ask it to you feed in a big piece of code and ask the agent to optimize it and make it more performant. You&#8217;re then going to eventually feed it into the GPU and the GPU is going to output a bunch of tokens, which is basically the modified code. And then you want to run that code and test it and see if it needs further modification. Or you could get more sophisticated and say, you know what, output five versions of this code for me, and then I&#8217;m going to run all five versions in a fan out manner and figure out which one works the best and then come back with the response. And all of that&#8217;s going to drive more CPU usage once the tokens are output.</p><p><strong>Sure. Yes, that makes sense. So, okay, then if I&#8217;m thinking about like a one-word kind of priority for these different sockets, if you will, like what&#8217;s the priority for the host node versus the agent server versus general purpose?</strong></p><p><strong>MR:</strong> So, the general purpose, of course, it depends because general purpose is such a broad category, but in general, when you&#8217;re running general purpose at scale, you are trying to hit some kind of latency SLA while maximizing capacity, right? So, you don&#8217;t want to over index on just single threaded performance, but you don&#8217;t want to over index on throughput either because if you over index on throughput, then each individual user may be unhappy because their responses are slower. If you over index on single threaded performance, a single user might be happy, but then you end up with not being able to serve enough users, right? So there is always this fine balance that hyperscalers have been doing for years to get the exact right SKU in their data center, meeting the balancing latency and throughput. Right? So that&#8217;s kind of how I think about the general purpose. Of course, enterprises have different constraints, right? For them, they need to run their workload, but they may have constraints such as, I&#8217;m in a data center with only 10 kilowatts of rack power, or I don&#8217;t have any more space in my IT closet. So there are other constraints that make enterprises pick different SKUs, or there&#8217;s software licensing involved. I have a core-based licensing model, so therefore I may pick a lower core count SKU with a very high frequency. So, I think there isn&#8217;t a single answer in general purpose, but that&#8217;s also the reason why we have a lot of SKUs there. Now going to the host node, And you can see us advancing PCIe generations in the new Helios. You might have seen us have an actual coherent interconnect between the CPU and the GPU. So those are the areas we&#8217;re innovating there. And on the agentic, it&#8217;s going to be a mix and I think anyone that says we know exactly what agents are going to do now, will probably have a surprise because that&#8217;s going to change. But what we&#8217;re largely seeing today with the coding agent is they&#8217;re very coarse grain. They want extreme concurrency. They want as many threads as possible. They aren&#8217;t necessarily valuing per thread performance as much today because they want to run as many tools as possible and then they&#8217;re waiting for an LLM call, which happens in latencies that are pretty large. So, concurrency seems to be the primary metric today for most of the deployments we&#8217;re seeing, but at the same time, there could be more fine-grained agents. There could be some kind of agents and orchestration tiers where you need that more performance. But that&#8217;s the reason we have a portfolio, right? We have parts that have high single threaded performance and parts that have extreme concurrency because what&#8217;s happening today with agents may not be what&#8217;s happening in two years. Today, for example, agentic AI has a lot of Python code running. They create lots of Python code on the fly. That drives a certain kind of compute need. Two years from now, the agents could decide they&#8217;re going to run something else. They could come up with their own language that&#8217;s more efficient, right? And that&#8217;s going to need a different type of compute. So I think the best way to make sure you&#8217;re the best agentic CPU is to make sure you have a portfolio of CPUs that serve a fairly broad set of your usages effectively.</p><p><strong>Sure. Yes, that makes sense. Oh, that&#8217;s helpful. So it feels like general purpose, it depends, but some sort of balance between throughput and latency. The host node, definitely it&#8217;s kind of what I hear you saying is speed. So things like clock speed, IPC, memory bandwidth, IO bandwidth and so on. And then, if what I felt like you were saying right now is for the agent server, it&#8217;s kind of it&#8217;s like about density and concurrency, but maybe, I felt like you were saying have an open mind because maybe they&#8217;ll maybe there&#8217;ll be some bifurcation where some of the workloads may need performance, some may So, do you foresee like a world where there&#8217;s like two different classes of parts for agentic AI?</strong></p><p><strong>MR:</strong> Yeah, I would never say never. I think data center operators try to keep as few SKUs as possible deployed at the data center level. And they usually try to pick the SKU that balances across multiple vectors the best. But if there is demand for any given workload that gets large enough to demand a certain set of CPU characteristics, they absolutely will demand unique will deploy unique infrastructure for that.</p><p><strong>Sure. And I guess, but I so tell me a little bit more like why try to minimize the number of SKUs? Because I&#8217;m thinking out loud like, I guess you would have to orchestrate the different agentic workloads to the different SKUs. So that already is complex.</strong></p><p><strong>MR:</strong> Correct. So because think about it this way, right? You get a let&#8217;s say you get a thousand jobs that come in. If all of the servers you have are the exact same type of server, you just have to figure out which server is free and load balance your workload there. If you now have some servers with more memory and some servers with less memory and maybe your different storage configs, now you have the added layer of, well, this workload needs more storage, so it&#8217;s going to go here. This workload needs more memory, so it needs to go here. And then you add the third layer of, well, now different CPU SKUs optimized for different workloads. That&#8217;s another layer of how you load balance and orchestrate it. And the more you fragment things, the more you risk having stranded capacity. So if you&#8217;re operating at scale, you end up figuring out what&#8217;s the baseline demand for the different types of workloads and you can afford to kind of deploy different infrastructure for those different workloads. But until you can figure out what that baseline is, it&#8217;s a hard choice to make because you may make a bet on the wrong side.</p><p><strong>Mhm. Mhm. I see. Yeah, that&#8217;s very interesting. Yeah, having stranded.</strong></p><p><strong>MR:</strong> And maybe just to add on to that, sorry. Yeah, go ahead. One more example I can think of is, for example, you could buy a 256 core Venice. And if there&#8217;s a workload that needs more memory bandwidth and a little higher single threaded performance, well, you could just keep 128 cores out of the 256 utilized and you can get that dynamic range to get that performance. So that&#8217;s one way to mitigate that is to say, well, I&#8217;ll deploy the higher core count CPUs, but if there are cases where I need to get a little more performance, a little more turbo frequency, that&#8217;s one way to say, you know what? I&#8217;ll modulate the utilization on that server to get more performance per core. So, there are different ways to handle this. Build a super set SKU and modulate the utilization, use two different SKUs. So, there are multiple ways I envision data center operators could is to could handle this situation.</p><p><strong>Yeah, okay, that&#8217;s really interesting. I like the idea of sort of dynamically taking a particular SKU like that has lots of cores and kind of making it flex and behave differently for a different workload. Interesting. Okay, so now tell me, in the life cycle of hyperscalers and kind of the people on the bleeding edge who are needing agentic racks because they are orchestrating lots of agents, like where are we as far as like where are agents actually running? So I assume, and correct me if I&#8217;m wrong, that early on it was like, okay, we&#8217;re just getting into this agentic world. There&#8217;s no such thing as an agentic rack yet. Are they were they running those on the head node or were they actually not so concerned with like the IO latency and just sending them out to a CPU rack somewhere? And then how is that kind of changing? Like where are we today?</strong></p><p><strong>MR:</strong> Yeah, so I think there are a few different models. I think of how I use agents every day for myself. And at work, I use Claude Code a lot. And in that case, my agents are running right on my laptop. And the API calls are being made somewhere in the cloud to run the large frontier models to get tokens, right? So in that model, I got about four or five different agents running on my laptop. All of the API calls to the LLMs are happening in the cloud. So it&#8217;s that&#8217;s one model. The other model could be, well, the agents on my laptop, I can&#8217;t share them. So I want to share them with a bunch of my colleagues because I want that agent to be useful to all of them. So one option is, well, I could run it on a on-prem server in behind the firewall, well access controlled and everything that my colleagues can access. So now you start having those shared agents. Or you can run them on any virtual machine in the cloud anywhere, right? So that&#8217;s the other way to run it. So I think most of the these types of agents tend to run a little decoupled from the LLM inference call itself. On the other hand, I could have like a Strix Halo where I have one box where I&#8217;m doing the agents and I&#8217;m doing the LLM and I&#8217;m doing everything else. So I think multiple models exist. But for at scale deployments today, I&#8217;m predominantly seeing these agents run on largely general purpose servers in a scale out manner, quite decoupled from the LLM, the GPU nodes that are running the LLMs themselves because when you&#8217;re running, The GPU could be serving another LLM inference call and that&#8217;s a much better use of the GPU resources. So, at least my view is I&#8217;d be hard pressed to see why you would run it that way at scale unless there&#8217;s a very specific use case where it&#8217;s extremely tightly coupled.</p><p><strong>Mm, okay. Yes. So, and I think you make a good point, which is there&#8217;s a lot of agentic use case today that&#8217;s happening that&#8217;s running on our laptop CPUs. So it&#8217;s not always even running in a VM on a server CPU. And so when it is, what I feel like I hear you saying is today it&#8217;s probably running on racks that are you mentioned scale out, but these would not be in the GPU data hall on the east-west network. They would be in a traditional data center somewhere.</strong></p><p><strong>MR:</strong> Yes.</p><p><strong>Yeah. Gotcha. Okay. Yeah, that&#8217;s helpful. So then, maybe just extending this train of thought, when we hear about like, oh, so and so&#8217;s going to deploy a gigawatt of AI compute, of course a person can kind of like calculate how many head node CPUs that is. And we&#8217;ve talked about going from four to one ratio in the head node world to now it&#8217;s more closer to one. But how should people think about like, is there kind of like a rule of thumb of like, oh, a gigawatt&#8217;s getting deployed, so how many agentic CPUs that is? And I&#8217;m sure it&#8217;s it obviously is the answer is it depends. But should we think like people are buying agentic CPUs greenfield when they&#8217;re deploying a gigawatt?</strong></p><p><strong>MR:</strong> So, I think when they&#8217;re deploying a gigawatt, they have to think about the predominant use cases and the ratios, right? We&#8217;ve heard multiple ratios in the industry, 20 to 30% on compute nodes, maybe 70% on GPU nodes. We&#8217;ve heard ratios that are more even 50/50. So we&#8217;ve seen numbers that range all over the place. We&#8217;ve done some of our own internal testing with coding agents and how long they spend on the CPU nodes and how long they are actually spending on the LLM GPU nodes. So we got some coarse grain numbers from those as well. Now, some use cases stress the CPUs a lot more, some use cases stress the GPUs a lot more, so there&#8217;s going to be a range here. But generally we are seeing this one is to one becoming much more common based on those power ratios I just spoke about. But what we expect to happen over time is that ratio shifts even more in favor of the CPU because if you just plot what has been happening to let&#8217;s use SPECint_rate as a metric for CPU performance and plot what&#8217;s going on at SPECint_rate say over the last five years. And maybe even project forward what it looks like for the next five years. Project GPU tops and how many tokens GPUs put out over the last five years and project forward for the next five years, you&#8217;ll see that one line looks like this, the GPU line kind of goes significantly higher. And the more tokens that are put out by the GPU, the more CPU resource you need resources you need to do useful work with those tokens. So we actually think one is to one will climb to even more on the CPU side as we move forward in time driven by the sheer capability of the GPUs and all of the software innovations happening on quantization and so on.</p><p><strong>Mm, that&#8217;s fascinating. So now, at the same time, you guys are always introducing like more and more cores per CPU, which maybe can kind of offset that a little bit. Like, but is there room to run with the number of cores that goes on a CPU or is like 256, 512, is that kind of going to be like the max?</strong></p><p><strong>MR:</strong> I have not seen that slowdown. We&#8217;ve always figured out how to put more cores and more importantly feed those cores, right? Because it&#8217;s not important it&#8217;s not enough to just have a lot of cores. You need enough memory bandwidth for those, you need enough fabric bandwidth, you need enough IO and networking. And you can see how quickly memory ecosystems been evolving, DDR5 from 6400 to 8,000 to MRDIMM 12800 and even faster speeds in the future. You can see the PCIe specs evolving fast with Gen 6 now and faster speeds coming soon. So we have managed to figure out how to squeeze in more cores and still feed the cores generation over generation and I envision that continues.</p><p><strong>Okay, fascinating. Okay, so advancing AI, you guys had a lot of like, Venice SKUs that you launched and you were kind of tying it into this these SKUs could be good for host node versus agentic versus general purpose. Like, high level in case anyone didn&#8217;t watch it, like, what should we have in our head when we think AMD head node, do we think Venice HF and what do we think for each socket?</strong></p><p><strong>MR:</strong> Yeah. So, maybe I&#8217;ll talk about both Turin and Venice. Turin&#8217;s out now. Venice is going to be available in volume really soon. If I think head node host CPU, for Turin we had the 64 core 5 GHz high frequency part. And in Venice, we&#8217;ll have a 96 core 5 GHz high frequency part. For general purpose, big range of SKUs, right? Going all the way up to 192 cores, 500 watts in Turin, going up to 256 cores, 600 watts in Venice. And for these agentic tools and sandboxes, we have a power optimized 256 core, which is going to be closer to 400 watts instead of 600 watts. Optimized for sandbox concurrency. And again, the rationale behind that part is you&#8217;re going to have a lot of sandboxes provisioned on a server, but sandboxes also tend to have very spurty, bursty, branchy code running on it. So they&#8217;re not all going to be running like a throughput workload on a server at any given time. So, this 256 cores at around 400 watts maximizes how many threads you can provision in any given data center. And each thread still has the dynamic range to run as fast as possible when a given sandbox is active. So that&#8217;s kind of the rationale. It maximizes threads per megawatt.</p><p><strong>Ah. Yeah. Maximizes threads per megawatt. That&#8217;s pretty that&#8217;s really interesting. Yeah, do you think, I mean, do you think there&#8217;s a point where people are talking about threads per megawatt as opposed to like CPUs per megawatt?</strong></p><p><strong>MR:</strong> I think they&#8217;re starting to get there sandboxes per megawatt, cores per megawatt because that is the unit of deployment and as you very well know, there&#8217;s a lot of power constraint in data centers. So if I have a 10 megawatt facility, how can I maximize how much sandbox capacity I can get there, right? So that is a critical metric and cores is a pretty good proxy for that as long as your cores perform well.</p><p><strong>Yeah. All right, super interesting. Okay, so now let&#8217;s talk kind of competitive positioning. I for anyone who&#8217;s been following this closely knows there&#8217;s been a lot of companies in the space saying they&#8217;ve got an agentic CPU rack now. And I maybe to remind listeners, Nvidia has Vera, which was originally a head node and now they also have a standalone CPU rack. That was 88 cores and 176 threads. Arm launched their AGI CPU. I think it was up to 136 cores and 136 threads, so no SMT or multi-threading. And interestingly, which set us up nicely for this conversation, on Arm&#8217;s most recent earnings call, there was a sell side analyst who specifically used the AMD framing and he said head node,</strong></p><p><strong>MR:</strong> Yeah. So, the one SKU fits all, I think, anyone in this industry long enough to know that the customer environments are diverse, the workloads are diverse, the constraints are diverse, that I don&#8217;t necessarily buy into that view of the world. Sometimes you&#8217;re different driven to different SKUs based on workload, and sometimes you&#8217;re driven to different SKUs based on constraints. Sometimes you&#8217;re driven to SKUs based on, hey, I already plumbed this data center for liquid cooling because I have a bunch of racks and GPUs. And now I&#8217;d like to slide in a compute rack that also takes advantage of the liquid cooling, right? Some of what are being called agentic racks, in my view are just dense liquid cooled racks that take advantage of what was provisioned in the data center for GPUs already. I don&#8217;t think there&#8217;s anything that makes them agentic specifically because again, as I mentioned, agentic is a diverse set of steps. Many of those steps are like gateways running nginx, which are things that we know and love for many years now. And some of them are like these agentic tools and orchestration. So you&#8217;re going to have an entire data center that&#8217;s running a end-to-end agentic workflow, and you&#8217;re going to have multiple parts to save those to serve those workloads, right? Your rag databases are going to need something different from what your storage nodes are going to need, or are going to be something different from what your gateway needs and what the tools need, right? So, I think you will continue to see those types of diverse deployments, and even if you consolidate down to say two or three SKUs for your entire data center, you&#8217;ll still have multiple configurations in the rack when it comes to memory provisioning and how much storage you put in it and so on.</p><p><strong>Yeah, that makes a lot of sense.</strong></p><p><strong>MR:</strong> Yeah, there&#8217;s no such thing as an agentic rack.</p><p><strong>Yeah, well, no, and I think that&#8217;s very fair based on what we talked about, where you&#8217;ve outlined there are different workloads and they have different characteristics and therefore they need something different from the compute that they&#8217;re deployed on. And I actually like the way at the end saying there&#8217;s no such thing as an agentic rack, stripping away the marketing and just saying like, a dense liquid cooled rack, a dense air cooled rack, a maybe less dense but high performance liquid cooled rack that&#8217;s really what we&#8217;re talking about at the end of the day.</strong></p><p><strong>MR:</strong> Yes. Exactly. 100% agree with that. That&#8217;s how I think about the world.</p><p><strong>Yeah, no, that is helpful. So thank you for sharing that. Okay, maybe one last question on the CPU front and agentic CPU front. A lot of hyperscalers have designed their own arm-based CPUs, Google Axion and Microsoft Cobalt and so forth. Traditionally those have been for general purpose and actually even a lot of them are for first-party workloads, but some are for third-party workloads as CSPs. Do you see those fitting into serving some of there&#8217;s obviously general purpose. Do you see them fitting into the agentic workloads at all?</strong></p><p><strong>MR:</strong> Yeah, I mean, they are general purpose CPUs and they can run a broad set of workloads, right? I mean, within the constraints of the Arm ecosystem, but if they wanted to run some Python code on a single thread, they could obviously do that really well. But our job is to make sure we run it even better than that and we give the concurrency and the PCO and the for what benefits that make it hard for them to choose their own CPU versus what we have to offer. So, that&#8217;s how I think about it is a competitive environment. We have other merchant silicon vendors that are really smart that we compete against, and then we have all of the hyperscalers who are building their own silicon with really smart people there, and we just have to do better than them and be generationally ahead, and that keeps justifying why our deployments continue to grow.</p><p><strong>Nice. Okay, that&#8217;s a good answer. All right, let&#8217;s move into the enterprise AI stuff. So, can you walk us through listeners in case they forgot. Obviously, there&#8217;s a lot of talk about Helios lately, which is rack scale, but talk to us about the sub rack scale options.</strong></p><p><strong>MR:</strong> Yeah, and maybe we should talk, we can start talking about tokenomics and then we can talk about some other motivations like regulatory reasons and some other things like that. But if you think tokenomics, I think every corporation is starting to use AI far more extensively, and that has put a lot of pressure on token budgets, right? Every everyone I know is using a lot more tokens and the bills are growing dramatically. Now, the question that we have to ask is not every AI related task needs to go to a PhD scientist equivalent AI. Some of them could go to an intern equivalent AI, right? So, I think that&#8217;s really where tokenomics comes in is route the highest value requests to the highest cost AI resource, and things that can be done with smaller models in lower cost hardware, route them to that and create a more balanced environment in terms of token spend. I think that&#8217;s something that every single IT department across enterprises is looking at. Our own IT department is looking at that. We are doing experiments with the MI350P PCIe card, which can go into a broad set of on-prem enterprise servers, air cooled, slotted into a PCIe slot and get it going. And we are looking to see if we can serve save like 40% on token costs by load balancing the right tokens to that on-prem resource, and then the right set of token requests to the really large frontier models in the cloud. So, I think that&#8217;s what&#8217;s motivating customers to think about it is minimizing cloud token cost spend while still not compromising on the output of the AI. So, where you need the frontier models, you don&#8217;t want to run it on a smaller model because the quality of the output will be wrong and you&#8217;ll just run it five times to get the right answer. So, you don&#8217;t want to do that. You have to know what to balance where, and I think that&#8217;s where a lot of work is going on to have intelligent model routers and so on.</p><p><strong>Yeah, okay, interesting. I mean, I can relate to this even from my own personal experience, which is I&#8217;ve got my $200 Claude plan and when I need Fable and I&#8217;m doing frontier model for coding, but then the stuff that I&#8217;m coding, I&#8217;m ultimately running and maybe have some lighter LLMs in it where I&#8217;m like, oh, dude, I&#8217;m totally not burning Fable tokens on some of this stuff. And in fact, and this is probably is actually very related to our conversation here. I have an AMD Ryzen AI Max, I think it&#8217;s 395 plus, it&#8217;s like the framework version, here in my And so even I have been thinking about like, how do I reduce my spend and orchestrate my tokens to running sort of from my own token generator, so free tokens if you will, versus paying for frontier tokens. And so you&#8217;re saying, a lot of enterprises are thinking the same way and MI350P is enough for their medium-sized models and maybe even the eight-way server nodes are enough to run even like 1 trillion plus parameter models.</strong></p><p><strong>MR:</strong> Yes. I think that&#8217;s what we&#8217;re seeing. You can do some medium to large reasoning on prem with those servers and then use the frontier models for the highest value stuff. And in some cases, they don&#8217;t even use those smaller GPUs on prem. They rent the smaller GPUs in the cloud. So that&#8217;s another model too.</p><p><strong>Oh, there you go.</strong></p><p><strong>MR:</strong> Yeah. So I think both models are valid. You I think the whole world is was going hybrid even before this happened. And now it&#8217;s a much more complicated hybrid.</p><p><strong>Sure, totally. So what about what are you seeing with customers like, are there concerns that to be able to run on premise or even on GPUs that you rent, you have to use an open source model, generally a little bit behind the frontier, versus the frontier models. Like is the open weight model that might be kind of six months behind a hindrance or not really?</strong></p><p><strong>MR:</strong> I think that&#8217;s where the model routers come in. If it&#8217;s a simple summarization task, I don&#8217;t think it&#8217;s necessarily as much of a hindrance. But if you&#8217;re asking it to do write the most complex piece of code, then it could be because it takes a few months for those models to catch up. So I think this is where how you route the models becomes important, right? Simpler use cases or use cases those models are good at, you route to those devices and the rest goes to the frontier models and that&#8217;s going to be a constant it&#8217;s a constantly shifting thing.</p><p><strong>Yes. Indeed. Indeed. So and then what about desk side inference, whether it&#8217;s a laptop or a small kind of mini PC that are kind of gaining steam. Do you see those also just plugging into it sits behind the orchestration and or how are you thinking about it?</strong></p><p><strong>MR:</strong> Oh, I think you&#8217;re going to see all kinds of models, right? I&#8217;m setting up one of those at home to run some personal agents and I&#8217;m not necessarily and right now I&#8217;m still at a point where I don&#8217;t let the agents go unrestrained. So I would like to control the access policies and what they can access and so on. So I think there&#8217;s going to be a whole bunch of people doing things like that. I already spoke about running a bunch of agents on my laptop. I think that model will continue as well. I can envision a day when you even have a couple of smaller agents running on your phone. Right? So I think there is going to be agents deployed across multiple places and maybe even talking to each other and doing the right thing in the right place. I think this has always been the panacea of edge computing for a while is, hey, let&#8217;s have a far edge and a near edge and a regional data center and a central data center and the problem was always orchestrating it. But when AI can figure out how to do that, maybe that actually becomes more reality.</p><p><strong>Sure, totally. Do you it feels like there&#8217;s probably opportunity for startups or companies to create this orchestration. I mean, is everyone doing their own bespoke orchestration or are there companies out there that can tackle this for customers?</strong></p><p><strong>MR:</strong> I think we are seeing some open source orchestration stacks pop up, but if I think about a lot of the biggest customers out there, it&#8217;s still a little bespoke. I think just like any rapidly evolving industry, it&#8217;ll start with things that are bespoke and people are modifying and then eventually become a little more standardized and open source. I think Kubernetes is a great example of that and I think we will see the same evolution here. I think eventually openness wins.</p><p><strong>Sure, totally. And what about security? This just popped in my mind. We see the Anthropic&#8217;s got models that are very advanced and they&#8217;re very good at cyber security, either defense or offense. If I&#8217;m buying my own token generator, running stuff on prem, do I need to get more sophisticated from a security perspective in like now my on premise models need to be able to prevent someone&#8217;s other models from attacking them or something or is this just like a networking firewall issue?</strong></p><p><strong>MR:</strong> My view on AI and security is especially the most advanced AI models are extremely useful tools to make your entire system more secure, right? So I think you have to leverage regardless of what how you are splitting your workload between on prem GPUs and advanced frontier models. I think you&#8217;d use the most advanced models to find all of the security attack surfaces and close them. So I think of this as an useful tool to secure the system, no matter how you compose the system.</p><p><strong>I like it. Yeah, that makes sense. AI is yet another tool to help you level up and do even better from a security perspective. Yes. Okay. So this has been great, Madhu. Let&#8217;s see, last question. I so I feel like what I&#8217;m hearing the theme for this whole conversation is AMD has a portfolio of compute for your different workloads, for your different size customers. We talked about CPUs and now we&#8217;ve talked about essentially sub rack scale GPUs. Anything else is that like the correct maybe vision and the way we should think about AMD going forward and oh, and is there anything that we&#8217;re under appreciating with that sort of portfolio and you&#8217;ve mentioned open before the open and the portfolio value.</strong></p><p><strong>MR:</strong> Yeah, and I think I should also add, right? There&#8217;s the CPUs, there&#8217;s the sub rack scale GPUs, there&#8217;s the Helios to run the most advanced frontier models. We have our Pensando DPUs to connect all of these together. Then we have the laptops and the we have the smaller AI boxes. So we got a portfolio that spans all of those and I think that&#8217;s important in this era because people are deploying this so many different ways.</p><p><strong>Yes. So the TAM will continue to grow from the desktop to the edge, to the far edge, to the cloud and AMD has offerings for all of it.</strong></p><p><strong>MR:</strong> Yes. And we will continue on the server CPU side to make sure we have a portfolio because anyone that says they know exactly what agents are going to do say three years from now, well, I would love them to give me some stock advice as well. But I think having a portfolio that serves a broad set of workloads is what is going to win the win at the end of the day.</p><p><strong>Yes, I agree with you. Yes, like you said earlier, maybe it runs Python today, maybe the agents decide to run rust or go tomorrow. Who knows? But</strong></p><p><strong>MR:</strong> Or they come up with their own new language too.</p><p><strong>Right? Yes, there&#8217;s probably something even more efficient out there. There you go. Awesome. Well, this was a great conversation. We covered a lot of ground. Thank you for your time and we&#8217;ll have to stay in touch.</strong></p><p><strong>MR:</strong> Yeah, absolutely. Thank you so much.</p>]]></content:encoded></item><item><title><![CDATA[The Hardest Computing Problem Ever, According to Arm's Physical AI Chief]]></title><description><![CDATA[Arm's Physical AI chief on why humanoid robots are the hardest computing problem yet, and why the physical economy still runs on a tenth of the digital side's compute.]]></description><link>https://newsletter.chipstrat.com/p/the-hardest-computing-problem-ever</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/the-hardest-computing-problem-ever</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Mon, 10 Aug 2026 13:03:12 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/210388207/9a8814c283eab7228a4ebe5076be4358.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><em>Drew Henry ran GeForce at Nvidia for 11 years. He launched Arm&#8217;s Neoverse infrastructure business. Now he leads Arm&#8217;s Physical AI unit. I sat down with him for the latest Bit by Bit conversation.</em></p><p><em>His main argument: robots, humanoids, and autonomous vehicles add up to one of the hardest systems-engineering problems in computing. And the physical half of the global economy still runs on a fraction of the compute the digital half already gets.</em></p><p>Physical AI comes down to two numbers: latency from a sensed photon to torque on an actuator, and system weight in grams. <em>No bolting on a bigger cooling system the way a data center would.</em></p><p>And Global GDP is about $115T. Roughly $70T is physical, $40T is digital. Almost all compute spend goes to the smaller, digital half. And yet <strong>the physical economy runs on a tenth of the compute per dollar of output.</strong></p><p>That&#8217;s pretty mind-blowing! More things stood out:</p><ul><li><p><strong>The vaue prop of Physical AI is time, not automation.</strong> An autonomous vehicle doesn&#8217;t remove the drive to your kid&#8217;s game. It gives back your attention during it.</p></li><li><p><strong>Every physical industry breaks into four verbs: grown or mined, manufactured, moved, operated.</strong> AI gains there flow into two-thirds of global GDP.</p></li><li><p><strong>A humanoid hand may be harder to build than an AI data center.</strong> Every joint shares the same latency and torque constraints, zero tolerance for error. <em>&#8220;The thing you never want to do is hallucinate during the middle of surgery.&#8221;</em></p></li><li><p><strong>A robot isn&#8217;t one compute problem, it&#8217;s four:</strong> locomotion (real-time), conversation (LLM-based), actuation orchestration, and cloud sync for retraining.</p></li><li><p><strong>Memory and model size at the edge have no settled answer yet.</strong> Henry expects a decade of trading precision for footprint before it converges.</p></li><li><p><strong>Braking stays local, coordination goes to the cloud.</strong> No cell-tower latency on a meters-scale decision, but a warehouse of robots needs central coordination.</p></li><li><p><strong>Fleet orchestration, not the chip, is becoming the real edge.</strong> Design the robot and the conveyor together, don&#8217;t bolt one onto the other.</p></li><li><p><strong>Arm co-designed Graviton with Amazon</strong>, once Amazon knew its workloads well enough to spec exactly what it needed. Same pattern now starting in AVs and robotics: go vertical on custom silicon, built on Arm.</p></li><li><p><strong>No ceiling on compute demand, every robotics company tells Henry the same thing.</strong> His comparison: PC modding culture from his GeForce days.</p><p></p></li></ul><p><em>Full transcript below (lightly edited for clarity).</em></p><div><hr></div><p><strong>Hello listeners. We have a special guest with us today, Drew Henry, EVP of Arm&#8217;s Physical AI Business Unit. Welcome, Drew.</strong></p><p>DH: Thank you, Austin. I&#8217;ve been looking forward to this all day.</p><p><strong>Awesome, good, me too, man. And of course, nice studio you&#8217;re in. That looks awesome.</strong></p><p>DH: Isn&#8217;t it nice? I know. I really like it.</p><p><strong>Yeah, it&#8217;s good. Everyone who&#8217;s watching take notes. This is how you do it for these Silicon companies out there. </strong></p><p><strong>Okay, so Drew, before we get into talking about physical AI, which is the topic here, I wanted to go through your background because you had an amazing journey. And so I&#8217;m just going to kind of read your LinkedIn profile for the listeners in case they aren&#8217;t familiar with you. And I&#8217;d love afterward for you to give me some color on maybe the through line of all your interesting experiences and how it helps you in the role that you&#8217;re in now. </strong></p><p><strong>So what I thought was awesome, you co-founded Animated Technologies in the early days of computer graphics and that got acquired. Then you went to Silicon Graphics, which I think most folks have heard of, that&#8217;s Jim Clark&#8217;s company during its heyday. And from there, in the dotcom boom, you went to a streaming media startup where you were an exec through the IPO. So all already right there, very interesting. </strong></p><p><strong>And then you went to this little company called Nvidia, which I think most people are familiar with in 2001, and you were a GM there for 11 years. And then you were at SanDisk, which is also a very hot company right now. A couple more startups, and you&#8217;ve been at Arm almost nine years now, where first you were the founding GM of the infrastructure business and launched Neoverse. Again, very awesome. </strong></p><p><strong>Quite a career there, but now you&#8217;re taking Arm into physical AI. And also, I noticed you&#8217;re on the board of Teradyne. So quite the expansive background from startups to Fortune 500 companies. </strong></p><p><strong>So, yes, tell us, what&#8217;s the through line and how does it apply to what you&#8217;re doing now in physical AI?</strong></p><p>DH: Yeah, thanks. The through line for me is just I just always enjoyed working on really interesting things. And so from early days in studying engineering, both my undergrad and graduate degrees. I just was you just there&#8217;s just a curiosity that you have and that you build from something like that. And so that&#8217;s kind of the thing. I just always wanted to work on very interesting things through my career and I had the opportunity to do that being here in Silicon Valley. I&#8217;m one of the few people, by the way, that&#8217;s actually born in Silicon Valley and works in Silicon Valley. The but some of those things have been really informative to me. </p><p>So for instance, working at Silicon Graphics helped me deeply appreciate what it means to be disrupted and not being a disruptor. So I was through, I went through that entire cycle with that company. I was there in the very early days and when I kind of finished that cycle, I realized, man, I just I really I want to be on where technology is disrupting, where technology is on it on its growth side, where it&#8217;s where the company is really well positioned and stuff like that. And so every decision I made from that is based upon being on the disruptor side of things and not being on the disrupted side of things. </p><p>And when I joined Nvidia was literally a few hundred people at the time that I joined. Very early days. And so that was the fun time of trying to figure out what you&#8217;re going to do to create markets and stuff like that. So that&#8217;s that through line. The through line was learning painfully early on in my career, some of the lessons of being disrupted and then just figuring out how to be on the side of being a disruptor.</p><p><strong>Sure, interesting. Really fascinating. I love that chasing your curiosity and then yes, trying to think how can you be on the front line of the disruption. But I think very cool for you as you&#8217;re a GM to think there&#8217;ll be a natural evolution to what you&#8217;re working on and how can you probably prevent the disruption in the long run.</strong></p><p>DH: Yes.</p><p><strong>That I wake up and go to bed at night thinking about.</strong></p><p>DH: Great, that&#8217;s great.</p><p><strong>Okay, so now you&#8217;ve moved into physical AI. So tell us more. I&#8217;m curious, like, how does Arm even define physical AI? I think it has a broad definition out there. People think of many different things. How do you define physical AI?</strong></p><p>DH: Yeah, for us, it is we think of it as how AI, the latest in AI technologies are being embodied into physical devices. Devices that can sense what&#8217;s going on in the world, can take can make a decision about what to do based upon what it&#8217;s sensing and then can go off and take action in it. But do that safely in a way in which you&#8217;re really helping your helping solve problems. And the and so that&#8217;s how we think about it. When I turn it into kind of technical terms, I think about it as the key metric being the latency between a sensed photon and torque being applied to an actuator. That metric is key. And then the second metric that we always think about is that in this physical AI space, every gram matters. So you can&#8217;t have a massively huge computing system with a massively huge cooling system. Like in my GeForce days, we make huge cooling systems we put on those things. These things are everything is about a total system design. So that&#8217;s how we think about it. That&#8217;s how we kind of think about what&#8217;s most important for it.</p><p><strong>Hm, I like it. Sense, decide and act. I like that at a high level because it&#8217;s not focused on a particular embodiment, humanoids or automotive or something, but it&#8217;s a little bit more abstract, higher level. And then, yeah, the engineering definitions there are interesting. Photon and torque applied that does helps get people in the mindset of this is sensing the world and taking action in the world. So yeah, very cool. Now, okay, so why should a normal person care about physical AI? Is this about productivity? What is it?</strong></p><p>DH: Yeah, I&#8217;ll put it in terms I think that a lot of people can appreciate. One of the greatest benefits of what we&#8217;re trying to work on in physical AI is actually to give people time back. So, for instance, when my I&#8217;ve I have two daughters and the and of course, my wife and I, like anyone, very involved in our daughter&#8217;s lives and they were volleyball players. So we carting them around to volleyball tournaments and stuff like that. And I was myself or my wife were driving all the time. And the and there&#8217;s this amazing amount of time that you have with your kid when you&#8217;re driving them around all over the place, but you&#8217;re driving, right? So there&#8217;s only so much presence or interaction that you get with your kid at during that time. And being able to build an autonomous vehicle where, yeah, I want to be with my kid as they go to the this volleyball tournament or something like that. But man, there&#8217;s an opportunity to be more present with them, right? We can tell stories together, we can watch movie together, we can do whatever during that particular time. And so to me, that&#8217;s what this benefit of physical AI is about giving people time back or really at the core, just creating productivity. And there and applying that across this whole world of the of physical oriented industries is going to probably result in one of the biggest growth in global GDP in the history of the world.</p><p><strong>Sure, totally. Really cool. Yeah, I like the framing around time as the KPI, if you will, or the value prop. Obviously, time is such a precious resource for everyone. And then, of course, when you think about enterprises, time matters and you can usually quantify that. But I think even for consumers, framing things around time is really interesting. Of course, we see robots dancing or robots making a cup of coffee, and it&#8217;s not often a problem that I can relate to, but the problem of like, I don&#8217;t have enough time. I think that is pretty universal and very sort of user-centric. So that as a recovering product manager, that resonates with me.</strong></p><p>DH: I agree. It&#8217;s funny when you I was listening to something recently and somebody framed something in a really interesting way about how to think about the physical AI. And if you just kind of look around the world, things are either grown or mined. They&#8217;re then manufactured. They are then moved and then they&#8217;re operated in some fashion, right? Yeah. So grown or mined, manufactured and then moved in some fashion or other operated. And the and when you look across all of those things, you think about, wow, if I can improve the productivity of any of those things or whatever metrics matter, that is what the benefit of physical AI actually is to all of those different industries. So that&#8217;s transportation, that&#8217;s logistics, that&#8217;s manufacturing, that&#8217;s mining, that&#8217;s agriculture. All those things are things that will benefit substantially, if we appropriately apply</p><p><strong>Yeah. And that, makes your argument that this could be sort of the biggest impact to GDP and obviously a huge TAM because really when you talk about it&#8217;s grown or mined, it&#8217;s manufactured, it&#8217;s moved. I think supply chains and just anything you can think about from fishing to farming to making silicon wafers and chips you can see to making tractors, you can see how physical AI could be a substrate underneath all of that.</strong></p><p>DH: And it&#8217;s interesting. I and I agree with you and I like the way that you described it, Austin. The interesting thing for us when I really think about this is I if you take a step back and you just kind of look at global GDP, gross domestic product. This is just what the world spends on creating valued goods. It&#8217;s about $115 trillion last year, ish, right? $70 trillion, are related to things that are manufacturing, transportation, construction, these things that are very physically oriented. A lot of physical labor, all those things I was just talking about. And then the other $40 trillion, approximately speaking, is digital, kind of the digital side of the economy. It is things like finance and entertainment, communications, things like that are all kind of digitally oriented. And the very interesting thing when you look at both those things, is you actually recognize the vast majority that we spend today as a world on compute is actually applied to that digital side of the economy.</p><p><strong>Yeah.</strong></p><p>DH: And matter of fact, if you really dig into it&#8217;s about a 10x difference between the right the way that compute is used on the digital side of the economy than it is on the physical side of the economy. And so on the physical side of the economy, if we can bring AI the benefit is really coming from the world of AI. If we can bring AI into the physical economy, into transportation, logistics manufacturing, all the areas I just talked about, and just increase productivity small amounts, that&#8217;s a massive benefit to what is the largest portion, two-thirds of global GDP.</p><p><strong>Fascinating. Yeah, I&#8217;ve never thought about that way of just how such a percentage of GDP is really physical. And I can think back to previous jobs working like for a grocery retailer and thinking about what was spent on compute and obviously a lot of it was just on like the let&#8217;s make a web app or a mobile app or the IT side of things. But really very little on anything supply chain, physical movement of goods and everything. And so yeah, I can totally see how we&#8217;re in 40% of that GDP, we&#8217;re investing in compute, but what about that other 60% or whatever the numbers were, what it&#8217;s just like untouched. What could happen? And of course, not only if we apply compute, but now, obviously, in this era, there&#8217;s very interesting AI that can help and we can get into that. Not make interesting value happen physically, not just how do you apply AI to solving the problem of how should we, restock our groceries, but what about something more physical? Could you have something restocking groceries at night, for example?</strong></p><p>DH: Exactly. Exactly. And do it in a way where there&#8217;s an intelligent system that&#8217;s behind it that is constantly working on optimizing it and then transferring the optimizations that just recognized there right in the place where the goods are transacting people are walking into a market to buy a to buy something and then up and then upstreaming all that intelligence going, hey, listen, this is what&#8217;s hot right now because of whatever demand is interesting in the marketplace. Let&#8217;s make sure that we are keeping this well stocked and let&#8217;s make sure the entire supply chain is prepared for that. That&#8217;s the interesting thing that we&#8217;re on the verge of being able to add to this physical side of the economy.</p><p><strong>Nice, amazing. So now going back, you mentioned every gram matters and that got me thinking about constraints. So talk to me about constraints in physical AI because it&#8217;s obviously it&#8217;s different than data centers or a laptop on your desk. We&#8217;re talking you&#8217;ve got batteries because things are moving, but yet there&#8217;s actuators, you&#8217;re like applying torque somewhere in the system. So there&#8217;s like, there&#8217;s definitely power constraints, there&#8217;s weights, there&#8217;s probably thermals, there&#8217;s probably connectivity if it&#8217;s mobile. So I&#8217;m wondering how physical AI system designers balance all of those tradeoffs. They&#8217;re a little bit different. But then at the same time, I&#8217;m thinking like, well, this does feel a lot like mobile in some respects, weight, thermals, connectivity. So how are you seeing, the similarities and differences in designing around the constraints of physical AI?</strong></p><p>DH: It&#8217;s such an interesting question, Austin, and I really like the way you framed it because it is a it&#8217;s a total systems engineering problem. And I actually believe that it is the most complicated computing problem ever. This issue on physical AI, just particularly if you look at like a humanoid. And the and this issue, if you just spend a little bit of time thinking about it, you can appreciate the complexity of it. So, imagine you&#8217;re trying to build a humanoid robot and what you care about is the precise movement of this little tiny joint at the end of a finger, okay? Just trying to move that thing. And you want to move that thing and you and it&#8217;s got to grab something safely, it can&#8217;t crush it&#8217;s but it&#8217;s got to grip it appropriately. It has to actually get in touch. It&#8217;s got to make touch with it.</p><p><strong>Yeah.</strong></p><p>DH: The computing system and then the communication, the computing system through all the actuation that exists through that robotic arm, all the way out to this little tiny actuator that&#8217;s moving around the very end of the appendage. All of that has to be coordinated extraordinarily well. And if you&#8217;re and if each one of those joints is off by even the smallest amount of tolerances, you&#8217;re never going to grab what you want to grab. So you&#8217;ve got to think about it from the from that endpoint and bring that design thinking all the way through. This is what I&#8217;m trying to accomplish and can I accomplish this in a way in which I actually get this task done. And so these issues of latency I see something, I sense something, I touch something to act to informing actuation, the action that you have to take. And then you have to do that all safely because you don&#8217;t want to do anything wrong. I tell people all the time, the thing you never want to do is hallucinate during the middle of surgery.</p><p><strong>Right. Totally. Yes.</strong></p><p>DH: So these are the constraints, which is what makes it a systems problem and with all due respect to my friends that are building these incredibly massive AI data centers for AI factory kind of data centers. I think the humanoid robotic platform is more complicated because of all those things.</p><p><strong>Sure. Yeah. No, that is fascinating of thinking about the systems approach. Lots of actuators, lots of sensors, lots of communication, lots of control, latency budgets, to your point. And then I guess maybe even more broadly, because there is an AI model that&#8217;s running on the humanoid in this example, it is even bigger systems problem, which is you got the humanoid, but then you do have the like training of the AI model and the deployment of the AI model. So yeah, maybe say a little bit about that briefly, just it doesn&#8217;t feel like just embedded computing anymore. It&#8217;s got that AI model piece applied to it.</strong></p><p>DH: Yeah, embedded computing for those that aren&#8217;t familiar with it. You&#8217;re literally writing exactly what you want everything to do. Literally coding it. You&#8217;re just writing, please do this, then do that. Whereas AI is very different. AI is about, hey, I&#8217;m creating a model that kind of understands all the different things that this device could do and how do I get that device to do the things that I really want it to do by exploiting AI. The interesting thing to me about this is as we have just gotten continuously deeper and deeper into us. And by the way, we&#8217;ve been working in these industries at from arm for a very long time. I remind people that just in the last 12 months, we&#8217;ve shipped two billion devices in this in these areas that relate to these physical AI application areas. So it&#8217;s not we&#8217;re not new to us. But the computing has changed quite a bit. And to your point, that question that you&#8217;re asking about Austin about different AI models, there if you look at a autonomous any type of an autonomous system, there&#8217;s actually four different compute that compute layers that you need to think about. There&#8217;s the compute layer that is the layer of motion. In robotics, it&#8217;s referred to as locomotion; in autonomous vehicles, it&#8217;s about the autonomous motion through the world. And that is very real-time based. I&#8217;m driving down the road, I&#8217;ve got to break in that in a very fixed period of time, so there&#8217;s actually latency matters. There&#8217;s a very fast clock that&#8217;s going, so there&#8217;s a real-time nature to the way these things work. It&#8217;s actually real-time system. That is different than when you&#8217;re interacting. So I&#8217;m sitting in an autonomous vehicle and I&#8217;m having a conversation with that vehicle, right? That is a that conversation with that vehicle or I might be asking that vehicle, hey, change my destination, I want to go somewhere else. That&#8217;s not quite the same type of real-time system. It&#8217;s more LLM-based than what the other one is based. And so there&#8217;s even different AI models that you run inside these different types of systems. And then all of that has to hit the compute plane that manages all the actuation. And actually orchestrates all the actuation. It&#8217;s got to do that safely, it&#8217;s got to do that reliably and repeatably. And then eventually, the thing does connect to that to connect to the cloud because you want to update it, train it bring up new models and stuff like that. So there&#8217;s all these different layers of compute. So when you when we talk to people about how we the computing platforms that you want to design that are built on top of arm, a lot of questions come up on, well, exactly which plane of this computer are you trying to solve for? Because models are different, training is different.</p><p><strong>Yes. Oh man, that&#8217;s that is complex. I can see your argument for it&#8217;s the most complex system because yeah, real-time systems, latency budgets, you have to your point of like, you got to get photon and make a decision, hit the brakes if you need to. You&#8217;re going 60 miles an hour, it&#8217;s got to be so fast because otherwise it&#8217;s literally turning into meters until maybe you strike something. But then on the other hand, yeah, I can see like talking to the vehicle is a totally different ball game. So which gets me thinking like, okay, GPUs, CPUs, NPUs, FPGA, like what real-time operating systems, like what are all like the building blocks here?</strong></p><p>DH: Yeah, those all of the above. I guess is the best way to describe it. We as a as I tell people, we don&#8217;t have specifically dog in the hunt on that. Meaning that you&#8217;ve got to design a system, the system has to be able to accomplish something. It&#8217;s it and it&#8217;s an engineering problem and it&#8217;s got to do things in a certain amount of time and there are portions of that workload that you execute on the CPU subsystem. There&#8217;s portions of that workload that you actually execute on the real-time safety subsystem that underlies this whole thing. So if the thing goes haywire, it has a safe space to fall back to. There&#8217;s workload that you, hey, I&#8217;ve got this workload. I&#8217;ve got to do it in a fraction of a second and so I need a dedicated acceleration unit to do that could be any one of the things that you talked about. It&#8217;s a system design and our experience is and as we all the customers that we talk to in this area, and we talk to the world&#8217;s leading, there&#8217;s just not a robotics or autonomous vehicle company that of that I haven&#8217;t talked to in the last 90 days. All of them have just different views about how they want to go about doing this stuff. And that&#8217;s the power of what we offer is that we don&#8217;t walk in and say, here you go boom, follow this thing go to our the to this specific conference that we&#8217;re going to do and we&#8217;re going to tell you exactly what you need to build, you&#8217;ve got to use this XX platform to use it. Our approach is to enable people to build what is the right solution for the market, while also bringing the benefit of arm, which is that if you just if you do it on top of arm, you&#8217;re just going to get the software compatibility across all the stuff that you&#8217;re trying to design. So you get the benefit out of open source contributions and the like because there&#8217;s just people are just building on top of the arm ecosystem today. So the we offer that flexibility people because it is just a systems engineering problem as you&#8217;re talking about earlier, Austin, and you&#8217;ve got to have that design flexibility to design across all of it. And we just like I said, we don&#8217;t have a dog in a hunt, you just decide what you want.</p><p><strong>Yeah. So we talked about compute and how your goal at arm is really to support customers&#8217; needs, however they see fit with whatever compute blocks they need, and of course, you&#8217;ve got the software compatibility on top of that. I&#8217;m curious, the entire system, also memory is like a big thing that&#8217;s different these days than AI of the past, before maybe it was convolutional neural networks, working on some of that perception, not as big of models. And now we&#8217;re in a world where some folks at the frontier are using like vision language action models and things like this that are these end-to-end models and it&#8217;s like taking perception, planning and control and to just like a big one big model and it has like big memory requirements. So obviously, compute is part of it, but how are you guys seeing memory constraints changing at the edge? I know at the same time there&#8217;s innovations to like make KD cache smaller and stuff like this. But yeah, what do you see when you talk to all these companies? Are they concerned about memory? Do they have enough?</strong></p><p>DH: Well, I think the world wish it had more memory right now. Well, true. The simple thing. You can&#8217;t listen to anybody&#8217;s earning reports today and not know that there&#8217;s questions about memory access and the like. But those are constraints that exist at the moment. The so to get to your question, this is what I think this is what I just like the physical AI space and it&#8217;s and what&#8217;s so interesting to me about it with the kind of the background that I&#8217;ve had and things I&#8217;ve worked on in my career, is it is all new, right? This isn&#8217;t you&#8217;re really it&#8217;s not a disruption market. This is we&#8217;re actually going in and bringing into really well understood areas of manufacturing, mining, all the things I talked about before. And we&#8217;re bringing the power of this into that area. And so it&#8217;s all new. And so designs are varying. The specific issue about memory is fascinating because it&#8217;s this constant tradeoff that people have about, well, how big do I want my model to be based upon the constraints that you talked about and how do I distill that into as small as possible? So that again, I can trade off that this the robot can&#8217;t walk around with something a compute system that weighs a ton, right? It&#8217;s got to it&#8217;s measured in grams, how much it can weigh. And including the memory system subsystem and the like. So you&#8217;re constantly making those tradeoffs. So what we find is that what people care about, they really want, they want absolutely as the maximize the value, the performance that they have. Your system has a certain amount of performance in it. How do I maximize that? How do I make it so super fast, which is just architectures that we&#8217;re good at. And if memory is a perceived problem for me, well, how do I then reduce my need for it, right? So you can actually you can change data types. Some people use FP floating point data types with lots of bits and those are relatively big for a model. But if you can distill that into another lower resolution data type, but that gives you still sufficient precision to what you want to do, that can substantially change the size of the models that you have. But there&#8217;s tradeoffs if you do that. So what we find is that there is no right answer today. This isn&#8217;t like this is a world where 20 years of optimizations have existed. This is a world where new models are being invented, new architectures are being explored, new inventions are being decided upon. And there and we will eventually find what are the most efficient of those over the next decade, but it&#8217;s going to be a decade to figure all that out.</p><p><strong>Mhm. Mhm. Nice. I see. Yeah, so it sounds much more nuanced than just like, you got to have 32 gigs of RAM, otherwise you&#8217;re not serious. But it&#8217;s much it sounds much more like, this is the system architecture that we think makes sense and so how can we do more if memory is a constraint? How can we distill down into a point where we feel like we have enough headroom, it&#8217;s not a constraint or whatever. It does seem like I can see some back and forth between like the ML engineers and maybe like the system designers a little ping-pong of like, hey, I want a little bit higher precision versus like, okay, all the tradeoffs. That&#8217;s engineering, right? Like, okay, fine, it&#8217;s going to the bomb&#8217;s going to go off.</strong></p><p>DH: That&#8217;s my day.</p><p><strong>I bet. I bet. Every day, that&#8217;s awesome. Okay, I have another question for you, Drew. So, when we&#8217;re talking about these systems, so take the car, for example, there&#8217;s obviously some real-time stuff that has to happen. It has to happen there at the edge. And then there&#8217;s some LLM inference happening. Like, how do you see like edge to cloud kind of hybrid playing out across physical AI? Like, will we see some stuff runs in the cloud, some stuff runs locally? Like, how are companies thinking about this?</strong></p><p>DH: Yeah, I think it is going to eventually it&#8217;ll distill into kind of both those things. It&#8217;ll be very use case dependent. So, let&#8217;s go back to the car example. I don&#8217;t see anytime soon or ever that a the braking system of a car is going to be something that&#8217;s figured out in the cloud.</p><p><strong>Mm, right.</strong></p><p>DH: Right? It&#8217;s just it&#8217;s too there the needs are too instantaneous. And so why would why in the world would you say, hey, listen, I&#8217;m going to add a requirement of latency, ping time to a cell tower and then the actual time back to the computer. I&#8217;m not you&#8217;re just not going to do that. So there are things that without doubt will remain local. Then you kind of get into, well, but every really the most advanced autonomous cars today that we see on the road, everything from what Tesla&#8217;s doing to what Waymo&#8217;s doing, to what the Wave guys are doing, the Rivian team these are all just world-class teams. They are looking at, okay, there&#8217;s a portion of this work that I&#8217;m going to do in the cloud. Like, for instance, I&#8217;m training all my models in the cloud. So I&#8217;ve got a new model. I&#8217;m going to download that to the car off offline. So that there&#8217;s a connection that exists. And then when you start thinking about robotic systems, you start thinking about, well a humanoid robot will be doing work, but it&#8217;s probably let&#8217;s project it into the work in into enterprise and into businesses. It is probably part of a fleet. So that fleet then has fleet management that&#8217;s associated. Okay, well, this robot needs to go is going to go over there. This robot&#8217;s going to go over there. It&#8217;s going to do these different tasks. And so there&#8217;s a coordination that ends up happening where you want to actually have the system be able to do all the autonomous work that it needs to be able to do, but then recognize that, hey, I want to there&#8217;s coordination that needs to happen. And so there&#8217;s intercommunication that needs to happen. That intercommunication is going to be cloud-based. And so that&#8217;s where I think it becomes a tradeoff. The interesting is when you actually start thinking about logistics centers and factories and how they evolve because in a if you&#8217;ve got if I&#8217;m in a logistics center like the center that&#8217;s shipping packages, goods in, goods out kind of thing. That&#8217;s the four walls that make up that building are under the control of the systems designer that&#8217;s designed that whole system. And they&#8217;ve got robots that are all over the place inside those things today. But you may make decisions that, hey, listen, because I&#8217;ve got a whole bunch of robots that are there and I might have central compute that&#8217;s available to you. There are some things that I might actually allow the decisions to be done in a compute system because it&#8217;s close enough from a latency standpoint. And that&#8217;s where that&#8217;s the thing that just boggles my mind about this world of physical AI is it&#8217;s not just about robots. It&#8217;s not just about autonomous vehicles. It&#8217;s about how all of that then interacts in those things I talked about before, growing crops, mining, manufacturing, moving, logistics centers, all that stuff. There&#8217;s a level of designing it as a at a systems level that is as complicated as just dividing the designing the system the robotic platform itself.</p><p><strong>Yeah, fascinating.</strong></p><p>DH: Right?</p><p><strong>Really interesting to think through. I know it&#8217;s like kind of mind-blowing to think through like the fleet level orchestration. Obviously, LLMs and AI will be very helpful there. And then you yeah, you start to get to this like, how do you define the area of responsibility? Like, what is the robot allowed to do for itself? Even making that work is awesome. It can walk over and pick up stuff. And what decisions can it make versus what decisions does it need to send up to like the central planning brain? And then of course, and then this is interesting because then now some organizations their like differentiation and how they compete with others could be not just the platform that they&#8217;re using, but that higher level orchestration. Like, are they really sophisticated about getting goods in and out moving, making interesting decisions?</strong></p><p>DH: Yeah. Yeah. For instance it&#8217;s you can just think about a conveyor system, right? A conveyor system is moving goods down the line. And then maybe there&#8217;s a robot there that&#8217;s sorting goods and things like that. At some point, someone&#8217;s going to go, hey, you know what? Let&#8217;s design those two things together as a system. Let&#8217;s not just stick a robot in front of a conveyor system that&#8217;s there. Let&#8217;s actually, let&#8217;s take a step back and think about all the coordination of those things that happen. That&#8217;s this orchestration that occurs.</p><p><strong>Interesting. So interesting. Okay, so let me ask you a question then. In where my mind is going is like, okay, let&#8217;s say we&#8217;re in this world where there&#8217;s robots, but there&#8217;s also like a data center based model, maybe doing some orchestration. Like, is there a benefit? Does architectural consistency matter? Is there a benefit if I&#8217;m running arm in the robot, running arm in the cloud? Or how do you guys think about that?</strong></p><p>DH: Yeah. So first off, the answer to the question is, do we see that happening? The simple answer is yes, we see it happening all over the place. And it&#8217;s following the same pattern that we saw even in the as people build large cloud data centers if you take a step back and just kind of look, just squint and look at that large cloud data center, it&#8217;s just a big computing system, right? And so the Amazon guys when they when we worked with them to build Graviton, they just knew the workloads that they were building that they were running. They knew the power constraints that they had. They knew the IO requirements that they had. Everything when you&#8217;re at designing at that level, not just trying to piece something out of what&#8217;s available through a bunch of supply chain points. When you&#8217;re really designing that, you really care then about how your compute plane looks. So you actually start looking at, well, how do I customize that to my particular needs? And so they were they all voted with their feet. Every major cloud company&#8217;s already decided to do that. We&#8217;re seeing the same thing now with the happening with the autonomous drive space. Tesla early first mover on that, right? Tesla is super famous for the fact that they build everything themselves. Rivian if you&#8217;re a AI first company like Rivian is, you go, listen, I care about latency. I care about the software load that I have. I care about the data types that I care about. I really want to have something that&#8217;s specific to my designs. And so RJ and that team, they just said, we&#8217;re going to go after it. We&#8217;re going to build our own platform. And now as we see this in the just as the physical AI space into other types of form factors, not just cars, autonomous cars, which are robots on wheels. Other things are happening that other seeing the same thing in other areas around the world, other companies around the world, just going, I just I need to make a decision about when something is so unique to what I&#8217;m doing, I want to make sure it&#8217;s part of my vertical supply chain. Now, there&#8217;s a lot of people that are building really powerful things using technologies that from folks like Nvidia, which is amazing. Their technologies are amazing. And they&#8217;re really satisfied with that because it solves the problem that they need. Nvidia does great work. And then there&#8217;s people like the Rivian guys who are going, listen, I just I want to do something that&#8217;s a little more specific to what I need and I&#8217;m going to design my own thing. So we&#8217;re seeing it all over the place.</p><p><strong>Yeah, interesting. Okay, so maybe last thing, as you&#8217;re talking to all these leading robotics players, is there something that comes up in every conversation that maybe you&#8217;re surprised by or that we haven&#8217;t been thinking about us in the audience?</strong></p><p>DH: What&#8217;s interesting to me is and surprising and I guess somewhat validating is that everyone I&#8217;ve talked to has said, Drew, I&#8217;m convinced that we will have you will not be able to satisfy the amount of compute that I need at any time. So please just keep building more and more compute systems. Because they just take a step back and they go, okay, well, maybe we solve the problem about how a car drives around the world, but then I got to put people in the car and they and we want them to have this natural experience being able to talk to the car and give the car direction, and that it could take action on and have it actually infer through hand gestures and the like. So there&#8217;s these everyone comes back and going, it&#8217;s a decade or more of high performance. Matter of fact this is I was joking recently and I think this is serious. I was talking to a friend of mine. I was thinking back to the days when I was doing the GeForce business and we had this world of computer enthusiasts, right? The modders, the ones that are building their own computing platforms and the like and tweaking it and overclocking and all that stuff. Austin, I&#8217;m sure you were one of them. That is I think it&#8217;s going to happen to robotics. Sure. I think we&#8217;re going to have people that are going to get robotic platforms and they&#8217;re going to want to overclock it. They&#8217;re going to want to tweak it. They&#8217;re going to want to do their own optimizations on it. We&#8217;re in that type of a world where people are just going to want more computers than they can possibly get it and they&#8217;re going to even try to figure out how to hack it for themselves.</p><p><strong>Yeah, that makes a lot of sense. And speaking of GeForce, it does kind of remind me of like computer graphics where it&#8217;s like never enough. It&#8217;s like, wow, this looks amazing. Can we make it look even more real? Can you give us even more compute? And the same thing. I don&#8217;t think people are going to be like, okay, this robotics platform is useful enough you can stop here.</strong></p><p>DH: Yeah, right. No, I think the analogy is perfect. And where are we now in that? Everything&#8217;s neural graphics based, right? We&#8217;re using AI technologies now to in to through AI generate frames that don&#8217;t exist. It&#8217;s not rendered, it&#8217;s actually completely AI generated the frames that insert between other rendered frames. And so the so you&#8217;re absolutely right. This is one of those spaces where people will thirst for as much computing as they can possibly get.</p><p><strong>Awesome. I love it. Well, Drew, thank you for this. This got me fired up about physical AI. I look forward to seeing what you guys at Arm do over the next couple years and I&#8217;m sure the listeners walked away learning something. So thank you.</strong></p><p>DH: Austin, it&#8217;s always good to see you, man.</p>]]></content:encoded></item><item><title><![CDATA[Why Did AMD Buy Taalas?]]></title><description><![CDATA[Does it move AMD stock? Speculation. Helios plus Taalas cluster, what about long-context, who uses it, and why buying a one-model chip is less crazy than it sounds.]]></description><link>https://newsletter.chipstrat.com/p/amd-bought-a-chip-that-runs-one-model</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/amd-bought-a-chip-that-runs-one-model</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Fri, 07 Aug 2026 21:40:07 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a6b4b93b-29ab-42e0-a5a0-0520e8ddf94d_2258x1170.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We got a big surprise yesterday after the market closed. <a href="https://ir.amd.com/news-events/press-releases/detail/1296/amd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market">AMD agreed to acquire Taalas</a>, a 25-person chip startup in Toronto.</p><p><em>I was pumped to see it. This is right up my alley! It&#8217;s AI accelerator startups + a strategy puzzle of sorts... and requires us to get into the technical nitty-gritty a bit. Does it matter to AMD&#8217;s competitiveness? Stock price? </em></p><p>Taalas builds what it calls model-specific silicon. Rather than loading weights from memory into a programmable accelerator, it hard-codes a specific model&#8217;s weights directly into the chip as read-only memory. The result is fast, cheap, and completely inflexible. <em>Which is the point!</em> Its first chip, the HC1, runs Llama 3.1 8B and nothing else. <em>And that&#8217;s the point...</em></p><p><em>Also note the HC1 is referred to as just a tech demonstration</em>.</p><p><em>Btw, Bajic spent 10 years at AMD, then later co-founded <a href="https://tenstorrent.com/en">Tenstorrent</a> and ran it as CEO and CTO before starting Taalas in 2023.</em></p><p>So AMD paid an undisclosed amount for a company with no revenue, no named customers, and one chip that runs one open-weight model. But <s>weight</s> wait! The idea has merit once you dig into it!!! </p><p>Let&#8217;s unpack it.</p><p><em>If you&#8217;re new here, skim this <strong>useful background from the Chipstrat archive</strong> so you can follow the narrative arc:</em></p><ul><li><p><em><a href="https://www.chipstrat.com/p/the-rise-of-groq-slow-then-fast">The Rise of Groq: Slow, then Fast</a> (Jun 2024). How an SRAM-based accelerator found its market, and why ultra-fast inference turned out to be a business rather than a demo.</em></p></li><li><p><em><a href="https://www.chipstrat.com/p/right-tool-for-the-job">Right Tool for the Job</a> (Dec 2025). Why Nvidia bought Groq to fill an ultra-low-latency gap in its portfolio.</em></p></li><li><p><em><a href="https://www.chipstrat.com/p/the-multi-silicon-era-is-here">The Multi-Silicon Era Is Here</a> (Mar 2026). Nvidia formally embracing disaggregation, and what it opens up for multi-vendor datacenters.</em></p></li><li><p><em><a href="https://www.chipstrat.com/p/right-systems-for-agentic-workloads">Right Systems for Agentic Workloads</a> (Jan 2026). Why low-latency accelerators and long-context accelerators are different machines.</em></p></li><li><p><em><a href="https://www.chipstrat.com/p/right-sized-ai-infrastructure-marvell">Right-Sized AI Infrastructure. Marvell XPUs.</a> (Oct 2025). Right-sizing infrastructure per workload instead of buying one monolithic platform.</em></p></li></ul><p>The overarching question we&#8217;ll attack: <strong>does this change how we should perceive AMD&#8217;s competitiveness?</strong> And it ultimately comes down to whether you believe ultra-high interactivity is a real and profitable product tier. <em>Jensen does....</em></p><p>Of course the whole story is a bit juicy too. Recently at Advancing AI (July 23rd), AMD announced a partnership with Cerebras that is also a very-fast-tokens play. But then AMD agreed to buy Taalas fourteen days later!? <em>What to make of that?</em></p><p>Let&#8217;s talk through:</p><ul><li><p>Why agents want fast tokens</p></li><li><p>The interactivity revenue curve</p></li><li><p>Why GPUs can&#8217;t get to high interactivity</p></li><li><p>What Taalas is</p></li><li><p>&#128272; Speculative ways to use Taalas. </p></li><li><p>&#128272; The IP AMD is buying</p></li><li><p>&#128272; The semi-custom inference cluster</p></li><li><p>&#128272; How much context actually fits in HC1</p></li><li><p>&#128272; Enterprise and edge?</p></li><li><p>&#128272; What this means for Cerebras</p></li><li><p>&#128272; So is AMD cheap here?</p></li></ul><h2>Interactivity <em>is</em> a product tier, and it has a premium price</h2><p>Remember when Groq CEO Jonathan Ross did the <a href="https://www.cnn.com/videos/world/2024/02/14/exp-groqs-record-breaking-ai-chip-jonathan-ross-intv-021410aseg2-cnni-world.cnn">real-time LLM demo on CNN</a> on Valentine&#8217;s Day in 2024? That was awesome. <em>Demo is like 2 minutes into the video. Watch to the end.</em></p><p>Real-time voice was a fun demo and shows the power of super-low-latency tokens. But agents are the reason the fast token market is blossoming. </p><p>In the chatbot days, humans were the bottleneck; we can only read and interact so fast. But agents operate much faster! And as a result, they generate way more tokens. <em>And even though the machine will run for days, of course we humans still want to see all that agentic work done in a reasonable time.</em></p><p>To illustrate, a 50-step agent chain generating 500 tokens per step is 25,000 tokens of sequential work. At 200 tokens per second per user, that&#8217;s roughly two minutes. But at 1,200 tps that&#8217;s only ~20 seconds. And at 12,000 tps, it&#8217;s ~2 seconds. Which one commands a premium? <em>Not the slow two minutes of tokens...</em></p><p>To pique your interest, Taalas claims 16,960 tokens per second for a particular use case! </p><h2>Nvidia says customers will pay a lot for fast tokens</h2><p>Look at Nvidia&#8217;s modeling of the value of fast tokens:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MhNN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MhNN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 424w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 848w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 1272w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MhNN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png" width="1456" height="831" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:831,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Nvidia's interactivity curve&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Nvidia's interactivity curve" title="Nvidia's interactivity curve" srcset="https://substackcdn.com/image/fetch/$s_!MhNN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 424w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 848w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 1272w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Throughput per megawatt on the y-axis, tokens per second per user on the x-axis, with the price per million tokens for each service tier underneath. S<a href="https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/">ource</a></em></figcaption></figure></div><p>The x-axis is interactivity aka tokens per second per user. <em>Faster is better.</em> </p><p>Notice Nvidia shows five service tiers of sorts in dollars per million tokens. <br>Free $0. Medium $3. High $6. Premium $45. Ultra $150.</p><p>So the same million tokens sells for <strong>25x</strong> more at the Ultra tier than at High. <br><em>That&#8217;s the premium for speed!</em></p><p><strong>So why isn&#8217;t everyone selling Ultra tier tokens? Because GPUs can&#8217;t get there.</strong></p><p>Blackwell NVL72, in all its liquid-cooled, state-of-the-art scale-up domain glory, can&#8217;t hit 600+ tokens per second per user. <em>It. just. can&#8217;t.</em></p><h2>GPUs can&#8217;t do super fast tokens</h2><p>Why? Communication latency. <em>Collectives latency.</em></p><p>Recall that decode is memory-bound; every token means streaming the activated weights plus the KV cache out of HBM before doing a little math on it. Of course, useful models don&#8217;t generally fit on a single GPU&#8217;s HBM, so we shard it across more GPUs. </p><p>Tensor parallelism helps here, since each GPU reads a slice of every matrix and you get that bandwidth back in parallel. <strong>The tradeoff is latency</strong>, because now the GPUs have to talk back and forth for <em>every layer, every token.</em></p><p>Take a big MoE like <a href="https://arxiv.org/abs/2412.19437">DeepSeek-V3</a> at 671B parameters. Its <a href="https://huggingface.co/deepseek-ai/DeepSeek-V3/blob/main/config.json">config</a> tells us it&#8217;s 61 layers, with the first 3 dense, so 58 MoE layers, 256 routed experts, 8 active per token.</p><p>Tensor parallelism requires at least one all-reduce per layer. Then each MoE layer adds two all-to-alls on top; one to dispatch each token to its experts and one to gather the results back. That&#8217;s something like 116 all-to-alls plus 61 all-reduces, so call it ~175 collectives per token. <em>Don&#8217;t forget that a collective means all the GPUs wait for the slowest one to respond...hence needing low jitter&#8230;</em></p><p><strong>So how long does it take a collective to process?</strong> <a href="https://arxiv.org/abs/2606.05951">Measured small-message collectives over NVLink</a> land in the single-digit microseconds, roughly 4 to 9 &#181;s depending on the protocol.</p><div class="captioned-image-container"><figure><div class="image-link image2" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qv0T!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png 424w, https://substackcdn.com/image/fetch/$s_!Qv0T!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png 848w, https://substackcdn.com/image/fetch/$s_!Qv0T!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png 1272w, https://substackcdn.com/image/fetch/$s_!Qv0T!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qv0T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png" width="490" height="217.26415094339623" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:470,&quot;width&quot;:1060,&quot;resizeWidth&quot;:490,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Measured collective latency over NVLink&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Measured collective latency over NVLink" title="Measured collective latency over NVLink" srcset="https://substackcdn.com/image/fetch/$s_!Qv0T!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png 424w, https://substackcdn.com/image/fetch/$s_!Qv0T!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png 848w, https://substackcdn.com/image/fetch/$s_!Qv0T!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png 1272w, https://substackcdn.com/image/fetch/$s_!Qv0T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb393dc23-62d8-4f85-b6cd-f4b036b08e76_1060x470.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></div><figcaption class="image-caption"><em><a href="https://arxiv.org/abs/2606.05951">Source</a></em></figcaption></figure></div><p><strong>So 175 collectives at 4 &#181;s is ~700 &#181;s per token.</strong> <em>And that&#8217;s just the GPUs talking to each other&#8230; before any math happens.</em></p><p>Granted, engines <a href="https://arxiv.org/abs/2511.09557">overlap some of this with compute</a>, but a collective requires synchronization between all GPUs so there&#8217;s a floor on how much you can overlap. </p><p><em>To be fair, this is single-user decode napkin math and tricks like speculative decode improve interactivity. But still, there are limits to a GPU-only path unlocking the &#8220;Ultra&#8221; tier above.</em></p><p>Well then, what to do?</p><p><strong>Nvidia showed the way.</strong> Prefill on Vera Rubin NVL72 and part of decode on Groq&#8217;s SRAM-based LPUs. </p><p>(<em>See more in <a href="https://www.chipstrat.com/p/right-tool-for-the-job">Right Tool for the Job</a> and <a href="https://www.chipstrat.com/p/the-multi-silicon-era-is-here">The Multi-Silicon Era Is Here</a>. Disaggregation makes this possible.)</em></p><p>Which unlocks that blue circle:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MhNN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MhNN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 424w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 848w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 1272w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MhNN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png" width="1456" height="831" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:831,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Nvidia's interactivity curve&quot;,&quot;title&quot;:&quot;Nvidia's interactivity curve&quot;,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Nvidia's interactivity curve" title="Nvidia's interactivity curve" srcset="https://substackcdn.com/image/fetch/$s_!MhNN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 424w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 848w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 1272w, https://substackcdn.com/image/fetch/$s_!MhNN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e7f90e6-3a73-410c-8c13-ffa4ee3a6a5c_2168x1238.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Note that Blackwell stops&#8230; so would Helios</em></figcaption></figure></div><p>Yet AMD&#8217;s pareto curve would be like the Blackwell NVL72 curve. It would stop short of the blue circle.</p><p>So AMD addressed that. But not with a $20B deal like Nvidia. Rather, a partnership, a disaggregated system with Cerebras: Helios does prefill, the Wafer-Scale Engine does decode. </p><p><em>Why not buy Cerebras? AMD doesn&#8217;t have that kind of money, and Cerebras is public and too expensive anyway. And maybe there&#8217;s better long term solutions anyway&#8230;</em></p><p>But then fourteen days later AMD announced an agreement to buy Taalas!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.chipstrat.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.chipstrat.com/subscribe?"><span>Subscribe now</span></a></p><h2>What Taalas is</h2><p>Taalas <a href="https://taalas.com/the-path-to-ubiquitous-ai/">came out of stealth</a> in February 2026 with the <a href="https://taalas.com/products/">HC1</a>. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Erc6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Erc6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png 424w, https://substackcdn.com/image/fetch/$s_!Erc6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png 848w, https://substackcdn.com/image/fetch/$s_!Erc6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png 1272w, https://substackcdn.com/image/fetch/$s_!Erc6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Erc6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Taalas HC1 accelerator card&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Taalas HC1 accelerator card" title="Taalas HC1 accelerator card" srcset="https://substackcdn.com/image/fetch/$s_!Erc6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png 424w, https://substackcdn.com/image/fetch/$s_!Erc6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png 848w, https://substackcdn.com/image/fetch/$s_!Erc6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png 1272w, https://substackcdn.com/image/fetch/$s_!Erc6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e6a1126-10ae-4cfd-906b-d58e7d2371e8_1532x758.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The HC1 on a PCIe card. Taalas calls it a technology demonstrator, not a production part.  <a href="https://taalas.com/products/">Source</a></em></figcaption></figure></div><p>TSMC N6, 815 mm&#178;, 53 billion transistors, and no HBM or external DRAM. Everything the chip needs sits on the die. <em>Sounds good when memory prices are so dang high right?</em></p><p>The model weights live in a mask-ROM recall fabric burned into the silicon. A smaller SRAM holds the KV cache and any fine-tuned weights. <em>Hmm, what about long contexts?</em></p><p>No advanced packaging either. <em>No CoWoS bottlenecks. Cheaper.</em></p><p>About 250W in a PCIe card, in racks that draw 12 to 15 kW and run on air. <em>Deploy anywhere.</em></p><p>It runs Llama 3.1 8B. <em>Hmm, edge model? Draft model? Hmm&#8230; </em></p><p>Changing models means changing two mask layers, which carry both the weights and the dataflow, on a two-month turnaround. </p><p>The idea stems from earlier tech. From <a href="https://www.eetimes.com/taalas-specializes-to-extremes-for-extraordinary-token-speed/">Sally&#8217;s article</a>:</p><blockquote><p><em>Taalas is borrowing some ideas from the structured ASICs of the early 2000s to make its hardwired model-specific chips. <a href="https://www.eetimes.com/how-hybrid-structured-asics-provide-low-cost-solutions-for-mid-range-applications/">Structured ASICs</a> used gate arrays and hardened IP blocks, changing only the interconnect layers to adapt the chip to a specific workload. At the time, this was seen as a more cost-effective alternative to a full-custom ASIC that was more performant than an FPGA.</em></p><p><em>&#8220;There are definitely parallels,&#8221; Bajic said. &#8220;It&#8217;s a similar idea to <a href="https://www.eetimes.com/is-it-an-asic-is-it-an-fpga-no-its-easic/">eASIC</a> and gate arrays, but the underlying technology looks really different.&#8221;</em></p></blockquote><p>Of course, the eye-candy: 16,960 tokens per second per user at 0.75 cents per million tokens. <em>OK! That&#8217;s FAST!</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HVoW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HVoW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png 424w, https://substackcdn.com/image/fetch/$s_!HVoW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png 848w, https://substackcdn.com/image/fetch/$s_!HVoW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png 1272w, https://substackcdn.com/image/fetch/$s_!HVoW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HVoW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png" width="598" height="507.03661971830985" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1204,&quot;width&quot;:1420,&quot;resizeWidth&quot;:598,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Tokens per second per user, Llama 3.1 8B&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Tokens per second per user, Llama 3.1 8B" title="Tokens per second per user, Llama 3.1 8B" srcset="https://substackcdn.com/image/fetch/$s_!HVoW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png 424w, https://substackcdn.com/image/fetch/$s_!HVoW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png 848w, https://substackcdn.com/image/fetch/$s_!HVoW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png 1272w, https://substackcdn.com/image/fetch/$s_!HVoW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a7df240-e8cd-473a-88de-6f1539845abb_1420x1204.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Taalas&#8217;s own benchmark. Note the footnote: input sequence length 1k/1k. <a href="https://taalas.com/products/">Source</a></em></figcaption></figure></div><p>Of course the figure is self-reported, the model is quantized <em>&#8220;aggressively&#8221;</em> by Taalas&#8217;s own description, and the benchmark point is 1k/1k input/output. <em>That&#8217;s small input/output, an important detail, we&#8217;ll dig into it below.</em></p><p>The customers? Bajic from EETimes:</p><blockquote><p><em>&#8220;The big thing at the root of this idea is the assumption that <strong>the customer is willing to commit to this [chip/model] for a year.</strong> There will definitely be a lot of people who won&#8217;t, but some people will.&#8221;</em></p></blockquote><p>Did you notice how Nvidia&#8217;s chart with our blue circle, it topped out at 1,000 tokens per second per user and put $150 on that tier? Taalas claims 17K tokens per second on an 8B model at 0.75 cents per million tokens. How much is that worth? <em>Well, there are lots of asterisks.</em></p><p>So is Taalas a Cerebras replacement, or something else? </p><p>Is there a job an 8B chip does better than anything else on the market? </p><p>What about long contexts? Is this dead-on-arrival?</p><p>What about AFD?</p><p>Hmm, did AMD buy a product, or a hiring pipeline? </p><p>And does any of this change what AMD is worth today?</p><p>Let&#8217;s get into it.</p><h2>So what do I make of it?</h2>
      <p>
          <a href="https://newsletter.chipstrat.com/p/amd-bought-a-chip-that-runs-one-model">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Will ASML Finally Charge for Its Value?]]></title><description><![CDATA[Management hints ASML starts charging for value in 2027. Consensus says it won't.]]></description><link>https://newsletter.chipstrat.com/p/will-asml-finally-charge-for-its</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/will-asml-finally-charge-for-its</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Fri, 31 Jul 2026 22:26:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!pE4C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388f4a0c-8551-43e3-8242-3333c54e1061_2200x1206.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>ASML is an earned monopoly. So is TSMC.</p><p>That makes pricing complicated; it&#8217;s not simply &#8220;charge whatever you want&#8221;. You have to know what your product is worth to the customer (that alone can be hard to measure) and then decide how much of that to take. Take too little, and you&#8217;re leaving shareholders&#8217; money on the table. Take too much, and you&#8217;ll anger customers and incentivize competition. <em>And regulators start paying attention...</em></p><p>TSMC historically played it safe, erring on the side of undercharging. In return, they had happy customers with no reason to seek an alternative. That&#8217;s changing.</p><p>TSMC is raising prices. It&#8217;s part of their strategy; colloquially, &#8220;selling our value&#8221;. This &#8220;selling our value&#8221; is listed among their six profitability factors. Of course, TSMC prefers not to explicitly talk about raising prices; for example, the last call pointed to improved utilization and reduced costs for margin gains. <em>We are better at manufacturing, and that&#8217;s why margins are up.</em> Of course, gross margin rose 9.1 points, and manufacturing performance explained 3.2 of those points. The other 5.9 points? Pricing. <em><a href="https://www.chipstrat.com/p/tsmc-could-spend-10b-more-on-n2-capacityhttps://www.chipstrat.com/p/tsmc-could-spend-10b-more-on-n2-capacity">Highly recommend you read more here</a>.</em></p><p><strong>So what about ASML?</strong> </p><p>As an earned monopolist, do they approach pricing like TSMC?</p><p>Yes and no. Both call it value-based pricing, and ASML says so in those words. But ASML measures that value with a much narrower yardstick, and unlike TSMC, it&#8217;s happy to discuss the whole thing on an earnings call.</p><p><em>So, buy or sell? </em></p><p>Well, will ASML start charging for value? </p><p>Management spent this quarter hinting that it will. Consensus assumes it won&#8217;t &#8212; and the way the street has 2027 modeled, ASML&#8217;s customers end up paying <em>less</em> per wafer of capacity than they do today. <em>That has never happened&#8230;</em></p><p>Here&#8217;s what this article will cover.</p><ol><li><p><strong>How to model ASML.</strong> </p></li><li><p><strong>Does ASML charge for its value?</strong> </p></li><li><p><strong>What&#8217;s changing.</strong> April: <em>&#8220;not the way we do business.&#8221;</em> July: <em>&#8220;more flexibility for pricing.&#8221;</em></p></li><li><p><strong>ASML vs TSMC.</strong> </p></li><li><p><strong>Four things that lift 2027 revenue.</strong> </p></li><li><p><strong>Price per unit of capacity. </strong>The 2027 decline consensus is implicitly forecasting.</p></li><li><p><strong>So is it cheap?</strong> What 2027 looks like if the ASP moves, what it looks like if it doesn&#8217;t, and what today&#8217;s price already assumes. </p></li><li><p><strong>What breaks it.</strong> Three risks.</p></li></ol><p>So: does ASML start charging for that value, or doesn&#8217;t it? Below is how I&#8217;d answer that, and what each answer is worth at today&#8217;s price.</p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/will-asml-finally-charge-for-its">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[TSMC could spend $10B more on N2 capacity. Why won’t they?]]></title><description><![CDATA[TSMC has the cash... Why not build more N2? Implications of not doing so? And more]]></description><link>https://newsletter.chipstrat.com/p/tsmc-could-spend-10b-more-on-n2-capacity</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/tsmc-could-spend-10b-more-on-n2-capacity</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Fri, 24 Jul 2026 22:51:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0Uol!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>TSMC had a great quarter. But some things stood out. For example, TSMC says AI demand is unbounded. Why, then, doesn&#8217;t CapEx match that sentiment?</p><p>In the pandemic-era supercycle, TSMC spent 53% of revenue on CapEx. In the AI supercycle, with roughly $100 billion a year of operating cash flow, they&#8217;re spending only 34-36%:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0Uol!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0Uol!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png 424w, https://substackcdn.com/image/fetch/$s_!0Uol!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png 848w, https://substackcdn.com/image/fetch/$s_!0Uol!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!0Uol!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0Uol!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png" width="608" height="316.5274725274725" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:758,&quot;width&quot;:1456,&quot;resizeWidth&quot;:608,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;tsmc-capital-intensity.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="tsmc-capital-intensity.png" title="tsmc-capital-intensity.png" srcset="https://substackcdn.com/image/fetch/$s_!0Uol!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png 424w, https://substackcdn.com/image/fetch/$s_!0Uol!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png 848w, https://substackcdn.com/image/fetch/$s_!0Uol!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!0Uol!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca3391e-974e-420b-97b5-37b08acd52a2_1920x1000.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Matching the 2021 playbook today would mean another $30 billion a year. <em>What to make of that?</em></p><p>Maybe TSMC got burned by 2021-2023 and ~35% is the preferred comfort level for management. Or it could be a lack of faith in the long-term demand?</p><p>TSMC has arguably the best ability to ascertain true demand given the breadth of customers, and talking to customers&#8217; customers... So why aren&#8217;t they investing more? <em>What did Wei see?! ... jk</em></p><p>Let&#8217;s pull out a few interesting questions from TSMC&#8217;s great quarter.</p><h2>What&#8217;s in this piece</h2><ul><li><p>N2 has arrived</p></li><li><p>Gross margin peaks, gross profit keeps growing</p></li><li><p><strong>&#128272; </strong>Digging into the CapEx gap</p></li><li><p><strong>&#128272; </strong>What happens if they underinvest</p></li><li><p><strong>&#128272; </strong>Pricing&#8230; no talk, but it&#8217;s in the margins</p></li><li><p><strong>&#128272; </strong>What to watch going forward</p></li></ul><h2>N2 has arrived!</h2><p>N2 debuted at 3% of wafer revenue this quarter:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ngFh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ngFh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png 424w, https://substackcdn.com/image/fetch/$s_!ngFh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png 848w, https://substackcdn.com/image/fetch/$s_!ngFh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png 1272w, https://substackcdn.com/image/fetch/$s_!ngFh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ngFh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png" width="552" height="435.1833084947839" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1058,&quot;width&quot;:1342,&quot;resizeWidth&quot;:552,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;tsmc-7nm-and-below-revenue.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="tsmc-7nm-and-below-revenue.png" title="tsmc-7nm-and-below-revenue.png" srcset="https://substackcdn.com/image/fetch/$s_!ngFh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png 424w, https://substackcdn.com/image/fetch/$s_!ngFh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png 848w, https://substackcdn.com/image/fetch/$s_!ngFh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png 1272w, https://substackcdn.com/image/fetch/$s_!ngFh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6083aa5c-4673-40f2-a648-c4b3ff46d0a8_1342x1058.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">2nm on the scene! Source: TSMC.</figcaption></figure></div><p><em>Only 3%? Look at that baby little red bar in 2Q26...</em> But interestingly every node debuts at about a billion dollars of revenue:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vDPj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vDPj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png 424w, https://substackcdn.com/image/fetch/$s_!vDPj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png 848w, https://substackcdn.com/image/fetch/$s_!vDPj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png 1272w, https://substackcdn.com/image/fetch/$s_!vDPj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vDPj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png" width="618" height="313.2445054945055" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ebcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:738,&quot;width&quot;:1456,&quot;resizeWidth&quot;:618,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;tsmc-node-debut-table.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="tsmc-node-debut-table.png" title="tsmc-node-debut-table.png" srcset="https://substackcdn.com/image/fetch/$s_!vDPj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png 424w, https://substackcdn.com/image/fetch/$s_!vDPj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png 848w, https://substackcdn.com/image/fetch/$s_!vDPj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png 1272w, https://substackcdn.com/image/fetch/$s_!vDPj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febcd3bd8-39be-4f15-b93c-68129976a887_1491x756.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Mix share shrinks because TSMC&#8217;s revenue keeps growing.</p><p>Given that wafer prices rise every generation, flat debut dollars implies shrinking debut volume (<em>wafer prices are rough, just for illustration</em>):</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7ygx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7ygx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png 424w, https://substackcdn.com/image/fetch/$s_!7ygx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png 848w, https://substackcdn.com/image/fetch/$s_!7ygx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png 1272w, https://substackcdn.com/image/fetch/$s_!7ygx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7ygx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png" width="574" height="290.9423076923077" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:738,&quot;width&quot;:1456,&quot;resizeWidth&quot;:574,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;tsmc-implied-wafers-table.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="tsmc-implied-wafers-table.png" title="tsmc-implied-wafers-table.png" srcset="https://substackcdn.com/image/fetch/$s_!7ygx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png 424w, https://substackcdn.com/image/fetch/$s_!7ygx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png 848w, https://substackcdn.com/image/fetch/$s_!7ygx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png 1272w, https://substackcdn.com/image/fetch/$s_!7ygx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98de4fdf-2335-411c-a593-9aab0be9c363_1491x756.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Volume shrinks because each new node has more layers and longer cycle times, so the same capacity ships fewer wafers in its first ninety days.</p><p>So N2&#8217;s debut is the fewest, most expensive wafers TSMC has ever sold. Customers pay the highest price per wafer ever, and these wafers cost TSMC the most to make, with yields still climbing and a new fab&#8217;s depreciation spread over small volume.</p><p>Interestingly, N5 is still the biggest node in the portfolio (33%), even though it&#8217;s six years old:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_0UB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_0UB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png 424w, https://substackcdn.com/image/fetch/$s_!_0UB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png 848w, https://substackcdn.com/image/fetch/$s_!_0UB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png 1272w, https://substackcdn.com/image/fetch/$s_!_0UB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_0UB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png" width="534" height="416.6275167785235" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:930,&quot;width&quot;:1192,&quot;resizeWidth&quot;:534,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;tsmc-wafer-mix-2q26.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="tsmc-wafer-mix-2q26.png" title="tsmc-wafer-mix-2q26.png" srcset="https://substackcdn.com/image/fetch/$s_!_0UB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png 424w, https://substackcdn.com/image/fetch/$s_!_0UB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png 848w, https://substackcdn.com/image/fetch/$s_!_0UB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png 1272w, https://substackcdn.com/image/fetch/$s_!_0UB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4eef25a-68bf-48ec-b867-ac5ebba4ff19_1192x930.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: TSMC</figcaption></figure></div><p>Why is such an old node still so dominant? AI. <em>The answer is always AI.</em> &#128514;</p><p>N5 is the AI node. Chips ship on whatever node was mature when they were designed two or three years earlier, and today&#8217;s AI fleet (Blackwell on N4P, Google&#8217;s TPUs, Trainium, most custom XPUs) was designed in 2022-24, when N5 was the obvious choice for big dies needing proven yield.</p><p>It&#8217;s interesting how long each node family hangs on:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vfta!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vfta!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png 424w, https://substackcdn.com/image/fetch/$s_!vfta!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png 848w, https://substackcdn.com/image/fetch/$s_!vfta!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!vfta!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vfta!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;tsmc-node-age-curves.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="tsmc-node-age-curves.png" title="tsmc-node-age-curves.png" srcset="https://substackcdn.com/image/fetch/$s_!vfta!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png 424w, https://substackcdn.com/image/fetch/$s_!vfta!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png 848w, https://substackcdn.com/image/fetch/$s_!vfta!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!vfta!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F212eadce-13ee-4a59-9ece-1dc17bb570d8_2000x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Of course, TSMC&#8217;s &#8220;5nm&#8221; node isn&#8217;t one process, it&#8217;s a node family. N5, then N5P, N4, N4P, N4X. <em>So what&#8217;s six years old and 33% is the family, not the original process.</em></p><p>On this earnings call, C.C. said A14 will be &#8220;an even larger and long-lasting node for TSMC than N2, just like a 2-nanometer technology is a larger and longer-lasting node than 3-nanometer&#8221;. Each node is a decade-long product line, each bigger than the last.</p><p>Rubin-class accelerators move to N3 through 2027, hence N3 getting three new fabs. TSMC even approved three greenfield fabs for N3, a four-year-old node. One in Taiwan, one in Arizona, one in Japan. <em>New fabs for kind of old nodes.</em></p><p>The longer a node family hangs on, the better the margins. N3 is &#8220;very tight,&#8221; its gross margin set to &#8220;cross over to the corporate average in second half 2026&#8221;, and its 2022 tools depreciate off in 2027 (equipment runs on a five-year schedule). <em>Sold out, above average, depreciation-free.</em></p><p><em>Given that fabs are a ten-year+ asset, today&#8217;s CapEx is TSMC&#8217;s opinion on demand for the next decade...</em></p><h2>Gross margin peaks, gross profit keeps growing</h2><p>Gross margin printed 67.7%, a fifth straight record. But Q3 is guided down to 66.0%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aB9W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aB9W!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png 424w, https://substackcdn.com/image/fetch/$s_!aB9W!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png 848w, https://substackcdn.com/image/fetch/$s_!aB9W!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png 1272w, https://substackcdn.com/image/fetch/$s_!aB9W!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aB9W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png" width="1456" height="887" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:887,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;tsmc-gm-vs-gross-profit-dollars.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="tsmc-gm-vs-gross-profit-dollars.png" title="tsmc-gm-vs-gross-profit-dollars.png" srcset="https://substackcdn.com/image/fetch/$s_!aB9W!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png 424w, https://substackcdn.com/image/fetch/$s_!aB9W!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png 848w, https://substackcdn.com/image/fetch/$s_!aB9W!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png 1272w, https://substackcdn.com/image/fetch/$s_!aB9W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79964020-cf70-4657-bce3-6f03dcb24bf5_2100x1280.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Reminds me of an MBA professor that I had, who told our class you bank profit dollars, not margin percentages. <em>i.e. don&#8217;t sweat the margins too much</em></p><p>And the dollars keep climbing. Gross profit was roughly $27.2 billion this quarter; the Q3 guide implies about $29.8 billion. Annualized, that&#8217;s about $120 billion a year of gross profit.</p><p>Why the margin decline? N2 ramp. TSMC guided Q3 margin down &#8220;primarily as we expect the steep ramp-up of our 2-nanometer technology to dilute our gross margin by about 3 to 4 percentage points&#8221;. That&#8217;s worse than the 2 to 3 points guided in April, because the ramp got steeper. But steep ramp and more dilution sooner means more premium wafers sooner! So it&#8217;s fine.</p><p>The 67.7% gross margin might be the high-water mark until roughly 2028. Through 2027 there&#8217;s more margin drag as N2 dilutes for about eight quarters, A16 ramps behind it, and Arizona Fab 2 arrives 2H27. But if Q3 prints margins above 66.5%, well then, it would mean N2 pricing is covering its own ramp. <em>One way to get pricing hints...</em></p><p>The last margin peak, 62.2% in 4Q22, was largely an FX bump at a cyclical top, and it stood for eleven quarters because demand then collapsed and margins fell. This peak is different! It comes from loading a new node as fast as physically possible while demand is spiking. <em>So we&#8217;ll see another little mountaintop on the margins chart, but for an opposite and much better cause.</em></p><h2>What paid subscribers get in the rest of this piece</h2><p>First we&#8217;ll digging into the CapEx gap to figure out how much they could invest but aren&#8217;t. Then we&#8217;ll talk through the implications of TSMC underinvesting.</p><p>Also, pricing&#8230; TSMC doesn&#8217;t credit it to margins verbally, but it&#8217;s there. We dig in.</p><p>And what to watch going forward.</p><p>Thoughts, numbers and charts. Carry on:</p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/tsmc-could-spend-10b-more-on-n2-capacity">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[🎙️ PicoJool’s Al Yuen: The Case for GaAs VCSELs in Scale-Up Interconnects]]></title><description><![CDATA[GaAs VCSELs at 200G/lane, unconstrained supply vs. indium phosphide, the 8&#215;200 / 16&#215;100 / 32&#215;50 flavors, WIN Semiconductors, the roadmap to 12.8T, and more]]></description><link>https://newsletter.chipstrat.com/p/picojools-al-yuen-the-case-for-gaas</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/picojools-al-yuen-the-case-for-gaas</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Thu, 16 Jul 2026 20:37:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/Ywh2BHKjeHM" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Austin talks with Al Yuen, CEO of PicoJool, about why a 25-year-old technology is the pragmatic answer for scaling up optical interconnects. Al makes the case for GaAs VCSELs over InP single-mode solutions, explains why the VCSEL supply is &#8220;unconstrained&#8221;, and walks through PicoJool&#8217;s newly announced 200-gigabit-per-lane device and its path to 12.8T. </p><p><strong>Things we cover:</strong></p><ul><li><p>Inventing the active optical cable with Mellanox</p></li><li><p>Why hyperscale AI&#8217;s error-free requirement changed the game</p></li><li><p><strong>GaAs (unconstrained) vs. InP (substrate-constrained)</strong></p></li><li><p>The three 1.6T flavors: 8&#215;200, 16&#215;100 LPO, 32&#215;50 NRZ</p></li><li><p>The roadmap to 3.2T and 12.8T via bi-di and 2-D arrays</p></li><li><p>The fabless model and WIN Semiconductors capacity</p></li></ul><div id="youtube2-Ywh2BHKjeHM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;Ywh2BHKjeHM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/Ywh2BHKjeHM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><em>This podcast is lightly edited for clarity.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.chipstrat.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.chipstrat.com/subscribe?"><span>Subscribe now</span></a></p><h2><strong>From HP Labs to the First 10 Gigabit Ethernet</strong></h2><p><strong>Austin:</strong> Hello everyone. Today we have a special guest, Al Yuen, CEO of PicoJool. PicoJool is an optical connectivity company and we&#8217;ll get into all the interesting details. But first I wanted to introduce you guys to Al. So Al, tell us about you and your background &#8212; I know you&#8217;ve been in the industry for a long time.</p><p><strong>Al:</strong> Yes. So after grad school at UC Santa Barbara &#8212; where a lot of the photonics folks have originated &#8212; I went into HP, HP Labs, where we focused on photonics research. And then around &#8216;99 I left HP and started my first company, called Alvesta, and we created the world&#8217;s first 10 gigabit Ethernet. At the time, 10 gigabit &#8212; which is 10 billion bits per second &#8212; we were actually trying to figure out applications, how people would use this in &#8216;99. We would make up things like, people will want to stream video someday in all the rooms in their house. But today, of course, we&#8217;re doing 1600 gigabit, or 1.6 terabits. So since then I&#8217;ve gone to various companies. I ran a division for Coherent, then started some other solar and clean-tech companies, and finally ended up at Lumentum in my last gig. And about two years ago, through Playground Global &#8212; which is our funder &#8212; we started PicoJool. So far so good on what we&#8217;re doing, which is basically in the interconnect space, and we&#8217;ll talk more about that today.</p><p><strong>Austin:</strong> Awesome. Wow, what a great background. I love that you guys invented early 10 gigabit Ethernet and then had to create ideas to sell people &#8212; to convince people that yes, this is useful, people will want to use it. That&#8217;s awesome. Now, remind me &#8212; you also helped invent the active optical cable. Is that right?</p><p><strong>Al:</strong> Sure. Very interesting story. Back in Alvesta, my first startup &#8212; another small startup at the time called Mellanox, which of course now is inside Nvidia and created this whole hyperscale and the whole InfiniBand ecosystem &#8212; they approached us. They had these very bulky copper cables, even back then, 25 years ago. They said, the copper cables are very bulky, they could only reach tens of meters &#8212; now it&#8217;s even shorter. We&#8217;d like an optical option, but we don&#8217;t really want to commit to a full optical solution. So could you put the optics inside the connector? And we said, why not? So we took that same transceiver that&#8217;s typically on a board and literally embedded it right into the connector. Then we thought, well, that&#8217;s not very efficient and clean &#8212; so we embedded the whole thing inside; the optics went inside the connector. And this is the world&#8217;s first demo of an active optical cable, meaning electrical to electrical: electrical comes in, electrical goes out, but inside there&#8217;s the electrical-to-optical transition. Electrons come in, photons carry the information, and electrons go back out at point B. So that&#8217;s how the whole active optical cable concept came through Mellanox. And since then, for the last 25 years, the AOC has been the standard workhorse in many, many data centers.</p><h2><strong>The Engineer&#8217;s Mindset and Why Optics Must Exit the Rack</strong></h2><p><strong>Austin:</strong> Amazing. So having invented that and then watched it become widely adopted &#8212; produced in the millions of cables &#8212; how has that impacted the way you think about what&#8217;s possible as an entrepreneur in this space?</p><p><strong>Al:</strong> I like to think of ourselves more as engineers. There&#8217;s a difference between scientists or researchers &#8212; what people call R&amp;D &#8212; and engineering, or product development. I always tell people I&#8217;m more of an engineer. Engineers solve problems and want to create products that meet specifications. If you need a Prius, you certainly don&#8217;t design a Ferrari &#8212; that&#8217;s overkill. The product fits the need: spec, cost, reliability, and in our case reach, or amount of power. So what we do is really look at solving the specific problem. Today, copper has shrunk to about 3 to 4 meters of reach at 200 gigabits per second per lane. And therefore they can&#8217;t get the information off each rack. Racks and racks of GPUs, CPUs &#8212; lots of compute power &#8212; and these racks are getting very hot because they&#8217;re packing more and more GPUs per rack, because they can&#8217;t exit the rack. To exit the rack, you need an optical solution. There are many technologies, and VCSEL-based active optical cables are one of them &#8212; that&#8217;s what we&#8217;re focused on. Low cost, highly reliable &#8212; going way back to our roots, replacing a copper cable with an active optical cable. So the problem really hasn&#8217;t changed; it&#8217;s just that the speed and aggregate bandwidth is now 1600 times more than the old 1 gigabit Ethernet. It&#8217;s exciting. It&#8217;s been a long journey, but we&#8217;re still at it, and we see a long future for VCSEL-based technology going forward.</p><p><strong>Austin:</strong> Nice. So you have this history of thinking like an engineer &#8212; what is the problem right ahead of us in the industry? Not &#8220;we need to invent new physics,&#8221; but how can we be thoughtful about it. Same cable form factor in the AOC case, electrons in, electrons out, but we could use optics here &#8212; could we put the transceiver on the end of the cable? Thinking very pragmatically. And now, fast forward 25 years, in 2024 you started PicoJool. It sounds like you&#8217;re tackling the problem again of communicating lots of information &#8212; this time GPUs talking to each other, rack to rack, where it&#8217;s such high bandwidth that copper is shrinking to 3 meters or so. So you&#8217;re again thinking about how to use active optical cables, but this time with VCSELs. So tell us &#8212; why VCSELs, and what about the problem made you want to start PicoJool and actually get in the game?</p><h2><strong>Why VCSELs &#8212; 25 Years of Shipping in the Millions</strong></h2><p><strong>Al:</strong> Great question. VCSELs, for one thing, have been around since &#8216;96 &#8212; a technology that&#8217;s been in product, in the field, in data centers since 1996, starting with the first gigabit Ethernet. I&#8217;m more of a historian today: here&#8217;s a 1 gigabit Ethernet that HP created, and at the time Honeywell &#8212; very early companies that did 1 gigabit. And from 1 gigabit all the way to today &#8212; 1600 gigabit, what they call 1.6 terabits &#8212; it&#8217;s all been VCSEL-based. The practicality of these solutions has to meet capacity, demand, cost &#8212; all of it. If you&#8217;re trying to get into a hyperscaler today, you need to meet the whole checklist. You can&#8217;t say, I meet everything but it&#8217;s three times the cost of your target; or everything&#8217;s great but I can&#8217;t get the reach. They want all of it. And right now the only thing that meets it is copper &#8212; mainly driven by cost. Cost for copper is very minimal. Even active copper, where you have some signal integrity and signal processing built in, like an AEC &#8212; active electrical cable &#8212; the cost is still quite minimal compared to more elegant, longer-reach technologies.</p><p>Now, what&#8217;s important in a data center: when you say &#8220;fiber optic communication,&#8221; everybody hears it and goes, oh, isn&#8217;t that the over-the-ocean stuff, the transatlantic subsea cables? Absolutely &#8212; there are many flavors of optical communication; it&#8217;s a very big umbrella. What we&#8217;re talking about today for data centers is a very short reach &#8212; today they call it scale-up. It&#8217;s basically a row, typically 25 to 30 meters, or in English terms 75 feet. Very short &#8212; within your house, from one end to the other, or potentially even shorter. And in that short reach, the shorter the reach, typically the higher the volume. In the data center you&#8217;re connecting all these GPUs, CPUs, and ASICs together &#8212; millions of interconnections, miles of fiber inside a data center that may be a football-field length. But when you leave that data center and go between cities, buildings, countries, you have fewer fibers but need longer distance. So back to the data center: you have to have millions-per-month capacity to meet the volume demand.</p><p>That&#8217;s one of the boxes &#8212; you&#8217;ve got to check off all the boxes. If you demonstrate one or two racks, or go to a show and demonstrate just the technology &#8212; technology, science, R&amp;D show the capability &#8212; that&#8217;s a demonstration. But to ship in volume, millions per month, you need this whole ecosystem: the connector, the transceiver, the sockets &#8212; everything has to be millions per month. And any one of those bill-of-material parts, if it&#8217;s missing, then you can&#8217;t ship in millions per month. That&#8217;s what&#8217;s happening with a lot of these new technologies that are fantastic from a demonstration and capability standpoint going forward. But VCSEL&#8217;s domination is that for the last 25 years we&#8217;ve been shipping in the millions per month. So there&#8217;s no invention of a technology, capacity, or foundries needed &#8212; we already have that. So we design the latest VCSEL &#8212; today it&#8217;s 200 gigabit, which we just announced &#8212; put it into the whole ecosystem, they package it together, and voila, we can build millions per month without waiting for machines, or even buildings, to be built up, then new machines, then new processes. All of that has existed for the last 25 years to service this short-reach data center application.</p><p><strong>Austin:</strong> Okay, let me reflect it back to you. The problem you&#8217;re solving is scale-up with optical interconnects. But the key insight &#8212; with your pragmatic hat on &#8212; is: how can we use technology and a supply chain that already exists and can already ship millions of cables, components, and parts per month? That&#8217;s one reason VCSELs are so attractive &#8212; they&#8217;re not a new technology, the supply chain is not new. So even though we hear about all these other interesting things &#8212; the Broadcoms, Coherents, Lumentums; silicon photonics, EMLs, micro-LEDs &#8212; you&#8217;re saying, hey, don&#8217;t rule out VCSELs. Even though they&#8217;re not new, there are advantages. But if VCSELs have been around so long, why aren&#8217;t other folks trying to take them to 200 gig per lane, times eight lanes, 1.6T and beyond?</p><h2><strong>The Hyperscale Shift to Error-Free</strong></h2><p><strong>Al:</strong> Yeah, like you said, there are multiple technologies &#8212; silicon photonics, EML, even micro-LEDs &#8212; all vying for this 200 gigabit electrical signal coming in. From the ASIC, GPU, or CPU you have 200G per lane of electrical signal. When you go to the electrical-to-photonic (E-to-O) transition, it can go to different lanes through a different IC that may be a gearbox. You don&#8217;t necessarily have to run at 200G straight through &#8212; which we can &#8212; and that&#8217;s the most elegant: straight through without an intermediary gearbox saves power, cost, and latency. So 200 gigabit would be ideal.</p><p>But what&#8217;s changed in hyperscale is different from Ethernet. Ethernet information is sent in packets. Whatever information &#8212; &#8220;where should I go in Italy, I&#8217;m going on vacation&#8221; &#8212; gets divided up by your search engine, sent in packets, and comes back to you. We&#8217;re not very cognizant if there&#8217;s an error drop or some delay, because those are in the hundreds of nanoseconds or milliseconds &#8212; to us it&#8217;s a thousandth of a second, we don&#8217;t notice. But for hyperscale systems, which are literally thousands of GPUs acting as one brain &#8212; one supercomputer, one high-performance cluster &#8212; that latency is super critical, because the GPU notices anything in the tens or hundreds of nanoseconds. So the typical Ethernet bit error rate needs to drop from 10&#8315;&#8310; to below 10&#8315;&#185;&#8304;, what we call error-free, because any errors slow down the whole training and inference &#8212; all the AI infrastructure. That&#8217;s changed, and it&#8217;s allowed these higher-cost single-mode solutions &#8212; single-mode being longer distance, longer reach, very high performance &#8212; to come into the data center and hyperscale systems, because of this requirement for very low bit error rate. And now VCSELs have to raise the bar from 10&#8315;&#8310; (typical Ethernet the last 25 years) to 10&#8315;&#185;&#8304;, 10&#8315;&#185;&#178;. And we&#8217;ve done that &#8212; we&#8217;ve pushed our VCSELs to 200 gigabit, and even used at 100 or 50 gigabit NRZ you can leverage that to very low bit error rate. So the answer is: things changed for hyperscalers to require error-free, and that&#8217;s allowed these high-end single-mode solutions to compete directly with VCSELs, because of that additional spec that&#8217;s new to hyperscale AI systems.</p><p><strong>Austin:</strong> Okay. So because the capacity for sustaining errors is much lower &#8212; we don&#8217;t want all these GPUs just waiting &#8212; that changes the game from the cloud/SaaS days to now, when everything&#8217;s acting as one big computer. So the bit error rate has to be much lower. And VCSELs, which are short reach, have to come down to meet that; or, what you&#8217;re saying is, these other technologies that are longer reach and higher power were already closer to the necessary bit error rate, so people say, why don&#8217;t we bring those into short reach? But surely that has power and cost tradeoffs &#8212; taking something that talks at low error rate over a long distance and trying to bring it in. So you guys must be taking a different tack: no, no, let&#8217;s just make VCSELs error-free. And presumably that&#8217;s a cost or power tradeoff you&#8217;d rather make &#8212; or is it back to the manufacturing supply chain capacity?</p><h2><strong>One Pizza Oven vs. 5,000 Pizzas &#8212; The Capacity Argument</strong></h2><p><strong>Al:</strong> Yeah, it&#8217;s again a complicated matrix of items you have to check off. Those shipping silicon photonics and EML for long reach &#8212; or DFBs &#8212; single-mode, high-performance devices, have been around a similar time, 25 years. They&#8217;ve been used for long reach because of that super-high performance, very few fibers, and what we call WDM &#8212; wavelength division multiplexing. Fibers are very expensive when you&#8217;re going hundreds of kilometers, so you want to use very few fibers but pass more information through more wavelengths, more colors, in that same fiber. For short reach &#8212; tens of meters &#8212; you&#8217;re not as locked into the cost of the fiber; it comes down, because you&#8217;re only 10 meters versus 10 kilometers. So the volume supply chain of the high-performance, low-bit-rate device would seem a natural fit: hyperscalers want low bit rate, let&#8217;s go with the Ferrari &#8212; the super-high-speed, high-performance single mode. But those players have been used to building in maybe 100K &#8212; 100,000 &#8212; volumes, because you don&#8217;t need as many between cities and countries. And all of a sudden they come into the data center. Even though their performance is excellent, the cost is a little higher because of the single-mode packaging. And the infrastructure to build millions per month is 10, 20, 50&#215; what exists. That&#8217;s brick and mortar. It&#8217;s like: I only have one pizza oven, and I&#8217;ve been used to a small clientele &#8212; maybe 20 an hour. Someone comes in and says, I&#8217;d like to order 5,000 pizzas, and I need them in an hour. You&#8217;re going, I need 100 pizza ovens. That&#8217;s the exact same problem the high-performance single-mode long-reach players are coming into. Silicon photonics, EML &#8212; excellent technology, very good bit rate &#8212; but the infrastructure needs to be built up. That&#8217;s what you&#8217;re seeing: a lot of announcements with people holding shovels, saying we&#8217;re investing in the next supply chain, the buildings. That&#8217;s great &#8212; bringing more manufacturing to the world, and back to the US, all great for the industry. But VCSELs have been around for 25 years and shipping in the millions. So we&#8217;re not building, we&#8217;re not putting shovel to ground &#8212; we&#8217;re just changing the actual VCSEL.</p><p>We&#8217;re just changing the actual VCSEL performance &#8212; the chip &#8212; and leveraging the existing infrastructure. And so for VCSEL capacity, we call it unconstrained. Constrained means they&#8217;re sold out: we&#8217;re sold out through next year; if you&#8217;re going to place an order, it&#8217;s going to be an eight-month, 18-month lead time &#8212; a year and a half from now we can get it to you. For us, unconstrained just means we have a certain lead time that&#8217;s limited only by our cycle time of building through the factory. A VCSEL run and then a packaging run may be four, eight, twelve weeks &#8212; but that&#8217;s limited just by the fact that we have to build it and ship it, not by the constraint of the supply and ecosystem of machines, or pizza ovens. We have plenty of pizza ovens: place the order, and we&#8217;ll get you your order in the cycle time we commit to.</p><p><strong>Austin:</strong> Gotcha. Okay, that&#8217;s very interesting, and a great competitive advantage for you. So on the constrained side &#8212; is that what listeners hear about when they hear about indium phosphide being a bottleneck? Where in the supply chain is it constrained? And for you, what&#8217;s different about VCSELs that makes them unconstrained?</p><h2><strong>InP vs. GaAs and the Fabless Split</strong></h2><p><strong>Al:</strong> Yeah. So a lot of silicon photonics and EMLs are based on this material, indium phosphide. For VCSELs, our technology has always been gallium arsenide. For the listeners it may be, okay, one III-V compound versus another &#8212; what&#8217;s the difference? Indium phosphide is material-constrained from the very beginning. You can&#8217;t even get a base substrate &#8212; even before you process it, just the substrates for indium phosphide are limited, before you get it made into a product, whether that&#8217;s EML, silicon photonics, or VCSELs. The indium phosphide bare material is already limited. Gallium arsenide: unconstrained. So we start with that. And then you go through the fabs &#8212; fabrication depends on foundries, usually very large companies. Companies like PicoJool and others don&#8217;t own large clean-room factories. We design the individual VCSEL chip device, and then we use foundries to manufacture it. Those foundries are available &#8212; but they can&#8217;t get enough indium phosphide starting material. Now, after that, once you get to the chip level &#8212; you dice it up, okay, great, I&#8217;ve got the laser, I&#8217;m ready to go &#8212; to get from the chip to a pluggable device, an actual optical engine, there&#8217;s a ton of stuff that happens. You&#8217;ve got laser drivers, you&#8217;ve got boards, and for single mode you have to have all the machines that align that silicon photonics or EML to a very, very small-core single-mode fiber. Those machines have to be readily available. VCSELs, again, have that millions-per-month volume infrastructure &#8212; it doesn&#8217;t need to be built up. So from the very beginning: indium phosphide material constraint, then you have to build it into lasers, then finally package it into transceivers &#8212; and all along that supply chain, it&#8217;s not used to building millions per month. All of that has to be built up: hardware, alignment machines, testers. VCSELs have all of that infrastructure existing already.</p><p><strong>Austin:</strong> I see &#8212; yeah, that makes a ton of sense. And I love your props, by the way &#8212; I liked the little VCSEL you held up. So tell us more: what are we looking at, and walk us through exactly what you design, and where it gets handed off, built, and packaged &#8212; where your responsibilities end.</p><h2><strong>The VCSEL Tree of Knowledge and the Handoff to WIN</strong></h2><p><strong>Al:</strong> Yeah &#8212; so background is super important here. There&#8217;s what we call a tree of knowledge of VCSELs. You have this line from Honeywell through Finisar, and Finisar goes into II-VI, which goes into Coherent &#8212; so there&#8217;s the Coherent line. Then for us, obviously, HP went into Avago, went into Broadcom &#8212; the HP&#8211;Broadcom line. And finally there&#8217;s Picolight and E2O &#8212; I&#8217;m throwing in some old names from 25, 30 years ago &#8212; which go into JDSU, another big name from the dot-com time, and JDSU spins off Lumentum and Viavi. We&#8217;re from the Lumentum arm. All three of these major arms have a lot of VCSEL knowledge &#8212; and the strength of PicoJool is that we&#8217;ve tapped into designers from all three branches. Imagine all that know-how &#8212; again, not patented; know-how, recipes. I use this example: you can hand three different chefs a recipe for a souffl&#233; and you&#8217;ll most likely get three different souffl&#233;s, because it&#8217;s really difficult to get a perfect one. It&#8217;s not just crack the eggs, beat the eggs. Same thing in the VCSEL.</p><p>So to answer your question &#8212; what do we do? We take all that know-how and design the epi layers. All these little lines you see &#8212; these are epi layers designing this vertical cavity. It&#8217;s not edge-emitting, where the light comes out of the edge of a flat chip; VCSELs are vertical-cavity, surface-emitting &#8212; the light comes out of the surface. We design the internal cavity of the laser, all the dopings, all the process. After we design, we hand it to an epi foundry that grows the material and gives us back an unprocessed epi wafer. We&#8217;ve been working with WIN Semiconductors in Taiwan &#8212; that&#8217;s our foundry. Once the epi wafer is ready, we hand it to WIN, they run it through their clean-room process, and they make the final VCSEL device. And here&#8217;s another beauty of VCSELs versus an edge emitter: at the wafer level you can start testing and probing each one &#8212; 100% tested, what they call known good die &#8212; before you have to singulate and dice it into arrays. That advantage is huge, because the more work you add before you yield the device, the more value you lose downstream. You always want to yield upstream. Wafer-level testing for VCSELs is a real advantage versus edge-emitting technologies.</p><p>And WIN is only the wafer processing. The wafer comes out and is diced into individual VCSELs, and we work with our partners to build those into either active optical cables or transceivers. That&#8217;s another foundry, if you will, but a packaging company. That said &#8212; companies like TSMC are now going into co-packaging: after they make their silicon wafer, they&#8217;ll package the optics directly on top of it. There&#8217;s a whole emerging field called CPO, where the traditional wafer-processing foundries are stacking technologies together &#8212; 3-D wafer-level packaging. For us: we only do the wafer at WIN, and then that WIN VCSEL goes to our module-integrator partners, who build it up into the active optical cables or transceivers.</p><p><strong>Austin:</strong> Gotcha, that&#8217;s helpful. So take us back to your roadmap. I know you mentioned a 50G version, 100G, 200G &#8212; and a recent launch. Tell us more about your roadmap, what you launched, what you announced.</p><h2><strong>200G/Lane and the Three Flavors of 1.6T</strong></h2><p><strong>Al:</strong> As I mentioned early on, 200 gigabit per lane is the benchmark &#8212; the bar you have to clear. EMLs and silicon photonics have all done that, and now VCSELs have reached it. PicoJool just announced our 200 gigabit; we&#8217;ll start sampling next quarter. That&#8217;s for a very simple transceiver where you have eight channels of 200 gigabit coming in &#8212; the aggregate bandwidth of 8&#215;200 is 1600 gigabit, or 1.6 terabits.</p><p>So that&#8217;s the standard ramping today. There are 800 gigabit transceivers as well, that&#8217;s also shipping &#8212; that&#8217;s 8&#215;100. And the next generation, today&#8217;s generation, is 8&#215;200 &#8212; the 200 gigabit VCSEL we announced. But there are many flavors of that because of the specification. Aggregate bandwidth 1.6T &#8212; check. But there are different ways to get there if you want very low power or very low bit error rate.</p><p>Running 200G, I liken it to a Ferrari &#8212; it can give you 200 miles per hour, but it&#8217;s very high-end, relatively expensive, because you need certain signal integrity and signal processing. So now I say, I want very low cost, low power, but I still want low bit error rate. And what people have done is: let&#8217;s slow it down. Let&#8217;s use the 200G-performance VCSEL, but run it at 100 gigabit.</p><p>So now you have an excellent VCSEL that gives you a lot more performance, and if you use it at half the speed, you get much better bit error rates &#8212; the signal-to-noise, or relative intensity noise, drops as well. That&#8217;s 100G. But you need more lanes &#8212; to get to 1600 you need 16 lanes of 100. And recently something came out called micro-VCSELs.</p><p>And that goes even slower, down to 50G &#8212; they call it NRZ. So instead of PAM4, which has four levels (0, 1, 2, 3), we go back to the original NRZ, which is 0 and 1. Now you use the entire 0-to-1 signal-to-noise, which again reduces your bit error rate. But you need more channels &#8212; 32 channels at 50G to get to 1.6T.</p><p>But we&#8217;re shipping all three, and all three are in demand depending on whether customers want what they call fast and narrow (8&#215;200); or an LPO &#8212; linear drive, no DSP, low power &#8212; which is 16&#215;100G; or, if they want really low bit error rate and very low power, the 32&#215;50G NRZ, which they call slow and wide.</p><p>We tend to call it &#8220;fast and wide&#8221; and &#8220;faster and narrow.&#8221; When you&#8217;re in high-speed interconnect, we try not to use &#8220;slow&#8221; in any of our marketing.</p><p><strong>Austin:</strong> That&#8217;s good, that&#8217;s good. Okay, interesting &#8212; this is really cool. So you&#8217;re saying there are many ways to get to 1.6T. You could have 8&#215;200, which uses PAM4 and requires a lot of DSP and power, but it&#8217;s definitely possible. Or 16&#215;100 or 32&#215;50, and you need less DSP for each of those &#8212; the 32&#215;50 has much less because it&#8217;s NRZ. Yeah, this is all fascinating. So will that approach still hold once you move to 3.2T? Is it going to be that same combination of possibilities?</p><h2><strong>The Roadmap to 3.2T and 12.8T</strong></h2><p><strong>Al:</strong> Great question &#8212; people always ask, what&#8217;s the roadmap ahead, what&#8217;s the future, is this the end of the road? Okay, 1.6T, VCSELs can do it &#8212; but is there a 3.2T? A 6.4T? A 12.8T? I mean, Andy Bechtolsheim &#8212; notorious &#8212; he&#8217;s created an XPO that&#8217;s literally going to give you 12.8T in a big pluggable today.</p><p>So they&#8217;re planning way ahead, because no one has ever told us in the last 30 years, &#8220;whoa, we have way too much bandwidth.&#8221; We have to have that bandwidth. And exactly like you said &#8212; what&#8217;s the future? Number one, we can use something called bi-di, bidirectional. Meaning we just add another wavelength &#8212; not the complexity of WDM where you have eight or 16 wavelengths like single mode.</p><p>We basically just add another wavelength to our existing one. So two wavelengths, passing them in both directions &#8212; bidirectional. That doubles the bandwidth without changing anything else, except you add another laser at a different wavelength, and you leverage the entire ecosystem. So from 1.6 to 3.2, we could add another wavelength. The other way, of course, is to double the speed. Can we do 100G NRZ?</p><p>That&#8217;s in the works. We&#8217;re developing 100G NRZ &#8212; today is 50G NRZ, but we&#8217;re developing 100G NRZ, leveraging our 200 gigabit VCSEL running at 100G NRZ. And in the future, we can go to more channels. The beauty of VCSELs again &#8212; surface emitting. Edge emitters can only have a one-dimensional array &#8212; a 1&#215;4, 1&#215;8, 1&#215;12 &#8212; it just makes a long bar.</p><p>But for surface emitting, we can have a two-dimensional array &#8212; 2&#215;4, 2&#215;12, 2&#215;16 &#8212; and couple all the light very elegantly with an optical fiber bundle. In that case I&#8217;m kind of unlimited. I can go up to 64 channels today in a 4&#215;16 connector that&#8217;s the size of &#8212; let me show you. A 4&#215;16 fiber is this size &#8212; here&#8217;s my finger. And that has 64 channels in it. If I run them at 200, that gets me to 12.8T. So in essence, the technology of today &#8212; without having to go to 400G per lane, which we&#8217;re also looking at &#8212; but at 200G, with more channels, with more colors (another color for bi-di), you double, you triple, by size. So that roadmap to 12.8T, we believe, is very solid and very clear &#8212; without even having to invent any new technology. And then with new technology, it just gets easier, if you can do 400G per lane.</p><p><strong>Austin:</strong> Sure. Fascinating. So it&#8217;s just the same 200G VCSEL over and over &#8212; put it in an array and you get more of those. Or are you having to invent a new VCSEL to get the 100G version to run at NRZ?</p><p><strong>Al:</strong> We&#8217;re just starting tests. We believe the 200G VCSEL has the capability to go to 100G NRZ. So it&#8217;s not a new VCSEL &#8212; it&#8217;s just a different coding on the signal coming in, running at NRZ instead of PAM4, to take advantage of the 0-to-1, using the whole signal-to-noise ratio as one bit as opposed to four levels. So that&#8217;s the difference.</p><h2><strong>Lead Times, Yields, and WIN&#8217;s Capacity</strong></h2><p><strong>Austin:</strong> Gotcha. Yeah, this goes back to your engineering, pragmatic mindset &#8212; taking a Lego block and figuring out different ways to place it, different ways to use it, to unlock this whole roadmap. That&#8217;s pretty cool. So you talked about unconstrained gallium arsenide and working with WIN &#8212; they&#8217;re used to making this stuff, they&#8217;ve got all the pizza ovens they need. So if a big hyperscaler comes to PicoJool and says, we want a million of your pizzas &#8212; what does that look like? How does that actually happen?</p><p><strong>Al:</strong> Right. If they want a million VCSELs, we&#8217;d typically give them an eight- to ten-week lead time. That&#8217;s the basic. We can accelerate it &#8212; push and have engineering carry certain wafers &#8212; but typically eight to ten weeks on the VCSEL side. You&#8217;ll get a wafer, or individually diced VCSELs. If you want transceivers, then that eight weeks tags on a certain number of weeks to package it all into the transceiver.</p><p>Those are constrained strictly by the packaging process &#8212; not by ordering equipment or building capacity. That ecosystem works, and it leverages the typical cycle time of building out. Typical cycle time: eight weeks for the VCSEL device, then four to six weeks for the module after that. So you&#8217;re looking at anywhere from 12 to 16 weeks to get to the full module, starting from an epi reactor growing epi and going all the way through.</p><p>And we&#8217;ll continue to drive that lower &#8212; lead time or cycle time. But also yields, which I didn&#8217;t mention much. Yields are: if I make a VCSEL, how many known-good die can I get out of a wafer? The higher the yield &#8212; up to 100% &#8212; the fewer wafers I have to run, and the higher the capacity of the factory. If every wafer goes through and I get 100%, I need fewer wafers, and therefore less capacity for the demand. Obviously getting to 100% is very hard, but VCSELs have been perfecting that process for many years, for many decades &#8212; and now we&#8217;re leveraging all of that. It&#8217;s not new, it&#8217;s not something that has to be established, it&#8217;s not based on new technology. That&#8217;s very important. WIN has been doing it for about 10 years, since we transferred that for a consumer electronics application back in 2016. They have a capacity of up to a thousand of these wafers per week. And there are about 240,000 VCSELs on each wafer, because they&#8217;re really tiny. So that adds up to a million yielded from maybe 10 wafers. The capacity is huge for datacom. So I think we have no worries &#8212; once we get to the 200G, the product specs are met with the customer, and reliability qualification is done, then we just ramp readily with WIN, and they&#8217;re ready to go.</p><h2><strong>The Fabless Model &#8212; Stay Small, Leverage the Supply Chain</strong></h2><p><strong>Austin:</strong> Nice. Amazing. It&#8217;s quite compelling. Normally when people hear &#8220;there&#8217;s a startup trying to compete in a space with these huge incumbents,&#8221; the question is: how&#8217;s this startup going to compete? How do they get to market, find customers, build up supply chain? But what I hear you saying is that you&#8217;re taking industry veterans with process know-how &#8212; similar ways of thinking as competitors, you&#8217;ve been in the game a long time &#8212; and tapping into an existing supply chain. And at the end of the day, it&#8217;s not like you have to win 50 customers; there&#8217;s a handful of big customers that would really make a difference for PicoJool if they said yes. But the most important point is the unconstrained gallium arsenide &#8212; being able to make a million, 10 wafers with 240,000 on each, whatever you yield, we&#8217;re talking millions of VCSELs very quickly. Because across all of semiconductors &#8212; memory, CPUs, AI accelerators &#8212; there&#8217;s so much demand and such constrained supply that, yes, you want to compete on cost and engineering performance, but there&#8217;s also a bit of: if it&#8217;s good enough and it&#8217;s in production and you can install it into my data center, game on. So it feels like you have a strategy that lets you deliver shipped VCSELs as soon as possible.</p><p><strong>Al:</strong> Yeah, the model has been around for decades in silicon. I&#8217;m in Palo Alto, in Silicon Valley. Most chip companies designing integrated circuits, CPUs, GPUs, don&#8217;t have their own foundry. A lot of people use Intel; even AMD uses TSMC. AMD is a huge chip company and they don&#8217;t have their own foundries today. TSMC, Intel, Global Foundries &#8212; many foundries are the factory floor, the clean rooms, of all these startups. So just because of our small startup size doesn&#8217;t mean we can&#8217;t ship in the millions per month and compete directly with the very large presence of other optical suppliers and competitors. That&#8217;s the beauty &#8212; we can stay very lean. Our motto is &#8220;stay small,&#8221; and that has to do with many things &#8212; our name is PicoJool, right? Very, very low power, and that&#8217;s one of the things driving us. We&#8217;re a well-experienced small team, but we get huge benefit by leveraging WIN Semiconductors and working with our supply chain. They&#8217;ve got the factories and the clean rooms, and we can ramp very quickly by providing our designs and our unique specialty, then partnering with these large companies that are already shipping &#8212; and they just drop-ship. So we don&#8217;t need a big company to ship in the millions per month.</p><p><strong>Austin:</strong> That&#8217;s amazing. Definitely punching above your weight &#8212; that is awesome. So when is your high-volume ramp targeted for? If I recall, the press release said something &#8212;</p><h2><strong>Ramp Timing &#8212; Sampling Now, HVM in Early 2027</strong></h2><p><strong>Al:</strong> Yes, the press release said we&#8217;re starting to sample next quarter. And that&#8217;s in different flavors &#8212; we&#8217;ve got customers for the 50G NRZ, the 100G LPO, and the 200G. So all of those begin to sample. And the process from our device to actual product shipment &#8212; revenue, in our case &#8212; is a period they call qualification, or reliability testing. Everything has to not only meet spec at zero hour, but be predicted to last 10 years, or a number of years, through what they call accelerated aging &#8212; they test it at higher temperature and higher bias conditions, then estimate back to normal operating conditions: can it last 10 years in the field? For very small, tier-two customers, that can be as short as three months. But for tier one &#8212; because they have much more to lose if there&#8217;s an issue with their connection &#8212; it takes more than six months to get up and running. So ramping at WIN is available today; but we have to go through this qualification cycle with our customers before they give the orders and everything is approved. So we&#8217;ve also said we&#8217;ll most likely start ramping in early 2027, next year.</p><p><strong>Austin:</strong> Okay, got it. Thank you for the education here. So &#8212; sampling, then the qualification process, then ramping. And probably, as you&#8217;ve been saying throughout, you&#8217;re not concerned about the ramping. You had to build the product and let customers kick the tires, and once they say let&#8217;s go, it&#8217;s off to the races. Okay, awesome. Well, we&#8217;ve covered so much &#8212; this has been amazing. I&#8217;ve learned a lot, and I know the listeners will have too. Is there anything else, any last things about PicoJool, or anything we didn&#8217;t talk about that you were hoping to cover?</p><h2><strong>Mentoring the Next Generation of Photonics Engineers</strong></h2><p><strong>Al:</strong> One thing that&#8217;s very interesting to me &#8212; as you can see from the image, I&#8217;m a very experienced, elderly startup person. One thing to point out: very few people went into hardware and photonics in our space, because young people over the last 25 years &#8212; literally since dot-com &#8212; came into the workforce more on the application side, the software side. People wanted to go into computer science. So the aging, experienced hardware folks in photonics need to transfer all this knowledge. That&#8217;s one of my passions &#8212; to bring on the next generation, and the generation after that, for photonics. Because just like VCSELs and other technology, we see many decades ahead, and as I said, no one&#8217;s saying &#8220;way too much bandwidth.&#8221; We see bandwidth demand increasing with robotics and autonomous vehicles and what have you &#8212; everything is going to be bit- and connectivity-constrained. So we want to spend time to educate and train. We&#8217;re trying to hire hardware engineers, and train young folks &#8212; maybe without the experience &#8212; to be the VCSEL designers and transceiver designers of the future. That&#8217;s really exciting for us, because we&#8217;ve got all this knowledge, 10, 20, 30, 40 years, and it feels wonderful to have the hardware excitement again &#8212; not only in the markets, but the investment community. Silicon Valley is booming with photonics and hardware. So we don&#8217;t take it for granted. It&#8217;s a great opportunity, and we definitely want to take young entrepreneurs, young engineers, folks interested in this space, along for the ride &#8212; and then they take it from there.</p><p><strong>Austin:</strong> I love it. Very inspiring, very cool. It&#8217;s never been a better time for interconnects, photonics, and optics folks. I love that you industry veterans want to bring up the next generation and give them the opportunity to learn from folks like you &#8212; to revitalize, rebuild, make sure we have a reinvigorated workforce, so that for my generation and the generation of my children, they can keep having more and more data moved around, faster and faster.</p><p><strong>Al:</strong> That&#8217;s right. Yeah, really enjoyed our conversation, Austin. Thank you.</p><p><strong>Austin:</strong> Awesome. Thank you, Al. Appreciate it.</p>]]></content:encoded></item><item><title><![CDATA[The Next Trillion-Dollar Chip Company]]></title><description><![CDATA[It might be simpler than you think. Etched, Fractile, MatX, Positron, Rebellions, SambaNova, Tensordyne, Tenstorrent, and more.]]></description><link>https://newsletter.chipstrat.com/p/the-next-trillion-dollar-chip-company</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/the-next-trillion-dollar-chip-company</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Wed, 15 Jul 2026 22:20:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!IJnR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Two AI chip startups had huge financial outcomes in the past year. Groq was acquired for $20B, and Cerebras IPO&#8217;d to a $40B+ market cap. <em>Huge, but not trillion-dollar huge.</em> And in this environment, an inference-first system that runs frontier models and scales to gigawatts of compute has a real shot at a trillion-dollar market cap if it can scale, grow, and IPO.</p><p>But Groq and Cerebras run architectures you can poke holes in. Both were designed before LLMs went mainstream, and their early design choices don&#8217;t suit today&#8217;s models. <em>I call them &#8220;preGPT&#8221; accelerators.</em></p><p>The famous example (<em>well-covered ground already</em>) is that they&#8217;re SRAM-only. It takes many Groq racks to serve even one smallish model, and even Cerebras&#8217; wafer-scale marvel (<em>900,000 cores on one piece of silicon</em>) can&#8217;t hold a frontier model&#8217;s weights on a single wafer, so you have to expand to many wafer-level systems. <em>Tough economics, a KV cache that doesn&#8217;t scale gracefully, and so on.</em></p><p>Admittedly, I looked at these engineering details and figured these companies were dead in the water when it came to LLMs. <em>Well, to their credit, who could have predicted such thicc models 10 years ago? But no HBM to support long-context... how will they survive? More context is better&#8230;. failure to thrive?</em></p><p>What&#8217;s important is that they both had production silicon available, and, regardless of the shortcomings of SRAM-only architecture, they outperformed GPUs on a certain KPI (<em>interactivity</em>). As Jensen nicely drew for us, SRAM chips for decode unlocked new possibilities on the Pareto frontier <em>for 2026</em>:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!V92f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!V92f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png 424w, https://substackcdn.com/image/fetch/$s_!V92f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png 848w, https://substackcdn.com/image/fetch/$s_!V92f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png 1272w, https://substackcdn.com/image/fetch/$s_!V92f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!V92f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png" width="1456" height="843" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/890fac5a-d8af-4fe7-b820-138084397609_1456x843.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:843,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Nvidia's Pareto frontier: SRAM decode extends throughput at high interactivity&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Nvidia's Pareto frontier: SRAM decode extends throughput at high interactivity" title="Nvidia's Pareto frontier: SRAM decode extends throughput at high interactivity" srcset="https://substackcdn.com/image/fetch/$s_!V92f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png 424w, https://substackcdn.com/image/fetch/$s_!V92f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png 848w, https://substackcdn.com/image/fetch/$s_!V92f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png 1272w, https://substackcdn.com/image/fetch/$s_!V92f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F890fac5a-d8af-4fe7-b820-138084397609_1456x843.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><a href="https://www.chipstrat.com/p/the-multi-silicon-era-is-here">Source</a></figcaption></figure></div><p>And that&#8217;s why they had eleven figure outcomes. <em>$XX Billion, kind of nuts right! </em></p><p>I was so wrong about architectural shortcomings preventing success.</p><p><strong>Looking back, what mattered most?<mark data-color="rgb(214, 240, 204)" style="background-color: rgb(214, 240, 204); color: rgb(0, 0, 0);"> Timing.</mark></strong></p><p>Groq was in production right when Nvidia came calling. Cerebras, right when OpenAI did. Forget the architectural shortcomings. <mark data-color="rgb(214, 240, 204)" style="background-color: rgb(214, 240, 204); color: rgb(0, 0, 0);">If you can ship and unlock a new Pareto frontier, good things happen.</mark> In fact, I thought the fundamental problem for Groq and Cerebras was being too early, making design decisions before transformers took off. <em>Being too early is indistinguishable from being wrong...</em> </p><p>And yes, starting that early nearly killed them; both came close to running out of cash. But starting early was also the whole advantage, because when the unforeseen inference wave hit, they had silicon in production. <em>Well, they were both still default dead until Meta released Llama, the first useful open weights LLM, which allowed Groq to go show the world how awesome a super fast chatbot experience was.</em></p><p>Again, the combination of production silicon at the right moment and a Pareto frontier the GPU can&#8217;t reach is all that mattered. And then Nvidia&#8217;s Dynamo was icing on the cake, disaggregating the workload so the memory-bound decode could run on SRAM chips.</p><h2>The postGPT era</h2><p>If the first era of inference was GPUs, Groq and Nvidia ushered in the next era with the GPU+LPU inference. But another era is incoming. <mark data-color="rgb(214, 240, 204)" style="background-color: rgb(214, 240, 204); color: rgb(0, 0, 0);">The postGPT accelerators.</mark> Rack-scale systems specifically designed for LLM inference from day one. From matmuls to interconnects to memory hierarchy.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WoeQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WoeQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png 424w, https://substackcdn.com/image/fetch/$s_!WoeQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png 848w, https://substackcdn.com/image/fetch/$s_!WoeQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png 1272w, https://substackcdn.com/image/fetch/$s_!WoeQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WoeQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png" width="1456" height="558" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:558,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The inference compute stack, in three eras&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The inference compute stack, in three eras" title="The inference compute stack, in three eras" srcset="https://substackcdn.com/image/fetch/$s_!WoeQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png 424w, https://substackcdn.com/image/fetch/$s_!WoeQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png 848w, https://substackcdn.com/image/fetch/$s_!WoeQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png 1272w, https://substackcdn.com/image/fetch/$s_!WoeQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d851bbb-8298-4ca4-a540-1ff9423eeb72_2400x920.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Inference compute in three eras</em></figcaption></figure></div><p>There are many startups in this space, not to mention custom silicon XPUs (<em>Maia, MTIA, Trainium, etc</em>). So which of these could be the next Groq/Cerebras with a whale of a customer and an acquisition/IPO?</p><h2>What it takes to be the next trillion-dollar chip company</h2><p>If we&#8217;ve learned anything so far, it&#8217;s that we shouldn&#8217;t overindex on engineering architectures. Sure they matter in the long run, but in this demand environment, contenders simply need to hit these four criteria:</p><ul><li><p><strong>Runs frontier 1T+ models</strong></p></li><li><p><strong>Ships rack-scale</strong></p></li><li><p><strong>Beats the incumbent on one KPI</strong></p></li><li><p><strong>Lands a frontier anchor</strong></p></li></ul><p><strong>Runs frontier 1T+ models.</strong> Frontier models are where the value gets captured and where scale happens, so frontier customers are the ones most willing and incentivized to deploy a postGPT accelerator at volume. <em>A system that can&#8217;t run a 1T-param model is playing a sub-frontier game. I&#8217;m bullish on the usefulness of sub-frontier, especially in the enterprise, but the SAM is small. They won&#8217;t mint the next Groq/Cerebras.</em></p><p><strong>Ships rack-scale.</strong> The race to gigawatt-scale inference runs on racks. To win, one must design and deploy racks of accelerators, quickly and reliably.</p><p><strong>Beats the incumbent, clearly, on one KPI.</strong> The challenger doesn&#8217;t need to be best at everything. But they must be an order of magnitude better at something.</p><p><strong>Lands a frontier anchor.</strong> The challenger needs someone serious about gigawatt-scale and buying into a roadmap and building a long-term relationship. A model lab or hyperscaler is best. Could be a big neocloud. Probably not a sovereign.</p><p>Of course, software support is table stakes.</p><p>But otherwise, I think it&#8217;s that simple.</p><p>Yes, the technical details matter; quantization, memory hierarchy, scale-up domain size, etc. But truly, right now they matter only insofar as they lead to differentiated performance (&#8221;beats the incumbent on one KPI&#8221;). The architecture behind each contender is in the teardowns; up here, judge them by the KPI it buys.</p><p>This is a unique time. Token demand FAR exceeds supply, and the shortage is worst for low-latency frontier tokens, where GPUs can&#8217;t keep up or can only do so at uneconomical throughput. This high-interactivity token scarcity turned &#8220;good enough architecture in production&#8221; into a winning hand for Groq and Cerebras.</p><p>But it&#8217;s time-bound; there was a window, and Groq/Cerebras grabbed it. And being first matters.</p><p><strong>Which raises the next question: what&#8217;s the next window, and who will get there soonest?</strong></p><p>Picking the next winner could be as simple as asking <strong><mark data-color="rgb(214, 240, 204)" style="background-color: rgb(214, 240, 204); color: rgb(0, 0, 0);">who can deploy a gigawatt of postGPT inference accelerators soonest?</mark></strong></p><p>To answer that, here&#8217;s the rest of the piece.</p><ul><li><p><strong>The frontier race</strong>, all eleven ranked by when their frontier rack ships, as a scoreboard and a timeline</p></li><li><p><strong>A full teardown of Tenstorrent</strong>, the only one serving a frontier-class rack today, a free sample of the diligence</p></li><li><p>[Paid] <strong>The other ten teardowns</strong>, each scored and sourced, the frontier contenders, the decode add-in, and the sub-frontier plays</p></li><li><p>[Paid] <strong>The receipts</strong>, only two of the eight frontier names have a clean third-party number</p></li><li><p>[Paid] <strong>Who actually looks like the next trillion-dollar chip company</strong>, and the argument that reframes the whole list</p></li></ul><h2>The frontier race: who reaches rack-scale, soonest</h2><p>Well then: &#8220;Which startups clear those four hurdles, sorted by ship date soonest?&#8221;</p><p>Let&#8217;s take a look. I&#8217;ve sorted them by when they <em>claim</em> their frontier-class generation hits rack scale. <em>Take it with a grain of salt, and remember this isn&#8217;t yet anything about performance, but just when they say they&#8217;ll have silicon deployed.</em></p><p><em>Also note that many of these companies have shipped products already, but the earlier products were sub-frontier as they pivoted their way toward rack scale for modern frontier LLMs. So I&#8217;m only focusing on when they are shipping frontier-capable products according to their roadmap</em></p><p>If we scour all the public sources (press releases, blogs, podcasts), we get the following:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RRn1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RRn1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png 424w, https://substackcdn.com/image/fetch/$s_!RRn1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png 848w, https://substackcdn.com/image/fetch/$s_!RRn1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png 1272w, https://substackcdn.com/image/fetch/$s_!RRn1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RRn1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png" width="1456" height="846" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da982695-94c2-4371-90ef-49e6a926e257_2220x1290.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:846,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The frontier race: first frontier rack, and to whom&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The frontier race: first frontier rack, and to whom" title="The frontier race: first frontier rack, and to whom" srcset="https://substackcdn.com/image/fetch/$s_!RRn1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png 424w, https://substackcdn.com/image/fetch/$s_!RRn1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png 848w, https://substackcdn.com/image/fetch/$s_!RRn1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png 1272w, https://substackcdn.com/image/fetch/$s_!RRn1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda982695-94c2-4371-90ef-49e6a926e257_2220x1290.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Visualized:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IJnR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IJnR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png 424w, https://substackcdn.com/image/fetch/$s_!IJnR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png 848w, https://substackcdn.com/image/fetch/$s_!IJnR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!IJnR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IJnR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;When each competitor lands&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="When each competitor lands" title="When each competitor lands" srcset="https://substackcdn.com/image/fetch/$s_!IJnR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png 424w, https://substackcdn.com/image/fetch/$s_!IJnR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png 848w, https://substackcdn.com/image/fetch/$s_!IJnR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!IJnR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf83ccbe-1bfc-433c-8dfe-193068449542_2000x1120.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A 2026 cohort, and a 2027+ group.</figcaption></figure></div><p>Caveat though. Every entry above is a first-rack or first-sample milestone, not a gigawatt figure. <em>IMO first rack matters, but time to first gigawatt matters most.</em></p><p>First rack proves the whole shebang works. But gigawatt scale proves the <em>company</em> works.</p><h3>Who is shipping in 2026, and who isn&#8217;t?</h3><p>Tenstorrent, Etched, and SambaNova ship production racks in 2026.</p><p>A year back sit Positron, Tensordyne, and Rebellions. Rebellions claims 2026, but its frontier REBEL is still only sampling, so its real rack is more like 2027.</p><p>Fractile and MatX are two years back. Both have 2027 milestones on their roadmaps, but those are samples and footholds, not scaled production shipments. Production racks at customers are realistically a 2028 story.</p><p><strong>In this supply-starved environment, a year is expensive.</strong> That&#8217;s not to say the 2027+ companies aren&#8217;t technically competitive, nor does it imply they won&#8217;t win serious customer share in the grand arc of time. After all, it only takes one of the few large customers to get to scale. <strong><mark data-color="rgb(244, 206, 211)" style="background-color: rgb(244, 206, 211); color: rgb(0, 0, 0);">But you&#8217;re not in the game until you&#8217;re on the field!</mark></strong> No matter how technically sound your architecture is, if you&#8217;re not on the field, you&#8217;re not in the game.</p><h3><strong>Who will be on the field first?</strong></h3><p><strong>Tenstorrent</strong> claims to be &#8220;shipping production silicon now&#8221;. <em>Promising</em>! But what about customers? Tenstorrent has neoclouds and sovereigns; no frontier lab or hyperscaler has been announced yet. Could those early customers accelerate Tenstorrent to gigawatt scale? We dig into that in its teardown below.</p><p>An acquisition could help Tenstorrent get to market, as big silicon incumbents already have relationships with hyperscalers. Qualcomm was <a href="https://www.theregister.com/systems/2026/06/16/qualcomm-said-to-be-circling-ai-chip-biz-tenstorrent-in-10b-risc-v-power-play/">reported</a> in mid-June to be circling a roughly $8&#8211;10B deal. But Keller supposedly denied any Qualcomm talks recently, so that path looks cold for now.</p><p><strong>Etched</strong> is a bigger contender than it looked just months ago. It claims rack-scale shipments this summer, well ahead of the other popular names like MatX and Fractile. I think Etched was left for dead, but it turns out they are alive and claim to be right at the front of the line. </p><p>Notably, Etched is talking gigawatts. It claims <em>&#8220;a path to gigawatt-scale in 2027,&#8221;</em> and it&#8217;s talking about building the company infrastructure to pull it off with a Taiwan factory, plus a data center, test house, and NPI prototyping lab at its San Jose HQ. From the <a href="https://www.globenewswire.com/news-release/2026/06/30/3319922/0/en/etched-emerges-from-stealth-with-working-chip-800m-raised-and-over-1b-in-customer-contracts.html">press release</a>, co-founder Rob Wachen said <em>&#8220;Our approach from the beginning has been to build for gigawatt-scale... Production is the product.&#8221;</em></p><p>Etched hasn&#8217;t disclosed specific customers but mentioned $1B+ in contracts booked. Etched&#8217;s talk about a &#8220;path to gigawatt-scale&#8221; and &#8220;building for gigawatt-scale&#8221; seems to imply they are targeting the type of companies that can support that scale.</p><p><em>Whether Etched actually hits gigawatt scale in 2027 remains to be seen, but it&#8217;s at least aimed at the right target if they want to be in the trillion-dollar race.</em></p><p><strong>SambaNova</strong> also claims a 2026 deployment on its frontier SN50 landing in the second half. <em>This is going to make for an exciting 2H26 and early 2027 if we&#8217;ve got SambaNova, Etched, and Tenstorrent all shipping competitive racks! </em></p><p>So far, SambaNova&#8217;s named customers are enterprises JPMorgan Chase and SoftBank, and these are explicitly NOT gigawatt installations. From <a href="https://www.eetimes.com/sambanova-raises-1-billion-signs-jpmorganchase-as-a-customer/">Sally Ward-Foxton</a></p><blockquote><p>SambaNova has a multi-year agreement in place with JPMorganChase.</p><p>&#8220;This is a breakthrough because most suppliers haven&#8217;t figured out how to get into the enterprise,&#8221; Liang said. &#8220;The enterprise market is going to be significant; it won&#8217;t be 100%, but it&#8217;s going to be a meaningful player in the overall market, and we&#8217;re starting to see that.&#8221;</p><p>Enterprise customers are looking to host private models and private data in their own environment, <strong><mark data-color="rgb(251, 226, 198)" style="background-color: rgb(251, 226, 198); color: rgb(0, 0, 0);">which is unlikely to be a gigawatt-scale data center, Liang said.</mark></strong></p></blockquote><p>Of course, if you have tens of 50-100MW deployments, well, that counts! </p><p><em>Maybe SambaNova&#8217;s plan is to demonstrate feasibility with reasonably sized enterprises and use that success to try to win hyperscalers? Or just win with enterprises, which should turn into a very large market on its own.</em></p><p>Behind the 2026 crew are five more competitors including Rebellions, Positron, and Tensordyne in 2027, with Fractile and MatX not reaching production racks until 2028.</p><p><strong>Rebellions</strong> reaches a rack soonest of this 2027+ tier, with REBEL sampling now. But its named customers, mostly Korean enterprises plus Saudi Aramco, sit on the sub-frontier ATOM, not the frontier REBEL, and none is a gigawatt buyer.</p><p><strong>Positron</strong> has a promising customer, Oracle. </p><p>Some consider OCI (<em>Oracle Cloud Infrastructure</em>) a hyperscaler and others call it a neocloud (<em><a href="https://www.clustermax.ai/cloudreview/oracle">ClusterMax</a> rates it as a gold tier neocloud</em>). Regardless, OCI can support GW scale installations.</p><p>So is a promising path to GW scale for Positron, and they are the first on this list who can publicly say they&#8217;re already deployed in a hyperscaler. The catch is that Oracle is currently running <a href="https://www.positron.ai/atlas">Atlas</a>, Positron&#8217;s sub-frontier first-gen server. <em>It&#8217;s not rackscale.</em> The frontier system, Asimov, doesn&#8217;t tape out until late 2026. </p><p>So Positron already has a customer who could scale to a gigawatt, just not yet on a frontier product. </p><p><strong>Tensordyne</strong> <a href="https://www.tensordyne.ai/stories/tensordyne-announces-breakthrough-inference-system-to-end-ais-speed-vs-cost-trade-off">recently announced</a> its Napier system has been taped out on TSMC 3nm and production is underway. And Tensordyne&#8217;s <a href="https://www.eetimes.com/podcasts/how-tensordyne-built-an-ai-accelerator-around-logarithmic-math/">logarithmic math approach</a> is super interesting.</p><p>But its customer book looks a lot like Tenstorrent&#8217;s. Cirrascale and BlueSky Compute are named, plus a dozen-plus letters of intent to <em>evaluate</em> beta systems, and roughly $200M of forecast, not booked, demand. Real interest, but neoclouds and LOIs, not a frontier lab ordering a gigawatt. </p><p><strong>Fractile</strong> is further out. Its roadmap targets a late-2026 commercial tape-out, customer samples in Q2 2027, and production silicon only at the very end of 2027, so racks in customer data centers are realistically a 2028 story. But it&#8217;s courting exactly the right customer. In May 2026, <a href="https://www.theinformation.com/articles/anthropic-talks-buy-ai-chips-u-k-startup">The Information reported</a> Anthropic was in early talks to buy Fractile chips once they ship, which would make a frontier lab its anchor, the kind that scales fastest. The talks were preliminary and the deal size wasn&#8217;t disclosed<em> </em>though.</p><p><strong>MatX</strong> is pre-silicon, on roughly the same clock as Fractile. It targets tape-out within a year of its <a href="https://matx.com/research/series_b">$500M Feb 2026 raise</a>, first chips a few months after that, and a 2027 foothold inside a frontier lab, routing ~1% of production traffic to its silicon as an A/B test. </p><p>That foothold is a validation experiment, not production racks; by Pope&#8217;s own tape-out-to-production math of one to two years, volume racks are a 2028-at-earliest story. </p><p>It hasn&#8217;t dated real rack scale, only that CEO Reiner Pope <em>&#8220;would like to be... shipping multiple gigawatts a year&#8221;</em> (<a href="https://cheekypint.substack.com/p/reiner-pope-of-matx-on-accelerating">Cheeky Pint</a>). So it&#8217;s got the right goal, just behind on the timeline.</p><p>Yet it may have the best customer list of anyone here. MatX says it <a href="https://www.chipstrat.com/p/an-interview-with-matx-ceo-reiner">sells only to frontier labs</a>, <em>&#8220;maybe five different customers in the world&#8221;</em> and its 1% trial would be a foothold with such a customer. If that little trial is a success, that&#8217;s the type of customer who would scale quickly to a gigawatt faster than any neocloud or enterprise mentioned.</p><h2>Deeper Teardown</h2><p>Of course, details matter. Here&#8217;s the due diligence on eleven companies with one question for each. </p><p><strong>Can it put a frontier-class rack in front of a whale of a customer before the window closes?</strong></p><p>Below is the summarized diligence on all eleven, the eight frontier contenders first, then the add-in and the sub-frontier plays.</p><p>For each we&#8217;ll hit on</p><ul><li><p><strong>KPI</strong> is the one outcome the chip beats a GPU on.</p></li><li><p><strong>Anchor</strong> is who is actually deploying it.</p></li><li><p><strong>Architecture</strong></p></li><li><p><strong>Verified?</strong></p></li></ul><p>Then a bunch of extra sourced details.</p><h3>Tenstorrent (Q2 2026)</h3><p>Jim Keller is arguing a contrarian narrative. One chip runs everything, no disaggregation. Keller believes that, in the long run, general-purpose can perform as well as specialized + disaggregation.</p><ul><li><p><strong>KPI:</strong> low cost at high interactivity. Runs ~any model (~90% of Hugging Face), prefill and decode on the same silicon.</p></li><li><p><strong>Anchor:</strong> five-plus neocloud colos serving live, plus sovereign and IP licensing (Japan, LG). <em>No frontier lab, no hyperscaler.</em></p></li><li><p><strong>Architecture:</strong> GDDR6, no HBM; RISC-V Tensix cores; Ethernet scale-out on-chip; 100% open-source software stack. One chip runs both prefill and decode, the deliberate counter to disaggregation.</p></li><li><p><strong>Verified?</strong> <strong>&#9888;&#65039;</strong> DeepSeek ~308 tok/s/user is a served-model number, but relayed by Keller, not an independently-published AA figure. The Prodia video claim is similar: AA independently ranks Prodia&#8217;s model at the top of its leaderboard, but that predates the Tenstorrent partnership, and the 10&#215; speed-on-Tenstorrent-hardware number is Tenstorrent/Prodia&#8217;s own benchmark, not AA-reproduced. No MLPerf; the 350&#8594;500 tok/s headline is partly forward-looking.</p></li></ul><h4>DETAILS</h4><p><strong>The anti-thesis, on record.</strong></p><ul><li><p>At TT-Deploy (May 2026), Jim Keller bet against specialization and disaggregation: <em>&#8220;the big fad is disaggregation, special purpose hardware, SRAM. Do you know how many people are going to be talking about that in 2 years? None.&#8221;</em></p></li><li><p>And: <em>&#8220;you don&#8217;t just run part of an LLM&#8230; that&#8217;s not a good business plan long run. You could be one-shotted by one model.&#8221;</em></p></li><li><p>Stan Sokorac&#8217;s (Sr. Fellow, Software) proof point: a pipeline that downloads random Hugging Face models achieves a <strong>~90% pass rate</strong>, so <em>&#8220;we can run ~2.5M of the 2.8M models on Hugging Face.&#8221;</em> (<a href="https://www.youtube.com/watch?v=8ZS9GvawgFs">TT-Deploy keynote</a>; <a href="https://www.youtube.com/watch?v=6IZEo5XmHxM">Sokorac talk</a>)</p></li></ul><p><strong>Architecture.</strong></p><ul><li><p><strong>Tensix core</strong> = matrix engine + vector unit + 5 &#8220;baby&#8221; RISC-V cores per tile, each with local L1 SRAM, over a 2D NoC mesh (no shared global memory); Blackhole adds <strong>16 &#8220;big&#8221; 64-bit RISC-V CPU cores</strong> and can run host-less (<a href="https://www.theregister.com/2024/08/27/tenstorrent_ai_blackhole/">The Register</a>).</p></li><li><p><strong>Anti-disaggregation by design:</strong> each chip has SRAM + DRAM + networking, so it runs prefill + decode + video on the same silicon (<a href="https://www.youtube.com/watch?v=QLQKzXADYcA">Vasiljevic talk</a>).</p></li><li><p><strong>Scale-out:</strong> standard Ethernet, <strong>10&#215;400 Gbps = 1 TB/s chip-to-chip</strong>, the anti-NVLink bet. Blackhole = TSMC 6nm.</p></li><li><p><strong>Memory:</strong> large on-chip SRAM + GDDR6, explicitly no HBM. Official Blackhole spec: <strong>180 MB</strong> on-chip SRAM, <strong>28&#8211;32 GB GDDR6 at 448&#8211;512 GB/s</strong> (<a href="https://tenstorrent.com/en/hardware/blackhole">tenstorrent.com/hardware/blackhole</a>).</p></li></ul><p><strong>Shipping status.</strong></p><ul><li><p>Three taped-out generations (Grayskull &#8594; Wormhole &#8594; Blackhole).</p></li><li><p>Shipping Blackhole cards: <strong>120 Tensix / 664 TFLOPS BLOCKFP8 / 300 W</strong>; p100a <strong>$999</strong>, p150a/b <strong>$1,399</strong>.</p></li><li><p>Blackhole <strong>Galaxy production servers reached GA Apr 28 2026</strong> (6U, 32 chips, 23 PFLOPS dense FP8, <strong>~$110K</strong>) (<a href="https://www.theregister.com/2026/04/28/tenstorrent_galaxy_blackhole_ai_servers_ga/">The Register GA</a>).</p></li><li><p>Yield caveat, now in the official spec: launched at 140 Tensix, a <strong>Jan 2026 firmware cut to 120</strong> (&#8776;1&#8211;2% perf hit on typical workloads; new cards ship Bin-3 parts with two disabled Tensix columns) (<a href="https://www.tomshardware.com/tech-industry/semiconductors/jim-kellers-tenstorrent-is-downgrading-blackhole-p150-cards-from-140-to-120-tensor-cores-via-firmware-update-will-ship-cards-with-120-tensor-cores-going-forward-company-claims-existing-users-should-expect-1-2-percent-performance-drop">Tom&#8217;s Hardware</a>).</p></li><li><p>&#8220;Installed in <strong>&#8805;5 neocloud colos</strong> outside Tenstorrent&#8221; goal hit (TT-Deploy).</p></li></ul><p><strong>Deployments / anchors.</strong></p><ul><li><p>Infra/colo/neocloud + finance, not hyperscalers: <strong>Equinix, Cirrascale, ai&amp;, Prodia Labs, Virtu Financial, Turiyam, OrionVM</strong> (<a href="https://tenstorrent.com/en/newsroom">Tenstorrent newsroom</a>).</p></li><li><p>The other half is RISC-V/chiplet IP licensing + sovereign design: <strong>Japan/LSTC</strong> (Ascalon IP, Rapidus), a <strong>~$50M</strong> Japan program to train 200 designers, <strong>LG</strong> co-developing SoCs.</p></li><li><p><strong>The gigawatt question.</strong> The July 2026 slate widened the book (Galaxy superclusters now anchor <strong>Equinix&#8217;s Distributed AI Hub</strong>, launched with OrionVM and BetterBrain, plus new ai&amp;, Virtu Financial, and Cirrascale deployments), but none of it is gigawatt-class. The headline neocloud, Cirrascale, is <strong>ClusterMAX Silver, running in the thousands of GPUs</strong>, and already lost OpenAI to Azure (<a href="https://www.clustermax.ai/cloudreview/cirrascale">ClusterMAX</a>). Every name here buys racks, yes, but not gigawatts. Aggregating many mid-size buyers may get there, but slower.</p></li></ul><p><strong>Performance / verification.</strong></p><ul><li><p>Absent from MLPerf v6.0.</p></li><li><p><strong>~308 tok/s/user</strong> on DeepSeek live is a served-model measurement, but per Keller&#8217;s own telling, not an independently published Artificial Analysis number; the roadmap to <strong>500 tok/s/user at $6/M TCO</strong> is forward-looking.</p></li><li><p>The <strong>Prodia video record</strong> (Wan 2.2, 2.4s, 10&#215; a GPU) needs the same unpacking: Prodia&#8217;s model already led the Artificial Analysis leaderboard <em>before</em> the Tenstorrent partnership, so that ranking is genuinely AA-verified &#8212; but the <strong>10&#215; speedup on Tenstorrent hardware</strong> (33.8 fps vs. 5.5 fps on Nvidia) is Tenstorrent and Prodia&#8217;s own collaboration benchmark, not a number AA independently re-ran (<a href="https://tenstorrent.com/en/newsroom/tenstorrent-enables-ai-at-scale-with-industry-leading-performance">Tenstorrent newsroom</a>; <a href="https://www.youtube.com/watch?v=QLQKzXADYcA">Vasiljevic talk</a>).</p></li><li><p>Stack is <strong>100% open source</strong> (TT-Forge, TT-Lang), the anti-CUDA-moat argument.</p></li></ul><p><strong>Team / funding.</strong></p><ul><li><p>Founded <strong>2016</strong> (Toronto) by Ljubisa Bajic, Ivan Hamer, and Milos Trajkovic; <strong>Jim Keller</strong> CEO since ~2023.</p></li><li><p><strong>Series D $693M+</strong> closed Dec 2024 at <strong>~$2.6&#8211;2.7B post</strong> (Samsung Securities + AFW lead; Bezos, Hyundai, LG) (<a href="https://tenstorrent.com/newsroom/tenstorrent-closes-693m-of-series-d-funding-led-by-samsung-securities-and-afw-partners">Series D</a>).</p></li><li><p>A reported <strong>~$800M / ~$3.2B</strong> Fidelity round (Nov 2025) is not company-confirmed.</p></li></ul><p>That&#8217;s the free read. Behind the paywall:</p><ul><li><p><strong>The other ten teardowns,</strong> each scored on the same four questions as Tenstorrent, its KPI, its anchor, its architecture, and what&#8217;s actually verified, then the full sourced details behind every line. The seven remaining frontier contenders (Etched, Rebellions, SambaNova, Positron, Tensordyne, Fractile, MatX), the decode add-in (d-Matrix), and the two sub-frontier plays (Furiosa, Taalas).</p></li><li><p><strong>The verification scorecard,</strong> which two of the eight frontier contenders have a clean independent number, which one has only a partner-run result, and which remain vendor-stated.</p></li><li><p><strong>Which &#8220;frontier&#8221; chips can&#8217;t yet run a frontier model,</strong> the sub-1T ceiling sitting under every verified number, and whose only MLPerf result is for the wrong chip.</p></li><li><p><strong>The memory split, company by company,</strong> who leans on HBM for raw bandwidth and who drops it for commodity LPDDR or GDDR to dodge the supply crunch, and what each choice costs them.</p></li><li><p><strong>Where the field lands on disaggregation,</strong> one chip running prefill and decode versus a split tier, and which way it&#8217;s converging.</p></li><li><p><strong>Who actually looks like the next trillion-dollar chip company</strong></p></li></ul><p>If you&#8217;re deciding who to watch, partner with, invest in, or compete against, this is the half that matters.</p><h3>Etched (H2 2026)</h3><p>A frontier inference system, sold as a full rack, has its real edge in low-voltage inference and cluster-scale memory. And more flexible than it&#8217;s remembered for: the system that launched in June runs DeepSeek, Qwen, Llama, and Mamba. <em>Mamba is an SSM, not a transformer&#8230;</em> <em>so clearly it&#8217;s more flexible than previous talk of &#8220;transformer-only ASIC&#8221; made it out to be&#8230; </em></p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/the-next-trillion-dollar-chip-company">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[High Bandwidth Flash: The Full Report]]></title><description><![CDATA[Why HBF for inference decode fits, what it does to the DRAM and NAND markets, a worked cost model, who moves first, and risks]]></description><link>https://newsletter.chipstrat.com/p/high-bandwidth-flash-the-full-report</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/high-bandwidth-flash-the-full-report</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Tue, 07 Jul 2026 20:30:12 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/28462db4-0d7e-4c08-a332-f93f590d1653_2016x1200.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There&#8217;s been a lot of chatter about High Bandwidth Flash (HBF) recently.</p><p>If you&#8217;re not up to speed, the basic idea is that HBF is a stack of NAND dies built the way an HBM stack is built; dies stacked vertically, wired together with through-silicon vias (TSVs), and sitting right next to the GPU on the package interposer:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WREa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WREa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png 424w, https://substackcdn.com/image/fetch/$s_!WREa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png 848w, https://substackcdn.com/image/fetch/$s_!WREa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png 1272w, https://substackcdn.com/image/fetch/$s_!WREa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WREa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png" width="411" height="344.86636363636364" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:923,&quot;width&quot;:1100,&quot;resizeWidth&quot;:411,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!WREa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png 424w, https://substackcdn.com/image/fetch/$s_!WREa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png 848w, https://substackcdn.com/image/fetch/$s_!WREa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png 1272w, https://substackcdn.com/image/fetch/$s_!WREa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb8d2e8d-7774-4c89-b2b4-e05126c9e63d_1100x923.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Looks a lot like HBM. Source: Sandisk.</figcaption></figure></div><p>What&#8217;s interesting is that it has the <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">same read bandwidth as an HBM4 stack, but with roughly 10x the capacity</mark>. And it&#8217;s made of NAND, not DRAM like HBM. (<em>NAND is the cheap stuff.)</em></p><p>The first samples of the memory itself are expected very soon from Sandisk, sometime in the second half of 2026. Samples of the first AI inference devices built with HBF follow in early 2027.</p><p>Everything has tradeoffs, flash too. We&#8217;ll look at why those tradeoffs aren&#8217;t so bad for inference decode workloads, and the impact HBF will have on the memory market.</p><p><strong>Table of contents</strong></p><ul><li><p>Flash is storage. But could it work for memory?</p></li><li><p>Weight memory is what drives GPU count</p></li><li><p>DRAM is no longer scaling well. But NAND is</p></li><li><p>CMOS directly Bonded to Array: how the HBF stack is built (for the device physics layer, see Vik&#8217;s <a href="https://www.viksnewsletter.com/p/high-bandwidth-flash-nands-bid-for-ai">High Bandwidth Flash: NAND&#8217;s Bid for AI Memory</a>)</p></li><li><p>[Paid] Inference decode fits this profile; training does not</p></li><li><p>[Paid] Four ways HBF attaches to a GPU, each with a different supply chain winner</p></li><li><p>[Paid] The disaggregated inference cost model</p></li><li><p>[Paid] What HBF does to the memory markets</p></li><li><p>[Paid] Nvidia is the certification gate; custom silicon moves first</p></li><li><p>[Paid] Competitive landscape: Sandisk, SK Hynix, Samsung, YMTC</p></li><li><p>[Paid] The patent: a processor bonded to NAND</p></li><li><p>[Paid] HBM is infeasible for edge devices; HBF is not</p></li><li><p>[Paid] Timeline and risks</p></li></ul><h2>Flash is storage. But could it work for memory?</h2><p>NAND flash has traditionally been used for information <em>storage</em>, i.e. where data lives when it&#8217;s not being used. <em>Cheap, dense, and non-volatile. But slooooooowwwww....</em></p><p>But <em>memory</em> is where the accelerator keeps information handy during computation, and it has to be fast enough that compute never waits. SRAM is used for memory and can be read very quickly, on the order of about a nanosecond. Next in the memory hierarchy is DRAM, which is slower to read at ~100 nanoseconds.</p><p>But storage like NAND takes about a hundred microseconds, roughly 1,000x slower than DRAM. By definition that&#8217;s <em>storage</em> and not <em>memory</em>, right? It takes way too long to read, but good for long term storage. That won&#8217;t work for computation... way too slow! Compute would just be sitting idle, waiting forever for data.</p><p><strong>So how could NAND flash possibly be used as memory?</strong></p><p><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">Well, could there be a way to get the right data to the accelerator in time for computation, even with NAND&#8217;s slow reads?</mark> <em>If you start the read it WAY before the accelerator needs the data, it could work?</em></p><p>In addition to intelligently scheduling when data is requested, one can also try to achieve high bandwidth from NAND. Bandwidth is the amount of data that arrives per unit of time. So if you know the read is going to take a long time, well, might as well have a bunch of reads running in parallel, if possible, right?</p><p>HBF unlocks high bandwidth by stacking NAND dies to increase the &#8220;width&#8221; or the parallelism. Each die is divided into many sub-arrays, which are small blocks of NAND that can each be read at the same time, independently of one another. An HBF stack places 16 of these dies behind a single interface, thousands of bits wide, so a huge number of sub-arrays are available to read at once. <mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">Any single read still takes thousands of nanoseconds, but thousands of reads can run in parallel, so a lot of data can be moved simultaneously.</mark></p><p>The 16 flash dies are stacked and connected with through-silicon vias (TSVs). A controller logic die is then bonded directly onto the NAND array. That logic die schedules the parallel sub-array reads and drives the results out over the wide interface to the GPU. The bonded die is called &#8220;CMOS directly Bonded to Array&#8221; or CBA:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fOUC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fOUC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png 424w, https://substackcdn.com/image/fetch/$s_!fOUC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png 848w, https://substackcdn.com/image/fetch/$s_!fOUC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png 1272w, https://substackcdn.com/image/fetch/$s_!fOUC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fOUC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png" width="608" height="311.18545454545455" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:563,&quot;width&quot;:1100,&quot;resizeWidth&quot;:608,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!fOUC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png 424w, https://substackcdn.com/image/fetch/$s_!fOUC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png 848w, https://substackcdn.com/image/fetch/$s_!fOUC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png 1272w, https://substackcdn.com/image/fetch/$s_!fOUC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2205fa-e11d-48c9-9e48-3f4f3db472af_1100x563.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Kioxia</figcaption></figure></div><p>The logic + HBF stack is on a high-bandwidth interposer. The resulting HBF delivers 1.6 TB/s of read bandwidth, which is the same as an HBM4 stack at the JEDEC spec&#8217;s 6.4 Gb/s operating point:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Mk9O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Mk9O!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png 424w, https://substackcdn.com/image/fetch/$s_!Mk9O!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png 848w, https://substackcdn.com/image/fetch/$s_!Mk9O!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png 1272w, https://substackcdn.com/image/fetch/$s_!Mk9O!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Mk9O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png" width="600" height="241.63636363636363" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:443,&quot;width&quot;:1100,&quot;resizeWidth&quot;:600,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!Mk9O!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png 424w, https://substackcdn.com/image/fetch/$s_!Mk9O!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png 848w, https://substackcdn.com/image/fetch/$s_!Mk9O!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png 1272w, https://substackcdn.com/image/fetch/$s_!Mk9O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe44bb4-e796-4ced-8a31-56a472bbde9e_1100x443.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>To underscore the importance of what the packaging is doing here, compare against the NAND you can buy today. Conventional flash ships about 14 GB/s behind a PCIe 5.0 NVMe controller. Packaged as HBF, the same material delivers 1.6 TB/s. Roughly 100x the bandwidth, from packaging alone. <em>Bullish advanced packaging!</em></p><p>Again, HBF&#8217;s latency is still 10-100x slower than HBM. But if the accelerator knows what data it needs in advance, it can prefetch it and avoid waiting on any single slow read. <mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">The argument for HBF is that inference decode is the perfect workload.</mark> <em>I&#8217;ll explain in more detail later for paid subscribers.</em></p><h2>Weight memory is what drives GPU count</h2><p><strong>What are the economic implications of HBF?</strong> </p><p>Inference cost scales with GPU count, and for today&#8217;s massive frontier models, GPU count is often driven by memory capacity per GPU. </p><p><em>Why?</em> </p><p>Well, you need enough HBM to hold all the weights, but HBM is co-packaged with the accelerator; a fixed amount of HBM is bonded into each GPU package. So you can&#8217;t add memory without adding GPUs. Hence, bigger models mean more GPUs.</p><p>Of course, weights aren&#8217;t the only thing the accelerator needs to store in memory and acccess quickly and often. The KV cache and activations sit in memory too, and both stay in HBM or DRAM. </p><p>But the KV cache takes new writes every token; NAND&#8217;s endurance can&#8217;t handle that. NAND has much lower write endurace. </p><p>And activations need low-latency random access that NAND is too slow to give.</p><p><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">So HBF is for storing model weights.</mark></p><p>And frontier weights are huge and always wanting to be even bigger. <em>Wouldn&#8217;t it be nice to hold the model in significantly cheaper memory than HBM?</em></p><p>Recall that a 70B parameter model at fp16 (2 bytes per parameter) requires 70 &#215; 10&#8313; &#215; 2 = 140 GB just for weights. A 1T parameter model thus needs 1,000 &#215; 10&#8313; &#215; 2 = 2 TB. But that&#8217;s a lot of HBM; today&#8217;s shipping HBM4 stacks hold 36 GB (12-Hi); 16-Hi parts push 48 GB, and <a href="https://www.jedec.org/news/pressreleases/jedec%C2%AE-and-industry-leaders-collaborate-release-jesd270-4-hbm4-standard-advancing">the JEDEC spec tops out at 64 GB</a>. So a large model needs many stacks, which means many GPUs. <em>Expensive!</em> </p><p>That also means more interconnect (to move data around all those GPUs), which requires more power and, ultimately, a higher cost per output token (watts and $).</p><p>But HBF can provide 512 GB of capacity per stack!</p><p>So 512 GB per stack versus 48 GB for HBM4 is roughly 10x the capacity at the same bandwidth. Sandisk states up to 8-16x across the HBM family, depending on which generation you compare against.</p><p>Interesting! Super promising. Consider the implications of storing frontier models in a few HBF stacks rather than spread across many HBM stacks across many GPUs in a server or rack. </p><p><em>New opportunities arise too! Could on-premises, air-cooled deployments at enterprises have frontier-model weights capacity without needing so many expensive GPUs? </em> </p><h2>CMOS directly Bonded to Array (CBA)</h2><p><strong>Let&#8217;s dig into the manufacturing of HBF a bit more.</strong> <em>Dang, I seriously didn&#8217;t mean to be punny with that bit.</em></p><p>Sandisk stacks 16 BiCS NAND dies (BiCS, or Bit Cost Scalable, is just its brand of 3D NAND) with TSVs (<em>the same way HBM is built</em>) and bonds the controller onto the array:</p><div class="captioned-image-container"><figure><div class="image-link image2" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yJBL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png 424w, https://substackcdn.com/image/fetch/$s_!yJBL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png 848w, https://substackcdn.com/image/fetch/$s_!yJBL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png 1272w, https://substackcdn.com/image/fetch/$s_!yJBL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yJBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png" width="456" height="228" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:330,&quot;width&quot;:660,&quot;resizeWidth&quot;:456,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!yJBL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png 424w, https://substackcdn.com/image/fetch/$s_!yJBL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png 848w, https://substackcdn.com/image/fetch/$s_!yJBL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png 1272w, https://substackcdn.com/image/fetch/$s_!yJBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25de952e-7035-450f-a556-85dc91d1b9fa_660x330.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></div><figcaption class="image-caption">Source: Kioxia</figcaption></figure></div><p>The result matches HBM4&#8217;s footprint, height, and power envelope, so it can drop into the same interposer slot an HBM stack sits in today. It speaks nearly the same electrical interface, too, though <a href="https://blocksandfiles.com/2025/02/12/sandisk-spills-its-technolgy-futures-beans/">the host memory controller has to change</a>, so it isn&#8217;t quite drop-in. And being NAND, it&#8217;s non-volatile (it holds data with the power off), so unlike DRAM it needs no constant refresh power.</p><p>One open question is the NAND cell type. NAND can pack more than one bit into each memory cell:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zWvc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zWvc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png 424w, https://substackcdn.com/image/fetch/$s_!zWvc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png 848w, https://substackcdn.com/image/fetch/$s_!zWvc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png 1272w, https://substackcdn.com/image/fetch/$s_!zWvc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zWvc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png" width="497" height="395.115" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:795,&quot;width&quot;:1000,&quot;resizeWidth&quot;:497,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!zWvc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png 424w, https://substackcdn.com/image/fetch/$s_!zWvc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png 848w, https://substackcdn.com/image/fetch/$s_!zWvc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png 1272w, https://substackcdn.com/image/fetch/$s_!zWvc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc621b38a-f0e9-4a85-957d-7a92f22a2d57_1000x795.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: SK Hynix via Vik&#8217;s Newsletter</figcaption></figure></div><p>Highly recommend this explainer from Vik:</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:173745369,&quot;url&quot;:&quot;https://www.viksnewsletter.com/p/high-bandwidth-flash-nands-bid-for-ai&quot;,&quot;publication_id&quot;:2065897,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Vik's Newsletter&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!9JlA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa409d69d-ca10-4bfe-a1fc-f8d291690566_185x185.png&quot;,&quot;title&quot;:&quot;High Bandwidth Flash: NAND&#8217;s Bid for AI Memory&quot;,&quot;truncated_body_text&quot;:&quot;Welcome to a &#128274; subscriber-only deep-dive edition &#128274; of my weekly newsletter. Each week, I help investors, professionals and students stay up-to-date on complex topics, and navigate the semiconductor industry.&quot;,&quot;date&quot;:&quot;2025-09-21T18:29:33.558Z&quot;,&quot;like_count&quot;:26,&quot;comment_count&quot;:4,&quot;bylines&quot;:[{&quot;id&quot;:124411709,&quot;name&quot;:&quot;Vikram Sekar&quot;,&quot;handle&quot;:&quot;vikramskr&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/afc78b68-c3cf-4c29-94f3-1422781e3e92_185x185.png&quot;,&quot;bio&quot;:null,&quot;profile_set_up_at&quot;:&quot;2023-01-21T13:11:05.033Z&quot;,&quot;reader_installed_at&quot;:&quot;2023-01-21T03:23:33.933Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:2068317,&quot;user_id&quot;:124411709,&quot;publication_id&quot;:2065897,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:2065897,&quot;name&quot;:&quot;Vik's Newsletter&quot;,&quot;subdomain&quot;:&quot;viksnewsletter&quot;,&quot;custom_domain&quot;:&quot;www.viksnewsletter.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;AI infrastructure research across photonics, memory, interconnects, power, and packaging. Engineering depth translated for professionals and investors.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a409d69d-ca10-4bfe-a1fc-f8d291690566_185x185.png&quot;,&quot;author_id&quot;:124411709,&quot;primary_user_id&quot;:124411709,&quot;theme_var_background_pop&quot;:&quot;#EA410B&quot;,&quot;created_at&quot;:&quot;2023-10-29T03:11:48.585Z&quot;,&quot;email_from_name&quot;:&quot;Vik's Newsletter&quot;,&quot;copyright&quot;:&quot;Vikram Sekar&quot;,&quot;founding_plan_name&quot;:&quot;\&quot;Expense it!\&quot;&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;magaziney&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}},{&quot;id&quot;:9006559,&quot;user_id&quot;:124411709,&quot;publication_id&quot;:8781267,&quot;role&quot;:&quot;contributor&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:8781267,&quot;name&quot;:&quot;Semi Doped&quot;,&quot;subdomain&quot;:&quot;semidoped&quot;,&quot;custom_domain&quot;:&quot;www.semidoped.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;The Daily Brew of Semiconductors. News and analysis from Vik Sekar and Austin Lyons.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/979b934f-dffb-48a2-8597-186738f44571_1024x1024.png&quot;,&quot;author_id&quot;:500274950,&quot;primary_user_id&quot;:500274950,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2026-04-23T14:07:09.663Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Semi Doped&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;twitter_screen_name&quot;:&quot;vikramskr&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:1000,&quot;status&quot;:{&quot;bestsellerTier&quot;:1000,&quot;subscriberTier&quot;:5,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;bestseller&quot;,&quot;tier&quot;:1000},&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.viksnewsletter.com/p/high-bandwidth-flash-nands-bid-for-ai?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!9JlA!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa409d69d-ca10-4bfe-a1fc-f8d291690566_185x185.png" loading="lazy"><span class="embedded-post-publication-name">Vik's Newsletter</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">High Bandwidth Flash: NAND&#8217;s Bid for AI Memory</div></div><div class="embedded-post-body">Welcome to a &#128274; subscriber-only deep-dive edition &#128274; of my weekly newsletter. Each week, I help investors, professionals and students stay up-to-date on complex topics, and navigate the semiconductor industry&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">a year ago &#183; 26 likes &#183; 4 comments &#183; Vikram Sekar</div></a></div><p><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">Note that the fewer bits each NAND cell packs, the faster and more durable the cell is, but at the tradeoff of density. </mark></p><p>QLC (quad-level cell) holds four bits, cheap and dense, but is the quickest to wear out. SLC (single-level cell) holds one, yet is far more durable. </p><p>Sandisk hasn&#8217;t said which HBF uses, but <a href="https://irrationalanalysis.substack.com/p/market-memo-hot-summer-topics">Irrational Analysis</a> reported recently that industry sources tell him it will be SLC, which would improve HBF&#8217;s write endurance by an order of magnitude, but at the cost of density.</p><p>Sandisk makes the NAND, but turning 16 dies into a stack that matches HBM&#8217;s footprint, thermals, and interface is advanced packaging work... the same thing the big 3 HBM folks have already perfected. So <mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">SK Hynix is the stacking partner</mark>. </p><p>Sandisk brings the flash, SK Hynix brings the stack.</p><p>The two <a href="https://www.sandisk.com/company/newsroom/press-releases/2026/2026-02-25-sandisk-and-sk-hynix-begin-global-standardization-of-next-generation-memory-solution-high-bandwidth-flash-hbf">began an OCP standardization effort in February 2026</a>.</p><p><em>Standardization?</em></p><p>You might expect Sandisk to want a monopoly on HBF, right? Actually, it wants other suppliers, so long as Sandisk is the biggest. After all, hyperscalers won&#8217;t design a sole-source part into a billion-dollar platform, so being the only supplier is a risk to HBF adoption. </p><p>A ratified standard with multiple suppliers removes that risk and gets HBF designed in. Again, Sandisk would rather own a big share of a market that exists than all of one that never does. <em>It reminds me of with Credo and active electrical cables. Grow the category into a real market, and win the largest slice of it.</em></p><p><strong>So how does HBF compare to HBM? </strong></p><p>HBF has the edge when it comes to capacity vs HBM4, but trails in most other respects:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Marm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Marm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png 424w, https://substackcdn.com/image/fetch/$s_!Marm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png 848w, https://substackcdn.com/image/fetch/$s_!Marm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png 1272w, https://substackcdn.com/image/fetch/$s_!Marm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Marm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png" width="1456" height="635" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:635,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!Marm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png 424w, https://substackcdn.com/image/fetch/$s_!Marm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png 848w, https://substackcdn.com/image/fetch/$s_!Marm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png 1272w, https://substackcdn.com/image/fetch/$s_!Marm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe18aa7c0-8eb4-4a48-a101-081b6d23b53f_1760x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Source: <a href="https://documents.sandisk.com/content/dam/asset-library/en_us/assets/public/sandisk/collateral/company/Sandisk-HBF-Fact-Sheet.pdf">Sandisk HBF Fact Sheet</a> (July 2025) and <a href="https://arxiv.org/abs/2601.05047">Ma &amp; Patterson, IEEE Computer, May 2026</a>. Capacity and bandwidth are Sandisk figures; the power, latency, and read-granularity rows for both HBF and HBM4 come from Ma &amp; Patterson&#8217;s Table 3, a self-described &#8220;ballpark comparison,&#8221; and the per-watt rows are derived from those figures. HBM4 shown at the 6400 operating point (48 GB, 1.6 TB/s); the <a href="https://www.jedec.org/news/pressreleases/jedec%C2%AE-and-industry-leaders-collaborate-release-jesd270-4-hbm4-standard-advancing">JEDEC base spec</a> runs to 2 TB/s and 64 GB.</em></figcaption></figure></div><p><em>Note the 512 GB capacity is 14x the 36 GB 12-Hi HBM4 shipping today, ~10.7x the 48 GB 16-Hi flagship, and 8x the 64 GB JEDEC max.</em></p><p>For the same bandwidth, HBF has up to 8-16x the capacity at roughly 2x the power. </p><p>Yes, HBF has worse bandwidth per watt, but for model weight storage one can argue the metrics that matter most are capacity and capacity per watt, where HBF wins. </p><p>SanDisk also claims a <a href="https://documents.sandisk.com/content/dam/asset-library/en_us/assets/public/sandisk/collateral/company/Sandisk-HBF-Fact-Sheet.pdf">similar cost to an HBM stack</a> despite 8-16x the capacity, which works out to roughly 10x lower cost per GB. </p><p><em>I thought NAND was way cheaper?!</em></p><p><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">The raw NAND is cheap per bit, but an HBF </mark><em><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">stack</mark></em><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);"> isn&#8217;t cheap.</mark> Most of its cost is the same advanced packaging that makes HBM expensive like the 16-high TSV stacking, the CBA logic die, and the interposer. So you get HBM-stack pricing with ~10x the bits, not SSD pricing. </p><p>Of course, HBF has weaknesses compared to HBM. <em>It&#8217;s all tradeoffs.</em> </p><p>Microsecond latency against HBM&#8217;s ~100 ns, a minimum read 128x larger, and write endurance that rules out anything write-heavy.</p><p>Objections to HBF are primarily those three issues. But <mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">the painfulness of those weaknesses depends on the workload. And for decode inference, HBF isn&#8217;t so bad! </mark><em>We&#8217;ll cover it more below.</em></p><p>Interestingly, John Carmack has been thinking about this too; yesterday&#8217;s X post has over 1M views:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/ID_AA_Carmack/status/2074248758422864226&quot;,&quot;full_text&quot;:&quot;Memory cost and capacity are significant issues for AI accelerators.\n\nUnlike game rendering, model inference can have a deterministic memory access pattern. You don&#8217;t need &#8220;random access memory&#8221; at all for model weights, and you could tolerate cold-start latencies in the multiple&quot;,&quot;username&quot;:&quot;ID_AA_Carmack&quot;,&quot;name&quot;:&quot;John Carmack&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1560764938083352577/B1X3m4NN_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-06T21:46:56.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:250,&quot;retweet_count&quot;:420,&quot;like_count&quot;:6346,&quot;impression_count&quot;:1407116,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Click on it and read the whole thing! Here&#8217;s an important snippet:</p><blockquote><p><em><span>&#8220;model inference can have a deterministic memory access pattern. You don&#8217;t need &#8216;random access memory&#8217; at all for model weights, and you could tolerate cold-start latencies in the multiple milliseconds, as long as continuous reads were delivered at the necessary bandwidth.</span></em></p></blockquote><p>Clearly, there&#8217;s merit to flash for weights for inference. </p><p>There are lots of nuances too. </p><p>And many big questions. <em>How does HBF impact HBM demand? How does HBF impact the already hot NAND market?</em> I&#8217;ll address those and more.</p><p><strong>What paid subscribers get in the rest of this piece:</strong></p><ul><li><p><strong>Why decode doesn&#8217;t care about NAND&#8217;s weaknesses</strong></p></li><li><p><strong>The four ways HBF attaches to a GPU:</strong> direct HBM replacement, mixed HBM+HBF slots, HBM-as-cache, and disaggregated prefill/decode. Each implies a different GPU attachment rate and a different set of supply chain winners, and the earliest route to production isn&#8217;t the obvious one.</p></li><li><p><strong>The disaggregated cost model, worked:</strong> a table pricing the model weight tier for Llama 405B, DeepSeek-V3, and a 1T-class model. HBM vs HBF</p></li><li><p><strong>What HBF does to the DRAM and NAND markets</strong></p></li><li><p><strong>Route to market:</strong> Nvidia, hyperscaler custom silicon, what about ASIC startups and other merchant GPU vendors?</p></li><li><p><strong>A Sandisk patent that hints at the long term roadmap for HBF</strong></p></li><li><p><strong>Timeline and risks</strong></p></li></ul><p>And more!</p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/high-bandwidth-flash-the-full-report">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Micron's Blowout and the Case for a Memory Supercycle]]></title><description><![CDATA[Micron's GM hit 84.9%, it booked $10B of take-or-pay through 2030, and supply shortage stretched further out. Peak? When does supply increase? SK Hynix and Samsung too.]]></description><link>https://newsletter.chipstrat.com/p/microns-blowout-and-the-case-for</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/microns-blowout-and-the-case-for</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Fri, 26 Jun 2026 22:52:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!E2qY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As you probably heard, Micron crushed earnings this week. It got me thinking. Man, Micron has come a long way in the past year.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gwkr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gwkr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png 424w, https://substackcdn.com/image/fetch/$s_!gwkr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png 848w, https://substackcdn.com/image/fetch/$s_!gwkr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png 1272w, https://substackcdn.com/image/fetch/$s_!gwkr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gwkr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png" width="574" height="387.13461538461536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:982,&quot;width&quot;:1456,&quot;resizeWidth&quot;:574,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Micron has come a long way&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Micron has come a long way" title="Micron has come a long way" srcset="https://substackcdn.com/image/fetch/$s_!gwkr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png 424w, https://substackcdn.com/image/fetch/$s_!gwkr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png 848w, https://substackcdn.com/image/fetch/$s_!gwkr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png 1272w, https://substackcdn.com/image/fetch/$s_!gwkr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483d9e98-b6c0-4cb4-be70-0f47f90763d7_1750x1180.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Micron has come a long way</em></figcaption></figure></div><p>Just look at the past four quarters:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xsvm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xsvm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png 424w, https://substackcdn.com/image/fetch/$s_!xsvm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png 848w, https://substackcdn.com/image/fetch/$s_!xsvm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png 1272w, https://substackcdn.com/image/fetch/$s_!xsvm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xsvm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png" width="1390" height="602" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:602,&quot;width&quot;:1390,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:116557,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.chipstrat.com/i/203765473?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xsvm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png 424w, https://substackcdn.com/image/fetch/$s_!xsvm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png 848w, https://substackcdn.com/image/fetch/$s_!xsvm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png 1272w, https://substackcdn.com/image/fetch/$s_!xsvm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7df8fe5a-de4a-4263-b17a-2e16c018ffe1_1390x602.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Revenue went from $11.3 billion to $41.5 billion. Gross margin went from 45.7% to 84.9% (!!!)</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Sbre!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Sbre!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png 424w, https://substackcdn.com/image/fetch/$s_!Sbre!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png 848w, https://substackcdn.com/image/fetch/$s_!Sbre!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png 1272w, https://substackcdn.com/image/fetch/$s_!Sbre!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Sbre!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png" width="1257" height="753" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:753,&quot;width&quot;:1257,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Gross margin trend&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Gross margin trend" title="Gross margin trend" srcset="https://substackcdn.com/image/fetch/$s_!Sbre!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png 424w, https://substackcdn.com/image/fetch/$s_!Sbre!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png 848w, https://substackcdn.com/image/fetch/$s_!Sbre!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png 1272w, https://substackcdn.com/image/fetch/$s_!Sbre!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feda9ebab-fd7a-4a68-a630-bc1e52a0060d_1257x753.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Whaaaaat!</em></figcaption></figure></div><p>Each quarter beat the guidance Micron gave the quarter before, and the magnitude of the beat grew each time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iJno!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iJno!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png 424w, https://substackcdn.com/image/fetch/$s_!iJno!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png 848w, https://substackcdn.com/image/fetch/$s_!iJno!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png 1272w, https://substackcdn.com/image/fetch/$s_!iJno!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iJno!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png" width="1257" height="753" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:753,&quot;width&quot;:1257,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Guide vs actual: each quarter beat its prior-quarter guide, by a widening margin&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Guide vs actual: each quarter beat its prior-quarter guide, by a widening margin" title="Guide vs actual: each quarter beat its prior-quarter guide, by a widening margin" srcset="https://substackcdn.com/image/fetch/$s_!iJno!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png 424w, https://substackcdn.com/image/fetch/$s_!iJno!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png 848w, https://substackcdn.com/image/fetch/$s_!iJno!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png 1272w, https://substackcdn.com/image/fetch/$s_!iJno!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0d3347c-ef84-4fd0-8463-7e4a8dbbb4da_1257x753.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Guide vs actual: each quarter beat its prior-quarter guide, by a widening margin</em></figcaption></figure></div><p>And here&#8217;s a mind-bender. <mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">Micron&#8217;s fiscal Q3 revenue of $41.5 billion was larger than its revenue in any previous full year in company history.</mark> Bigger than the $37.4 billion it booked in all of FY2025. Bigger than the prior record of $30.8 billion in FY2022.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!E2qY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!E2qY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png 424w, https://substackcdn.com/image/fetch/$s_!E2qY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png 848w, https://substackcdn.com/image/fetch/$s_!E2qY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png 1272w, https://substackcdn.com/image/fetch/$s_!E2qY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!E2qY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png" width="1456" height="731" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:731,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;One quarter now tops any full year in Micron's history&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="One quarter now tops any full year in Micron's history" title="One quarter now tops any full year in Micron's history" srcset="https://substackcdn.com/image/fetch/$s_!E2qY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png 424w, https://substackcdn.com/image/fetch/$s_!E2qY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png 848w, https://substackcdn.com/image/fetch/$s_!E2qY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png 1272w, https://substackcdn.com/image/fetch/$s_!E2qY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97cae510-a391-4c60-ad10-1305668d7c4f_1557x782.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>One quarter now tops any full year in Micron&#8217;s history</em></figcaption></figure></div><p>Of course, what does that margin chart scream for folks who&#8217;ve been around the block? <em><strong>Peak</strong></em>. Memory is the most cyclical business in semiconductors. It&#8217;s like a sine wave:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K7Jo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K7Jo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png 424w, https://substackcdn.com/image/fetch/$s_!K7Jo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png 848w, https://substackcdn.com/image/fetch/$s_!K7Jo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png 1272w, https://substackcdn.com/image/fetch/$s_!K7Jo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K7Jo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png" width="1282" height="722" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:722,&quot;width&quot;:1282,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Micron quarterly revenue: the memory sine wave&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Micron quarterly revenue: the memory sine wave" title="Micron quarterly revenue: the memory sine wave" srcset="https://substackcdn.com/image/fetch/$s_!K7Jo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png 424w, https://substackcdn.com/image/fetch/$s_!K7Jo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png 848w, https://substackcdn.com/image/fetch/$s_!K7Jo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png 1272w, https://substackcdn.com/image/fetch/$s_!K7Jo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5881eb6e-1c11-4454-81ff-dead109b9635_1282x722.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Micron quarterly revenue: the memory sine wave. <a href="https://stockanalysis.com/stocks/mu/revenue/">Source</a></em></figcaption></figure></div><p>That nice little rolling wave there is the semiconductor cycle. </p><p>I love how Doug O&#8217;Laughlin visualized the cycle here with various markets moving around on a merry-go-round:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VFhU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VFhU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png 424w, https://substackcdn.com/image/fetch/$s_!VFhU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png 848w, https://substackcdn.com/image/fetch/$s_!VFhU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png 1272w, https://substackcdn.com/image/fetch/$s_!VFhU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VFhU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Doug O'Laughlin's semiconductor cycle map&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Doug O'Laughlin's semiconductor cycle map" title="Doug O'Laughlin's semiconductor cycle map" srcset="https://substackcdn.com/image/fetch/$s_!VFhU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png 424w, https://substackcdn.com/image/fetch/$s_!VFhU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png 848w, https://substackcdn.com/image/fetch/$s_!VFhU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png 1272w, https://substackcdn.com/image/fetch/$s_!VFhU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f33b82-992c-45c9-91ab-87ca20c93104_1993x1102.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Doug O&#8217;Laughlin&#8217;s semiconductor cycle map</em></figcaption></figure></div><p>Think of memory&#8217;s revenue wave as the weighted average of all the memory-consuming semis markets. Each one sits somewhere on the merry-go-round and carries weight equal to how much memory it buys, so the cycle you see in Micron&#8217;s revenues is really <mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">just the combined center of mass. </mark></p><p>Which raises the question&#8230; which quadrant is that center of mass in now, and what would tip it into the next one (top right)?</p><p>Well, the merry-go-round still exists, it&#8217;s just that the mass of datacenter AI is sooo big that all the other kids on the merry-go-round don&#8217;t meaningfully contribute to the center of gravity as much anymore. <em>AI is the dad who jumped on the merry-go-round with a couple of three-year-olds. Dad&#8217;s weight wins.</em></p><p>And yeah, I cropped that earlier Micron revenue chart for dramatic effect. Here&#8217;s what it looks like now:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WAeT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WAeT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png 424w, https://substackcdn.com/image/fetch/$s_!WAeT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png 848w, https://substackcdn.com/image/fetch/$s_!WAeT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png 1272w, https://substackcdn.com/image/fetch/$s_!WAeT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WAeT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png" width="1456" height="713" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:713,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Micron quarterly revenue: AI tips the wave&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Micron quarterly revenue: AI tips the wave" title="Micron quarterly revenue: AI tips the wave" srcset="https://substackcdn.com/image/fetch/$s_!WAeT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png 424w, https://substackcdn.com/image/fetch/$s_!WAeT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png 848w, https://substackcdn.com/image/fetch/$s_!WAeT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png 1272w, https://substackcdn.com/image/fetch/$s_!WAeT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22829bd7-1859-4b19-86b3-8ce51ccb7971_1540x754.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>YO! The amplitude of that sine wave is much bigger!</em></p><p>So let&#8217;s dig into that demand more and figure out if and when more supply is coming online. <mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">Can supply counterbalance, or will this memory supercycle persist?</mark></p><p>Behind the paywall:</p><ul><li><p><strong>Demand has two drivers&#8230; and don&#8217;t forget NAND:</strong> HBM is stacked DRAM, so it pulls many wafers off the DRAM market, right as server CPUs pull from the same pool. And AI is eating NAND too.</p></li><li><p><strong>The shortage goes out through 2027:</strong> CapEx up by half while the shortage horizon slid out a full year, and the construction-clock, node-shrink, and cleanroom limits behind it.</p></li><li><p><strong>What the contracts change:</strong> sixteen take-or-pay deals lock ~20% of DRAM and a third of NAND to 2030</p></li><li><p><strong>SK Hynix and Samsung say it on their own calls:</strong> the supply-demand gap widening into 2027, demand for the next three years above capacity, DRAM ASPs up ~60% in a quarter, HBM &#8220;sold out&#8221;.</p></li><li><p><strong>What still cuts the other way:</strong> ~80% of DRAM still reprices with the market, and Micron&#8217;s own guide already flags a &#8220;moderation in the rate of price increases&#8221;.</p></li><li><p><strong>The scorecard:</strong> the handful of variables that decide persist versus revert, each with a read and a date.</p></li></ul>
      <p>
          <a href="https://newsletter.chipstrat.com/p/microns-blowout-and-the-case-for">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The ISA Doesn't Matter Where It Counts]]></title><description><![CDATA[x86's lock-in is real. It's just furthest from the GPU.]]></description><link>https://newsletter.chipstrat.com/p/the-isa-doesnt-matter-where-it-counts</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/the-isa-doesnt-matter-where-it-counts</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Thu, 18 Jun 2026 22:40:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!wuD-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AMD, Intel, Nvidia, Arm, and Qualcomm are all selling datacenter CPUs into the AI buildout. The previous piece mapped them across five sockets orbiting the GPU and ranked those sockets by value: coherent host, standard host, thinker, doer, traditional cloud. </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;02a11ec5-3627-4183-8558-c663036df143&quot;,&quot;caption&quot;:&quot;AMD, Intel, Nvidia, and Arm are all selling datacenter CPUs into the AI buildout, and Qualcomm is trying to get in too. They are piling in because agentic AI turned the CPU from an afterthought into a fast-growing market, as the CPU-to-GPU ratio in AI infra has moved from ~1:4 toward 1:1.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Are Agentic CPUs a Commodity? It&#8217;s Complicated.&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8066776,&quot;name&quot;:&quot;Austin Lyons&quot;,&quot;bio&quot;:&quot;Chipstrat, Creative Strategies, Semi Doped. MSEE + MBA.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c180a750-7572-4aff-88e4-317aa435d533_1203x902.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-06-10T22:11:50.480Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ca431d67-f4a1-42c5-a055-a6ab42b449bd_1854x874.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.chipstrat.com/p/are-agentic-cpus-a-commodity-its&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:201514372,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:25,&quot;comment_count&quot;:0,&quot;publication_id&quot;:2003179,&quot;publication_name&quot;:&quot;Chipstrat&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!rCMl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27769444-42f3-4b43-9683-4fe7826c06b8_608x608.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>The coherent host is the most valuable. The traditional cloud CPU is the least.</p><p>Many readers asked if it matters whether the CPU is x86 or Arm.</p><p>Honestly, not as much as made out to be. But let&#8217;s go socket by socket.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wuD-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wuD-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png 424w, https://substackcdn.com/image/fetch/$s_!wuD-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png 848w, https://substackcdn.com/image/fetch/$s_!wuD-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!wuD-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wuD-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png" width="1456" height="968" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:968,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:167096,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.chipstrat.com/i/202647100?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wuD-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png 424w, https://substackcdn.com/image/fetch/$s_!wuD-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png 848w, https://substackcdn.com/image/fetch/$s_!wuD-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!wuD-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F230b5f07-3942-4526-8b93-08d3967ef323_1664x1106.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Quick context</h2><p>The ISA is the language a CPU speaks. Software gets compiled into that language, and a chip can only run code written for its dialect.</p><p>x86 has been the server default for decades. Yet Arm has been gaining in servers, first slowly, then quickly as Graviton, Axion, and Cobalt took hold in cloud, and now inside AI infrastructure as hyperscalers build Arm into their GPU server stacks.</p><p>Naturally, everyone asks which ISA is &#8220;better&#8221; for agentic AI; they&#8217;re both just fine.</p><p>The more interesting question at each socket is whether the software running there cares which ISA it runs on? Specifically, is the ISA a &#8220;moat&#8221; at any of the agentic sockets? Let&#8217;s see:</p><h2>1. Coherent host: ISA is irrelevant</h2><p><strong>The coherent host&#8217;s moat is the coherent link to the GPU, not its ISA. </strong></p><p>NVLink-C2C connects Nvidia&#8217;s Grace CPU to the Blackwell GPU at 900 GB/s, providing a shared address space in which the GPU reads CPU DRAM as if it were local. Vera doubles that to 1.8 TB/s with Rubin. Infinity Fabric ties AMD&#8217;s EPYC to the Instinct MI455X at comparable bandwidth. <strong>The coherent link is what makes this socket valuable.</strong> <em>It&#8217;s what no other CPU can replicate without a bilateral design agreement with the GPU vendor... like NVLink Fusion...</em></p><p>Before Grace, Nvidia GPU servers shipped with standard x86 hosts (Intel Xeon or AMD EPYC) connected over PCIe. Grace Hopper (2023) was Nvidia&#8217;s first coherent superchip: Grace CPU (Arm, Neoverse V2) connected to the Hopper GPU via NVLink-C2C at 900 GB/s &#8212; and Nvidia&#8217;s first deployment of the full datacenter CUDA stack on an Arm server CPU. <em>CUDA already ran on Arm through the Jetson embedded line, but this was the server-grade debut.</em></p><p>Grace Blackwell carried that forward; Vera Rubin extends it with a custom Arm CPU (88 Nvidia-designed cores) at 1.8 TB/s to Rubin. </p><p>So clearly, ISA isn&#8217;t a differentiator for the 800-lb gorilla. Host software runs on either. </p><p>What about AMD? ROCm is effectively x86-native. AMD&#8217;s coherent platform is built around EPYC, so an Arm port has naturally never been a priority.</p><p><strong>The main takeaway is that the ISA is baked into the accelerator platform choice.</strong></p><p>NVLink Fusion is Nvidia&#8217;s move to open the coherent-host socket to third-party CPUs. Previously, the only CPU that could claim a coherent seat on Nvidia&#8217;s backend was the one Nvidia built (Grace/Vera). NVLink Fusion allows other vendors to couple their processors to Blackwell GPUs over the same high-bandwidth coherent link Grace uses. <em>Note that no NVLink Fusion product has actually shipped yet, these are simply announced partnerships.</em> But the partner list includes <a href="https://nvidianews.nvidia.com/news/nvidia-nvlink-fusion-semi-custom-ai-infrastructure-partner-ecosystem">Qualcomm</a> (Arm), <a href="https://nvidianews.nvidia.com/news/nvidia-nvlink-fusion-semi-custom-ai-infrastructure-partner-ecosystem">Fujitsu</a>, <a href="https://newsroom.intel.com/artificial-intelligence/intel-and-nvidia-to-jointly-develop-ai-infrastructure-and-personal-computing-products">Intel</a> (x86), and <a href="https://www.sifive.com/press/sifive-nvidia-nvlinkfusion-datacenter">SiFive</a> (RISC-V). </p><p>If and when these ship, the coherent-host socket will be accessible to any ISA, so the moat is most definitely not the ISA. <em>RISC-V even... although lots of software porting required.</em></p><h2>2. Standard host: ISA nearly irrelevant, and eroding</h2><p>The standard host&#8217;s job is to keep the GPU fed: tokenize inputs, batch requests, stage data over PCIe, manage memory. The CPU needs to work as fast as possible and also move a lot of data. <em>PCIe can become a bottleneck here&#8230; hence the coherent host.</em></p><p>The hyperscalers started with x86 standard hosts paired with their XPUs, but that has moved toward Arm. AWS pairs <a href="https://aws.amazon.com/ai/machine-learning/trainium/">Graviton with Trainium</a>. Google pairs <a href="https://cloud.google.com/blog/products/compute/tpu-8t-and-tpu-8i-technical-deep-dive">Axion with its gen 8 TPUs</a>. </p><p>The feed-the-XPU stack runs on x86 or Arm interchangeably; ISA is not the moat.</p><p>Note that there is still an x86 standard-host business in smaller deployments, specifically enterprises and small neoclouds running DGX, Instinct MI355X, RTX Pro 6000 servers, and so on. </p><p>In these setups, the host often runs double duty with GPU feeding and application-tier workloads on the same box. That brings legacy x86 software dependencies back into the picture, and ISA does matter. Lower volume, but will grow.</p><p><strong>Takeaway: if the host is doing double duty as application processor, then ISA matters. Otherwise, nope.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TAAU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TAAU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png 424w, https://substackcdn.com/image/fetch/$s_!TAAU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png 848w, https://substackcdn.com/image/fetch/$s_!TAAU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png 1272w, https://substackcdn.com/image/fetch/$s_!TAAU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TAAU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png" width="1456" height="1298" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1298,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:200579,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.chipstrat.com/i/202647100?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TAAU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png 424w, https://substackcdn.com/image/fetch/$s_!TAAU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png 848w, https://substackcdn.com/image/fetch/$s_!TAAU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png 1272w, https://substackcdn.com/image/fetch/$s_!TAAU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F457ee094-1843-4304-a02b-a1263f10a056_1590x1418.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The two orbits closest to the GPU give the same answer: ISA does not matter there. The three that remain do not all agree. One has a real x86 lock-in story. One has a wrinkle. One&#8230; not so much. </em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.chipstrat.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Chipstrat is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div><hr></div><h2>3. Thinker: ISA irrelevant, and the category is still forming</h2><p>The <a href="https://www.chipstrat.com/p/are-agentic-cpus-a-commodity-its">thinker</a> is the GPU-coupled agent: reasoning-heavy work that runs close to the GPU on the scale-out backend, passes large contexts to the model, and needs per-core performance and low round-trip latency above all else. Real-time world model controllers, agents driving tight perception-action loops, anything where a single network hop eats the frame budget.</p><p>The thinker commands a higher ASP than the doer. That premium is earned on per-core performance and bandwidth, not protected by a proprietary link. The scale-out backend runs over open InfiniBand or Ethernet, reachable by any NIC-equipped CPU. AMD Venice Classic (Zen 6, high-frequency cores), Intel Diamond Rapids (P-cores, 2027), Nvidia Vera, Qualcomm Oryon, and Arm AGI in its thinker-flex configuration can all reach it. <strong>Every major ISA in the same fight.</strong></p><p>The software stack for thinker agents consists of Python, orchestration frameworks, and custom agent runtimes built over the last two years, running in containers and targeting both architectures from day one. There is no installed base of thinker software to be locked into, because the thinker socket barely exists at volume yet. Most agentic work today is coding agents, and that is a doer workload. <strong>So when the thinker software stack gets written at scale, it will be ISA-agnostic by default.</strong> </p><p>Hence, the ISA question is already settled before the volume arrives.</p><h2>4. Doer: ISA matters, with a specific wrinkle</h2><p><strong>The doer is the most valuable orbit where ISA matters.</strong></p><p>The doer is the action agent: coding, tool use, web search, API calls, compiling, running tests, filing pull requests. The metric is threads per watt &#8212; run as many agents as possible per rack. The agent&#8217;s own orchestration code is ISA-agnostic. <em>Python runs anywhere.</em> The frameworks that drive agents are all multi-architecture.</p><p>Yet the ISA matters in the execution environment the agent spins up.</p><p>If the target codebase compiles x86 binaries, e.g. a C++ service, a Go application, a Rust binary targeting x86 Linux and runs x86 test suites, you need x86. <em>Yes technically you could use an x86 CPU emulator on Arm but that would kill performance</em>.</p><p>So the doer&#8217;s ISA dependency is a dependency on what the <em>target codebase</em> compiles to, not on the agent runtime itself. A coding agent deployed to assist a Python web shop has no meaningful ISA dependency, but a coding agent deployed against a legacy enterprise C++ codebase does!</p><p>This matters for companies running agents against their own internal codebases. Their internal build pipelines, CI systems, and test suites were likely written assuming x86. <em>Years of accumulated toolchain dependencies.</em> The long tail of enterprise repos still carries enough x86-specific build logic that moving to Arm would require real migration work. That&#8217;s where ISA creates genuine friction today. <em>LLMs make it easier than ever, but it&#8217;s still work. Maybe Fable can one shot it though...</em></p><p>So <em>doer</em> agentic CPUs, especially in bigger enterprise deployments, have an ISA dependence.</p><h2>5. Traditional cloud: the strongest moat</h2><p>Agentic AI is also <a href="https://www.chipstrat.com/p/are-agentic-cpus-a-commodity-its">driving traditional cloud workloads</a>. Agents ultimately query ERPs, hit databases, and call APIs, all of which run on general-purpose cloud CPUs.</p><p>This is the orbit with real x86 legacy lock-in. Proprietary enterprise software from the 2000s and 2010s shipped as x86 binaries, sometimes with no source code and no porting path.</p><p>Definitely an ISA moat here. But the least value capture in the agentic AI path. Moreover, this ISA moat is shrinking over time. </p><p>But it&#8217;s real, especially for enterprises.</p><h2>Where it lands</h2><p>x86&#8217;s moat lives at the outer orbits: the traditional cloud socket, where old software licenses and ISV certifications built a real moat, and the doer socket, where it depends on what the agent actually executes. </p><p>The playing field is level at the inner orbits where the software is new, the frameworks were built for containers, and competing on power efficiency matters more than installed base.</p><p>The most valuable socket, the coherent host, does not care about the ISA. It cares about which accelerator agreed to share a coherent address space with you.</p><p><strong>The closer you get to the GPU, the newer the software and the less the ISA matters.</strong></p>]]></content:encoded></item><item><title><![CDATA[Are Agentic CPUs a Commodity? It’s Complicated.]]></title><description><![CDATA[Five sockets, different economics. A model for telling them apart, and where each competitor fits. AMD, Intel, Nvidia, Arm, Qualcomm.]]></description><link>https://newsletter.chipstrat.com/p/are-agentic-cpus-a-commodity-its</link><guid isPermaLink="false">https://newsletter.chipstrat.com/p/are-agentic-cpus-a-commodity-its</guid><dc:creator><![CDATA[Austin Lyons]]></dc:creator><pubDate>Wed, 10 Jun 2026 22:11:50 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ca431d67-f4a1-42c5-a055-a6ab42b449bd_1854x874.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AMD, Intel, Nvidia, and Arm are all selling datacenter CPUs into the AI buildout, and Qualcomm is trying to get in too. They are piling in because agentic AI turned the CPU from an afterthought into a fast-growing market, as the CPU-to-GPU ratio in AI infra has moved from ~1:4 toward 1:1.</p><p>So, who wins?</p><p><strong>Every one of them will tell you it&#8217;s them.</strong> I&#8217;ve spent a lot of time listening to executives at all of these companies, and each frames the comparison around the socket where its own chip happens to win. The market, meanwhile, sees the crowd of launches and weighs them all the same. <em>Up and to the right!</em> But these are really several different CPUs competing for different jobs, and they don&#8217;t all carry the same ASP or capture equal value; some sockets are near-monopolies, others are headed for a price war.</p><p><strong>To cut through the positioning, I needed a mental model to help answer my questions.</strong> <em>Which sockets matter? Which specs matter? Who competes where?</em> </p><p>That model is what follows.</p><h2>Framing: The CPU Sockets that Orbit the GPU</h2><p>Let&#8217;s start with the GPU at the center of our model. <em>Don&#8217;t let anyone tell you otherwise. CPUs are not the center of the AI universe; GPUs are. Also, I&#8217;m using GPU interchangeably for AI accelerator / XPU / GPU.</em></p><p>In this model, the orbits represent the CPU&#8217;s jobs to be done. Each job creates a socket a CPU can fill, and the closer the job is to the GPU, the more valuable that socket is. <em>I know you have lots of questions. But aren&#8217;t CPUs general-purpose? Does &#8220;closeness&#8221; to GPU truly matter? We&#8217;ll get to those.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ePxE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ePxE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png 424w, https://substackcdn.com/image/fetch/$s_!ePxE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png 848w, https://substackcdn.com/image/fetch/$s_!ePxE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png 1272w, https://substackcdn.com/image/fetch/$s_!ePxE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ePxE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png" width="558" height="534.4080882352941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1042,&quot;width&quot;:1088,&quot;resizeWidth&quot;:558,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Pasted image 20260609152402.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Pasted image 20260609152402.png" title="Pasted image 20260609152402.png" srcset="https://substackcdn.com/image/fetch/$s_!ePxE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png 424w, https://substackcdn.com/image/fetch/$s_!ePxE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png 848w, https://substackcdn.com/image/fetch/$s_!ePxE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png 1272w, https://substackcdn.com/image/fetch/$s_!ePxE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0893db0-fb0b-485f-87e9-a998ff3423c5_1088x1042.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The CPU sockets that orbit the GPU. Closer is more valuable.</figcaption></figure></div><h3>1) First orbit: the Host CPU</h3><p>Every GPU server is built around a host CPU, usually one or two sockets, that the GPUs attach to. It&#8217;s also called the head node, and its job is to run the GPU driver, launch kernels, tokenize and stage data, manage memory, and keep the accelerator fed. <em>Really, it&#8217;s just &#8220;do whatever&#8217;s necessary to keep the accelerator fed.&#8221; Some people joke the head node is a glorified memory controller.</em></p><p>The host CPU sits in the token path. Thus, the host CPU must NEVER be the bottleneck. Stalled GPUs are insanely expensive, and the host exists to make sure that never happens.</p><p>The specs that matter for a host CPU follow from that fact: high per-core performance (so it can keep up with the GPU and decide what to do next quickly), high bandwidth to the GPU (so it can move data without stalling it), and enough memory bandwidth and capacity to stage what the GPU needs. Core count is secondary. </p><h4>Nuance: coherent host vs. standard host</h4><p>Let&#8217;s get a bit nuanced.</p><p>The host orbit splits in two on the question, &#8220;Does the GPU need to use the CPU&#8217;s memory as an extension of its own?&#8221;</p><p>Early in the training era, the answer was no, and the link was PCIe. The host CPU stages data, the GPU copies it across the bus, the two keep separate memory pools. <em>PCIe Gen5 moves about 64 GB/s per direction, Gen6 about 128 unidirectional. That&#8217;s fine when the host is just feeding tokens.</em></p><p>Reasoning models changed the calculus. When a model thinks before it answers, its KV cache (the attention state it holds while generating) balloons, and when it spills out of GPU memory, the GPU has to tier it into CPU DRAM and read it back fast, inside the generation loop. <em>PCIe bandwidth is no longer enough to keep up.</em></p><p>So Nvidia built a coherent link, <a href="https://www.nvidia.com/en-us/data-center/nvlink-c2c/">NVLink-C2C</a>, that presents a shared address space, letting the GPU read CPU memory as if it were local. That is Grace Blackwell: a Grace CPU and a Blackwell GPU fused into one coherent module, sharing data at 900 GB/s, about seven times what PCIe Gen5 could move (Vera doubles it to 1.8 TB/s). </p><p><strong>The coherent host is the head node rebuilt for the reasoning era</strong>, and as context grows, more of its value moves across that coherence line. </p><p>So if we split the host orbit to account for the nuance, it looks like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!v5Nn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!v5Nn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png 424w, https://substackcdn.com/image/fetch/$s_!v5Nn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png 848w, https://substackcdn.com/image/fetch/$s_!v5Nn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png 1272w, https://substackcdn.com/image/fetch/$s_!v5Nn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!v5Nn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png" width="474" height="469.76785714285717" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/70db6620-3241-4614-96e4-6511e6100b0c_896x888.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:888,&quot;width&quot;:896,&quot;resizeWidth&quot;:474,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Pasted image 20260609153234.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Pasted image 20260609153234.png" title="Pasted image 20260609153234.png" srcset="https://substackcdn.com/image/fetch/$s_!v5Nn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png 424w, https://substackcdn.com/image/fetch/$s_!v5Nn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png 848w, https://substackcdn.com/image/fetch/$s_!v5Nn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png 1272w, https://substackcdn.com/image/fetch/$s_!v5Nn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70db6620-3241-4614-96e4-6511e6100b0c_896x888.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Coherent hosts are more valuable and higher volume right now, but also proprietary&#8230;</figcaption></figure></div><p>Closest in is the coherent host, tied to the GPU over a coherent link and built for KV-cache tiering and long-context reasoning. It is the most valuable seat and the most proprietary because that link exists only between an accelerator and the CPU it was designed with.</p><p>A step further out sits the standard host, tied to the GPU over ordinary PCIe and handling the tokenizing, batching, and feeding. It is less tightly coupled, but it has one thing the coherent host does not. It can sit in front of <em>any</em> accelerator, since PCIe is universal. That makes the standard host open, modular, multi-vendor territory.</p><h3>2) Second orbit: the Agents CPU</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!g0iF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!g0iF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png 424w, https://substackcdn.com/image/fetch/$s_!g0iF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png 848w, https://substackcdn.com/image/fetch/$s_!g0iF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png 1272w, https://substackcdn.com/image/fetch/$s_!g0iF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!g0iF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png" width="520" height="488.22810590631366" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:922,&quot;width&quot;:982,&quot;resizeWidth&quot;:520,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Pasted image 20260609154052.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Pasted image 20260609154052.png" title="Pasted image 20260609154052.png" srcset="https://substackcdn.com/image/fetch/$s_!g0iF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png 424w, https://substackcdn.com/image/fetch/$s_!g0iF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png 848w, https://substackcdn.com/image/fetch/$s_!g0iF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png 1272w, https://substackcdn.com/image/fetch/$s_!g0iF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F735b4be5-85d4-49f2-ad91-0a728c4598f5_982x922.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Let&#8217;s talk agentic CPUs. They serve a different purpose than hosts, and therefore need different specs</figcaption></figure></div><p>Agents are the new workload, and today&#8217;s agents are mostly <em>not</em> a GPU job. An agent is given a task, and it loops. It asks the model what to do, gets back an instruction, executes it (run some code, query a database, search the web, call an API), observes the result, and asks again. </p><p>The model&#8217;s thinking is one step in that loop (on the GPU). </p><p>Everything around it, the parsing, the tool calls, the code execution, the state management, runs on a CPU. That work would overwhelm the host CPU if you let it. </p><p>The host is in the token path and cannot afford to be busy doing anything else, and there are now thousands or millions of agents pinging the model, not a handful of humans. So the agent work gets offloaded to standalone CPUs dedicated to running agents. The goal there is different from the host, namely to run as many agents as possible in a given power footprint. </p><p><strong>The metric is threads per watt, not raw per-core speed.</strong> Lots of cores, lots of threads, low power, enough cache to keep each agent&#8217;s small working set on-chip.</p><p><strong>This orbit is where the CPU demand is exploding.</strong> It is already most of the reason the CPU-to-GPU ratio in AI infrastructure has moved from roughly 1:4 toward 1:1.</p><h4>Nuance: thinkers vs. doers (GPU-coupled vs. CPU-bound).</h4><p>Not all agent work is the same; some should run in the same datacenter as the GPUs, but much needn&#8217;t. Vik&#8217;s <a href="https://www.viksnewsletter.com/p/the-ai-datacenter-cpu-yellow-pages">Yellow Pages</a> splits CPUs on exactly this axis, reasoning vs. action, and the same split applies to the work itself.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zo6z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zo6z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png 424w, https://substackcdn.com/image/fetch/$s_!zo6z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png 848w, https://substackcdn.com/image/fetch/$s_!zo6z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png 1272w, https://substackcdn.com/image/fetch/$s_!zo6z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zo6z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png" width="516" height="499.0422535211268" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1236,&quot;width&quot;:1278,&quot;resizeWidth&quot;:516,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Pasted image 20260610140634.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Pasted image 20260610140634.png" title="Pasted image 20260610140634.png" srcset="https://substackcdn.com/image/fetch/$s_!zo6z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png 424w, https://substackcdn.com/image/fetch/$s_!zo6z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png 848w, https://substackcdn.com/image/fetch/$s_!zo6z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png 1272w, https://substackcdn.com/image/fetch/$s_!zo6z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a91a37-39f6-465e-b6e7-48dc626ec387_1278x1236.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">NOT ALL AGENTIC CPUS ARE THE SAME!!! </figcaption></figure></div><p><strong>The doers (CPU-bound, action agents).</strong> Here, the heavy lifting is on the CPU, and the GPU sits idle while it waits for the CPU. Coding is the most popular and a great example. Agents write code, compile it, run it, open a pull request, etc. Compiling and running can takes seconds, which dwarfs any network latency. So the extra hop for the GPU to talk over the front-end network to a CPU rack in another datacenter is no big deal. Hence, the doers can leave the GPU data hall and run on cheaper CPUs somewhere else.</p><p><strong>The thinkers (GPU-coupled, reasoning agents).</strong> Here, the heavy lifting is the model&#8217;s thinking, and per step the agent is mostly waiting on the GPU. A reasoning-heavy task that ships a large context to the model every step, then waits, is GPU-coupled. This work aims to stay close to the GPU, on the scale-out backend network inside the data hall, so round trips and large context transfers remain fast. And because each agent mostly waits on the GPU, one fast core can host many of them, <strong>so this socket prizes per-core speed first</strong> and core count only second.</p><p>Take, for example, agents that drive real-time generation and world models. <a href="https://www.worldlabs.ai/blog/taxonomy-of-world-models">Fei-Fei Li&#8217;s World Labs</a> frames a world model as a renderer, a planner, and the real-time loop connecting them, the same perception-action cycle that drives an embodied agent. The GPU renders a frame, a CPU reads the user&#8217;s or agent&#8217;s input and conditions the next step, and around again, all inside a frame budget of 16 to 33 milliseconds (30 to 60 FPS). A single network hop eats that budget, so the controlling CPU has to sit on the backend, next to the GPU. It is a reasoning job, not an action one, so the CPU does little computing but has to be fast and close. Thus, <strong>we need high per-core performance and low latency</strong>, not necessarily high thread count. </p><p>And as the world&#8217;s persistent state grows the way a long-context KV cache does, it pushes toward the coherent host. As AI moves <a href="https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence">from words to worlds</a>, more of the agentic workload lands in this tightest, most GPU-coupled corner.</p><p><strong>Side note:</strong> The strongest case for keeping even the doers on the backend is the agent&#8217;s growing context. A long session accumulates a large context, and every turn the model attends over all of it through the KV cache, the attention state it builds from those tokens. Recomputing that cache each turn is wasteful, so the inference system keeps and reuses it, tiering it out of GPU memory into a dedicated context-memory rack as it grows (Nvidia builds this with BlueField DPUs, fast storage, and KV-aware routing in Dynamo). But that pulls the KV cache onto the backend, not the doer. The agent&#8217;s memory is really just tokens, and its durable memory (files, a vector database, retrieval stores) never touches the GPU at all. The doer ships a prompt and a tool result, text in and text out, and the inference system turns that into KV and caches it. So the context-memory rack is real backend infrastructure, and the doer&#8217;s compute can still leave.</p><p><strong>One caveat:</strong> if the work produces a very large artifact the model then has to reason over, you keep it near the backend so the big payload does not crawl across a slow link to reach the GPU. </p><h3>3) Third orbit: traditional cloud CPUs</h3><p>The outer orbit is the CPU as we have always known it, the general-purpose server CPU running web services, databases, microservices, and VMs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NT86!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NT86!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png 424w, https://substackcdn.com/image/fetch/$s_!NT86!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png 848w, https://substackcdn.com/image/fetch/$s_!NT86!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png 1272w, https://substackcdn.com/image/fetch/$s_!NT86!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NT86!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png" width="566" height="520.9065934065934" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1340,&quot;width&quot;:1456,&quot;resizeWidth&quot;:566,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Pasted image 20260610140944.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Pasted image 20260610140944.png" title="Pasted image 20260610140944.png" srcset="https://substackcdn.com/image/fetch/$s_!NT86!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png 424w, https://substackcdn.com/image/fetch/$s_!NT86!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png 848w, https://substackcdn.com/image/fetch/$s_!NT86!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png 1272w, https://substackcdn.com/image/fetch/$s_!NT86!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33a561f4-e60f-4c60-bcbd-498897b9f865_1500x1380.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">This is the mental model.</figcaption></figure></div><p>Cloud CPUs are not in the AI loop directly. But agents often reach out to all of it; the ERPs, the CRMs, the databases, the web servers, and so on. And the more that agents multiply, the more ordinary cloud-CPU demand they pull along behind them.</p><h4>Nuance: there is a portfolio here too</h4><p><em>I know you all know this, but to be pedantic: </em></p><p>Traditional cloud buyers have always chosen cloud CPU SKUs by workload. General-purpose, compute-optimized, memory-optimized, storage-optimized, etc. A per-core-licensed database wants fewer, faster cores, because the license is priced per core. A web tier wants density and throughput per watt. That same split, fast cores vs. dense cores, runs straight through every vendor&#8217;s lineup, and it is why &#8220;the cloud CPU&#8221; was never one SKU.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VzNs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VzNs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png 424w, https://substackcdn.com/image/fetch/$s_!VzNs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png 848w, https://substackcdn.com/image/fetch/$s_!VzNs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png 1272w, https://substackcdn.com/image/fetch/$s_!VzNs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VzNs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png" width="566" height="526.3489010989011" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1354,&quot;width&quot;:1456,&quot;resizeWidth&quot;:566,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Pasted image 20260610142330.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Pasted image 20260610142330.png" title="Pasted image 20260610142330.png" srcset="https://substackcdn.com/image/fetch/$s_!VzNs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png 424w, https://substackcdn.com/image/fetch/$s_!VzNs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png 848w, https://substackcdn.com/image/fetch/$s_!VzNs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png 1272w, https://substackcdn.com/image/fetch/$s_!VzNs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe32599f9-6974-447c-af3c-1a77dfad7028_1622x1508.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Lots of different SKUs and ASPs in the cloud CPU portfolio</figcaption></figure></div><h2>Where is the Value and Who Competes at Each Socket</h2><p>Put it together and you have a spectrum of CPUs that all benefit from agentic AI, but not in the same place or to the same degree. Map them onto the orbits above, and each lands in a different socket: coherent host, standard host, thinker, doer, or traditional cloud, and some in more than one.</p><p><strong>Those sockets are not worth the same. One is close to a monopoly; the rest will be in a price war.</strong></p><p>This is what current head-to-head comparisons miss. Every vendor counter-positions, framing the whole &#8220;agentic CPU&#8221; question around the socket where its own part wins. Nvidia leans on the coherent host and thinker, AMD and Arm on the doer&#8217;s rack-scale density, and so on. The marketing then reads as if everyone leads, because each is measuring a different socket.</p><p>To know who actually captures value, you have to ask two things of each chip. Which socket is it really built for, and what is that socket worth?</p><p>Behind the paywall, I answer both. </p><p>First, the money. What each of the five sockets is worth, which one is a near-monopoly, and which are headed for a price war. </p><p>Then the field. Where Nvidia&#8217;s Vera, AMD&#8217;s EPYC, Intel&#8217;s Xeons, Arm&#8217;s AGI rack, and Qualcomm&#8217;s CPU each land, who owns the moat, who is the most complete, who is shut out of the one socket that actually pays, and who is renting a way in. </p>
      <p>
          <a href="https://newsletter.chipstrat.com/p/are-agentic-cpus-a-commodity-its">
              Read more
          </a>
      </p>
   ]]></content:encoded></item></channel></rss>