← Writing

The AI Inference Stack

AI entering the public markets

Anthropic, the creator of Claude, will soon IPO at a $2 trillion valuation, bigger than the blockbuster SpaceX IPO in May. Until now, Anthropic and OpenAI have been private companies, inaccessible to retail investors. Buying shares of a pricey IPO like Anthropic is a risky play: over 35% of SpaceX's publicly available stock is being sold short.[1]
Short-selling occurs when hedge funds borrow shares from a broker, sell them, and expect to buy them back later at a lower price.

What is AI inference?

The massive market capitalizations of Anthropic, OpenAI, and SpaceXAI do not lie. AI is transforming every industry, and these companies are seeing revenues in the hundreds of billions. The key driver of that growth is AI inference: the service of running models on physical silicon. Including Anthropic and OpenAI's models, AI inference as a market could be as large as electricity, and there is no single player set to dominate — not even OpenAI or Anthropic. In reality, models are coming from an ecosystem of vendors including NVIDIA, and unlike Claude or ChatGPT, many of them are open source, which means anyone can run them and make money.

Which companies are in the stack?

There are three primary layers to AI inference:

  1. Application layer. Companies that provide a specific service, like AI accounting, to their customers. They may serve their own open-source model or use the frontier models from OpenAI and Anthropic.
  2. Cloud layer. The "cloud" is the management of server hardware by a third-party vendor. In the AI age, "neo-cloud" businesses purchase computing hardware and provide an interface for application-layer companies to build on top of.
  3. Hardware layer. Think of AMD or NVIDIA. These businesses provide the physical chips that cloud companies purchase and manage. They benefit the most when customers plan for large inference spend and can commit to buying chips up front.[2]

Where are the best opportunities to invest? Most application-layer companies are still private, which restricts retail investors. The business model for cloud companies like Oracle Cloud Infrastructure (ORCL) involves buying a lot of hardware up front and renting it to the application layer. That is why CoreWeave (CRWV) and Nebius (NBIS) have already gone public, even at far lower market capitalizations than the application-layer companies, to seek extra investment. Being a cloud company, however, offers little room for innovation.

Can AI hardware be reinvented?

Besides NVIDIA and AMD, an outcropping of hardware companies is popping up to serve the inference workload. NVIDIA started by creating graphics processing units for consumer gamers and later pivoted that hardware to machine learning. Those chips were not designed for the specialized requirements of AI inference.

A useful analogy is the airplane. Imagine you needed to serve both fighter pilots and commercial travelers. A Boeing 737 could have the seats ripped out and missiles added, but the plane is still not designed for high-speed flight. Specialized AI chips are more like fighter jets: equipped with the speed and technology the workload actually requires.

The space of specialized hardware

The hardware layer covers a large surface area. The focus here is companies that list AI co-designed processors as their product. Within that category:

  1. Axelera develops two processors, Europa and Metis, which focus on increased energy efficiency for a smaller number of TOPS (trillion operations per second).
    The focus on energy efficiency benefits industries such as robotics or edge devices, where data-center-level power is not feasible.
  2. Rebellions AI also develops a processor that is more energy efficient than competitors like NVIDIA.
  3. Tendrils is designing a hardware architecture for CPUs that uses interaction nets to massively speed up computation for dependency-graph problems, such as finding the longest subsequence of two strings.
  4. Etched designed and already shipped a custom ASIC for AI inference, with specialized components for the prefill and decode stages of an LLM.
  5. Groq develops racks of LPU chips with 128GB of SRAM across the chips. Compared with NVIDIA chips, Groq chips have ultra-fast memory access because the memory sits close to the silicon, which helps LLM customer experience. NVIDIA entered a $20 billion licensing agreement with Groq in 2025.

Who are the hardware suppliers for these chips?

Below are the public stock suppliers for each of these chip companies, sourced from vendor relationships disclosed in company blogs and third-party articles. None of these vendors have been validated in a bill of materials.

  1. Axelera
    1. Samsung (SSNLF). A vice president of the Foundry Technology Planning Team is quoted saying that Samsung manufactures Axelera's Europa chip with its 5nm semiconductor technology.[3]
    2. TSMC (TSM). This competing foundry manufactured Axelera's 12nm Metis chip, the slightly older chip in Axelera's lineup.[3]
    3. Andes Technology (TWSE: 6533). Andes, a Taiwanese manufacturer of 32-bit and 64-bit processor cores, published a press release stating they are providing the AX65 processor for Axelera's Europa chip.[4]
    4. Advantech (TPE: 2395). Advantech sells embedded systems that package AI chips like Europa. Axelera lists Advantech as one of its partner systems.[5]
  2. Rebellions
    1. Samsung (SSNLF). A different vice president of the Foundry Technology Planning Team said in a Rebellions press release that they are "proud to support Rebellions with our 4nm process."[6] Samsung is also said to supply Rebellions with HBM memory.[7]
    2. SK Hynix (SKHY). In the same reporting, SK Hynix is described not only as a provider of HBM memory but also as an investor in Rebellions, along with its sister company, SK Telecom.[7]
    3. Qualcomm (QCOM). Rebellions is reported to have "licensed its UCI-Express-A controller from Alphawave Semi."[7] That interconnect technology connects multiple chips together in order to scale total processing capability. Alphawave was acquired by Qualcomm in 2025.
    4. Arm (ARM). Rebellions uses pre-built Arm-designed components for its REBEL chip.[8]
    5. Marvell (MRVL). Marvell provides Rebellions with "signaling SerDes, chip-to-chip interconnects, and advanced packaging."[7]
  3. Tendrils
    Tendrils is early enough that little is public about their hardware. Their job board lists a role requiring experience with FPGA prototyping: custom hardware circuits aimed at massively parallelizing operations. Vendors for this kind of production include:
    1. Synopsys (SNPS). Provides EDA (electronic design automation) software products and AI workflow optimization.[9]
    2. Cadence (CDNS). Also provides EDA products and is working on an AI chip design tool.[10]
    3. Siemens (SIEGY). Same category: EDA tools, now with an AI-design push.[11]
  4. Etched
    1. TSMC (TSM). Etched is said to have manufactured their chip with TSMC's N4P technology, an improved 5nm process.[12]
  5. Groq
    1. Samsung (SSNLF). Manufactures Groq's LPU chip. NVIDIA, notably, works exclusively with TSMC.[13]

Other companies in the thesis

American Turbines is building mass-manufacturable, compact gas turbines to power data centers. Karman Industries is developing modular heating units that use excess heat from the data center to cool their racks. Both are betting on a rapid increase in data-center energy demand and on the need for more compact, modular infrastructure. That energy thesis lines up with the chip startups aiming to specialize their silicon and cut energy consumption.
More chip-related public and private companies
  1. Astera Labs
  2. Companies listed on Arm's Total Design partner wall

This is not investment advice.

References

  1. “SpaceX’s $101 Billion Unlock Heaps Pressure on Battered Shares.” Bloomberg, 5 Aug. 2026, bloomberg.com.
  2. Etched. etched.com.
  3. “The Edge of Tomorrow.” Data Center Dynamics, 20 Aug. 2026, datacenterdynamics.com.
  4. “Axelera AI and Andes Technology Partner to Power Next-Generation Europa AI Platform with High-Performance RISC-V AX65 Cores.” Andes Technology, 1 June 2026, andestech.com.
  5. “Advantech.” Axelera AI, axelera.ai/systems/advantech.
  6. “Rebellions Debuts REBEL-QUAD at Hot Chips 2025, Breaking AI’s Energy Tax with High-Performance Chiplet Innovation.” Rebellions, rebellions.ai.
  7. “Rebellions AI Puts Together an HBM and Arm Alliance to Take on NVIDIA.” The Next Platform, 23 Dec. 2025, nextplatform.com.
  8. “Rebellions Joins Arm Total Design to Drive Next-Gen AI Infrastructure Solutions.” Rebellions, rebellions.ai.
  9. “AI.” Synopsys, synopsys.com/ai.html.
  10. “AI Overview.” Cadence, cadence.com.
  11. “EDA AI.” Siemens, siemens.com.
  12. “Inference Chip Startup Etched Emerges from Stealth with $800m Funding, Unveils Working Chip.” Data Center Dynamics, datacenterdynamics.com.
  13. “Nvidia Says Groq Racks Will Be Online This Year After $20 Billion Deal.” CNBC, 24 Aug. 2026, cnbc.com.