Back to Learn
    blog 7 min read

    Cerebras and the Coming IPO: The Wafer-Scale Challenger to NVIDIA

    Cerebras builds the world's largest AI chip and is heading to the public markets. Here is what they actually do, how their wafer-scale architecture compares to NVIDIA GPUs, and what their IPO means for the AI infrastructure race.

    88

    88 Labs AI

    Editorial Team

    Cerebras and the Coming IPO: The Wafer-Scale Challenger to NVIDIA
    Share:

    # Cerebras and the Coming IPO: The Wafer-Scale Challenger to NVIDIA


    For most of the last decade, the conversation about AI hardware has had exactly one main character: NVIDIA. Every training run, every inference benchmark, every earnings call seemed to orbit Jensen Huang's GPUs. But a quieter contender has been building something radically different — a single chip the size of a dinner plate — and it is finally heading for the public markets.


    Cerebras Systems is preparing for an IPO, and it is the most serious architectural challenge to NVIDIA's dominance in years.


    What Cerebras Actually Does


    Cerebras designs and builds AI accelerators, but they do not look anything like a GPU. Their core product is the Wafer-Scale Engine (WSE) — a processor that occupies an entire 12-inch silicon wafer instead of being cut into hundreds of small chips.


    The current generation, WSE-3, packs:


  1. 900,000 AI-optimized cores on a single piece of silicon
  2. 44 GB of on-chip SRAM (yes, on the chip, not attached over a bus)
  3. 4 trillion transistors
  4. 21 PB/s of memory bandwidth — orders of magnitude beyond any GPU

  5. That chip lives inside the CS-3 system, a refrigerator-sized box that behaves, from a software perspective, like one giant accelerator. Cerebras also runs Cerebras Inference, a hosted API that serves open models like Llama 3 and Qwen at speeds the GPU world genuinely cannot match — frequently quoted at 1,000–2,000+ tokens per second for 70B-class models, where high-end GPU stacks land in the low hundreds.


    The customer list is no longer experimental: G42, Mayo Clinic, AstraZeneca, Meta (as a partner for Llama API), the U.S. Department of Energy, and a growing roster of sovereign AI projects in the Middle East.


    How It Compares to NVIDIA


    NVIDIA's playbook is scale-out: take a great GPU (H100, B200), pack eight of them into a server, then network thousands of those servers together with NVLink and InfiniBand. The chip is small, the cluster is enormous, and the software stack — CUDA — is the real moat.


    Cerebras's playbook is scale-up: make the chip itself as big as physics allows, so that the work that normally requires thousands of GPUs talking over a network can happen inside a single piece of silicon.


    Here is the honest side-by-side:


    | | Cerebras WSE-3 / CS-3 | NVIDIA B200 / GB200 NVL72 |

    |---|---|---|

    | Form factor | One wafer-scale chip | Many discrete GPUs networked together |

    | On-chip memory | 44 GB SRAM on die | ~192 GB HBM3e per GPU (off-die) |

    | Memory bandwidth | ~21 PB/s on-chip | ~8 TB/s per GPU |

    | Best at | Long-context inference, large models without sharding, tightly-coupled training | General-purpose AI + HPC, the broadest software ecosystem |

    | Software | PyTorch via Cerebras SDK | CUDA, the industry default |

    | Deployment model | Buy a CS-3, or use Cerebras Inference API | Buy GPUs, rent from every cloud, or use any inference provider |

    | Ecosystem | Narrow but deep | The entire AI industry |


    The short version: NVIDIA wins on ecosystem, optionality, and sheer ubiquity. Cerebras wins on raw speed-per-model and on workloads where moving data between chips is the bottleneck — which, increasingly, is most of frontier AI.


    Why The IPO Matters


    Cerebras filed confidentially in 2024 and has been working through a regulatory review tied to its largest customer, G42 in the UAE. The pieces have been falling into place: revenue has scaled into the hundreds of millions, the inference business is growing fast, and the appetite for "anything but NVIDIA" among hyperscalers and sovereign buyers is the strongest it has been since the AI boom started.


    A successful IPO would do three things:


    1. Give Cerebras the capital to build more CS-3 fleets and start shipping the next generation, which is critical because their scale-up approach demands enormous up-front fab and packaging investment.

    2. Create a public market benchmark for "non-NVIDIA AI compute," the way Snowflake created one for cloud data warehouses. Investors want a way to bet on the picks-and-shovels story without being 100% long NVIDIA.

    3. Pressure NVIDIA's pricing and roadmap. NVIDIA is not going to lose its lead, but a credible, public, well-capitalized rival changes the negotiating dynamic for every large GPU buyer.


    It also raises real questions: customer concentration (G42 is a huge share of revenue), the difficulty of competing with CUDA's lock-in, and whether wafer-scale economics work outside of a handful of premium use cases.


    What This Means If You Are Building With AI


    Most teams reading this are not buying a CS-3. But the Cerebras story matters anyway, for two reasons.


    First, inference is becoming a commodity with very different speed tiers. A year ago, "fast inference" meant 60 tokens/second. Today, Cerebras and a few others are serving the same open models at 10–20× that rate. If your product depends on agent loops, real-time voice, long-context reasoning, or anything where latency is UX, you should be benchmarking across providers — not assuming the GPU API you used last year is still the right one.


    Second, the hardware layer is finally diversifying. NVIDIA, AMD MI300X, Google TPUs, AWS Trainium, Groq, and Cerebras are all viable for different workloads. The right answer is increasingly "use the chip that fits the job," not "default to whatever your cloud provider stocks."


    Where 88 Labs AI Comes In


    Picking the right model and the right inference provider is a small but real part of what we do when we deploy custom AI agents in 14 days. The bigger work is making sure the agent is wired into your business correctly: the tools it can call, the data it can see, the guardrails on what it can do, and the workflow it actually replaces.


    If you want an agent that takes advantage of the new generation of fast, cheap inference — without having to learn the difference between a wafer and a GPU — see your free demo. We will show you exactly what we would build, on your real workflow, before you spend a dollar.


    Ready to see this in action?

    Get a free, personalized demo of an AI agent built for YOUR business.

    Get Your Free Demo