Cerebras just put 750 PFLOPs into a rack that runs frontier models 30x faster per user than GPU clusters, at 10x better throughput per watt than its own prior generation. The interesting part is not the speed number. It is that Wafer Scale Engine architecture has crossed the threshold where it out-executes GPU chiplet stacks on the metric that actually determines datacenter economics: tokens per second per dollar of power.
WSE-3T solves a bottleneck GPU clusters cannot solve by adding more chips: inter-chip coordination overhead. GPU clusters spend bandwidth on NVLink, InfiniBand, and all the fabric tax between dies. The CS-4 puts 160.5 PB/s of compute fabric bandwidth inside a three-wafer rack with two-microsecond wafer-to-wafer latency. At that latency, a model with 50 trillion parameters that would require a full GPU cluster becomes a single-rack problem. The 129.6 PB/s memory bandwidth number is more directly load-bearing than the PFLOP count: it determines how fast the weights can be streamed through the compute, which is the actual constraint on token rate for frontier-scale models.
Hardware teams designing inference silicon against GPU-cluster reference architectures now have a different competitive bar to clear. 30x tokens-per-second-per-user is not a benchmark footnote; it is a product experience floor that AI application teams will start to treat as the baseline. Teams still building toward GPU-cluster-competitive inference architectures should model whether they can close within 5x of this on total cost of ownership within 18 months. If they cannot, WSE-native inference is the architecture to partner against, not the one to outbuild.