BTC $84,838.05 -2.26%
ETH $2,681.92 -2.62%
BNB $770.64 -1.18%
XRP $1.49 -3.85%
SOL $119.35 -2.38%
TRX $0.3363 +0.41%
DOGE $0.0927 -4.34%
ADA $0.2450 -4.81%
BCH $311.48 -1.46%
LINK $13.89 -3.61%
HYPE $87.93 -3.43%
AAVE $181.10 -1.28%
SUI $1.18 -0.42%
XLM $0.2148 -4.80%
ZEC $1,308.79 -5.82%
AAPL $333.26 +0.31%
AMZN $251.80 +0.41%
GOOGL $342.85 +0.38%
MSFT $517.40 -0.33%
META $728.51 -0.68%
NVDA $234.26 -0.72%
TSLA $370.93 +3.77%
SNDK $1,716.56 -3.28%
INTC $118.19 -4.41%
SPCX $159.09 +6.24%
MU $1,069.35 -3.29%
AMD $632.27 +0.07%
BTC $84,838.05 -2.26%
ETH $2,681.92 -2.62%
BNB $770.64 -1.18%
XRP $1.49 -3.85%
SOL $119.35 -2.38%
TRX $0.3363 +0.41%
DOGE $0.0927 -4.34%
ADA $0.2450 -4.81%
BCH $311.48 -1.46%
LINK $13.89 -3.61%
HYPE $87.93 -3.43%
AAVE $181.10 -1.28%
SUI $1.18 -0.42%
XLM $0.2148 -4.80%
ZEC $1,308.79 -5.82%
AAPL $333.26 +0.31%
AMZN $251.80 +0.41%
GOOGL $342.85 +0.38%
MSFT $517.40 -0.33%
META $728.51 -0.68%
NVDA $234.26 -0.72%
TSLA $370.93 +3.77%
SNDK $1,716.56 -3.28%
INTC $118.19 -4.41%
SPCX $159.09 +6.24%
MU $1,069.35 -3.29%
AMD $632.27 +0.07%
hot_img

Cerebras releases the fourth generation AI inference system CS-4: performance doubled, power consumption doubled, more flexible deployment

2026-08-19 10:38:29

Cerebras released its fourth-generation AI inference system CS-4 this week, based on the same 5nm WSE-3 wafer, achieving double the performance by doubling the clock frequency and power consumption. A single CS-4 cabinet accommodates 3 wafers (CS-3 has 2), featuring a modular "backpack" design that simplifies manufacturing and deployment, with a TDP of approximately 125 to 135kW. The CS-4 can provide an inference speed of nearly 4000 tokens/second/user, about twice that of the CS-3, and supports decomposed inference with heterogeneous systems such as AMD and AWS Trainium.

Cerebras claims that the CS-4 offers about 2000 times the on-chip memory bandwidth of NVIDIA's Rubin (43PB/s), but the 44GB SRAM capacity remains unchanged, and long-context inference still requires multi-wafer stacking. For example, with the DeepSeek V4 Pro (1.6T parameters), approximately 20 systems are needed for a 1M context window, and about 40 systems are required for 256 concurrent users, corresponding to a CAPEX exceeding 20 million USD. Cerebras is collaborating with clients such as OpenAI and plans to achieve approximately double performance improvements each year, aiming for a 20-fold throughput increase by 2027. The "backpack" cabinet design of the CS-4 will continue into the next-generation "Nexus" platform.

app_icon
ChainCatcher Building the Web3 world with innovations.