The inference engine for the agentic era.
GPU-class inference with the simplicity of a CPU, on commodity memory. Agents think and act on one chip, with all the memory they need.
Think and act on one chip
More tokens per dollar
Abundant memory holds every user's context, so one read of the model serves them all, and the cost of each token falls.
Performance that scales at every level
One matrix engine, repeated from the tensor unit up to the tray. Each level reuses data from the one below, so performance grows without a new design.
Open at every layer
Open hardware, an open software stack, one unified Ethernet fabric and air cooling. Nothing proprietary to adopt, and nothing to be locked into.
Running today
Open models run end to end on our RISC-V prototype, under Linux.
What is the capital of France? Answer in one word.
Paris