NVIDIA was refreshingly open about how performance of the technology has improved, and the numbers are impressive. They started out needing dual RTX 5090s, and have made the model five times faster since, through a combination of kernel optimization, hardware co-design, and model-side research. Edward's summary was that the model shown today is smaller, faster, more realistic and more controllable than the one at GTC.
Asked whether the model is NVFP4 quantized, given that it is Blackwell-exclusive anyway, the answer was "not yet," so improved performance seems on the horizon.