GLM-5.3-FlashX launches on B.AI with API and Web Chat support
The B.AI platform announced on Sept. 21 the launch of GLM-5.3-FlashX, a high-speed multimodal model developed by Z.AI. The model features a sparse architecture with 320B total parameters and 18B active parameters, achieving generation speeds of up to 200 tokens per second, which is approximately five times faster than the GLM-5.3-Flash version. It supports a 1M token context window and native video and file understanding. Developers can access the model via the B.AI platform or the web chat interface.
Summaries are written by AI from the original article. Not investment advice.