Reports
AI-generated structured vendor updates
NVIDIA
Other
1970-01-01
SGLang 0.5.13 Delivers 25x MoE Inference Speedup via Predictive Routing and Sparse KV Cache
SGLang 0.5.13 introduces two-stage MoE routing prediction and sparse KV cache, achieving a 25x inference speedup on NVIDIA GB300 NVL72. Benchmarks on A100 show 65% throughput gain, 40% latency reduction, and 62% lower routing overhead. This optimization directly attacks the core bottleneck of MoE inference, potentially reshaping AI inference economics.
Anthropic
Other
1970-01-01
US Export Controls Force Anthropic Global Shutdown: AI Model Deployment Hits Compliance Architecture Gap
Anthropic globally pulls Fable 5 and Mythos 5 due to inability to filter users by nationality under US export controls. White House talks fail, jeopardizing $965B IPO. Highlights compliance architecture gaps in AI model deployment.