Reports
AI-generated structured vendor updates
NVIDIA and Wistron Open US Factory for GB300 and Vera Rubin AI Superchips
Wistron opens its first US manufacturing facility in Fort Worth, producing NVIDIA GB300 Grace Blackwell Ultra and Vera Rubin superchips. The $700M plant aims for tens of thousands of boards monthly, marking NVIDIA's strategic shift to domestic AI hardware production.
NVIDIA Vera Rubin Platform Specs Revealed: 10x Tokens per Watt, Monolithic CPU+GPU Design
NVIDIA unveiled Vera Rubin platform specs with a monolithic design pairing 2 Rubin GPUs with 1 Vera CPU, flagship NVL72 integrating 36 CPUs and 72 GPUs. Claims 10x tokens per watt and 3x memory bandwidth over Grace Blackwell. Vera CPU sold standalone. First customers: Microsoft, OpenAI, Oracle. Mass production H2 2026. Performance claims await independent verification.
Google's Frozen v2 Chip Hardwires Gemini Architecture for 6-10x TPU Efficiency, Set for 2028
Google is developing Frozen v2, a dedicated AI chip that hardwires the Gemini model architecture into silicon for 6-10x energy efficiency per token over current TPUs. It is a new product line, planned for 2028, with weight update flexibility but a frozen architecture. This validates the industry shift from general-purpose GPUs to dedicated ASICs for AI inference.
Meta Launches Muse Spark 1.1 API at 25% Competitor Price, Ends Open-Source Era
Meta releases Muse Spark 1.1, a multimodal reasoning model with 1M token context window, and launches its first paid API at 25% of competitors' price. This ends the Llama open-source era, signaling a strategic shift to proprietary API monetization and aggressive market share capture.
Moonshot AI Launches Kimi K3: 2.8T Parameter Open-Source MoE Model at $3/$15
Moonshot AI unveils Kimi K3, a 2.8T parameter open-source MoE model with 896 experts (16 active), native vision, and 1M context. API pricing at $3/$15 input/output per million tokens undercuts rivals. Open weights release July 27. GPU capacity exhausted within 2 days. Arena score 1679 tops Fable 5.
NVIDIA Agent Toolkit Shifts AI Agent Control from Cloud to Local DGX Station
NVIDIA launches Agent Toolkit for DGX Station, comprising NemoClaw, Nemotron 3 Ultra, Omniverse Libraries, and OpenShell. It enables local AI agent deployment in 30 minutes, locking developers into NVIDIA's hardware-software stack and shifting control from cloud services to on-premises hardware.
测试来源:36氪
...
Alibaba Launches 2.4T Parameter Qwen3.8-Max MoE Model with 0.2x Pricing
Alibaba released Qwen3.8-Max-Preview, a 2.4 trillion parameter MoE multimodal model with 1M context window. It launched Qoder platform and Token Plan with aggressive discounts up to 0.2x, significantly reducing inference cost. The company claims it is second only to Anthropic Fable 5, marking China's AI entry into dual-track of parameter arms race and open-source competition.
PPIO Launches Agentic Cloud, Intelligent Model Gateway Becomes New Control Point
PPIO unveiled Agentic Cloud and Intelligent Model Gateway at WAIC 2026, targeting AI agent workloads with semantic routing and cost-aware scheduling. With over 1.2 trillion daily tokens and sub-200ms sandbox cold start, it signals the emergence of dedicated agent infrastructure.
Palo Alto Networks Launches AI Gateway as Centralized Control Plane for Enterprise AI
Palo Alto Networks announces general availability of AI Gateway, integrating Portkey technology, positioned as the enterprise AI control plane. It unifies LLM, MCP, and A2A gateway execution, processing over 68 trillion tokens with sub-millisecond latency and 99.999% availability.
AMD与OpenAI达成6GW算力供应历史性协议 1.6亿认股权证可获10%股权 股价盘前涨35%
...
Microsoft Replaces OpenAI/Anthropic with In-House MAI Models to Cut Costs and Reduce Dependency
Microsoft has started replacing OpenAI and Anthropic AI calls in Excel and Outlook with its in-house MAI models, handling tens of thousands of prompts weekly. The move aims to cut costs and reduce dependency on Anthropic, signaling a strategic shift toward internal AI models and impacting the AI vendor ecosystem.
SANS Identifies Distributed Scanning of MCP Servers and AI Assistant Configs
SANS Internet Storm Center reports systematic scanning of MCP servers, AI assistant configs, and local LLM endpoints. 49 IPs targeted MCP handshakes, exploiting CVEs in MCP SDKs, signaling AI infrastructure as a new attack vector.
Huawei Ascend 10K-Card Cluster Goes Live, UnifiedBus Protocol Pools All Resources
Huawei launched an Ascend 10,000-card AI cluster in Shaoguan, Guangdong, and showcased the Atlas 950 SuperPoD with its proprietary UnifiedBus interconnect supporting 8,192 NPUs at 16.3 PB/s. Huawei Cloud also entered the Gartner 2026 Cloud AI Infrastructure Leaders quadrant, reinforcing its push for a self-contained AI ecosystem.
AMD's Experimental Topological Ghost Protocol Boosts MI300X Inference 10x
AMD introduces experimental Topological Ghost Protocol (TGP) on MI300X GPUs, achieving 431 tokens/sec with 100% success in high-concurrency inference, 10x improvement over standard vLLM. TGP uses KV-cache recycling and segmented state management, still experimental but potentially redefining AI inference benchmarks.
Google Gemini 3.5 Pro Rebuilds from Scratch: 2M Token Context Window Reshapes AI Frontier
Google DeepMind targets July 17 for Gemini 3.5 Pro, a full architectural rewrite of its pretraining stack to overcome deficits in math reasoning, SVG generation, and image quality. Specs include a 2M token context window, Deep Think reasoning layer, and multi-step autonomous workflows, though unconfirmed by Google.
Anthropic Launches Claude Sonnet 5, Closing Gap to Opus, Targets Enterprise Workflows
Anthropic launches Claude Sonnet 5, a mid-tier model that nearly matches flagship Opus 4.8 on SWE-bench Pro (63.2% vs 69.2%) and surpasses it on GDPval-AA v2 (1618 vs 1615). Priced at 60% of the flagship, it is paired with Claude Science, a research workbench integrating 60+ scientific databases, aiming to deepen enterprise lock-in through tooling and cost-performance.
Qualcomm Enters AI Inference with Dragonfly C1000 CPU and HBC Near-Memory Compute
Qualcomm unveils Dragonfly roadmap with Oryon-based C1000 CPU and AI300 inference accelerator featuring HBC near-memory compute. Meta and Microsoft are early adopters. The strategy targets AI inference TCO reduction and memory wall breakthrough, bypassing Nvidia's training dominance.
Anthropic Launches Sonnet 5: 40% Cost for Near-Opus Performance, Reshaping AI Inference Economics
Anthropic launches Claude Sonnet 5, a mid-range flagship model priced at 40% of Opus 4.8. It scores 63.2% on SWE-bench Pro, approaching Opus's 69.2%, and surpasses Opus on GDPval-AA v2. With native 1M token context and 48B average activated parameters, Sonnet 5 targets high-volume API revenue growth.
Announcing the Monetization Gateway: charge for any resource behind Cloudflare via x402
...