Reports
AI-generated structured vendor updates
OpenAI and Anthropic Jointly Push 30-Day Federal Review for Frontier AI Models
OpenAI and Anthropic propose a 30-day federal review for frontier AI models before release, citing security risks. The August 1 deadline defines coverage thresholds, sparking a policy battle between closed-source advocates and open-source supporters including NVIDIA, Meta, and Microsoft. Chinese open models GLM-5.2 and Kimi K3 are central to the debate.
ASML股价下跌因中国DUV光刻机量产报道冲击芯片设备市场
...
Meta Develops Switchboard AI Model Router to Control Inference Costs and Ecosystem
Meta's internal incubator AAI Labs is developing Switchboard, an AI model router that analyzes task complexity and routes to the most suitable model to minimize inference costs. Initially applied to internal AI coding agents, it may later become a commercial service, positioning Meta as a control layer for AI inference.
Microsoft's Project Perception Automates Vulnerability Remediation with Multi-Model AI Orchestration
In response to competitors like Anthropic and Palo Alto, Microsoft's Project Perception leverages multi-model AI orchestration to automate vulnerability discovery and remediation. This product signifies a transition from manual security operations to autonomous self-healing systems, with control shifting from security analysts to an AI-driven platform.
PrismML's 1-bit Compression: 27B Qwen Model Runs Fully on iPhone 17 Pro in 4GB
PrismML compressed a 27B-parameter dense LLM (Qwen 3.6) to 4GB, running fully on iPhone 17 Pro. Using native 1-bit quantization (weights as {-1, +1}), it achieves >92% compression, 8x faster inference, and 75-80% energy reduction. This challenges Apple's sparse architecture, potentially shifting edge AI from cloud-reliant to device-native.
Towards Feature Complete Triton Support in JAX-Triton â ROCm Blogs
...
微软重构Copilot应用,消费者与企业版合并为单一产品
...
OpenAI Slashes Inference Costs 50%, Runs ChatGPT on Hundreds of GPUs via System-Level Optimization
OpenAI reduces AI inference costs by over 50% through system-level optimizations: model quantization (FP16 to INT4/INT8), KV-Cache optimization, dynamic batching, and speculative decoding. Using only hundreds of NVIDIA GPUs to serve ChatGPT's unlogged-in traffic, inference gross margin jumps from 38% to 65%, nearing breakeven.
NVIDIA AI Compute Partnership: Revenue Share and Credit Backstop to Lock Cloud Providers into DSX AI Factories
NVIDIA launches AI Compute Partnership with revenue sharing and credit backstop, shifting from hardware sales to recurring service revenue. Initial projects include 40K GB300 chips for Sharon AI and 170K GPUs for Firmus, totaling 200K+ high-end chips. NVIDIA is becoming the 'central bank' of AI compute, squeezing cloud brokers.
AWS and Anthropic Ink Token-Based Pricing, Reshaping AI Cloud Economics
Amazon AWS and Anthropic have agreed to a new token-based pricing model, shifting from compute-centric to usage-centric billing for running Anthropic models on AWS. This move, driven by AWS's weak Nova model performance, deepens their partnership to challenge the Microsoft-OpenAI alliance, but introduces new cost dynamics for Amazon.
Cisco Acquires WideField: Injecting Identity Session Intel into Splunk’s Agentic SOC to Win the AI Agent Security Control Plane
Cisco announces intent to acquire WideField Security to embed identity and session intelligence into Splunk's Agentic SOC. The move targets the new security risks from AI agents and non-human identities operating at machine speed, using deterministic data pipelines and session-level signals for evidence-backed autonomous response, strengthening the trust layer within the Cisco Data Fabric.
AMD MLPerf 6.0: MI350 GPUs Achieve 3.5x Leap with MXFP4, Debut Multi-Node Training
AMD submitted its most comprehensive MLPerf Training 6.0 results, including first multi-node training (FLUX.1 on 512 GPUs) and MXFP4 training recipe. MI355X delivers 3.5x generational leap over MI300X on Llama 2-70B, within 5% of NVIDIA B200. 10 ecosystem partners validated reproducibility.
NVIDIA Optimizes Google's DiffusionGemma for 1,000 tok/s Parallel Text Generation
NVIDIA optimizes Google DeepMind's DiffusionGemma, a diffusion-based text model generating 256 tokens per step in parallel. On a single H100, it achieves 1,000 tok/s, with deployment via NIM and NeMo. This breaks the sequential token bottleneck, slashing serving costs and latency for real-time AI.
OpenAI Pivots to Codex: From Chatbot to Agentic Control Plane for Enterprise Automation
OpenAI plans its biggest ChatGPT overhaul, integrating Codex, AI agents, and third-party apps into a super-app. This marks a strategic pivot from a Q&A chatbot to an agentic execution platform, with Codex as the new control plane, aiming to boost enterprise monetization and counter Anthropic's competitive threat.
Google TPU 8t/8i Enables Cross-Datacenter Training, Gemini 3.5 Flash 4x Faster
Google unveils TPU 8t (training) and TPU 8i (inference) with 3x raw compute and 2x perf-per-watt. JAX/Pathways enable distributed training across 1M+ TPUs across sites. Gemini 3.5 Flash delivers 4x output tokens per second vs frontier models. SynthID adopted by OpenAI, Nvidia, Kakao, Eleven Labs.
Cisco Acquires Astrix Security to Strengthen Non-Human Identity and AI Agent Security Control Plane
Cisco announces its intent to acquire Astrix Security, a Non-Human Identity (NHI) security specialist. The goal is to integrate AI agent and credential (API keys, service accounts) security management deeply into Cisco's Identity Intelligence platform and Zero Trust Access solutions. This move signals a shift in the security control plane from traditional human-machine interactions towards securing automated AI agent workloads, addressing the new attack surface created by AI agents abusing credentials.
Cisco Integrates AI with Networking via Vision Portal to Enhance Physical Security Incident Response
Cisco has introduced new software features in its Meraki Vision portal, leveraging AI and cross-camera tracking to deeply integrate smart cameras into the enterprise network management plane. This move aims to transform physical security incident response from passive monitoring to proactive, rapid investigation through a unified cloud management interface.
Cisco Optimizes Developer Portals via Product Sprints, Focusing on AI Agent Workflow Data
Cisco's DevNet team detailed its practice of optimizing developer portals and content through product sprints, focusing on establishing measurable product-market fit indicators. Notably, the newly added analytics events specifically track how developer content is consumed by AI coding assistants or agents, such as copying Markdown and downloading OpenAPI/SDK/MCP documents.
Cisco Announces Galileo Acquisition to Strengthen AI Agent Observability
Cisco plans to acquire Galileo, a startup specializing in AI observability. The move aims to integrate Galileo's AI quality evaluation, failure detection, and guardrail technology into the Splunk Observability Cloud, providing enterprises with full lifecycle visibility and security for their AI agent systems.
Samsung Re-Architects Bixby as an LLM-Core Device Agent
Samsung has re-architected its voice assistant Bixby, shifting from a command-based model to an agentic paradigm with an LLM at its core. The new Bixby understands device context and user intent to autonomously orchestrate device functions and APIs for complex tasks, aiming to become the primary interface for all Samsung products.